The papers the index reads
Every paper a search can return, ranked two ways: by the field's own upvotes, and by when it was published. Each row says who published it and how many open models cite it, which is the one measure of a paper that a citation count does not give you.
read 2026-09-18- 1BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
A 150M-parameter reasoning model using recurrent latent reasoning and in-context learning achieves a new cost-accuracy frontier on ARC-AGI-1.
- 2Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
Vidu S2 introduces real-time interactive avatar and video editing models that support high-resolution spatial video generation and dynamic reference updates.
- 3Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing
Swarm sAmpling Policy Optimization (SAPO) is a decentralized and asynchronous RL algorithm that enhances post-training language models without supervised fine-tuning, achieving significant reward gains and scalability across diverse hardware.
- 4GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning
GrandCode is a multi-agent reinforcement learning system that outperforms human competitors in competitive programming challenges by orchestrating specialized agent modules and employing novel reward policy optimization techniques.
- 5The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
A 1-bit LLM variant, BitNet b1.58, achieves comparable performance to full-precision models with reduced computational costs and introduces new scaling laws and hardware design opportunities.
- 6The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
BDH, a biologically inspired Large Language Model, combines scale-free network architecture with Hebbian learning to achieve Transformer-like performance while maintaining interpretability.
- 7Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
DisCo is a research agent that distills operational knowledge into reusable skills, significantly improving autonomous ML research performance across benchmarks.
- 8A Very Big Video Reasoning Suite
A large-scale video reasoning dataset and benchmark are introduced to study video intelligence capabilities beyond visual quality, enabling systematic analysis of spatiotemporal reasoning and generalization across diverse tasks.
- 9Kimi K3: Open Frontier Intelligence
Kimi K3 is a large-scale mixture-of-experts model with native vision and long-context capabilities that improves scaling efficiency and achieves strong performance across coding, reasoning, and agentic tasks.
- 10Less is More: Recursive Reasoning with Tiny Networks
Tiny Recursive Model (TRM) achieves high generalization on complex puzzle tasks using a small, two-layer network with minimal parameters, outperforming larger language models.
- 11Adam's Law: Textual Frequency Law on Large Language Models
A novel framework for improving large language model performance through textual frequency analysis, including laws, distillation, and curriculum training approaches.
- 12Orca: The World is in Your Mind
Orca establishes a unified world latent space through next-state-prediction modeling using multimodal data and demonstrates superior performance in downstream tasks compared to specialized baselines.
- 13Continued domain-specific pre-training of protein language models for pMHC-I binding prediction
Continued pre-training of protein language models on HLA-associated peptides improves pMHC-I binding affinity prediction, especially for underrepresented alleles.
- 14StudentSim: Training LLM-based Student Simulators
StudentSim trains personalized student simulators from sparse data to mirror learner responses and adapt to tutor guidance, outperforming existing models across chess, writing, and math.
- 15ABot-Earth 0.5: Generative 3D Earth Model
ABot-Earth 0.5 generates realistic 3D environments from satellite imagery using 3D Gaussian Splatting representation, enabling fast synthesis and real-time visualization for Embodied AI applications.
- 16Looped World Models
Looped World Models introduce iterative latent state refinement through shared transformer blocks, achieving 100x parameter efficiency while adapting computational depth to prediction complexity.
- 17DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-R1-Zero and DeepSeek-R1 utilize reinforcement learning and multi-stage training to enhance reasoning capabilities, with DeepSeek-R1 achieving performance comparable to OpenAI-o1-1217.
- 18Scaling Automatic Research Agents via World Models
World Model RL replaces costly environment execution with a learned world model and applies debiasing and denoising to accelerate post-training of autonomous research agents.
- 19StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weights.
- 20Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
A generative multi-agent world model is presented that uses simplex rotary agent encoding and sparse hub attention to enable scalable, permutation-symmetric interaction between multiple agents in interactive video generation.
- 21AI Can Learn Scientific Taste
Great scientists have strong judgement and foresight, closely tied to what we call scientific taste. Here, we use the term to refer to the capacity to judge and propose research ideas with high potential impact. However, most relative research focuses on improving an AI scientist's executive capability, while enhancing an AI's scientific taste remains underexplored. In this work, we propose Reinforcement Learning from Community Feedback (RLCF), a training paradigm that uses large-scale community signals as supervision, and formulate scientific taste learning as a preference modeling and alignment problem. For preference modeling, we train Scientific Judge on 700K field- and time-matched pairs of high- vs. low-citation papers to judge ideas. For preference alignment, using Scientific Judge as a reward model, we train a policy model, Scientific Thinker, to propose research ideas with high potential impact. Experiments show Scientific Judge outperforms SOTA LLMs (e.g., GPT-5.2, Gemini 3 Pro) and generalizes to future-year test, unseen fields, and peer-review preference. Furthermore, Scientific Thinker proposes research ideas with higher potential impact than baselines. Our findings show that AI can learn scientific taste, marking a key step toward reaching human-level AI scientists.
- 22Atria Dawn: The Dawn of Agentic Superintelligence
Atria Dawn Preview is a foundation agentic language model trained through verified tool interactions that achieves strong benchmark results and demonstrates a shift toward human-AI project-level collaboration in scientific research.
- 23NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
NeoHorse-1 uses agentic post-training with intelligent routing, structured feedback loops, and curriculum-based distillation to improve model capabilities across agent benchmarks.
- 24Agents' Last Exam
Agents' Last Exam (ALE) is a benchmark for evaluating AI agents on long-term, economically valuable real-world tasks across 13 industry clusters with 1K+ tasks, revealing significant gaps between benchmark performance and practical deployment.
- 25Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
Qwen-Drive-1.0 is a vision-language foundation model for autonomous driving that unifies 3D perception, visual question answering, and motion planning via shared representations and staged training.
- 26Compile by Training: Turning Natural-Language Specifications into Local Neural Functions
Compile by training converts natural-language specifications into reusable neural functions by distilling teacher-generated examples into small adapters, enabling efficient deployment without remote model dependencies.
- 27Qwen2.5 Technical Report
Qwen2.5, an enhanced series of large language models, demonstrates superior performance across various benchmarks and use cases through extensive pre-training and advanced post-training techniques.
- 28Demystifing Video Reasoning
Diffusion-based video models demonstrate reasoning capabilities through denoising steps rather than frame sequences, exhibiting behaviors like working memory, self-correction, and perception-before-action within specialized transformer layers.
- 29DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
DataFlex is a unified framework for dynamic data-centric training of large language models that supports sample selection, domain mixture adjustment, and sample reweighting while maintaining compatibility with standard training workflows and enabling efficient large-scale deployment.
- 30MolmoAct2: Action Reasoning Models for Real-world Deployment
MolmoAct2 presents an open-action reasoning model for robotics that improves upon previous systems through specialized vision-language-model backbones, new datasets, open-weight action tokenizers, architectural redesign for continuous-action prediction, and adaptive reasoning for reduced latency.
- 31OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
OPUS is a dynamic data selection framework that improves pre-training efficiency by scoring data candidates based on optimizer-induced update projections in a stable proxy-derived target space, achieving superior performance with reduced computational overhead.
- 32FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization
FIPO enhances reinforcement learning for language models by using discounted future-KL divergence to improve credit assignment and extend reasoning chains, achieving better mathematical problem-solving performance.
- 33A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
A.S.E is a repository-level benchmark for evaluating the security of AI-generated code, highlighting challenges in secure coding and the limitations of LLMs in real-world scenarios.
- 34Qwen3 Technical Report
Qwen3, a unified series of large language models, integrates thinking and non-thinking modes, reduces computational resources, and achieves state-of-the-art performance across various tasks and languages.
- 35CARLA-Air: Fly Drones Inside a CARLA World -- A Unified Infrastructure for Air-Ground Embodied Intelligence
CARLA-Air integrates high-fidelity driving and multirotor flight simulation within a unified Unreal Engine framework, supporting joint air-ground agent modeling with photorealistic environments and multi-modal sensing capabilities.
- 36Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
Macaron-V1 is an open agent-model family that uses a Mixture-of-LoRA architecture and recursive self-improvement to enable continual learning and collaboration across specialized tasks.
- 37HarnessEval-W: Agentifying the Evaluation of Visual Worlds
HarnessEval-W uses hierarchical sub-agents to decompose world-model evaluations into verifiable reasoning chains that justify scores with transparent evidence.
- 38mHC: Manifold-Constrained Hyper-Connections
Manifold-Constrained Hyper-Connections (mHC) stabilize and scale residual connection architectures by restoring identity mapping properties through manifold projection and infrastructure optimization.
- 39Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
Supervised finetuning and reinforcement learning exhibit conditional cross-domain generalization in reasoning tasks, influenced by optimization dynamics, data quality, and model capability, with asymmetric outcomes between reasoning improvement and safety degradation.
- 40HRM-Text: Efficient Pretraining Beyond Scaling
A Hierarchical Recurrent Model architecture with specialized training on instruction-response pairs achieves competitive language modeling performance with significantly reduced computational requirements compared to traditional Transformer-based approaches.