BigHugger
Research · papers

The papers the index reads

Every paper a search can return, ranked two ways: by the field's own upvotes, and by when it was published. Each row says who published it and how many open models cite it, which is the one measure of a paper that a citation count does not give you.

read 2026-09-18
  1. 1
    BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

    Pathway2026-08-10782 upvotes

    A 150M-parameter reasoning model using recurrent latent reasoning and in-context learning achieves a new cost-accuracy frontier on ARC-AGI-1.

  2. 2
    Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

    Tsinghua University2026-09-10694 upvotes

    Vidu S2 introduces real-time interactive avatar and video editing models that support high-resolution spatial video generation and dynamic reference updates.

  3. 3
    Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing

    Gensyn2025-09-10664 upvotes

    Swarm sAmpling Policy Optimization (SAPO) is a decentralized and asynchronous RL algorithm that enhances post-training language models without supervised fine-tuning, achieving significant reward gains and scalability across diverse hardware.

  4. 4
    GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning

    Ornith2026-04-03641 upvotes4 models cite it

    GrandCode is a multi-agent reinforcement learning system that outperforms human competitors in competitive programming challenges by orchestrating specialized agent modules and employing novel reward policy optimization techniques.

  5. 5
    The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

    Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang2024-02-27630 upvotes56 models cite it

    A 1-bit LLM variant, BitNet b1.58, achieves comparable performance to full-precision models with reduced computational costs and introduces new scaling laws and hardware design opportunities.

  6. 6
    The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain

    Pathway2025-09-30555 upvotes9 models cite it

    BDH, a biologically inspired Large Language Model, combines scale-free network architecture with Hebbian learning to achieve Transformer-like performance while maintaining interpretability.

  7. 7
    Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

    Beijing Academy of Artificial Intelligence2026-09-02553 upvotes

    DisCo is a research agent that distills operational knowledge into reusable skills, significantly improving autonomous ML research performance across benchmarks.

  8. 8
    A Very Big Video Reasoning Suite

    Video-Reason2026-02-23528 upvotes11 models cite it

    A large-scale video reasoning dataset and benchmark are introduced to study video intelligence capabilities beyond visual quality, enabling systematic analysis of spatiotemporal reasoning and generalization across diverse tasks.

  9. 9
    Kimi K3: Open Frontier Intelligence

    Moonshot AI2026-07-27523 upvotes2 models cite it

    Kimi K3 is a large-scale mixture-of-experts model with native vision and long-context capabilities that improves scaling efficiency and achieves strong performance across coding, reasoning, and agentic tasks.

  10. 10
    Less is More: Recursive Reasoning with Tiny Networks

    Samsung AI Lab (SAIL) Montreal2025-10-06521 upvotes8 models cite it

    Tiny Recursive Model (TRM) achieves high generalization on complex puzzle tasks using a small, two-layer network with minimal parameters, outperforming larger language models.

  11. 11
    Adam's Law: Textual Frequency Law on Large Language Models

    FaceMind2026-04-02511 upvotes

    A novel framework for improving large language model performance through textual frequency analysis, including laws, distillation, and curriculum training approaches.

  12. 12
    Orca: The World is in Your Mind

    Yihao Wang, Yuheng Ji, Mingyu Cao, Yanqing Shen2026-06-29508 upvotes1 models cite it

    Orca establishes a unified world latent space through next-state-prediction modeling using multimodal data and demonstrates superior performance in downstream tasks compared to specialized baselines.

  13. 13
    Continued domain-specific pre-training of protein language models for pMHC-I binding prediction

    Sergio E. Mares, Ariel Espinoza Weinberger, Nilah M. Ioannidis2025-07-16504 upvotes1 models cite it

    Continued pre-training of protein language models on HLA-associated peptides improves pMHC-I binding affinity prediction, especially for underrepresented alleles.

  14. 14
    StudentSim: Training LLM-based Student Simulators

    Microsoft Research2026-09-01492 upvotes

    StudentSim trains personalized student simulators from sparse data to mirror learner responses and adapt to tutor guidance, outperforming existing models across chess, writing, and math.

  15. 15
    ABot-Earth 0.5: Generative 3D Earth Model

    Alibaba AMAP CV Lab2026-06-08489 upvotes

    ABot-Earth 0.5 generates realistic 3D environments from satellite imagery using 3D Gaussian Splatting representation, enabling fast synthesis and real-time visualization for Embodied AI applications.

  16. 16
    Looped World Models

    FaceMind2026-06-16488 upvotes

    Looped World Models introduce iterative latent state refinement through shared transformer blocks, achieving 100x parameter efficiency while adapting computational depth to prediction complexity.

  17. 17
    DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

    DeepSeek2025-01-22463 upvotes524 models cite it

    DeepSeek-R1-Zero and DeepSeek-R1 utilize reinforcement learning and multi-stage training to enhance reasoning capabilities, with DeepSeek-R1 achieving performance comparable to OpenAI-o1-1217.

  18. 18
    Scaling Automatic Research Agents via World Models

    University of Illinois at Urbana-Champaign2026-08-29456 upvotes

    World Model RL replaces costly environment execution with a learned world model and applies debiasing and denoising to accelerate post-training of autonomous research agents.

  19. 19
    StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

    Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang, Kai Wang2026-08-15447 upvotes

    StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weights.

  20. 20
    Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players

    NVIDIA2026-05-27434 upvotes1 models cite it

    A generative multi-agent world model is presented that uses simplex rotary agent encoding and sparse hub attention to enable scalable, permutation-symmetric interaction between multiple agents in interactive video generation.

  21. 21
    AI Can Learn Scientific Taste

    OpenMOSS2026-03-15432 upvotes6 models cite it

    Great scientists have strong judgement and foresight, closely tied to what we call scientific taste. Here, we use the term to refer to the capacity to judge and propose research ideas with high potential impact. However, most relative research focuses on improving an AI scientist's executive capability, while enhancing an AI's scientific taste remains underexplored. In this work, we propose Reinforcement Learning from Community Feedback (RLCF), a training paradigm that uses large-scale community signals as supervision, and formulate scientific taste learning as a preference modeling and alignment problem. For preference modeling, we train Scientific Judge on 700K field- and time-matched pairs of high- vs. low-citation papers to judge ideas. For preference alignment, using Scientific Judge as a reward model, we train a policy model, Scientific Thinker, to propose research ideas with high potential impact. Experiments show Scientific Judge outperforms SOTA LLMs (e.g., GPT-5.2, Gemini 3 Pro) and generalizes to future-year test, unseen fields, and peer-review preference. Furthermore, Scientific Thinker proposes research ideas with higher potential impact than baselines. Our findings show that AI can learn scientific taste, marking a key step toward reaching human-level AI scientists.

  22. 22
    Atria Dawn: The Dawn of Agentic Superintelligence

    Intern Large Models2026-09-14422 upvotes4 models cite it

    Atria Dawn Preview is a foundation agentic language model trained through verified tool interactions that achieves strong benchmark results and demonstrates a shift toward human-AI project-level collaboration in scientific research.

  23. 23
    NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

    TokenRhythm2026-09-08418 upvotes16 models cite it

    NeoHorse-1 uses agentic post-training with intelligent routing, structured feedback loops, and curriculum-based distillation to improve model capabilities across agent benchmarks.

  24. 24
    Agents' Last Exam

    UC Berkeley2026-06-03392 upvotes

    Agents' Last Exam (ALE) is a benchmark for evaluating AI agents on long-term, economically valuable real-world tasks across 13 industry clusters with 1K+ tasks, revealing significant gaps between benchmark performance and practical deployment.

  25. 25
    Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

    Qwen2026-08-31386 upvotes6 models cite it

    Qwen-Drive-1.0 is a vision-language foundation model for autonomous driving that unifies 3D perception, visual question answering, and motion planning via shared representations and staged training.

  26. 26
    Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

    University of Waterloo2026-09-03385 upvotes

    Compile by training converts natural-language specifications into reusable neural functions by distilling teacher-generated examples into small adapters, enabling efficient deployment without remote model dependencies.

  27. 27
    Qwen2.5 Technical Report

    Qwen, An Yang, Baosong Yang, Beichen Zhang2024-12-19381 upvotes124 models cite it

    Qwen2.5, an enhanced series of large language models, demonstrates superior performance across various benchmarks and use cases through extensive pre-training and advanced post-training techniques.

  28. 28
    Demystifing Video Reasoning

    SenseNova2026-03-17373 upvotes

    Diffusion-based video models demonstrate reasoning capabilities through denoising steps rather than frame sequences, exhibiting behaviors like working memory, self-correction, and perception-before-action within specialized transformer layers.

  29. 29
    DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models

    Peking University2026-03-27366 upvotes

    DataFlex is a unified framework for dynamic data-centric training of large language models that supports sample selection, domain mixture adjustment, and sample reweighting while maintaining compatibility with standard training workflows and enabling efficient large-scale deployment.

  30. 30
    MolmoAct2: Action Reasoning Models for Real-world Deployment

    Ai22026-05-04357 upvotes34 models cite it

    MolmoAct2 presents an open-action reasoning model for robotics that improves upon previous systems through specialized vision-language-model backbones, new datasets, open-weight action tokenizers, architectural redesign for continuous-action prediction, and adaptive reasoning for reduced latency.

  31. 31
    OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration

    Qwen2026-02-05356 upvotes

    OPUS is a dynamic data selection framework that improves pre-training efficiency by scoring data candidates based on optimizer-induced update projections in a stable proxy-derived target space, achieving superior performance with reduced computational overhead.

  32. 32
    FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization

    Qwen2026-03-20354 upvotes1 models cite it

    FIPO enhances reinforcement learning for language models by using discounted future-KL divergence to improve credit assignment and extend reasoning chains, achieving better mathematical problem-solving performance.

  33. 33
    A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code

    Keke Lian, Bin Wang, Lei Zhang, Libo Chen2025-08-25350 upvotes

    A.S.E is a repository-level benchmark for evaluating the security of AI-generated code, highlighting challenges in secure coding and the limitations of LLMs in real-world scenarios.

  34. 34
    Qwen3 Technical Report

    Qwen2025-05-14346 upvotes2,400 models cite it

    Qwen3, a unified series of large language models, integrates thinking and non-thinking modes, reduces computational resources, and achieves state-of-the-art performance across various tasks and languages.

  35. 35
    CARLA-Air: Fly Drones Inside a CARLA World -- A Unified Infrastructure for Air-Ground Embodied Intelligence

    Tianle Zeng, Hanxuan Chen, Yanci Wen, Hong Zhang2026-03-30346 upvotes1 models cite it

    CARLA-Air integrates high-fidelity driving and multirotor flight simulation within a unified Unreal Engine framework, supporting joint air-ground agent modeling with photorealistic environments and multi-modal sensing capabilities.

  36. 36
    Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

    Mind Lab2026-08-10345 upvotes5 models cite it

    Macaron-V1 is an open agent-model family that uses a Mixture-of-LoRA architecture and recursive self-improvement to enable continual learning and collaboration across specialized tasks.

  37. 37
    HarnessEval-W: Agentifying the Evaluation of Visual Worlds

    MirroS2026-08-17341 upvotes

    HarnessEval-W uses hierarchical sub-agents to decompose world-model evaluations into verifiable reasoning chains that justify scores with transparent evidence.

  38. 38
    mHC: Manifold-Constrained Hyper-Connections

    DeepSeek2025-12-31336 upvotes2 models cite it

    Manifold-Constrained Hyper-Connections (mHC) stabilize and scale residual connection architectures by restoring identity mapping properties through manifold projection and infrastructure optimization.

  39. 39
    Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

    AI45Research2026-04-08330 upvotes61 models cite it

    Supervised finetuning and reinforcement learning exhibit conditional cross-domain generalization in reasoning tasks, influenced by optimization dynamics, data quality, and model capability, with asymmetric outcomes between reasoning improvement and safety degradation.

  40. 40
    HRM-Text: Efficient Pretraining Beyond Scaling

    Sapient AI2026-05-20325 upvotes12 models cite it

    A Hierarchical Recurrent Model architecture with specialized training on instruction-response pairs achieves competitive language modeling performance with significantly reduced computational requirements compared to traditional Transformer-based approaches.