Research

Towards rigorous, reliable, and capable AI scientists

  1. The Last Human-Written Paper: Agent-Native Research Artifacts

    What if a paper were something an AI could run and verify, not just read?

    NeurIPS · 2026

    NeurIPS · 2026ARA protocol

    The Last Human-Written Paper: Agent-Native Research Artifacts

    What if a paper were something an AI could run and verify, not just read?

  2. Open-Endedness Bench: Measuring Epistemic Process from Agent Records

    If every step of an agent's research is on record, what does reading the record reveal?

    arXiv · 2026

    Open-Endedness Bench: Measuring Epistemic Process from Agent Records

    If every step of an agent's research is on record, what does reading the record reveal?

  3. Beyond Natural Language: An Agent-Native Language for Autonomous Science

    What if every claim in a paper came with an argument a checker could accept, reject, or flag as missing a piece?

    arXiv · 2026

    arXiv · 2026

    Beyond Natural Language: An Agent-Native Language for Autonomous Science

    What if every claim in a paper came with an argument a checker could accept, reject, or flag as missing a piece?

  4. The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

    What has to be true before a system can turn its own experience into a better version of itself?

    arXiv · 2026

    arXiv · 2026

    The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

    What has to be true before a system can turn its own experience into a better version of itself?

  5. The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing

    Discovery is as sparse a signal as a crash, so what plays the role of coverage for an AI scientist?

    arXiv · 2026

    The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing

    Discovery is as sparse a signal as a crash, so what plays the role of coverage for an AI scientist?

  6. Sci-Reasoning: A Dataset Decoding AI Innovation Patterns

    How do the best researchers actually arrive at their breakthroughs, and could an AI learn to do the same?

    arXiv · 2026

    arXiv · 2026

    Sci-Reasoning: A Dataset Decoding AI Innovation Patterns

    How do the best researchers actually arrive at their breakthroughs, and could an AI learn to do the same?

  7. EXP-Bench: Can AI Conduct AI Research Experiments?

    Today's agents can design and code an experiment, so why can almost none of them finish one?

    ICLR · 2026

    ICLR · 2026

    EXP-Bench: Can AI Conduct AI Research Experiments?

    Today's agents can design and code an experiment, so why can almost none of them finish one?

  8. Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents

    How do you make an AI scientist's experiments rigorous enough to trust the results?

    arXiv · 2025

    arXiv · 2025

    Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents

    How do you make an AI scientist's experiments rigorous enough to trust the results?