Research
Towards rigorous, reliable, and capable AI scientists
The Last Human-Written Paper: Agent-Native Research Artifacts
What if a paper were something an AI could run and verify, not just read?
The Last Human-Written Paper: Agent-Native Research Artifacts
What if a paper were something an AI could run and verify, not just read?
Open-Endedness Bench: Measuring Epistemic Process from Agent Records
If every step of an agent's research is on record, what does reading the record reveal?
Open-Endedness Bench: Measuring Epistemic Process from Agent Records
If every step of an agent's research is on record, what does reading the record reveal?
Beyond Natural Language: An Agent-Native Language for Autonomous Science
What if every claim in a paper came with an argument a checker could accept, reject, or flag as missing a piece?
Beyond Natural Language: An Agent-Native Language for Autonomous Science
What if every claim in a paper came with an argument a checker could accept, reject, or flag as missing a piece?
The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
What has to be true before a system can turn its own experience into a better version of itself?
The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
What has to be true before a system can turn its own experience into a better version of itself?
The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing
Discovery is as sparse a signal as a crash, so what plays the role of coverage for an AI scientist?
The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing
Discovery is as sparse a signal as a crash, so what plays the role of coverage for an AI scientist?
Sci-Reasoning: A Dataset Decoding AI Innovation Patterns
How do the best researchers actually arrive at their breakthroughs, and could an AI learn to do the same?
Sci-Reasoning: A Dataset Decoding AI Innovation Patterns
How do the best researchers actually arrive at their breakthroughs, and could an AI learn to do the same?
EXP-Bench: Can AI Conduct AI Research Experiments?
Today's agents can design and code an experiment, so why can almost none of them finish one?
EXP-Bench: Can AI Conduct AI Research Experiments?
Today's agents can design and code an experiment, so why can almost none of them finish one?
Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents
How do you make an AI scientist's experiments rigorous enough to trust the results?
Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents
How do you make an AI scientist's experiments rigorous enough to trust the results?








