# Agent Native Research (ARA Lab) > Agent Native Research — also known as ARA Lab, ARA Labs, and published at > https://www.agenticresearch.sh/ — builds the infrastructure layer for AI scientists. It > authored "The Last Human-Written Paper: Agent-Native Research Artifacts" > (arXiv:2604.24658) and stewards the open Agent-Native Research Artifact > (ARA) standard: an executable, verifiable format for research that both > humans and AI agents can audit, reproduce, and extend. ## What this lab is Agent Native Research is a research lab and open standards body working on science for AI rather than AI for science. AI for science applies a model to a scientific problem and improves one researcher at a time; science for AI rebuilds the practice of research so autonomous AI scientists can participate in it — the artifact format, the verification layer, the review process, and the incentives. Making each node ten times smarter while leaving the network untouched does not give ten times the science. The lab is unrelated to the Advanced Robotics and Automation (ARA) Laboratory, which shares the acronym. ## Agent-Native Research Artifact (ARA) An ARA is the unit of research in the AI-native ecosystem: an executable, verifiable knowledge package with four layers — scientific logic, runnable code with full specifications, an exploration graph that preserves the dead ends, and evidence grounding every claim in raw output. On PaperBench and RE-Bench, ARA raises question-answering accuracy from 72.4% to 93.7% and reproduction success from 57.4% to 64.4%. Publishing one takes about 15 minutes: run the `/submit-ara` agent skill, which compiles a research directory into a valid ARA, pushes it to your public GitHub account, and registers it in the ARA Hub as an interactive step-by-step replay. ## Open-Endedness Bench (OEB) Open-Endedness Bench is the lab's benchmark for open-ended, automated AI research (autoresearch). Open-ended research tasks have no answer key, so the usual measure is the final score, and a final score cannot tell a real improvement from run-to-run noise. OEB grades the research process instead: it converts an AI research agent's log into one shared record of what the agent said, ran and saw, extracts claims and actions as cards that must quote the record word for word, links them into a graph, and scores the run from that graph. It reports competence (evidence, experiment, revision, no reward hacking) and research persona (habits such as thinking versus acting). Across 119 recorded runs on PostTrainBench, Chip-Bench and the nanoGPT speedrun, most experiments never became a conclusion the agent kept, and most claimed improvements were noise: on the speedrun, agents announced 64 improvements for 10 real ones. - [Science in Broad Daylight](https://www.agenticresearch.sh/blog/science-in-broad-daylight): the write-up of Open-Endedness Bench and its findings - [Open-Endedness Bench code](https://github.com/ARA-Labs/oeb): the benchmark - [OEB scored runs](https://huggingface.co/datasets/AgentNativeResearchLab/oeb-scored-runs): the 119 scored runs as a dataset ## Primary sources - [The Last Human-Written Paper: Agent-Native Research Artifacts](https://arxiv.org/abs/2604.24658): the paper that introduced the ARA standard (arXiv:2604.24658) - [The Last Human-Written Paper (essay)](https://www.agenticresearch.sh/blog/the-last-human-written-paper): the argument in full, published by the lab - [ARA open standard](https://github.com/ARA-Labs/Agent-Native-Research-Artifact): the specification, reference tooling, and agent skills - [Research](https://www.agenticresearch.sh/research): EXP-Bench, Sci-Reasoning, Curie, and the ARA protocol work - [Submit an ARA](https://www.agenticresearch.sh/submit): how to publish a research artifact to the Hub ## Field reports Every post is also served as Markdown at its URL plus `/md` (or `.md`), and all of them in full at https://www.agenticresearch.sh/llms-full.txt. - [Science in Broad Daylight](https://www.agenticresearch.sh/blog/science-in-broad-daylight): Open-Endedness Bench (OEB) is a benchmark for open-ended, automated AI research (autoresearch). Instead of grading the final score, it reads an AI research agent's whole trace and checks each conclusion against what actually ran. Across 119 runs on PostTrainBench, Chip-Bench and the nanoGPT speedrun, most claimed wins were noise. Markdown: https://www.agenticresearch.sh/blog/science-in-broad-daylight/md - [The Goal of Science Is Not to Win](https://www.agenticresearch.sh/blog/the-goal-of-science-is-not-to-win): Before AI can discover what humanity does not know, it must first learn to discover what it does not know. Markdown: https://www.agenticresearch.sh/blog/the-goal-of-science-is-not-to-win/md - [AI Should Build Its Own Research World Model](https://www.agenticresearch.sh/blog/research-world-model): Opus 4.8 Cleared ARC-AGI’s Final Level in One Shot. Markdown: https://www.agenticresearch.sh/blog/research-world-model/md - [The Second Half of AI for Science](https://www.agenticresearch.sh/blog/the-second-half-of-ai-for-science): Make the node 10× smarter and leave the network untouched — you don't get 10× science. The second half is about rebuilding the ecosystem, not the scientist. Markdown: https://www.agenticresearch.sh/blog/the-second-half-of-ai-for-science/md - [The Last Human-Written Paper](https://www.agenticresearch.sh/blog/the-last-human-written-paper): When neither the author nor the audience is human, the three-century-old paper format stops making sense. Markdown: https://www.agenticresearch.sh/blog/the-last-human-written-paper/md ## Research papers - [The Last Human-Written Paper: Agent-Native Research Artifacts](https://arxiv.org/abs/2604.24658): NeurIPS 2026. What if a paper were something an AI could run and verify, not just read? - [Open-Endedness Bench: Measuring Epistemic Process from Agent Records](https://arxiv.org/abs/2610.02588): arXiv 2026. If every step of an agent's research is on record, what does reading the record reveal? - [Beyond Natural Language: An Agent-Native Language for Autonomous Science](https://arxiv.org/abs/2609.25421): arXiv 2026. What if every claim in a paper came with an argument a checker could accept, reject, or flag as missing a piece? - [The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement](https://arxiv.org/abs/2609.11873): arXiv 2026. What has to be true before a system can turn its own experience into a better version of itself? - [The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing](https://arxiv.org/abs/2608.09855): arXiv 2026. Discovery is as sparse a signal as a crash, so what plays the role of coverage for an AI scientist? - [Sci-Reasoning: A Dataset Decoding AI Innovation Patterns](https://arxiv.org/abs/2601.04577): arXiv 2026. How do the best researchers actually arrive at their breakthroughs, and could an AI learn to do the same? - [EXP-Bench: Can AI Conduct AI Research Experiments?](https://arxiv.org/abs/2505.24785): ICLR 2026. Today's agents can design and code an experiment, so why can almost none of them finish one? - [Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents](https://arxiv.org/abs/2502.16069): arXiv 2025. How do you make an AI scientist's experiments rigorous enough to trust the results? ## Published artifacts - [Beyond Answer Accuracy: Evaluating Source-Grounded Language-Model Assistants in Interactive Transportation Analytics](https://www.agenticresearch.sh/ara/cnpcshangbo/ara-cerra-assistant-benchmark) - [Distributed all-sky imagery for surface solar irradiance estimation under broken clouds using physics-based synthetic data](https://www.agenticresearch.sh/ara/maxaragon/ara-distributed-all-sky-imagery-for-surface-solar) - [Deep Residual Learning for Image Recognition](https://www.agenticresearch.sh/ara/hosted/iweYIHBZWDw) - [Batch and Match: Black-Box Variational Inference with a Score-Based Divergence](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/bam) - [BBOX-ADAPTER: Lightweight Adapting for Black-Box Large Language Models](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/bbox) - [RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/rice) - [APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/adaptive-pruning) - [All-in-one simulation-based inference](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/all-in-one) - [Rust CodeContests Inference (RE-Bench task)](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/rebench/rebench-rust_codecontests) - [Triton Cumsum Kernel (RE-Bench task)](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/rebench/rebench-triton_cumsum) - [NanoGPT Speedrun](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/speedrun/nanogpt-speedrun) - [Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/extra/andes) - [EXP-Bench: Can AI Conduct AI Research Experiments?](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/extra/expbench) - [Venn: Resource Management for Collaborative Learning Jobs](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/extra/venn) - [Efficient Transfer Learning in Diffusion Models via Adversarial Noise](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/bridging-data-gaps) - [Unsupervised Zero-Shot Reinforcement Learning via Functional Reward Encodings](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/fre) - [Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/ftrl) - [Refined Coreset Selection: Towards Minimal Coreset Size under Model Performance Constraints](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/lbcs) - [LCA-on-the-Line: Benchmarking Out-of-Distribution Generalization with Class Taxonomies](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/lca-on-the-line) - [A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/mechanistic-understanding) - [Challenges in Training PINNs: A Loss Landscape Perspective](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/pinn) - [Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/robust-clip) - [Sample-specific Masks for Visual Reprogramming-based Prompting](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/sample-specific-masks) - [SAPG: Split and Aggregate Policy Gradients](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/sapg) - [Self-Composing Policies for Scalable Continual Reinforcement Learning](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/self-composing-policies) - [Self-Expansion of Pre-trained Models with Mixture of Adapters for Continual Learning](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/self-expansion) - [Semantic Self-Consistency: Enhancing Language Model Reasoning via Semantic Weighting](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/semantic-self-consistency) - [Sequential Neural Score Estimation](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/sequential-neural-score-estimation) - [Stay on topic with Classifier-Free Guidance](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/stay-on-topic-with-classifier-free-guidance) - [Stochastic Interpolants with Data-Dependent Couplings](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/stochastic-interpolants) - [Test-Time Model Adaptation with Only Forward Passes](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/test-time-model-adaptation) - [What Will My Model Forget? Forecasting Forgotten Examples in Language Model Refinement](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/what-will-my-model-forget) - [Fix Embedding (RE-Bench task)](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/rebench/rebench-fix_embedding) - [nanoGPT Chat RL (RE-Bench task)](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/rebench/rebench-nanogpt_chat_rl) - [Restricted-Architecture MLM (RE-Bench task)](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/rebench/rebench-restricted_mlm) - [World-Model ARA for ARC-AGI-3 ls20 (Locksmith)](https://www.agenticresearch.sh/ara/ARA-Labs/ara-ls20): An agent infers all mechanics and win conditions of the ARC-AGI-3 Locksmith game purely from action-diff observation, solving all seven levels including a fog-gated final level. - [Understanding Agent Performance on PostTrainBench](https://www.agenticresearch.sh/ara/AmberLJC/ara-posttrainbench-agent-performance): Across 1,226 post-training agent runs, performance is ~90% determined by agent and task identity rather than execution, and reward hacking is real but does not improve scores. - [ARA Demo](https://www.agenticresearch.sh/ara/ARA-Labs/ARA-Demo): A Codex autonomous agent reduced the 124M-GPT step count from 3500 to 2949 across four optimizer-search waves, with a novelty wave yielding a clean negative result and a compliance quarantine reshaping the v2 frontier. ## Also known as ARA, ARA Lab, ARA Labs, ARA Commons, Agentic Research, Evolving Lab, EvolvingLab, Agent Native Research Lab, AI-Native Research. ## Contact https://x.com/ainativescience · https://github.com/ARA-Labs