Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Aug 28, 2026 07:06 AM - 2 hours ago 4

Terminal-Bench-Science evaluates AI agents connected workflows from researchers' ain work. Scientists, not exemplary developers aliases information vendors, group the barroom for technological capacity successful AI.

Terminal-Bench-Science is simply a benchmark led by researchers astatine Stanford University and built by the squad down Terminal-Bench successful collaboration pinch domain experts from a scope of technological disciplines and investigation institutions astir the world. It measures the AI supplier capabilities done a divers group of challenging, expert-curated workflows drawn from technological research.

Terminal-Bench-Science is simply a continuous benchmark that evolves alongside frontier AI, creating a feedback loop betwixt technological needs and AI development. Our first merchandise includes 70 tasks from the life, physical, Earth, mathematical, and engineering sciences. The strongest exemplary evaluated, Claude Opus 5, achieves a 30% solution complaint connected Terminal-Bench-Science 0.1.

Terminal-Bench-Science 0.1 Leaderboard

Terminal-Bench-ScienceTerminal-Bench-Science

0%25%50%75%100%Claude Opus 5Claude Code30.0%GPT-5.6 SolCodex22.4%Claude Fable 5Claude Code21.4%Claude Opus 4.8Claude Code10.5%GPT-5.6 TerraCodex8.6%GLM 5.3Claude Code8.1%Kimi K3Claude Code7.1%Grok 4.6Grok Build7.1%GPT-5.6 LunaCodex3.3%

Resolution rates crossed 70 technological workflow tasks connected Terminal-Bench-Science 0.1

While Terminal-Bench has driven advancement successful AI agents for package engineering, Terminal-Bench-Science brings the aforesaid ambition to science. Our extremity is to thrust the improvement of agents pinch technological capabilities that make them useful investigation assistants. These agents should execute technically demanding and time-consuming workflows, freeing scientists to attraction much of their clip connected the parts of subject wherever quality judgement matters most: defining investigation questions, forming hypotheses, interpreting and validating results, and communicating findings. In this role, AI agents tin widen what researchers execute and thief accelerate technological discovery.

Achieving this requires benchmarks that bespeak existent technological practice, supply verifiable grounds of capability, and germinate alongside the AI frontier.

We request benchmarks drawn from existent technological workflows. Scientific capacity should beryllium evaluated connected existent investigation practice, not textbook questions aliases standardized exercises, contributed by practicing scientists themselves. Terminal-Bench-Science gives scientists crossed domains a nonstop sound and a shared level to group the barroom for AI advancement connected the problems they attraction about. The stakes successful subject are excessively high, and its benchmarks must bespeak the technological community's priorities alternatively than extracurricular interests.

We request verifiable grounds of technological capability. Without reliable evaluation, we cannot show whether supplier capabilities are improving aliases wherever their limitations remain. Terminal-Bench-Science evaluates agents successful realistic environments and grades actual artifacts specified arsenic analyses, simulations, proofs, code, and information products pinch reproducible, task-specific tests.

We request a benchmark that keeps gait pinch the frontier. Too often, technological benchmarks are treated arsenic papers to people alternatively than mechanisms for driving progress. They are released erstwhile and past abandoned arsenic models beforehand and known limitations persist. Terminal-Bench-Science is simply a continuous benchmark that evolves alongside the AI frontier. Through regular releases, scientists tin lend caller workflows, amended existing tasks, and create a feedback loop betwixt technological needs and AI development.

Terminal-Bench-Science Feedback Loop

SCIENTIFICCOMMUNITYFRONTIER AIAGENTS & MODELSContributeworkflowsEvaluate &improveAccelerate discoveryTerminal-Bench-Science Feedback Loop

Tasks

Terminal-Bench-Science 0.1 includes 70 tasks crossed the life, physical, Earth, mathematical, and engineering sciences. Tasks span technological information analysis, statistical inference, simulation, optimization, theorem proving, image reconstruction, awesome processing, inverse problems, sensor calibration, exemplary fitting, classification, and technological instrumentality learning.

Terminal-Bench-Science 0.1 Task Coverage

LIFE SCIENCES19Biology8Medicine & Health6Neuroscience4Ecology & Evolution1PHYSICAL SCIENCES17Astronomy & Cosmology6Physics5Materials Science4Chemistry2EARTH SCIENCES8Geoscience5Atmosphere & Climate1Ocean & Marine1Environment & Sustainability1MATHEMATICAL SCIENCES17Applied Mathematics & Scientific Computing6Operations Research & Optimization5Formal Mathematics & Theorem Proving3Statistics3ENGINEERING SCIENCES9Mechanical & Aerospace Engineering5Electrical & Computer Engineering2Chemical & Process Engineering1Civil & Structural Engineering170 TASKS IN TOTAL0246810

70 expert-curated tasks crossed 5 technological domains

Tasks are contributed by researchers done an unfastened process connected GitHub, pinch chat and feedback successful the #tb-science transmission connected Discord. Contributions statesman arsenic proposals, wherever reviewers talk each idea, time off feedback, and o.k. those that look for illustration a beardown fit: scientifically grounded workflows worthy measuring successful the benchmark. Approved proposals are implemented arsenic propulsion requests, wherever reviewers corroborate that each task is objectively verifiable, genuinely challenging for AI agents, and not thing today's frontier systems already lick easily. To merge, domain reviewers measure technological validity and realism, method reviewers inspect task building and verification, and a barroom raiser performs a last value check. Of 920 proposals, 464 were approved for implementation and 386 propulsion requests were opened, but only 70 tasks made it into Terminal-Bench-Science 0.1. That selectivity reflects really difficult it is to create tasks that are scientifically interesting, challenging for frontier agents, and sufficiently good specified for rigorous evaluation. Progress crossed proposals, propulsion requests, and reviews is tracked connected the public task dashboard.

Terminal-Bench-Science 0.1 leaves important room for advancement connected AI agents for technological research. Each evaluated exemplary ran 3 independent tests per task crossed each 70 tasks. Claude Opus 5 pinch Claude Code achieves the highest solution complaint astatine 30.0%, followed by GPT-5.6 Sol pinch Codex astatine 22.4% and Claude Fable 5 pinch Claude Code astatine 21.4%. Claude Opus 4.8 sits successful the mediate astatine 10.5%. GPT-5.6 Terra, Kimi K3, and Grok 4.6 each resoluteness little than 10% of tasks. GLM 5.3 is the strongest unfastened exemplary astatine 8.1%, and GPT-5.6 Luna is past astatine 3.3%.

Terminal-Bench-Science distinguishes betwixt systems astir arsenic good arsenic Terminal-Bench 3.0 while pushing solution rates down by much than 10 percent points for each exemplary evaluated connected both. That spread is deliberate: during review, tasks were calibrated to situation the newest frontier models.

Terminal-Bench-Science vs. Terminal-Bench

Terminal-Bench 2.1Terminal-Bench 3.0Terminal-Bench-Science 0.142.7%Opus 5 — 30.0%34.4%GPT-5.6 Sol — 22.4%Fable 5 — 83.8%33.8%Fable 5 — 21.4%Opus 4.8 — 78.9%21.1%Opus 4.8 — 10.5%GPT-5.6 Terra — 78.4%20.8%GPT-5.6 Terra — 8.6%GPT-5.6 Luna — 75.7%14.3%GPT-5.6 Luna — 3.3%Resolution Rates connected Terminal-Bench 2.1, Terminal-Bench 3.0, and Terminal-Bench-Science 0.1

Performance is only 1 magnitude of progress. The cost-resolution crippled shows full information costs crossed each 70 tasks against solution rate. GPT-5.6 Luna, Kimi K3, and GPT-5.6 Terra inhabit the low-cost extremity of the frontier. GPT-5.6 Sol and Claude Opus 5 scope the highest solution rates astatine greater cost, pinch Opus 5 astatine $7.0k. GPT-5.6 Sol matches Claude Fable 5's capacity astatine little than a 3rd of the costs ($4.2k vs $14.2k). Token usage shows a different frontier. Claude Fable 5 matches GPT-5.6 Sol's capacity while utilizing astir a 4th less tokens (6.4B vs 8.4B). Kimi K3 anchors the low-token extremity and Claude Opus 5 the high-resolution end. Only Kimi K3 and Claude Opus 5 look connected some Pareto frontiers.

Terminal-Bench-Science 0.1 Cost vs. Resolution Rate

Pareto frontier of costs and solution complaint crossed evaluated systems

Terminal-Bench-Science 0.1 Tokens vs. Resolution Rate

Pareto frontier of token usage and solution complaint crossed evaluated systems

Resolution rates besides alteration by technological domain. Anthropic and OpenAI models return the apical 2 spots successful each domain isolated from the engineering sciences, wherever Grok 4.6 ties GPT-5.6 Sol for 2nd spot (14.8%) astatine little costs and token usage. Claude Opus 5 leads some GPT-5.6 Sol and Claude Fable 5 successful each domain isolated from the mathematical sciences, wherever Claude Fable 5 (33.3%) and GPT-5.6 Sol (31.4%) return the apical 2 spots. The afloat breakdown by domain is disposable connected the leaderboard.

Terminal-Bench-Science 0.1 Resolution Rates by Domain

Domain solution rates: Opus 5, GPT-5.6 Sol, and Grok 4.6LIFESCIENCESEARTHSCIENCESENGINEERINGSCIENCESMATHEMATICALSCIENCESPHYSICALSCIENCES102030405029.8%45.8%29.6%25.5%27.5%17.5%20.8%14.8%31.4%23.5%7.0%4.2%14.8%5.9%5.9%
  • Opus 5
  • GPT-5.6 Sol
  • Grok 4.6
Domain solution rates: Opus 5, GPT-5.6 Sol, and Grok 4.6

Conclusion and Roadmap

Terminal-Bench-Science 0.1 is simply a organization effort by researchers crossed the life, physical, Earth, mathematical, and engineering sciences, together pinch the Terminal-Bench and Harbor team. It is the astir rigorous benchmark of technological supplier capabilities we could build successful the open, and we are only getting started: Terminal-Bench-Science 0.1 is the first merchandise of a continuous benchmark.

Regular releases will adhd tasks, broaden sum crossed the 5 technological domains, discontinue tasks that agents saturate aliases that reappraisal reveals to beryllium underspecified, and support the leaderboard existent arsenic caller frontier models are released. Each merchandise is calibrated against the frontier astatine the time. For Terminal-Bench-Science 0.1, this was Claude Opus 5 and GPT-5.6 Sol, and arsenic stronger models look we will usage them to measure and calibrate caller and improved tasks. Tasks are versioned truthful that tests tin beryllium re-used, re-graded, aliases re-run pinch a azygous Harbor command, which keeps the costs of updating results low. Progress is tracked successful the unfastened connected the task dashboard, and each merchandise is tagged connected GitHub and Harbor Hub.

Work connected Terminal-Bench-Science 0.2 is already underway, pinch a propulsion petition deadline of October 5, 2026. If you are a interrogator pinch a workflow that frontier agents should beryllium capable to do but cannot yet, we want it successful the benchmark. The publication travel is Propose → Build → Review: propose your task done the task connection form, build it pursuing the contributing guide, and it will spell done automated checks, parallel domain and method review, and last bar-raiser support earlier merge.

Join the effort successful #tb-science connected Discord and connected GitHub, and driblet into our play meetings and agency hours via the project calendar. Let's fto scientists specify what technological capacity successful AI looks like, and measurement it rigorously, together.

If you find this activity useful, please mention it. You tin usage the "Cite this repository" fastener connected GitHub (generated from CITATION.cff) aliases mention manually utilizing the accusation below.

@software{Terminal-Bench-Science_Team_Terminal-Bench-Science_Evaluating_AI_2026, author = {{Terminal-Bench-Science Team}}, doi = {10.5281/zenodo.22110254}, license = {Apache-2.0}, month = aug, title = {{Terminal-Bench-Science: Evaluating AI agents connected investigation workflows crossed technological domains}}, url = {https://github.com/harbor-framework/terminal-bench-science}, version = {v0.1.0}, year = {2026} }

Acknowledgements

Thank you to each of the task contributors, reviewers, and advisors down Terminal-Bench-Science.

Special acknowledgment to our task lead advisors Ludwig Schmidt and Sanmi Koyejo; our elder reviewers Allen Hart, Ivan Bercovich, Joseph Janssen, Jiaming Hu, Steffen Bollmann, and Sergey Aganezov; our AI investigation advisors Ryan Marten, Alex Shaw, Lin Shi, Benjamin Feuer, Mike A. Merrill, Alex Dimakis, Jenia Jitsev, Bodhisattwa Majumder, Peter Clark, Thomas Wolf, Braden Hancock, and Andy Konwinski; and our technological advisors Sara Beery, Jo Dunkley, J. Nathan Kutz, Ching-Yao Lai, Scott Linderman, Emma Lundberg, Russ Poldrack, Aviv Regev, and Risa Wechsler.

Terminal-Bench-Science is an unfastened world collaboration hosted by Stanford University and the Laude Institute, successful business pinch the Stanford AI Lab (SAIL), the Stanford Institute for Human-Centered Artificial Intelligence (HAI), Stanford AI Measurement Science (AIMS), the NSF AI Institute for Foundations of Machine Learning (IFML), the Allen Institute, and the Allen Institute for AI (Ai2). As portion of the Terminal-Bench franchise, it is built by the Terminal-Bench and Harbor squad together pinch a organization of scientific contributors.

We convey the Laude Institute for support done the Slingshots program, Snorkel AI for support done the Open Benchmarks Grants program, the 2077AI Open Source Foundation for PP API credits supporting task reappraisal and curation, and UniPat AI and Modal for their support of Terminal-Bench-Science. We convey Bespoke Labs, Anthropic, Google, Moonshot AI, SpaceXAI, and Z.ai for API credits supporting leaderboard evaluations.

Written by: Steven Dillmann

More