← Retour au blog
tech 28 August 2026

Terminal-Bench-Science: Evaluating AI Agents on Scientific Research Workflows

Terminal-Bench-Science 0.1 establishes a new benchmark for evaluating AI agent capabilities in scientific research workflows. Developed by Stanford, this benchmark involves scientists in setting evaluation criteria, encompassing 70 tasks from various scientific fields.

Article inspired by the original source
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows ↗ www.terminal-bench-science.ai

Introduction

In the realm of scientific research, AI is becoming an indispensable tool. However, to truly revolutionize the field, it's crucial to have benchmarks that accurately assess AI agent capabilities in real-world contexts. Terminal-Bench-Science 0.1, developed by researchers at Stanford University, addresses this need by providing a platform where scientists themselves can define evaluation criteria.

A Benchmark Built by Scientists for Scientists

Unlike other benchmarks often devised by model developers or data vendors, Terminal-Bench-Science relies on the expertise of researchers themselves. This approach ensures that the tasks and evaluation criteria genuinely reflect the challenges scientists face in their daily work. With 70 tasks spanning life sciences, physical sciences, Earth sciences, mathematics, and engineering, this benchmark is both diverse and rigorous.

Current Model Performances

In its initial release, Terminal-Bench-Science 0.1 highlights several AI models, with the Claude Opus 5 model achieving a resolution rate of 30% across all tasks. While this figure might seem modest, it represents a significant step toward enhancing AI agent capabilities in complex and specialized contexts.

The Importance of Continuous Evaluation

One of the most innovative aspects of Terminal-Bench-Science is its evolving nature. The benchmark is designed to adapt to AI advancements, creating a feedback loop between scientific needs and AI development. This not only ensures the benchmark's continued relevance but also promotes the continuous improvement of AI agents.

Use Cases and Potential Impacts

AI agents evaluated through Terminal-Bench-Science can transform scientific research by automating technically demanding and time-consuming processes. This frees researchers to focus on tasks where human judgment is crucial, such as defining research questions and interpreting results. For instance, in life sciences, automating complex data analyses can accelerate drug discovery.

Conclusion

Terminal-Bench-Science 0.1 marks a significant milestone in evaluating AI agents for scientific research. By directly involving scientists and incorporating real-world tasks, it sets a standard that could transform how we use AI in science. Let's discuss your project in 15 minutes.

References

  • [Official Terminal-Bench-Science website](https://www.terminal-bench-science.ai/announcement)

Keywords

  • Scientific AI
  • AI Benchmark
  • Terminal-Bench-Science
  • AI Evaluation
  • Scientific Research
IA scientifique Benchmark IA Terminal-Bench-Science Évaluation IA Recherche scientifique
Deepthix newsletter · 100% AI · every Monday 8am

An AI agent reads tech for you.

Our AI agent scans ~200 sources per week and ships the best articles to your inbox Monday 8am. Free. One click to unsubscribe.

Visit the newsletter page →

Want to automate your operations?

Let's talk about your project in 15 minutes.

Book a call