SciTra: Cost-Aware Science Trajectory
Evaluating the paths to scalable scientific discovery.
Scientific agents can reach the same outcome through very different sequences of decisions — and at very different costs. SciTra evaluates long-horizon scientific agents through their outcomes, resource use, and execution trajectories.

Across science and engineering
SciTra is being built across a growing range of scientific and engineering domains.
- Physics
- Materials Science
- Mechanical and Aerospace Engineering
- Medicine and Health
- Biology
- Chemistry
- Computer Science
- Robotics
- Electrical Engineering
- Astronomy
- Neuroscience
- Other
Why SciTra?
Long-horizon science is a resource-allocation problem.
Scientific agents do not simply produce answers.
They repeatedly choose what to investigate, compute, simulate, run, and revise. As these loops grow longer, inefficient decisions compound into substantial costs.
The same outcome can come at very different costs.
Two agents may reach the same valid outcome while consuming radically different amounts of time, compute, money, and other scientific resources. Measuring success alone cannot reveal that difference.
An agent that can solve a task once is not necessarily an agent that can scale scientific discovery.
Scaling automated science requires more than measuring outcomes. It requires measuring the paths and resources that produce them.
Long horizon
Real cost
Trajectory
From evaluation to improvement
Evaluation today. Learning environments tomorrow.
Scalable scientific autonomy requires systems that can learn from repeated interaction with scientific environments. By preserving resource-aware execution trajectories, SciTra aims to make evaluation useful not only for comparing agents, but eventually for improving how they explore, decide, and allocate resources.
Workflow
Collect
Run locally
Scale
Toward the full scientific loop
SciTra begins with computational science deliberately.
Computational and simulated environments provide a lower-cost, repeatable setting in which resource-aware evaluation can be developed and validated before moving into physical experimentation.
This sequencing matters.
In a real laboratory, inefficient exploration can consume scarce instrument time, materials, experimental throughput, and other physical resources. A framework for evaluating long-horizon decisions should therefore prove useful in computational settings before those decisions are allowed to incur substantially higher real-world costs.
The longer-term direction is to extend these principles to software-controlled laboratory workflows.
We are exploring LabVIEW-based systems as an initial path toward real experimental interaction.
Computational workflows
Physical experiments
The goal is to evaluate the complete loop — from deciding what to do, to executing an experiment, observing the result, and choosing what to do next — while accounting for the real resources consumed along the way.
Join SciTra
SciTra is built around expert-in-the-loop evaluation. Domain experts do not just contribute tasks — they help define what meaningful outcomes, trajectories, and resource trade-offs look like in their field.
Define with experts
Co-design the evaluation
Scale validated evaluations
Resource-aware evaluation cannot be one-size-fits-all: the costs and trade-offs that matter differ across scientific domains.
Have a task, evaluation idea, or infrastructure contribution? Apply to contribute (opens in a new tab)
Contributors
Resources
BenchFlow
Large-scale evaluation infrastructure.
Scientific Eval Environments
A curated knowledge base of papers, benchmarks, and related work on scientific agent evaluation.
Progress
- 39
- Ongoing
- Ongoing



