SciTra: Cost-Aware Science Trajectory
An open community exploring the paths to scalable scientific discovery.
Scientific agents can reach the same outcome through very different sequences of decisions — and at very different costs. SciTra is an open, long-term community of domain experts and AI researchers studying these agents together: where they fall short, how to evaluate them, and how to improve them.

Across science and engineering
SciTra’s community spans a growing range of scientific and engineering domains.
- Physics
- Computer Science
- Biology
- Chemistry
- Materials Science
- Mechanical and Aerospace Engineering
- Medicine and Health
- Robotics
- Electrical Engineering
- Neuroscience
- Astronomy
- Earth Science
- Other
Why SciTra?
Long-horizon science is a resource-allocation problem.
01
Cost efficiency is a precondition, not an extra.
02
Evaluation cannot be separated from scientists.
Long horizon
Real cost
Trajectory
What we’re building
Infrastructure and methodology, built together.
Infrastructure
SciTra Runtime
Methodology
Evaluation methodology
Trajectories have structure
In autonomous driving, route planning, lane changes, and steering are trajectories at different levels. Science and engineering have levels too. These levels should not be invented. They should grow from real human-agent trajectories and from domain knowledge.
| Driving | Science and engineering (example) |
|---|---|
| Route planning | Research strategy. Which question to pursue, and how. |
| Lane change | Scientific action. Run a simulation. Fit a model. |
| Steering and braking | Execution. Write code, set parameters, call tools. |
From evaluation to improvement
Evaluation today. Learning environments tomorrow.
Scalable scientific autonomy requires systems that can learn from repeated interaction with scientific environments. By preserving resource-aware execution trajectories, SciTra aims to make evaluation useful not only for comparing agents, but eventually for improving how they explore, decide, and allocate resources.
Workflow
Collect
SciTra
Run locally
SciTra Runtime
Scale
SciTra Runtime
Progress
- 55
- 32
- Ongoing
Toward the full scientific loop
SciTra begins with computational science deliberately.
Computational and simulated environments provide a lower-cost, repeatable setting in which resource-aware evaluation can be developed and validated before moving into physical experimentation.
This sequencing matters.
In a real laboratory, inefficient exploration can consume scarce instrument time, materials, experimental throughput, and other physical resources. A framework for evaluating long-horizon decisions should therefore prove useful in computational settings before those decisions are allowed to incur substantially higher real-world costs.
The longer-term direction is to extend these principles to software-controlled laboratory workflows.
We are exploring LabVIEW-based systems as an initial path toward real experimental interaction.
Current
Computational workflows
Future
Physical experiments
The goal is to evaluate the complete loop — from deciding what to do, to executing an experiment, observing the result, and choosing what to do next — while accounting for the real resources consumed along the way.
Scientists and AI, evolving together
Scientific AI is not a zero-sum game.
We are scientists, and we should grow together with the tools we build. We should help create the next generation of scientific tools. We should be builders, not just a source of training data.
As AI becomes more capable, we will learn to work with it, challenge it, and shape it. We do not decide in advance what we should do and what AI should do. That division of labor should emerge from real human-agent trajectories.
Benchmark scores are not the final measure. The real test is progress on problems that matter: understanding nature, developing new materials, improving medicine, and building new technologies.
Join SciTra
An open community, built for the long term.
SciTra has a theme — cost-aware scientific agents — but it is not a race to assemble a benchmark. The point is to learn and explore together, and to build collaborations that last well beyond any single task or paper.
01
Explore limitations
02
Study evaluation
03
Improve agents
The tasks members bring are the material for this research; the research itself is done together. Because the costs and trade-offs that matter differ across fields, it needs people from many domains.
Curious how scientific agents succeed — and fall short — in your field? Join us (opens in a new tab) Or come say hello on Discord (opens in a new tab).
Collaborators from
Resources
Living knowledge base
Evaluating AI Scientists
Benchmark observatory
Leaderboard of Benchmarks



