SciTra: Cost-Aware Science Trajectory

An open community exploring the paths to scalable scientific discovery.

Scientific agents can reach the same outcome through very different sequences of decisions — and at very different costs. SciTra is an open, long-term community of domain experts and AI researchers studying these agents together: where they fall short, how to evaluate them, and how to improve them.

SciTra — Cost-Aware Science Trajectory

Across science and engineering

SciTra’s community spans a growing range of scientific and engineering domains.

Explore current coverage

Why SciTra?

Long-horizon science is a resource-allocation problem.

  1. 01

    Cost efficiency is a precondition, not an extra.

    Solving a problem is not enough. If it takes unreasonable compute, time, experimental resources, or human attention, it will not accelerate science. Efficiency is what makes a capability usable.

  2. 02

    Evaluation cannot be separated from scientists.

    Style and answers can be learned from data. Strategic judgment cannot. It comes from scientists. In SciTra, we help build the agents for our own fields. We shape how trajectories are abstracted and how evaluation is designed. We do more than label data or write rubrics.

Outcome measures capability.Cost measures scalability.Trajectory connects the two.

These are open questions, and no single benchmark will settle them. SciTra exists to work on them together, over the long term.

What we’re building

Infrastructure and methodology, built together.

Trajectories have structure

In autonomous driving, route planning, lane changes, and steering are trajectories at different levels. Science and engineering have levels too. These levels should not be invented. They should grow from real human-agent trajectories and from domain knowledge.

DrivingScience and engineering (example)
Route planningResearch strategy. Which question to pursue, and how.
Lane changeScientific action. Run a simulation. Fit a model.
Steering and brakingExecution. Write code, set parameters, call tools.
The right column is an example. Each field’s structure comes from its own reviewed trajectories.

From evaluation to improvement

Evaluation today. Learning environments tomorrow.

Scalable scientific autonomy requires systems that can learn from repeated interaction with scientific environments. By preserving resource-aware execution trajectories, SciTra aims to make evaluation useful not only for comparing agents, but eventually for improving how they explore, decide, and allocate resources.

  1. Scientific task
  2. Agent interaction
  3. Trajectory + resource use
  4. Evaluation
  5. Learning signal
  6. Better scientific agent
The cycle repeats.

Workflow

  1. Collect

    SciTra

    Gather scientific tasks as shared material for the community's research.

  2. Run locally

    SciTra Runtime

    Use your agent as usual. SciTra Runtime automatically records the execution trajectory.

  3. Scale

    SciTra Runtime

    Scale validated evaluations to larger campaigns.

Progress

Ideas collected
55

View ideas

Local runs
32

View runs

Scaled tasks
Ongoing

View scale-up

These counts track shared research material, not a leaderboard. Public aggregate statistics will be synchronized as the work grows.

Toward the full scientific loop

SciTra begins with computational science deliberately.

Computational and simulated environments provide a lower-cost, repeatable setting in which resource-aware evaluation can be developed and validated before moving into physical experimentation.

This sequencing matters.

In a real laboratory, inefficient exploration can consume scarce instrument time, materials, experimental throughput, and other physical resources. A framework for evaluating long-horizon decisions should therefore prove useful in computational settings before those decisions are allowed to incur substantially higher real-world costs.

The longer-term direction is to extend these principles to software-controlled laboratory workflows.

We are exploring LabVIEW-based systems as an initial path toward real experimental interaction.

Physical experimentation should not be the place where a resource-aware evaluation framework is first invented.

Current

Computational workflows

  1. Data
  2. Code / Analysis
  3. Simulation
  4. Result
  5. Iteration

Future

Physical experiments

  1. Decision
  2. Instrument control
  3. Measurement
  4. Observation
  5. Next decision

The goal is to evaluate the complete loop — from deciding what to do, to executing an experiment, observing the result, and choosing what to do next — while accounting for the real resources consumed along the way.

Scientists and AI, evolving together

Scientific AI is not a zero-sum game.

We are scientists, and we should grow together with the tools we build. We should help create the next generation of scientific tools. We should be builders, not just a source of training data.

As AI becomes more capable, we will learn to work with it, challenge it, and shape it. We do not decide in advance what we should do and what AI should do. That division of labor should emerge from real human-agent trajectories.

Benchmark scores are not the final measure. The real test is progress on problems that matter: understanding nature, developing new materials, improving medicine, and building new technologies.

Our goal is not to replace scientists, but to expand what we can discover together with AI.

Join SciTra

An open community, built for the long term.

SciTra has a theme — cost-aware scientific agents — but it is not a race to assemble a benchmark. The point is to learn and explore together, and to build collaborations that last well beyond any single task or paper.

The tasks members bring are the material for this research; the research itself is done together. Because the costs and trade-offs that matter differ across fields, it needs people from many domains.

Curious how scientific agents succeed — and fall short — in your field? Join us (opens in a new tab) Or come say hello on Discord (opens in a new tab).

Collaborators from

Academic & Research Institutions

  • University of Pennsylvania

  • UC San Diego

  • UC Berkeley

  • The University of Texas at Austin

  • UCLA

  • Caltech

  • Northwestern University

  • Zhejiang University

  • EPFL

  • Boston University

  • University of Tennessee, Knoxville

  • University of Chicago

  • University of Florida

  • University of Illinois Urbana-Champaign

  • The Hang Seng University of Hong Kong

  • UC Santa Cruz

  • Washington State University

  • Georgia Tech

Industry & Infrastructure

  • KLA

  • Applied Materials

  • Amazon

  • BenchFlow

Resources