SciTra

SciTra: Cost-Aware Science Trajectory

Evaluating the paths to scalable scientific discovery.

Scientific agents can reach the same outcome through very different sequences of decisions — and at very different costs. SciTra evaluates long-horizon scientific agents through their outcomes, resource use, and execution trajectories.

SciTra — Cost-Aware Science Trajectory

Across science and engineering

SciTra is being built across a growing range of scientific and engineering domains.

Explore current coverage

Why SciTra?

Long-horizon science is a resource-allocation problem.

Scientific agents do not simply produce answers.

They repeatedly choose what to investigate, compute, simulate, run, and revise. As these loops grow longer, inefficient decisions compound into substantial costs.

The same outcome can come at very different costs.

Two agents may reach the same valid outcome while consuming radically different amounts of time, compute, money, and other scientific resources. Measuring success alone cannot reveal that difference.

An agent that can solve a task once is not necessarily an agent that can scale scientific discovery.

Scaling automated science requires more than measuring outcomes. It requires measuring the paths and resources that produce them.

Outcome measures capability.Cost measures scalability.Trajectory connects the two.

To our knowledge, SciTra is the first general-purpose scientific-agent benchmark to jointly treat long-horizon execution, heterogeneous resource costs, and full execution trajectories as first-class evaluation dimensions.

From evaluation to improvement

Evaluation today. Learning environments tomorrow.

Scalable scientific autonomy requires systems that can learn from repeated interaction with scientific environments. By preserving resource-aware execution trajectories, SciTra aims to make evaluation useful not only for comparing agents, but eventually for improving how they explore, decide, and allocate resources.

  1. Scientific task
  2. Agent interaction
  3. Trajectory + resource use
  4. Evaluation
  5. Learning signal
  6. Better scientific agent
The cycle repeats.

Workflow

  1. Collect

    SciTra

    Collect and organize scientific evaluation ideas.

  2. Run locally

    SciTra Runtime

    Use your agent as usual. SciTra Runtime automatically records the execution trajectory.

  3. Scale

    BenchFlow

    Scale validated evaluations to larger campaigns.

Toward the full scientific loop

SciTra begins with computational science deliberately.

Computational and simulated environments provide a lower-cost, repeatable setting in which resource-aware evaluation can be developed and validated before moving into physical experimentation.

This sequencing matters.

In a real laboratory, inefficient exploration can consume scarce instrument time, materials, experimental throughput, and other physical resources. A framework for evaluating long-horizon decisions should therefore prove useful in computational settings before those decisions are allowed to incur substantially higher real-world costs.

The longer-term direction is to extend these principles to software-controlled laboratory workflows.

We are exploring LabVIEW-based systems as an initial path toward real experimental interaction.

Physical experimentation should not be the place where a resource-aware evaluation framework is first invented.

Current

Computational workflows

  1. Data
  2. Code / Analysis
  3. Simulation
  4. Result
  5. Iteration

Future

Physical experiments

  1. Decision
  2. Instrument control
  3. Measurement
  4. Observation
  5. Next decision

The goal is to evaluate the complete loop — from deciding what to do, to executing an experiment, observing the result, and choosing what to do next — while accounting for the real resources consumed along the way.

Join SciTra

SciTra is built around expert-in-the-loop evaluation. Domain experts do not just contribute tasks — they help define what meaningful outcomes, trajectories, and resource trade-offs look like in their field.

Resource-aware evaluation cannot be one-size-fits-all: the costs and trade-offs that matter differ across scientific domains.

Have a task, evaluation idea, or infrastructure contribution? Apply to contribute (opens in a new tab)

Contributors

Academic & Research Institutions

  • University of Pennsylvania

  • UC San Diego

  • UC Berkeley

  • The University of Texas at Austin

  • UCLA

  • Caltech

  • Northwestern University

  • Zhejiang University

  • EPFL

  • Boston University

  • University of Tennessee, Knoxville

  • University of Chicago

  • University of Florida

  • University of Illinois Urbana-Champaign

  • The Hang Seng University of Hong Kong

  • UC Santa Cruz

  • Washington State University

  • Georgia Tech

Industry & Infrastructure

  • KLA

  • Applied Materials

  • Amazon

  • BenchFlow

Resources

Progress

Ideas collected
39

View ideas

Local runs
Ongoing

View runs

Scaled tasks
Ongoing

View scale-up

Public aggregate statistics will be synchronized as the benchmark grows.