Autonomous Science Network

Roadmap

A two-year community roadmap · July 2026

Toward trustworthy autonomous science

One year after the original AISLE roadmap, the field has moved faster than we anticipated. Multi-agent systems have produced experimentally validated hypotheses, self-driving laboratories have grown markedly more interoperable, and a national mobilization has placed autonomous experimentation at the center of U.S. federal science strategy. Progress has also met a sobering counter-current: a flagship autonomous-discovery result was corrected, benchmarks show agents completing only a small fraction of open-ended research, and no system has yet made and experimentally self-validated a genuinely novel discovery end to end.

Producing a candidate discovery is no longer the hard part. Verifying it is.

This asymmetry, more than raw model capability, is what limits autonomous science today. This roadmap therefore organizes the community agenda around seven dimensions — the five we revisit, and two we elevate from cross-cutting afterthoughts to first-class concerns — and scopes the path forward to a deliberately short two-year horizon.

Toward Trustworthy Autonomous Science report cover

What changed in a year

Seven shifts in the landscape bear directly on the original roadmap. We read them through two lenses: an autonomy ladder that grades systems as tool, analyst, or scientist, and a persistent capability–reliability gap between what systems can propose and what can be trusted.

1. From isolated demonstrations to validated discoveries

Multi-agent systems have generated hypotheses later confirmed in the laboratory — but humans still selected the problems and ran the experiments.

2. Self-driving laboratories enter a second generation

SDL 2.0 aims to be interoperable, orchestrated, safe, and capable of hypothesis generation. Interoperability is the concrete advance so far.

3. The interface layer begins to consolidate

MCP connects one agent to many tools; A2A connects agents to one another. Both moved under neutral governance, joined by an emerging skills layer.

4. Reasoning and foundation models as a new substrate

Reasoning-trained LLMs approach expert performance on graduate-level science, and domain foundation models produce corroborated designs.

5. Benchmarks expose the capability–reliability gap

Agents that rival experts on closed-ended questions complete only a small fraction of open-ended research. Reliability is the named obstacle in production.

6. Trust and governance move to the foreground

A corrected discovery claim, a training-data leak, fabricated citations at leading venues, and unaddressed dual-use risk make these prerequisites.

7. Industry becomes a primary actor

Well-capitalized entrants build robotic discovery laboratories at a pace public programs cannot match — supplying capital, but owing the same open interfaces and verification.

Seven dimensions

The first five revisit and update the original AISLE dimensions in light of the past year. The last two are elevated here from cross-cutting concerns to first-class dimensions of the roadmap.

Instrument and cyberinfrastructure integration

Vendor-neutral access to increasingly interoperable self-driving laboratories, with AI inference moving onto the instruments themselves. Instruments, edge, HPC, cloud, and digital twins form a single distributed cyber-physical system.

  • M1. Common instrument interfaces and vendor-agnostic hardware abstraction.
  • M2. End-to-end workflows across institutions and facilities.
  • M3. Open compute fabric with fault tolerance and digital twins.
  • M4. Scalable national instrument framework.

Agent-driven data management

Provenance and quality enforced at the point of capture, where scarce data are made AI-ready by construction. As campaigns generate data that humans never inspect, curation becomes a prerequisite for credible discovery rather than after-the-fact bookkeeping.

  • M5. AI-driven metadata systems with automated annotation.
  • M6. Federated data mesh with FAIR governance.
  • M7. Near-real-time processing with AI provenance.

Agent-driven orchestration

Reasoning agents that plan and coordinate specialized scientific methods — and that explain why a result holds. The orchestrator is one component of the ecosystem, not a replacement for the established methods it coordinates.

  • M8. Hierarchical LLM orchestration with verification.
  • M9. Cross-facility knowledge integration.

Interoperable agent interfaces

The consolidating MCP and A2A stack, now joined by an emerging skills layer. The lesson of the past year is to standardize the interface while leaving the calling pattern free to evolve.

  • M10. Standardized cross-vendor agent interfaces.
  • M11. Zero-trust, sub-second agent coordination.
  • M12. Self-discovering agent networks.

Education and workforce development

The judgment to resist the homogenizing pull of automation. AI-engaged researchers publish more and are cited more, even as the range of topics the community collectively studies contracts.

  • M13. National autonomous science education consortium.
  • M14. Immersive virtual labs and human-AI collaboration assessment.

Elevated to first-class

Trust, verification, and reproducibility

Verification and validation across the lifecycle as a first-class requirement. Verification asks whether a workflow was built and executed correctly; validation asks whether its result corresponds to physical reality — the harder problem, where only measurement can settle the question.

  • M15. Verification and validation across the experiment lifecycle.
  • M16. Reproducibility and efficiency benchmark for autonomous workflows.

Elevated to first-class

Safety, security, integrity, and governance

Screening and accountability where digital design meets physical execution. The autonomy granted to an agent should match the cost, risk, and reversibility of the action it would take, and should be raised only as auditable evidence of reliability accumulates.

  • M17. Screening at the digital-to-physical interface.
  • M18. Cross-institution agent identity and governance.

Milestone scorecard

We assess each original AISLE milestone (M1–M14) against a year of evidence and add four new milestones (M15–M18) for the elevated dimensions. Statuses are provisional and subject to confirmation against the cited evidence. We intend this rescoring as a recurring practice rather than a one-off, so that the roadmap stays honest about what has and has not been achieved.

# Milestone Status By
M1Common instrument interfaces, HALPartialY1
M2End-to-end cross-institution workflowsPartialY2
M3Open compute fabric, fault tolerance, digital twinsOpenY2
M4Scalable national instrument frameworkOpenY2
M5AI-driven metadata and annotationPartialY1
M6Federated data mesh, FAIR governanceOpenY2
M7Near-real-time processing and AI provenancePartialY2
M8Hierarchical LLM orchestration, verificationPartialY1
M9Cross-facility knowledge integrationOpenY2
M10Standardized cross-vendor agent interfacesReframedY1
M11Zero-trust, sub-second agent coordinationOpenY2
M12Self-discovering agent networksPartialY2
M13National education consortiumOpenY1
M14Virtual labs, human-AI assessmentOpenY2
M15Verification and validation across the lifecycleNewY1
M16Reproducibility and efficiency benchmarkNewY1
M17Screening at digital-physical interfaceNewY1
M18Cross-institution agent identity, governanceNewY2

The pattern is consistent: the dimensions closest to raw model capability have advanced the fastest, while those that require cross-institutional infrastructure, verification, and governance remain largely open. Orchestration and agent interfaces moved the furthest. Data management, education, and the federated aspects of instrument integration moved the least, because they depend on coordination that no single model improvement can supply.

The two-year trajectory

The roadmap climbs the autonomy ladder from tool to analyst to scientist, so that trustworthy autonomy rises to meet, rather than outrun, frontier model capability. We keep the horizon short deliberately: at the current pace of change, a longer-range plan would be obsolete before it could be acted upon.

0–8 months

Interfaces and adoption

  • Vendor-agnostic interfaces (M1)
  • MCP/A2A adoption (M10)

8–16 months

Data and the scaffolding of verification

  • AI-driven metadata (M5)
  • Federated data mesh (M6)
  • Verification and validation (M15)

16–24 months

Federation and governance

  • Zero-trust communications (M11)
  • Self-discovering networks (M12)
  • Agent identity and governance (M18)
  • National framework (M4)

The national and global ecosystem

The original roadmap argued for a grassroots, bottom-up network on the premise that no coordinated national program existed to connect autonomous laboratories. That premise has partly changed. The Genesis Mission has mobilized U.S. national laboratories around an integrated platform coupling HPC, scientific foundation models, datasets, and automated laboratory systems, alongside the Trillion Parameter Consortium, the European strategy for AI in science, and the Acceleration Consortium.

We argue that top-down mobilization and bottom-up coordination are complementary rather than redundant. Large programs supply what a grassroots network cannot: compute at scale, foundation models, a security framing, and funding at scale. A grassroots network, in turn, supplies what large programs and commercial vendors tend to underweight — vendor-neutral interfaces and cross-institutional standards that keep each new program or proprietary platform from becoming another silo.

AISLE as interoperability fabric

An interconnected network such as AISLE is best understood not as an alternative to federal mobilization but as the interoperability fabric and community-standards layer that allows these initiatives to connect to one another and to the broader research community, rather than to interconnect only internally.

The milestones above should be read as community commitments that any of these programs can adopt, instrument, and report against, rather than as the agenda of a single institution.

Roadmap report (2026)

R. Ferreira da Silva, M. Abolhasani, P. Beaucage, L. Biven, M. Bussmann, K. Chard, R. Coffee, S. DeWitt, S. Dolas, C. Eckert, D. Elbert, I. T. Foster, T. Ghosal, A. Giannakou, T. Gibbs, L. Hamilton, G. Lockwood, T. Mayer, B. Mintz, R. Nazikian, S. Nimer, A. Randles, W. Shin, S. R. Sukumar, F. Suter, M. Taheri, M. Taufer, D. Vrabie. "Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap." Technical Report, ORNL/TM-2026/4663, Oak Ridge National Laboratory, July 2026. arXiv:2607.12113

Licensed under CC BY 4.0.

The original roadmap (2025)

R. Ferreira da Silva, M. Abolhasani, D. A. Antonopoulos, L. Biven, R. Coffee, I. T. Foster, L. Hamilton, S. Jha, T. Mayer, B. Mintz et al. "A Grassroots Network and Community Roadmap for Interconnected Autonomous Science Laboratories for Accelerated Discovery." Workshop Proceedings of the 54th International Conference on Parallel Processing, 2025, pp. 142–150. arXiv:2506.17510

Establishes the AISLE network and the original milestones M1–M14, scored above.