AI Research Note · DOE Genesis Mission2026-09-02 · AI for Science / Agentic AI
DOE Genesis Mission · Phase I/AI Advantage/2025–2026 Research Landscape

AI가 과학을 더 잘하는가가 아니라,
발견을 더 빨리 증명할 수 있는가

DOE Genesis Mission Phase I: Proving AI Advantage in Scientific Discovery

Genesis-style scientific discovery loopA conceptual flow from scientific objective through agents, foundation models, simulation and experiment to evidence and revision. OBJECTIVEAGENTIC REASONINGSCIENTIFIC MODELSIM / LABEVIDENCEREVISION · MEMORY · NEXT EXPERIMENT
Central thesis

Genesis Mission Phase I의 핵심은 거대한 모델을 하나 더 만드는 데 있지 않다. AI가 기존 과학 워크플로보다 실제로 더 빠르고, 더 정확하고, 더 검증 가능한 발견을 만드는지 정량적으로 입증하는 것이 핵심이다.

2025년 11월 24일 미국 행정부는 Executive Order 14363으로 Genesis Mission을 출범시켰다. DOE는 이어 2026년 3월 The Genesis Mission: Transforming Science and Energy with AI(DE-FOA-0003612)를 공고했고, 7월 22일 첫 프로젝트들을 발표했다. 2026년 9월 2일 현재 이 프로그램 자체를 장기간 평가한 동료평가 논문은 아직 거의 없다. 따라서 이 글은 DOE의 공식 Phase I 정의와 실제 선정 사례를 기준점으로 삼고, 2025–2026년 Scientific Foundation Model, AI Scientist, Multi-Agent Scientific Reasoning, Autonomous Laboratory, Digital Twin 연구를 연결해 그 학술적 의미를 해석한다.

\[ W_{\mathrm{AI}} = \{\mathrm{Data},\,\mathrm{Model},\,\mathrm{Agent},\,\mathrm{Simulation},\,\mathrm{Experiment},\,\mathrm{Feedback}\} \]Genesis-style AI scientific workflow

비교 대상은 전통적인 과학 워크플로 \(W_0\)이다. Phase I이 묻는 질문은 결국 \(\Delta=(\Delta\mathrm{Accuracy},\Delta\mathrm{Time},\Delta\mathrm{Cost},\Delta\mathrm{Yield},\Delta\mathrm{Insight},\Delta\mathrm{Reproducibility})\)가 의미 있는 양의 변화를 만드는가이다.

Part I · Definition & Problem

Phase I은 ‘AI 데모’가 아니라 과학적 proof-of-advantage다

프로그램의 용어를 정확히 고정해야 이후의 연구문제와 평가 기준도 흔들리지 않는다.

§1 · Definition

Genesis Mission과 Phase I의 정확한 의미

Genesis Mission은 연방 연구기관과 DOE National Laboratories가 보유한 과학 데이터, 고성능컴퓨팅, 실험시설, AI 모델, 로봇·생산설비를 하나의 통합 과학발견 시스템으로 연결하려는 국가 규모의 AI-for-Science 프로그램이다. Executive Order와 DOE 설명에서 반복되는 중심 표현은 closed-loop AI experimentation platform이다. 즉 과학적 가설을 만들고, 계산·시뮬레이션하고, 실험을 설계하고, 결과를 다시 모델과 다음 행동에 반영하는 순환 구조를 지향한다.

그중 RFA의 Phase I은 소규모 팀이 명확하고 구체적인 AI 통합 연구 워크플로를 설계·시연하고 AI advantage의 잠재력을 정량 평가하는 단계이다. DOE 자료에 따르면 수행기간은 9개월, 과제당 총비용은 50만–75만 달러이며, 통상 6개월 시점에 Phase II 진입 자격을 판단하는 go/no-go 평가가 있다. 계획한 실험을 마치기 위해 go/no-go 시점을 최대 3개월 연장할 수 있다.

Main RFA Phase I
DE-FOA-0003612

과학·에너지 연구 워크플로에서 AI advantage를 입증하는 9개월 연구 프로젝트. 21개 Topic, 99개 Focus Area를 대상으로 한다.

Do not confuse with
Genesis-linked SBIR/STTR Phase I

DOE가 2026년 7월 22일 별도로 연 중소기업 기술사업화 기회다. 초기 공고는 biotechnology, quantum systems, predictable materials, autonomous laboratories의 4개 영역과 약 40개 예상 award를 제시했다.

Scope note · 2026-09-02이 글에서 “Genesis Mission Phase I”은 별도 표기가 없는 한 DOE 본 RFA인 DE-FOA-0003612의 Phase I을 의미한다.
§2 · Problem Definition

예측 정확도에서 과학적 발견 속도로 연구단위가 이동한다

많은 AI for Science 연구는 “기존 모델보다 예측 오차가 줄었다”에서 끝난다. 하지만 실제 과학은 문헌 검색, 관측, 가설, 시뮬레이션, 실험계획, 실험, 분석, 반증, 수정이 반복되는 과정이다. 따라서 진짜 병목은 하나의 predictor가 아니라 end-to-end scientific workflow다.

\[\mathrm{Prediction\ AI}\;\rightarrow\;\mathrm{Reasoning\ AI}\;\rightarrow\;\mathrm{Experimenting\ AI}\;\rightarrow\;\mathrm{Self\!\!-improving\ Scientific\ System}\]

2026년 Nature의 Robin은 literature search agent와 data-analysis agent를 연결해 생물학 연구에서 가설 생성과 실험 데이터 분석을 하나의 반복 과정으로 묶었다. 2025년 AI Co-Scientist는 generate–debate–evolve 식의 멀티 에이전트 구조로 가설 생성을 다룬다. 이 흐름을 DOE의 Phase I과 연결하면 연구의 평가단위가 모델 성능에서 발견 주기(discovery cycle)의 생산성·신뢰성으로 이동한다.

Part II · Core Concepts & Platform

모델, 에이전트, 실험실, HPC를 하나의 과학 시스템으로 묶는다

Genesis의 신규성은 구성요소 각각보다 이들을 하나의 검증 가능한 운영체계로 엮는 데 있다.

§3 · Core Concepts

핵심 개념 9개

Concept의미Phase I / 최신 연구와의 연결
AI AdvantageAI가 기존 방법보다 시간·비용·정확도·실험효율·통찰·재현성에서 실제 우위를 만드는가Phase I의 정량적 go/no-go 논리
Scientific Foundation Model다양한 과학 문제와 조건에 재사용·적응하는 대규모 사전학습 모델MatterGen, Aurora, molecular/cell multimodal FMs
Agentic ScienceAI가 목표를 해석하고 데이터·도구·모델을 선택해 다단계 연구를 수행AI Co-Scientist, Robin, AI Scientist-v2
Closed-loop Discoverypredict → make/simulate → measure → learn의 반복autonomous laboratory와 Genesis EO의 핵심 구조
Physics-informed Digital Twin물리 모델, AI surrogate, 관측 데이터를 결합해 실제 시스템 상태를 지속 추정accelerator, fusion, grid, complex flow
AI-ready Scientific Dataprovenance, metadata, semantics, 품질 정보가 붙은 학습·추론 가능한 데이터AmSC/ModCon의 Data Cards, standards, reproducibility
Multimodal Reasoning텍스트·수식·영상·그래프·센서·시뮬레이션을 함께 추론molecular cell biology, materials, particle physics
Uncertainty Quantification예측 불확실성과 모델이 모르는 영역을 수치화안전한 scientific agents와 experiment selection의 핵심
Human–AI Co-Science전문가와 AI가 역할을 분담하며 결과를 상호검증현재 기술 수준에서 현실적인 운영모델

2026년 Scientific Foundation Model 리뷰는 단순한 대형 사전학습 모델이라는 정의를 넘어 domain/problem adaptation과 generalization을 구분하고, 물리 일관성·OOD 일반화·UQ·표준 benchmark를 주요 과제로 제시한다.

§4 · Introduction

American Science and Security Platform: 과학발견 운영체계의 기반

DOE는 American Science and Security Platform을 Genesis Mission의 핵심 기술 엔진으로 정의한다. 고성능컴퓨팅, 실험시설, 데이터 자원, 생산 능력을 하나의 조정된 AI 기반 발견시스템으로 통합한다. 연구자 입장에서 중요한 두 축은 American Science Cloud(AmSC)ModCon: Transformational AI and Data이다.

AmSC는 안전한 federated science-optimized 환경으로 컴퓨팅, 실험시설, 데이터, 고성능 네트워크를 연결한다. ModCon은 그 위에서 재사용 가능한 AI R&D, 과학 워크플로 모범사례, 데이터 broker·standards, model/agent/data cards, evaluation과 협업 도구를 제공한다. AmSC가 플랫폼이라면 ModCon은 과학 AI capability layer라고 볼 수 있다.

National challengesEnergy · Materials · Biology · Physics · Grid · Water · Quantum
Scientific intelligenceFoundation Models + Agentic Reasoning + Multimodal Interfaces
ExecutionHPC Codes + Digital Twins + Autonomous Laboratories + Instruments
Evidence layerMeasurement + UQ + Provenance + Replication + Falsification
InfrastructureAmerican Science Cloud + ModCon + Federated Compute/Data

이 관점에서 Genesis Mission은 “ChatGPT for Science”라기보다 Scientific Discovery Operating System에 가깝다. DOE infrastructure 자료가 수백 개 agent의 장시간 과학 태스크와 heterogeneous compute resource matching을 직접 언급한다는 점도 이 해석을 뒷받침한다.

Part III · Motivation & Background

왜 지금 ‘과학용 폐루프 AI’인가

데이터 폭증, 계산비용, 실험 병목, 가설생성 자동화가 같은 방향에서 만나고 있다.

§5 · Motivation

2025–2026 문헌이 보여준 네 가지 압력

1. 과학 데이터의 증가 속도가 인간의 통합 능력을 넘어섰다

2025년 Nature Perspective는 genomics, transcriptomics, epigenomics, proteomics, metabolomics, spatial profiling을 함께 학습하는 multimodal foundation model을 제안한다. 중요한 점은 데이터가 많다는 사실 자체가 아니라 서로 다른 측정 맥락을 하나의 representation으로 묶어 실험 설계와 perturbation prediction까지 연결하려는 방향이다.

2. 고정밀 시뮬레이션은 정확하지만 비싸다

Aurora는 100만 시간 이상의 geophysical data로 사전학습한 Earth-system foundation model이다. 대기오염, 파랑, 열대저기압 경로, 고해상도 날씨 등 여러 downstream task에서 운영 시스템과 경쟁하거나 앞서면서 계산비용을 크게 낮췄다. 이는 scientific FM이 전통 solver를 무조건 대체한다기보다 HPC simulation과 AI surrogate를 목적에 맞게 조합할 가능성을 보여준다.

3. 실제 병목은 prediction보다 experiment feedback이다

2025년 autonomous laboratory 문헌은 database, large intelligent model, automated experimental platform, management/decision system을 결합해 predict–make–measure loop를 닫는 구조를 강조한다. 계산 결과가 다시 실험으로 이어지지 않는다면 discovery latency는 크게 줄지 않는다.

4. AI가 가설 생성과 연구 수행 단계까지 진입했다

AI Co-Scientist는 멀티 에이전트로 후보 가설을 생성·토론·진화시키고, AI Scientist-v2는 agentic tree search를 이용해 가설 설정부터 실험 실행, 분석, 논문작성까지의 자동화를 탐색했다. Robin은 이를 생물학적 wet-lab feedback과 연결했다. 이 세 흐름은 ‘모델을 과학자가 호출하는 구조’에서 ‘AI가 연구과정을 조직하는 구조’로의 이동을 보여준다.

§6 · RFA scope

21개 Topic과 99개 Focus Area: 응용문제와 공통기술을 동시에 묶는다

No.DOE Genesis RFA TopicResearch interpretation
1Reenvisioning Advanced Manufacturing and Industrial ProductivityAI-guided design, process control, extreme manufacturing
2Scaling the Biotechnology Revolutionbiological FMs, protein/cell design, biomanufacturing
3Securing America’s Critical Minerals Supplydiscovery, recovery, separation, supply-chain science
4Delivering Nuclear Energy that is Faster, Safer, Cheaperreactor design, operations, materials, safety
5Accelerating Delivery of Fusion Energyplasma/material FMs, digital twins, HPC-agent workflows
6Transforming Nuclear Restoration and Revitalizationremote sensing, robotics, planning, decommissioning
7Discovering Quantum Algorithms with AIAI-guided quantum algorithm discovery
8Realizing Quantum Systems for Discoveryquantum control, devices, networking
9Recentering Microelectronics in Americamaterials/process/device co-design
10Securing U.S. Leadership in Data Centersthermal, power, scheduling, energy optimization
11Achieving AI-Driven Autonomous Laboratoriesclosed-loop experimental science
12Designing Materials with Predictable Functionalityinverse design, synthesis, characterization
13Enhancing Particle Accelerators for Discoveryself-evolving digital twins, controls
14Unifying Physics from Quarks to the Cosmosfoundation models for large experimental data
15Predicting U.S. Water for EnergyEarth/water FMs, forecasting and management
16Scaling the Grid to Power the American Economyagentic planning, simulation, resilience
17Unleashing Subsurface Strategic Energy Assetsgeoscience FMs, subsurface digital twins
18HPC Code Curation, Translation, and Development for Accelerated Scientific DiscoveriesAI-assisted scientific software and solver modernization
19AI for Scientific Reasoninghypothesis, tool use, evidence-grounded agents
20Cybersecurity for AI-driven Science Workflowssecure autonomous research infrastructure
21AI in Fluid Flow for Energy Components and Technologiesphysics-informed models for complex flow

DOE 2026 informational webinar는 21개 Topic과 99개 Focus Area를 명시한다. Topic 1–17은 national challenge, 18–21은 American Science and Security Platform의 cross-cutting needs에 해당한다.

Part IV · Challenges & Research Questions

좋은 답보다 믿을 수 있는 과학 행동이 더 어렵다

Scientific agent의 실패는 문장 오류에서 끝나지 않는다. 잘못된 실험, 잘못된 자원 배분, 잘못된 과학적 결론으로 증폭될 수 있다.

§7 · Challenges

Phase I 이후 반드시 넘어야 할 일곱 개의 장벽

Scientific hallucination

존재하지 않는 근거를 만드는 문제보다 더 위험한 것은 과학적으로 그럴듯하지만 틀린 가설이다. 폐루프 시스템에서는 오류가 다음 실험으로 증폭된다.

OOD generalization

새 물질, boundary condition, reactor regime, biological context에서 성능이 급락할 수 있다. zero-shot scientific generalization은 아직 드물다.

Uncertainty

정답값 하나보다 epistemic/aleatoric uncertainty와 abstention이 중요하다. UQ가 없다면 agent는 모르는 상황에서도 행동한다.

Agent reliability

SciAgentArena는 잘 정의된 분석 문제에서는 가능성을 보이지만 open-ended exploration과 novel insight에서 약함을 보고한다.

Sim-to-real gap

시뮬레이션에서 좋은 후보가 실제 실험에서도 유효하다는 보장은 없다. high-fidelity simulation과 physical experiment에 의한 falsification이 필요하다.

Provenance & reproducibility

데이터, 모델 버전, prompt, solver, experimental condition, 판단근거를 추적하지 않으면 결과를 재현할 수 없다.

HPC–agent orchestration

여러 agent가 HPC job, model inference, database, instrument를 동시에 호출한다. resource scheduling과 workflow correctness가 새로운 systems problem이 된다.

Evaluation gap

ReplicationBench에서 최고 수준 모델조차 paper-scale astrophysics replication에서 20% 미만을 기록한다. 과학 연구능력 평가가 모델 benchmark보다 훨씬 어렵다는 신호다.

§8 · Research Questions

Genesis-style AI for Science를 위한 10개의 연구질문

RQ핵심 질문왜 중요한가
RQ1AI advantage를 과학적으로 어떻게 정의하고 비교할 것인가?Phase I 성공/실패 판단의 출발점
RQ2SciFM이 새로운 물리조건·물질·실험으로 얼마나 generalize하는가?benchmark 성능과 실제 활용 사이의 간극
RQ3AI가 만든 hypothesis가 plausible한지 genuinely novel한지 어떻게 판별하는가?새 지식 생성의 기준
RQ4simulation–experiment–agent loop를 어떻게 안정화할 것인가?폐루프 오류 누적 방지
RQ5uncertainty와 counter-evidence를 행동선택에 어떻게 반영할 것인가?위험한 overconfidence 완화
RQ6실험 실패와 negative result를 어떻게 장기 기억할 것인가?동일 실패 반복 방지
RQ7heterogeneous DOE data를 provenance를 보존하며 FM 학습에 연결하는 방법은?신뢰 가능한 AI-ready data
RQ8human scientist와 autonomous agent의 권한 경계를 어디에 둘 것인가?안전성과 생산성의 균형
RQ9multi-agent scientific workflow의 end-to-end correctness를 어떻게 검증할 것인가?구성요소 성능만으로 전체 시스템을 보장할 수 없음
RQ10scientific discovery speedup과 단순 compute throughput 증가를 어떻게 구분할 것인가?“더 많이 계산함”과 “더 빨리 발견함”은 다름
100배 많은 후보를 계산했다는 사실은 10배 빨리 새로운 과학적 사실을 발견했다는 뜻이 아니다.Evaluation principle
Part V · Approaches & Methods

과학 에이전트는 답을 생성하는 것이 아니라 실험 가능한 행동을 조직해야 한다

가설, 시뮬레이션, 불확실성, 실험, 반증, 기억을 하나의 control loop로 결합한다.

§9 · Reference Architecture

가장 유망한 Genesis-style closed-loop architecture

01 · ObjectiveScientist-defined goal + constraints + success criteria무엇을 발견하거나 최적화하려는지, 어떤 증거가 성공을 의미하는지 명시한다.
02 · Planner & RetrievalScientific Planner + Literature/Data/KG Retrieval관련 지식과 데이터의 provenance를 유지한 채 evidence space를 만든다.
03 · Hypothesis populationGenerate + Debate + Critique + Rank하나의 답에 조기 수렴하지 않고 경쟁 가설을 유지한다.
04 · Model & simulationSciFM + Physics Solver + Digital Twinfast surrogate와 high-fidelity solver를 비용·위험에 따라 조합한다.
05 · GateUncertainty + Physical Constraints + Safety낮은 신뢰도 또는 물리 위반 후보는 보류·재검증한다.
06 · ExperimentDesign Agent → Robot / Facility / Instrument정보가치와 비용을 함께 고려해 다음 실험을 선택한다.
07 · EvidenceStatistical/Causal Analysis → Support or Falsification관측이 기존 가설을 지지하는지, 반증하는지, 조건을 바꾸는지 기록한다.
08 · Epistemic memoryClaim + Provenance + Conditions + Counter-evidence + Revision결과를 단순 memory가 아니라 과학적 믿음의 변화로 저장한다.
§10 · Method families

여섯 가지 핵심 방법론

Scientific Foundation Models

MatterGen은 결정구조의 원자종·좌표·주기격자를 diffusion으로 생성하고 원하는 화학조성, 대칭성, 기계·전자·자기 특성으로 조건화하는 inverse design을 보여준다. 실제 합성을 통한 proof-of-concept까지 포함했다. Genesis의 predictable materials는 바로 이 생성모델을 실험 loop와 연결하는 문제다.

Multi-Agent Scientific Reasoning

AI Co-Scientist의 generate–debate–evolve와 AI Scientist-v2의 progressive agentic tree search는 하나의 LLM에 연구 전체를 맡기기보다 후보 경로를 경쟁시키고 experiment manager가 탐색을 통제하는 방향을 보여준다.

Physics-informed Digital Twins

\[\mathrm{Digital\ Twin}=\mathrm{Physics\ Solver}+\mathrm{Neural\ Surrogate}+\mathrm{Real\ Observations}\]

FRIB Phase I의 “Towards Self-Evolving, Physics-Informed Digital Twins of Ion Accelerators and Isotope Separators”는 이 아이디어가 실제 Genesis award로 연결된 대표 사례다.

Autonomous Experimentation

다음 실험을 성능 최대화만으로 선택하기보다 정보이득으로 고르는 active learning/Bayesian optimization 구조가 중요하다.

\[x_{t+1}=\arg\max_x\;\mathbb{E}[\mathrm{Scientific\ Information\ Gain}(x)]\]

Retrieval + Tool-Augmented Agents

모든 지식을 모델 파라미터에 압축하지 않고 문헌 DB, scientific KG, simulator, HPC code, laboratory API를 필요할 때 호출한다. 이때 tool result provenance와 version을 추론상태에 함께 넣어야 한다.

Uncertainty-aware Experiment Selection

\[\mathrm{Utility}(x)=\mathrm{Expected\ Discovery\ Value}-\lambda_1\mathrm{Cost}-\lambda_2\mathrm{Risk}+\lambda_3\mathrm{Information\ Gain}\]

가장 높은 예측값이 아니라 가장 높은 과학적 가치의 다음 행동을 선택하는 것이 폐루프 discovery의 핵심이다.

Part VI · Key Applications

응용분야가 달라도 반복되는 구조는 같다

가설을 만들고, 모델링하고, 실제 세계와 부딪치고, 근거를 다시 학습하는 루프다.

§11 · Applications

Genesis와 최신 연구가 만나는 여덟 개의 응용축

A. Autonomous Scientific Discovery

Robin은 질병을 입력받아 문헌을 검색하고, 실험 가능한 치료 후보를 제안하며, 실제 실험 데이터를 분석해 다음 후보를 생성하는 반복주기를 보여준다. 완전 자동 wet lab은 아니지만 hypothesis → experiment → analysis → revised hypothesis의 지적 단계를 하나의 multi-agent workflow에 묶었다.

B. Critical Materials

SLAC–USC의 실제 Genesis Phase I은 폐 리튬이온전지에서 cobalt, nickel, manganese의 선택적 회수를 위한 multi-agent framework를 구축한다. 여러 분야의 지식을 바탕으로 분리 경로를 제안하고 실험으로 학습하며, 9개월 동안 여러 AI–experiment cycle을 실행하고 목표 금속 80% 이상 purity를 지향하면서 기존 literature-search/recovery 방법과 비교한다. AI advantage를 실험 metric으로 연결한 매우 전형적인 Phase I 설계다.

C. Materials Discovery

MatterGen류 generative model과 synthesis automation을 연결하면 Desired Property → Material Generation → DFT/Surrogate → Synthesis → Measurement의 inverse-design loop가 된다. 핵심은 생성 다양성이 아니라 실제 합성·특성 측정으로 조건부 생성이 검증되는가이다.

D. Particle & Nuclear Physics

MIT–Brookhaven–Iowa State 협력 Phase I은 300 PB가 넘는 high-energy nuclear-collision data를 활용해 foundation model을 만들고 sPHENIX detector에서 입자 궤적 재구성을 목표로 한다. 거대 scientific data를 범용 representation으로 바꾸는 SciFM의 좋은 사례다.

E. Particle Accelerators

FRIB–Argonne 프로젝트는 heavy-ion accelerator와 isotope separator를 위한 self-evolving physics-informed digital twin을 개발한다. 장기적으로는 beam tuning, failure prediction, operation optimization을 autonomous control 문제로 재구성할 수 있다.

F. Electrical Grid

PowerChain은 GridLAB-D 같은 structured tools와 expert-verified reasoning trajectories를 사용해 distribution-grid analysis를 자동화하고, build–verify–feedback–correct 구조를 제시한다. 논문은 2026년 온라인 공개되었고 2027년 저널 volume에 수록된다. 이는 grid에서 “LLM 답변”보다 검증 가능한 도구 실행이 중요하다는 점을 보여준다.

G. Earth, Climate & Water

Aurora는 weather뿐 아니라 air quality, ocean waves, tropical cyclone tracks를 하나의 pretrained representation에서 fine-tune해 처리한다. Genesis의 water/energy 관련 문제에서 foundation model이 멀티 downstream task를 공유하는 기술적 선례다.

H. Biotechnology

multimodal cell foundation model에 scientific agent와 robotic laboratory를 결합하면 genotype-to-phenotype prediction, in-silico perturbation, protein/cell design, bioprocess optimization, biomarker discovery로 확장할 수 있다. 다만 실제 wet-lab 검증과 데이터 조건성 기록이 빠지면 predictive biology에 머문다.

§12 · What Phase I really tests

9개월의 목적은 ‘거대한 완성품’이 아니라 다음 투자를 정당화할 증거다

DOE의 Phase I 목표는 작은 팀이 한 번에 전국 규모 과학 인프라를 완성하는 것이 아니다. AI가 해당 과학 워크플로의 병목을 실제로 바꿀 수 있다는 최소한의 정량 증거를 만드는 것이다. 성공한 방향은 Phase II의 더 큰 팀·3년 규모 연구로 확장될 수 있고, 실패한 방향은 종료된다.

2026년 7월 DOE는 첫 RFA 응답에서 미국 전역에 걸친 거의 300개 프로젝트를 선정했다고 발표했다. 이 숫자는 Genesis가 단일 거대 모델 프로젝트가 아니라 여러 과학 workflow에서 AI advantage를 동시에 탐색하는 portfolio 전략임을 보여준다.

Part VII · Open Problems & Future Directions

다음 승부처는 더 큰 모델이 아니라 ‘반증 가능한 과학자’다

생성 능력보다 evidence, uncertainty, counter-evidence, revision이 강한 시스템이 과학적 신뢰성을 결정한다.

§13 · Open Problems

모델, 에이전트, 과학시스템의 세 층

모델 수준에서는 out-of-distribution generalization, physical consistency, causal reasoning, calibrated uncertainty, multi-fidelity learning이 남아 있다. SciFM survey가 강조하듯 foundation model의 이름만으로 새로운 문제 일반화가 보장되지는 않는다.

에이전트 수준에서는 hallucination, tool misuse, premature hypothesis convergence, confirmation bias, unsupported novelty claim, experiment-selection bias가 핵심이다. SciAgentArena와 ReplicationBench는 현재 agent가 명확히 정의된 태스크에서는 도움이 되지만 open-ended science와 paper-scale replication에서는 여전히 불안정하다는 근거를 제공한다.

과학시스템 수준에서는 더 어려운 문제가 있다. 결과가 단순 log로 남는 것이 아니라 다음 관계로 추적되어야 한다.

\[\mathrm{Evidence}\rightarrow\mathrm{Claim}\rightarrow\mathrm{Experiment}\rightarrow\mathrm{Counter\!\!-evidence}\rightarrow\mathrm{Revision}\]

따라서 future scientific agent에게 필요한 것은 단순 memory가 아니라 epistemic memory다. 무엇을 알고 있는가뿐 아니라 왜 믿는지, 어느 조건에서 성립하는지, 무엇이 그 주장을 약화·반박했는지, 그 뒤 믿음을 어떻게 수정했는지를 보존해야 한다.

§14 · Future Directions

2027년 이후 특히 중요해질 일곱 방향

01 · Falsifiable Scientific Agents

가설과 함께 supporting evidence, counter-evidence, test, falsification criterion을 생성한다. “그럴듯함”을 “검증 가능성”으로 바꾼다.

02 · Epistemically Accountable Multi-Agent Science

Planner 외에 Evidence, Skeptic, Uncertainty, Replication agent를 둬 한 방향의 합의가 과학적 확신으로 오인되지 않게 한다.

03 · Scientific World Models

text + equations + graphs + images + simulation + sensor streams + experimental actions를 하나의 과학 world representation으로 연결한다.

04 · Agent–SciFM–Digital Twin Co-Evolution

Agent, foundation model, digital twin, experiment가 서로 데이터를 주고받으며 함께 개선되는 구조를 연구한다.

05 · Uncertainty-Driven Autonomous Science

최고 predicted performance보다 최대 expected scientific information gain을 목표로 다음 실험을 선택한다.

06 · Distributed Autonomous Laboratories

하나의 self-driving lab을 넘어 여러 시설과 simulation center가 networked experimental system으로 협력하는 방향이다.

07 · Open Scientific Foundation Models

DOE는 2026년 8월 Genesis Open Models Initiative를 시작했고 첫 open-weight 모델 계열로 Genesis-Science-1을 추진한다고 발표했다. auditable·extensible SciFM이 공공 연구 인프라가 되는 흐름이다.

Evaluation · AI Advantage Science

accuracy가 아니라 discovery time, experimental yield, resource use, novelty, reproducibility, calibrated uncertainty를 함께 측정하는 새로운 평가과학이 필요하다.

§15 · 2025–2026 papers

Genesis Phase I을 이해하는 핵심 문헌 지도

WorkYear / venue핵심 의미Genesis connection
Towards an AI co-scientist2025 · arXivmulti-agent hypothesis generationAI for Scientific Reasoning
The AI Scientist-v22025 · arXivagentic tree search 기반 end-to-end ML researchautonomous research workflow
MatterGen2025 · Natureproperty-conditioned inorganic material generationpredictable materials design
Aurora2025 · NatureEarth-system foundation modelwater/weather/energy modeling
Multimodal FMs in molecular cell biology2025 · Natureomics 통합 foundation model visionbiotechnology
Autonomous laboratories in China2025 · Digital Discoverypredict–make–measure loopautonomous laboratory
UQ for NNP foundation models2025 · npj Computational Materialsepistemic/aleatoric uncertainty methodstrustworthy SciFM
ReplicationBench2025 · arXivpaper-scale agent reproducibilityAI-advantage verification
Robin2026 · Naturehypothesis–experiment–analysis multi-agent loopautonomous scientific discovery
Automated Scientific Discovery2026 · Machine Learningequation discovery에서 autonomous systems까지 taxonomyconceptual framework
On Scientific Foundation Models2026 · Neural NetworksSciFM 정의·generalization·evaluationModCon / model layer
SciAgentArena2026 · arXiv실제 과학 태스크에서 agent benchmarkgo/no-go evaluation
AutoResearchBench2026 · arXivscientific literature discovery benchmarkevidence retrieval
PowerChainonline 2026 · EPSR 2027verifiable agentic grid analysisgrid scientific tool orchestration
§16 · Synthesis

Scientific Foundation Model보다 더 큰 연구기회

Genesis Mission Phase I의 학술적 의미를 한 문장으로 줄이면 다음과 같다.

AI 모델의 성능 경쟁에서 벗어나, AI가 실제 과학적 발견 과정을 얼마나 빠르고 정확하며 검증 가능하게 변화시키는지를 실험적으로 증명하는 단계다.Interpretation across DOE documents and 2025–2026 literature
\[\boxed{\mathrm{AI\!\!-ready\ Data}\rightarrow\mathrm{SciFM}\rightarrow\mathrm{Multi\!\!-Agent\ Reasoning}\rightarrow\mathrm{Simulation/Digital\ Twin}\rightarrow\mathrm{Autonomous\ Experiment}\rightarrow\mathrm{Evidence/Falsification}\rightarrow\mathrm{Self\!\!-Improvement}}\]

현재 문헌에서 가장 빠르게 성장한 부분은 가설 생성과 도구 사용이다. 반면 가설을 반증하고, uncertainty와 counter-evidence를 관리하고, 실패한 실험을 장기 기억하며, 실제 discovery speedup을 객관적으로 증명하는 능력은 여전히 약하다. 이것이 SciAgentArena, ReplicationBench 같은 평가 연구가 중요한 이유다.

따라서 향후 높은 학술적 신규성을 가질 가능성이 큰 축은 단순 SciFM 자체보다 Falsifiable + Epistemically Accountable + Closed-Loop Autonomous AI Scientist다. 모델이 얼마나 유창한지가 아니라, 어떤 근거에서 행동했고 무엇이 그 믿음을 바꿨으며 그 결과 과학적 발견이 실제로 얼마나 앞당겨졌는지를 설명할 수 있어야 한다.

References & Primary Sources

DOE / U.S. Government
01
THE WHITE HOUSE · 2025-11-24
Genesis Mission의 목적, American Science and Security Platform, national challenge 선정의 정책적 근거.
02
THE WHITE HOUSE · 2025-11-24
closed-loop AI experimentation platform, supercomputing, foundation models, robotic laboratories의 공식 요약.
03
DOE OFFICE OF SCIENCE · 2026
Phase I/II RFA의 공식 공고.
04
DOE · 2026-03-31
21 topics, 99 focus areas, Phase I AI advantage 평가, 9개월 수행기간, 6개월 go/no-go와 award 규모의 근거.
05
U.S. DEPARTMENT OF ENERGY · 2026
Genesis Mission의 핵심 기술 엔진과 통합 플랫폼 설명.
06
U.S. DEPARTMENT OF ENERGY · 2026
AmSC, AI/data capability layer, Agent/Data/Model Cards와 재현성·검증 서비스.
07
U.S. DEPARTMENT OF ENERGY · 2026
에너지, 발견과학, 국가안보를 포괄하는 challenge portfolio.
08
U.S. DEPARTMENT OF ENERGY · 2026
award list, challenge documents, AmSC, ModCon, open models의 공식 자료 허브.
09
U.S. DEPARTMENT OF ENERGY · 2026-08-07
open-weight scientific foundation models와 Genesis-Science-1 계획.
10
U.S. DEPARTMENT OF ENERGY · 2026
본 RFA Phase I과 구분되는 중소기업 대상 별도 Genesis 연계 Phase I.
Genesis Phase I Examples
11
SLAC NATIONAL ACCELERATOR LABORATORY · 2026-07-22
실험과 결합한 multi-agent metal recovery Phase I의 구체적 AI-advantage 설계.
12
FACILITY FOR RARE ISOTOPE BEAMS · 2026-07-23
heavy-ion accelerator와 isotope separator의 physics-informed digital twin.
13
IOWA STATE UNIVERSITY · 2026-07-23
300 PB+ nuclear-collision data와 sPHENIX track reconstruction을 다루는 Phase I.
2025–2026 Research Papers
14
GOTTWEIS ET AL. · 2025
Gemini 2.0 기반 multi-agent hypothesis generation과 scientific proposal reasoning.
15
YAMADA ET AL. · 2025
progressive agentic tree search와 experiment manager를 사용한 end-to-end research automation.
16
NATURE · 2025
property-conditioned diffusion 기반 inverse materials design과 합성 검증.
17
NATURE · 2025
100만 시간 이상의 geophysical data로 pretrain한 multi-task Earth-system FM.
18
NATURE · 2025
다양한 omics와 spatial modalities를 통합하는 foundation model vision.
19
DIGITAL DISCOVERY · 2025
AI, robotics, database, decision systems의 predict–make–measure closed loop.
20
NPJ COMPUTATIONAL MATERIALS · 2025
readout ensemble과 quantile regression을 통한 uncertainty estimation.
21
YE ET AL. · 2025
paper-scale, expert-validated replication benchmark. 최고 모델도 20% 미만 수준을 보고.
22
NATURE · 2026
문헌 기반 가설 생성과 실제 생물학 실험 데이터 분석을 연결한 반복적 therapeutic discovery.
23
MACHINE LEARNING · 2026
symbolic/equation discovery에서 autonomous discovery systems까지의 리뷰.
24
NEURAL NETWORKS · 2026
SciFM의 정의, adaptation/generalization, applications, evaluation challenges.
25
LIU ET AL. · 2026
약 200개 real-world scientific tasks에서 agent의 강점과 open-ended reasoning 한계를 평가.
26
XIONG ET AL. · 2026
Deep Research와 Wide Research로 scientific literature discovery를 평가.
27
ELECTRIC POWER SYSTEMS RESEARCH · ONLINE 2026 / VOL. 262, 2027
structured power tools, expert-verified trajectories, feedback-driven correction을 결합한 grid agent.