AI Research NotesDrug Discovery · AI Co-ScientistMultimodal FM · Multi-Agent · Falsification · Lab-in-the-Loop
2025–2026 Research Trends/Drug Discovery AI Co-Scientist/Survey date 2026.08.23

모델 하나가 아니라,
과학의 폐루프를 설계한다

Multimodal Multi-Agent AI Co-Scientists for Drug Discovery — scientific foundation models as instruments, hypotheses as competing beliefs, experiments as decisions, and evidence as memory.

Drug-discovery AI Co-Scientist closed loop멀티모달 과학 파운데이션 모델이 과학 세계를 표현하고, 에이전트들이 가설을 생성·비판·선택하며, 전문 도구와 실험을 수행하고, 결과와 반증 증거가 scientific memory를 갱신해 가설을 수정하는 폐루프를 나타낸다. WORLD MODELSMolecule · Protein3D · Omics · ImageAssay · LiteratureMAMMAL · PhenoModel · TxGemma HYPOTHESIS SOCIETYGenerateRetrieve EvidenceSkeptic / FalsifyRank / Evolvescientific decision, not chat INSTRUMENTSBoltz-2 · DockingMD · FEP · ADMETSearch · Code · KGchoose what to trust LABDesignMake / TestObservepositive + negative MEMEvidenceProven.Uncert. counter-evidence → belief revision → next hypothesismeta-scientific reasoning selects models, evidence and experiments
Survey Thesis

2025년 이후 신약개발 AI의 질문이 달라졌다. “이 모델이 affinity를 얼마나 잘 맞히는가”에서 “여러 모델과 증거와 실험을 이용해 연구 과정 자체를 얼마나 잘 운영하는가”로 이동한다. 이 변화의 중심에는 하나의 거대한 모델이 아니라, 과학 모델을 도구처럼 고르고 가설을 경쟁시키며 실험 결과로 믿음을 수정하는 멀티에이전트 AI Co-Scientist가 있다.

첨부 자료의 핵심 결론은 분명하다. 미래의 Co-Scientist는 multimodal foundation model, multi-agent reasoning, 전문 계산 도구, wet-lab automation, provenance memory를 따로 쓰는 시스템이 아니다. 이들을 Prediction → Decision → Experiment → Revision의 하나의 폐루프로 연결하는 시스템이다.

Central Claim

Foundation Model은 과학적 세계를 표현한다. Agent는 다음 행동을 결정한다. Co-Scientist는 가설·증거·도구·실험·반증·수정을 하나의 과학적 의사결정 과정으로 운영한다. 이 세 층위를 구분해야 연구의 신규성과 한계를 정확히 볼 수 있다.

Sequential scientific decision problem
\[(h^*,x^*,e^*)=\arg\max_{h,x,e}\;\mathbb E\!\left[U_{\mathrm{scientific}}+U_{\mathrm{therapeutic}}-\lambda C_{\mathrm{experiment}}-\mu R_{\mathrm{safety}}\mid D,M,T\right].\]
3 multimodalitiesrepresentation, evidence, action을 함께 연결해야 Co-Scientist의 multimodality가 완성된다.
6 core functionsrepresentation, hypothesis society, scientific instruments, evidence grounding, falsification, lab-in-the-loop.
10 research questions표현부터 실험 선택, cross-model verification, accountability, human–AI division까지 이어진다.
L3 → L4/L5현재 가장 큰 연구 공백은 multi-agent co-scientist에서 closed-loop experimental scientist로 넘어가는 경계다.

이 글은 첨부된 2025–2026 연구 동향 문서의 전체 내용을 웹 읽기 흐름으로 재구성한 것이다. 논문별 수치와 사례는 첨부 자료의 서술에 근거하며, 외부 문헌을 별도로 재검증해 수치를 추가하거나 교정하지 않았다.

Part I · Definition & Transition

2025년의 전환:
Model-centric AI에서 Science-centric AI로

The unit of progress shifts from a predictor to a scientific organization.

Scientific Multimodal Foundation Model은 단순한 VLM이 아니다

신약개발의 과학특화 멀티모달 파운데이션 모델은 molecule, protein, cell, tissue, disease를 서로 다른 표상으로 다루는 범용 사전학습 계층이다. 입력과 증거의 공간은 SMILES, molecular graph, 3D structure, protein sequence, DNA/RNA, omics, cell image, assay, text를 가로지른다.

Scientific modalities
\[\mathcal M=\{\text{SMILES, Molecular Graph, 3D Structure, Protein Sequence, DNA/RNA, Omics, Cell Image, Assay, Text}\}.\]

첨부 자료는 multimodality를 세 층위로 나눈다. 첫째는 molecule–protein–3D–omics–image를 함께 표현하는 Representation Multimodality다. 둘째는 paper–assay–clinical–toxicity–omics처럼 서로 다른 증거를 통합하는 Evidence Multimodality다. 셋째는 search–code–docking–simulation–robotic experiment처럼 서로 다른 행동을 수행하는 Action Multimodality다.

Representation

세계를 표현한다

MAMMAL은 protein, small molecule, gene, transcriptomic information을 cross-modal하게 다루고, PhenoModel은 molecular structure와 cell morphology를 연결한다.

Evidence

증거를 묶는다

문헌, assay, 임상, 독성, omics가 같은 결론을 지지하거나 반박하는 근거망으로 연결된다.

Action

과학 행동을 선택한다

검색, 코드 실행, docking, simulation, 실험 장비 실행이 동일한 scientific workflow 안에 들어온다.

AI Co-Scientist는 연구자용 챗봇과 다르다

첨부 자료의 정의를 압축하면 AI Co-Scientist는 주어진 과학적 목표에 대해 문헌과 데이터를 탐색하고, 새로운 가설을 만들고, 경쟁 가설을 비판·순위화하고, 필요한 계산이나 실험을 설계·실행하며, 결과에 따라 기존 가설을 수정하거나 폐기할 수 있는 자율적 또는 반자율적 과학추론 시스템이다.

Google AI Co-Scientist는 Generate → Debate → Evolve 구조와 Generation, Reflection, Ranking, Evolution 같은 역할 분리를 보여 주었다. Robin은 literature → hypothesis → experimental proposal → wet-lab result → autonomous data analysis → revised hypothesis라는 연속적인 feedback loop를 실험생물학에 연결했다.

Foundation Model과 Co-Scientist는 같은 것이 아니다. 전자는 세계를 표현하고, agent는 행동을 선택하며, Co-Scientist는 과학의 전체 과정을 운영한다.

왜 2025년이 전환점인가

첨부 자료는 2023–2024년의 중심 질문을 “AI가 분자나 단백질을 얼마나 잘 예측하거나 생성하는가?”로 요약하고, 2025년 이후 질문이 “AI가 여러 모델과 데이터를 이용해 과학자의 연구과정 자체를 수행할 수 있는가?”로 바뀌었다고 본다.

System / ModelYearSource document’s key advance
AI Co-Scientist2025Generate–Debate–Evolve 기반 scientific hypothesis society.
TxGemma / Agentic-Tx2025therapeutic foundation model을 더 큰 agentic workflow의 도구로 사용.
BioDiscoveryAgentICLR 2025실험 결과를 보고 다음 genetic perturbation을 선택.
PharmAgents2025target → lead → preclinical의 virtual-pharma multi-agent pipeline.
RobinNature 2026hypothesis–experiment–analysis–revision 폐루프의 초기 실증.
BiomniScience 2026biomedical tools·DB·code·wet-lab를 통합하는 범용 agent.
DiscoVerse2026장기 제약 R&D archive에 근거한 traceable co-scientist.
ChatInvent2026실제 discovery pipeline의 multi-agent molecular design과 synthesis planning.
MAMMAL2026molecule–protein–gene 등 cross-modal biomedical foundation model.
PhenoModel2025/26molecular structure–cell morphology multimodal foundation model.

이 흐름을 첨부 자료는 Model-centric AI → Agent-centric AI → Science-centric AI로 정리한다.

Part II · Problem & Motivation

예측 문제를 넘어
Sequential Scientific Decision Problem으로

Drug discovery is not a single prediction; it is a sequence of decisions under uncertainty.

기존 AI 신약개발은 흔히 \((\text{protein},\text{ligand})\rightarrow\text{binding affinity}\) 또는 \(\text{molecule}\rightarrow\text{toxicity}\) 같은 정적 prediction을 풀었다. 실제 연구자의 질문은 훨씬 길다. disease mechanism을 이해하고, 아직 충분히 탐색되지 않은 target을 고르고, binding affinity·selectivity·ADMET·synthesizability를 함께 고려해 candidate를 고르며, 현재 uncertainty를 가장 잘 줄일 next assay까지 선택해야 한다.

즉 최적화 대상은 candidate 하나가 아니다. hypothesis \(h\), intervention \(x\), next experiment \(e\)의 묶음이다. scientific utility와 therapeutic utility를 키우되 experiment cost와 safety risk를 함께 고려해야 한다. 그리고 모든 결론에는 provenance와 uncertainty가 따라야 하며, 새로운 결과가 기존 믿음과 충돌하면 belief revision이 일어나야 한다.

01 · Predict구조·affinity·toxicity·phenotype 등 부분 모델의 예측을 얻는다.
02 · Decide가설과 candidate를 비교하고 우선순위를 정한다.
03 · Design어떤 계산 또는 실험이 uncertainty를 가장 줄일지 정한다.
04 · Executedocking, MD/FEP, assay, omics, wet-lab을 수행한다.
05 · Observepositive와 negative result를 모두 evidence로 취급한다.
06 · Revise가설·candidate ranking·model trust를 갱신한다.

신약개발은 원래부터 multimodal problem이다

disease mechanism에는 genomics와 transcriptomics가 필요하고, target에는 protein sequence와 3D structure가 필요하며, compound에는 graph·SMILES·conformation이 필요하다. biological effect는 Cell Painting, assay, pathology로 관찰되고, translational decision에는 toxicity, PK/PD, clinical evidence와 문헌이 들어간다. 최근 foundation model은 원래 존재하던 이 heterogeneity를 계산적으로 한 구조 안에 놓기 시작한 것이다.

신약개발은 원래부터 multi-agent problem이다

실제 제약 연구조직은 medicinal chemist, structural biologist, disease biologist, bioinformatician, toxicologist, DMPK scientist, statistician, clinician의 협업으로 움직인다. 첨부 자료는 그래서 monolithic agent보다 전문 역할이 분화된 scientific organization architecture가 자연스럽다고 본다. PharmAgents의 “virtual pharma”, ChatInvent의 multi-agent 확장이 이 관점을 구체화한다.

가장 비싼 자원은 계산이 아니라 잘못된 실험이다

prediction accuracy를 조금 더 높이는 것보다 “다음에 어떤 실험을 해야 가장 많은 정보를 얻는가?”가 더 중요한 경우가 있다. 첨부 자료는 향후 metric이 AUROC 하나보다 다음과 같은 방향에 가까워질 수 있다고 본다.

Value of scientific information
\[\frac{\text{Scientific Information Gained}}{\text{Experiment Cost}+\text{Time}}\]

BioDiscoveryAgent가 genetic perturbation experiment selection을 직접 문제로 삼은 이유도 여기에 있다. 첨부 자료는 Claude 3.5 Sonnet 기반 시스템이 6개 dataset에서 기존 Bayesian-optimization baseline 대비 평균 21% 향상을 보고했다고 정리한다.

Part III · Core Concepts & Evidence

과학적 세계를 표현하고,
가설을 경쟁시키고, 증거로 묶는다

A Co-Scientist needs more than generation: it needs representation, instruments, evidence, falsification and laboratory feedback.

첨부 자료는 이 분야의 core concept를 여섯 가지 기능으로 정리한다. 이 여섯 기능이 따로 존재하는 것이 아니라 하나의 scientific loop 안에서 연결될 때 Co-Scientist가 된다.

01 · Representation

Multimodal Scientific Representation

MAMMAL과 PhenoModel 같은 모델이 molecule, protein, gene, phenotype을 richer world representation으로 연결한다.

02 · Society

Hypothesis Society

Generator, Evidence Retriever, Skeptic, Experiment Designer, Statistician, Safety Critic, Judge가 서로 다른 역할을 맡는다.

03 · Instruments

Foundation Models as Scientific Instruments

Boltz-2, TxGemma, MAMMAL, docking, MD, FEP, ADMET, search를 문제에 따라 선택한다.

04 · Evidence

Evidence-Grounded Reasoning

scientific claim을 문헌·assay·clinical archive의 source-linked evidence와 연결한다.

05 · Falsification

Falsification, Not Just Generation

가설을 많이 만드는 능력보다 틀린 가설을 제거하는 Skeptic/Counter-Evidence 역할을 중시한다.

06 · Loop

Lab-in-the-Loop

Hypothesis → Experiment → Observation → Revision을 반복해 실험결과가 다음 가설의 input이 되게 한다.

Evidence-grounding의 실제 사례: DiscoVerse

첨부 자료에 따르면 Roche의 DiscoVerse는 Preclinical, Clinical 등 역할이 다른 agent와 Supervisor를 사용하면서 제약회사 내부 자료를 source-linked evidence로 연결한다. 180개 molecule과 40년 이상의 R&D archive를 대상으로 평가했고, 7개 benchmark query에서 recall ≥ 0.99, precision 0.71–0.91을 보고했다.

이 수치의 의미는 단순 QA 성능보다 traceability에 있다. 과학적 conclusion이 어느 source에 기대고 있는지 보이지 않으면 Co-Scientist의 판단을 검증하거나 반박하기 어렵다.

Lab-in-the-loop의 실제 사례: Robin

첨부 자료는 Robin이 dry age-related macular degeneration 연구에서 RPE phagocytosis를 치료 전략으로 제안하고 ripasudil과 KL001을 실험적으로 확인한 뒤, RNA-seq 결과를 분석해 ABCA1이라는 후속 mechanistic hypothesis까지 생성했다고 정리한다. 결과가 report의 마지막 줄이 아니라 다음 hypothesis의 시작점이 된 것이다.

가설을 만드는 능력만으로는 과학자가 되지 못한다. 과학적 시스템은 틀릴 수 있어야 하고, 틀렸다는 증거를 보존해야 하며, 그 증거 때문에 다음 행동을 바꿀 수 있어야 한다.이 문장은 첨부 자료의 falsification·negative memory·lab-in-the-loop 논지를 종합한 해석이다.
Part IV · Architecture & Methods

하나의 거대 모델보다
계층적 과학 시스템이 유망하다

Supervisor, specialist agents, scientific foundation models, tools and wet-lab form a hierarchy of scientific control.

첨부 문헌을 종합한 가장 유망한 구조는 monolithic super-model이 아니라 계층적 orchestration이다. Research Objective를 Supervisor가 분해하고, Hypothesis Agent·Evidence Agent·Skeptic Agent가 경쟁·협력한다. Experiment Designer와 Model/Tool Selection Agent가 다음 행동을 설계하고, MAMMAL·Boltz-2·TxGemma·PhenoModel 같은 전문 모델과 docking·MD·FEP·ADMET·omics·search를 호출한다. wet-lab 결과는 counter-evidence와 uncertainty로 다시 올라와 hypothesis revision을 일으킨다.

Objectiveresearch question과 therapeutic goal을 정의한다.
Supervisor문제를 분해하고 전문 agent와 tool budget을 배분한다.
Hypothesis Societygenerate, retrieve, critique, rank, evolve를 수행한다.
Scientific InstrumentsFM·docking·MD/FEP·ADMET·omics·search를 호출한다.
Wet Labdecisive experiment를 수행하고 관측값을 만든다.
Revisioncounter-evidence와 uncertainty를 memory에 반영해 belief를 수정한다.

Method 1 — Generate–Debate–Evolve

Google AI Co-Scientist가 대표적이다. 여러 candidate hypothesis를 생성하고 pairwise comparison과 critique로 competition시킨 뒤 더 좋은 hypothesis를 진화시킨다. 핵심은 single-shot generation보다 selection pressure를 넣는 데 있다.

Method 2 — Hierarchical Supervisor Architecture

Supervisor가 연구문제를 decomposition하고 전문 agent에 할당한다. DiscoVerse와 enterprise multi-agent workflow에 특히 적합한 방식으로 첨부 자료는 설명한다.

Method 3 — Tool-Augmented ReAct

Reason → Tool Call → Observe → Revise를 반복한다. TxGemma/Agentic-Tx와 최근 drug-discovery agent 연구가 이 구조와 밀접하다고 자료는 설명한다.

Method 4 — Cross-modal Foundation Modeling

MAMMAL은 heterogeneous biomedical entity를 통합하고, PhenoModel은 molecule–phenotype image alignment를 수행한다. 첨부 자료는 이들을 agent가 호출할 수 있는 learned scientific simulator 혹은 representation instrument로 본다.

Method 5 — Retrieval + Provenance

단순 RAG에서 한 걸음 더 나아가 모든 scientific claim에 source를 연결한다. DiscoVerse는 이를 pharmaceutical archive 수준에서 구현한 사례로 제시된다.

Method 6 — Agentic Experimental Design

BioDiscoveryAgent는 이전 perturbation 결과를 다음 prompt에 넣고 next experiment를 선택한다. 첨부 자료는 이를 LLM agent를 acquisition function처럼 사용하는 중요한 방향으로 해석한다.

Method 7 — Lab-in-the-loop Hypothesis Revision

Robin의 경우 실험결과가 시스템의 final output이 아니라 다음 가설의 input이 된다. 이 차이가 “assistant”와 “closed-loop scientist”를 가르는 핵심 중 하나다.

Part V · Research Questions & Applications

무엇을 맞힐 것인가보다
무엇을 실험할 것인가

The frontier questions move from representation accuracy to causal choice, cross-model verification and prospective science.

10개의 핵심 Research Question

RQQuestion
RQ1 · Multimodal Representation분자–단백질–세포–질병–환자 수준을 하나의 representation 또는 interoperable latent space로 연결할 수 있는가?
RQ2 · Hypothesis Generation기존 논문의 재조합이 아니라 novel, plausible, testable한 hypothesis를 만들 수 있는가?
RQ3 · Causal Discoverycorrelation과 causal intervention을 구분해 druggable target을 선택할 수 있는가?
RQ4 · Agent Collaborationdebate, tournament, supervisor, swarm 중 어떤 협력 구조가 scientific validity를 가장 높이는가?
RQ5 · Experiment Selection한정된 비용에서 어떤 실험이 hypothesis space를 가장 빨리 줄이는가?
RQ6 · Cross-model VerificationBoltz-2, MAMMAL, TxGemma, docking, MD/FEP의 결과가 충돌할 때 어떻게 판단할 것인가?
RQ7 · Closed-loop Learningnegative result까지 long-term memory에 저장하고 이후 hypothesis generation을 바꿀 수 있는가?
RQ8 · Prospective EvaluationQA benchmark가 아니라 실제 target/hit과 검증 실험 성공률로 평가할 수 있는가?
RQ9 · Accountability각 claim에 evidence, provenance, uncertainty, counter-evidence를 자동 부착할 수 있는가?
RQ10 · Human–AI Division어디까지 autonomous하게 하고 어느 지점에서 scientist approval을 요구할 것인가?

첨부 자료는 특히 RQ5와 RQ7, 즉 experiment selection과 negative-result-aware closed-loop learning을 assistant와 co-scientist를 가르는 중요한 경계로 본다.

주요 응용 영역

Target

Target Identification & Validation

omics, literature, KG, perturbation evidence를 합쳐 disease-driving target을 찾고 검증한다. AI Co-Scientist의 liver fibrosis 사례가 예로 제시된다.

Repurposing

Drug Repurposing

drug–disease–target–phenotype evidence를 재조합해 새로운 indication을 찾는다. AML과 dAMD 사례가 언급된다.

Hit

Hit Discovery / Selection

Target → Virtual Library → Affinity → Phenotype → ADMET → Hit을 여러 모델로 평가한다. PhenoModel은 phenotype-aware hit discovery 가능성을 보여 준다.

Lead

Lead Optimization

Affinity, Selectivity, ADMET, PK, Toxicity, Synthesizability를 동시에 최적화한다. PharmAgents는 초기 virtual-pharma pipeline을 제시한다.

Phenotype

Phenotypic Drug Discovery

PhenoModel/PhenoScreen은 osteosarcoma와 rhabdomyosarcoma cell line에서 phenotypically active compound 탐색 사례를 제시한다.

Genetics

Genetic Perturbation

BioDiscoveryAgent가 어느 gene 또는 gene combination을 perturb할지 반복적으로 선택해 target biology와 validation을 연결한다.

Reverse

Pharmaceutical Reverse Translation

clinical observation → mechanism → target → next design으로 되돌아가며, 실패한 R&D program을 negative scientific memory로 활용한다.

Design

Real-world Molecular Design

ChatInvent는 actual discovery pipeline에서 molecular design과 synthesis planning을 지원하고 multi-agent architecture로 확장된다.

Part VI · Challenges & Open Problems

Multi-Agent라고 해서
증거가 여러 개가 되는 것은 아니다

Reliability depends on epistemic diversity, uncertainty propagation and prospective validation—not agent count.

현재 가장 큰 난제는 model size가 아니라 신뢰할 수 있는 과학적 폐루프다. 첨부 자료는 cross-modal alignment, missing modality, assay/domain shift, causal reasoning, hallucination, correlated agent failure, uncertainty propagation, long-horizon coherence, prospective validation, provenance/reproducibility, safety/dual use, human–AI governance를 주요 challenge로 든다.

ChallengeCore issue
Cross-modal alignmentmolecule, protein, image, omics는 semantic scale 자체가 다르다.
Incomplete multimodality모든 compound에 3D·omics·phenotype data가 존재하지 않는다.
Assay/domain shiftlaboratory, protocol, cell line에 따라 data distribution이 바뀐다.
Causal reasoningcorrelation에서 intervention target을 구별하기 어렵다.
Hallucination잘못된 reference나 mechanism이 scientific hypothesis처럼 보일 수 있다.
Correlated agent failure동일 foundation model을 쓰는 agent들이 같은 오류를 반복할 수 있다.
Uncertainty propagationdocking → ADMET → agent decision을 거치며 uncertainty가 누적된다.
Long-horizon coherence수십~수백 단계 연구에서 초기 assumption을 잃기 쉽다.
Prospective validationretrospective benchmark가 실제 discovery success를 보장하지 않는다.
Provenance / Reproducibilitydata·model·version·prompt가 conclusion을 만든 경로를 기록해야 한다.
Safety / Dual usebiological design 능력의 증가는 안전관리와 직접 연결된다.
Human–AI governanceAI 제안과 human approval의 경계를 정의해야 한다.
Agent count is not evidence diversity
\[5\times\text{Agent}\;\neq\;5\times\text{Independent Evidence}.\]

같은 model family를 쓰는 다섯 agent는 같은 latent bias를 공유할 수 있다. 그래서 첨부 자료는 Agent Diversity보다 Epistemic Diversity를 요구한다. 서로 다른 model family, evidence source, physics-based simulator, symbolic rule, experimental data가 서로를 검증해야 한다.

Boltz-2의 독립 평가 사례는 이 교훈을 선명하게 만든다. 첨부 자료에 따르면 2026년 대규모 독립 평가에서 일부 affinity 영역은 weak-to-moderate correlation을 보였고, top-ranked ligand 수준에서 physics-based free-energy 계산과의 일치가 충분하지 않았다고 보고됐다. 따라서 어떤 scientific foundation model의 output도 “사실”이 아니라 evidence 중 하나로 다뤄야 한다.

여섯 개의 열린 문제

  • 01FM–Agent coupling. MAMMAL·PhenoModel 같은 multimodal FM과 AI Co-Scientist·Robin·DiscoVerse 같은 agent architecture는 아직 대부분 “Agent → Scientific Model as Tool” 수준으로 느슨하게 연결돼 있다.
  • 02Novelty vs truth. 새로운 가설은 쉽게 만들 수 있지만 \(\text{Novel}\not\Rightarrow\text{Correct}\)다. falsifiability, causal plausibility, experimental validation이 함께 평가돼야 한다.
  • 03Debate is not independent verification. Generator와 Critic이 같은 foundation model이면 동일한 bias를 공유한다. LLM critic + causal model + KG reasoner + physics simulator + experimental critic처럼 epistemic mechanism을 달리해야 한다.
  • 04Chain-level uncertainty. affinity와 ADMET가 모두 불확실한데 마지막 agent가 candidate를 단정하면 안 된다. claim-level uncertainty graph가 필요하다.
  • 05Negative result memory. 실패한 실험을 기억하지 않으면 system은 같은 실패를 반복한다. 장기 R&D archive는 미래 Co-Scientist의 중요한 negative memory가 된다.
  • 06Retrospective-to-prospective gap. random split보다 scaffold/temporal split, distribution shift, activity cliff, calibration을 보아야 하며, 최종 평가는 “AI 없이 할 때보다 얼마나 적은 실험으로 검증된 새 지식을 얻었는가”에 가까워야 한다.
Part VII · Future Directions

다음 단계는 더 큰 LLM이 아니라
더 과학적인 의사결정 구조다

Federated scientific models, causal world models, falsification-first agents, value-of-information, epistemic memory and human-in-command.

1. One Giant Model → Federation of Scientific Foundation Models

첨부 자료는 모든 것을 하나의 model이 담당하는 구조보다 Orchestrator와 여러 Specialized FM이 협력하는 federation을 더 유망하게 본다. 한 연구문제에서 MAMMAL은 disease/molecule representation을, AlphaGenome-type model은 regulatory biology를, Boltz-type model은 protein–ligand structure를, PhenoModel은 cellular phenotype을, TxGemma는 therapeutic property reasoning을 맡을 수 있다.

Federated scientific intelligence
\[\boxed{\text{Co-Scientist}=\text{Orchestrator}+\sum_i\text{Specialized FM}_i}.\]

따라서 핵심 능력은 “무엇을 예측하는가”뿐 아니라 어떤 모델을 어느 상황에서 얼마나 믿을 것인가를 판단하는 Meta-Scientific Reasoning이 된다.

2. Multimodal Model → Multiscale Causal World Model

현재 multimodal FM은 representation alignment에 강하다. 다음 단계는 Molecule → Protein → Pathway → Cell → Tissue → Patient의 multiscale causal relation을 다루는 biomedical world model이다. 그렇게 되면 “이 ligand가 target에 붙는가?”에서 “이 perturbation이 downstream pathway와 어떤 환자군의 efficacy/toxicity를 어떻게 바꾸는가?”로 질문을 확장할 수 있다.

3. Generative Agent → Falsification-First Co-Scientist

향후 가장 중요한 agent는 Generator보다 Falsifier일 수 있다. hypothesis \(H\)가 만들어지면 supporting evidence, contradicting evidence, hidden confounder, alternative explanation, decisive experiment를 별도 mechanism이 찾는 구조다.

Scientific reasoning loop
\[\text{Generate}\rightarrow\text{Attack}\rightarrow\text{Falsify}\rightarrow\text{Experiment}\rightarrow\text{Revise}.\]

4. Accuracy Optimization → Value-of-Information Optimization

가장 정확한 candidate를 고르는 것보다 “어떤 experiment가 uncertainty를 가장 많이 줄이는가”를 계산해야 한다. 첨부 자료는 Bayesian experimental design과 AI agent를 결합한 다음 목적을 제안한다.

Value-of-information experiment selection
\[e^*=\arg\max_e\frac{\operatorname{Expected\ Information\ Gain}(e)}{\operatorname{Cost}(e)+\operatorname{Time}(e)+\operatorname{Risk}(e)}.\]

5. Scientific Memory → Evidence / Provenance Graph

단순 vector memory는 “무엇을 기억하는가”를 저장한다. Co-Scientist는 “왜 그것을 믿는가”까지 저장해야 한다. 첨부 자료가 제안하는 Scientific Epistemic Memory는 hypothesis가 어떤 paper와 assay의 지지를 받고, 어떤 assay에 반박되며, 어떤 assumption에 의존하고, 어떤 model version이 예측했고, 어떤 experiment가 검증했으며, 어떤 uncertainty를 가지고 어떤 hypothesis로 수정됐는지를 graph로 남기는 방식이다.

Epistemic edgeExample meaning
supportsPaper P31, Assay A12가 Hypothesis H17을 지지한다.
contradictsAssay A19가 H17을 반박한다.
depends_onH17이 Assumption C3에 의존한다.
predicted_byBoltz-2 특정 version이 prediction을 만들었다.
tested_byExperiment E8이 hypothesis를 직접 시험했다.
uncertaintyclaim-level uncertainty를 저장한다.
revised_toH17이 새로운 evidence 때문에 H24로 수정된다.

6. Human-in-the-Loop → Human-in-Command

완전한 human removal보다 AI가 탐색하고 인간이 research objective, safety boundary, decisive experiment를 통제하는 구조가 현실적이라고 첨부 자료는 본다. governance는 automation의 반대말이 아니라 scientific accountability의 일부다.

7. Autonomous AI → Self-Driving Drug-Discovery Laboratory

궁극적인 loop는 Disease → Hypothesis → Target → Design → Make → Test → Analyze → Revise다. 2026년 Nature Chemical Biology 논평도 agentic AI와 laboratory automation의 결합을 self-driving laboratory와 adaptive DMTA cycle의 방향으로 논의한다고 첨부 자료는 설명한다.

현재의 성숙도: L0에서 L5까지

LevelFormSource document’s status
L0Predictor이미 성숙
L1Scientific Assistant널리 가능
L2Tool-using Agent빠르게 성숙
L3Multi-Agent Co-ScientistAI Co-Scientist, DiscoVerse, ChatInvent
L4Closed-loop Experimental ScientistRobin, Biomni 등이 초기 실증
L5Autonomous Drug-Discovery Scientist아직 미해결

첨부 자료가 가장 높은 연구가치를 두는 곳은 L3에서 L4/L5로 넘어가는 경계다. 단순히 LLM agent 하나를 더 만드는 것이 아니라, 여러 scientific foundation model을 causal/provenance-aware scientific memory 위에서 협력시키고, 서로 독립적인 epistemic mechanism을 가진 agent가 가설을 생성·반증하며, Value-of-Information에 따라 다음 계산과 실험을 선택하고, positive/negative result에 따라 belief를 수정하는 시스템이다.

RAG는 “무엇을 알고 있는가”를 묻는다. 미래의 AI Co-Scientist는 “현재 증거를 바탕으로 무엇을 믿어야 하며, 다음에 무엇을 실험해야 하는가”를 결정해야 한다.첨부 자료가 2025–2026 연구 흐름에서 도출한 가장 중요한 경계다.

연구 신규성의 교차점

첨부 자료는 특히 MAMMAL / PhenoModel / Boltz-2 / TxGemma 같은 과학특화 FM + Hyper-Relational Knowledge Graph + Causal Inference + Falsification Agent + Lab-in-the-Loop의 통합을 유망한 연구 공백으로 본다. 각 축은 이미 따로 발전하고 있지만, 이를 하나의 검증 가능한 scientific decision system으로 묶는 문제는 아직 열려 있다는 판단이다.

References & Source Notes

01
Gottweis et al., “Towards an AI co-scientist,” 2025.

Generate–Debate–Evolve와 biomedical validation. arXiv

02
Wang et al., “TxGemma: Efficient and Agentic LLMs for Therapeutics,” 2025.

Therapeutic foundation model과 Agentic-Tx. arXiv

03
Roohani et al., “BioDiscoveryAgent,” ICLR 2025.

Genetic perturbation experiment design. ICLR Proceedings

04
Gao et al., “PharmAgents: Building a Virtual Pharma with Large Language Model Agents,” 2025.

Virtual-pharma multi-agent pipeline. arXiv

05
Cui et al., “Towards multimodal foundation models in molecular cell biology,” Nature 2025.

Multimodal foundation-model framework for molecular cell biology. Nature

06
Ghareeb et al., “A multi-agent system for automating scientific discovery,” Nature 2026 — Robin.

Hypothesis–experiment–analysis–revision loop. Nature

07
Huang et al., “Autonomous biomedical research with an artificial intelligence agent,” Science 2026 — Biomni.

Biomedical tools·DB·code·wet-lab 통합. PubMed

08
Shoshan et al., “MAMMAL – Molecular Aligned Multi-Modal Architecture and Language for biomedical discovery,” npj Drug Discovery 2026.

Cross-modal biomedical foundation model. npj Drug Discovery

09
Wang et al., “PhenoModel: A multimodal phenotypic drug design foundation model,” 2025/2026.

Molecular structure–cell morphology alignment과 phenotypic drug discovery. PubMed

10
Zheng et al., “DiscoVerse: multi-agent pharmaceutical co-scientist for traceable drug discovery and reverse translation,” 2026.

Source-linked evidence와 long-term pharmaceutical archive. Frontiers

11
He et al., “Democratising real-world drug discovery through agentic AI,” Drug Discovery Today 2026 — ChatInvent/AstraZeneca.

Real-world molecular design and synthesis planning. ScienceDirect

12
Pastore & De Rango, “Foundation and Multimodal Models for Drug Discovery in Molecular Informatics,” 2026.

Evaluation, distribution shift, uncertainty를 다루는 review. PubMed

13
Huynh et al., “AI agents in drug discovery: applications and case studies,” Drug Discovery Today 2026.

Drug-discovery AI agent review. PubMed

14
Vijayan et al., “Entering the agentic era of AI in drug discovery,” Nature Chemical Biology, 2026-08-19.

Agentic AI + laboratory automation + closed-loop DMTA. Nature Chemical Biology

15
Boltz-2 independent reliability evaluation cited in the source document.

Structure and binding-affinity prediction reliability에 대한 독립 평가. arXiv

Source boundary: 본문은 첨부된 “신약개발용 과학특화 멀티모달 파운데이션 모델 기반 Multi-Agent AI Co-Scientist: 2025–2026 연구 동향” 문서의 논지·사례·수치·참고문헌을 재구성했다. 자료가 미래 연구방향으로 제안한 항목은 established result가 아니라 source-level analysis/inference로 구분해 서술했다.