2025년 이후 신약개발 AI의 질문이 달라졌다. “이 모델이 affinity를 얼마나 잘 맞히는가”에서 “여러 모델과 증거와 실험을 이용해 연구 과정 자체를 얼마나 잘 운영하는가”로 이동한다. 이 변화의 중심에는 하나의 거대한 모델이 아니라, 과학 모델을 도구처럼 고르고 가설을 경쟁시키며 실험 결과로 믿음을 수정하는 멀티에이전트 AI Co-Scientist가 있다.
첨부 자료의 핵심 결론은 분명하다. 미래의 Co-Scientist는 multimodal foundation model, multi-agent reasoning, 전문 계산 도구, wet-lab automation, provenance memory를 따로 쓰는 시스템이 아니다. 이들을 Prediction → Decision → Experiment → Revision의 하나의 폐루프로 연결하는 시스템이다.
Foundation Model은 과학적 세계를 표현한다. Agent는 다음 행동을 결정한다. Co-Scientist는 가설·증거·도구·실험·반증·수정을 하나의 과학적 의사결정 과정으로 운영한다. 이 세 층위를 구분해야 연구의 신규성과 한계를 정확히 볼 수 있다.
이 글은 첨부된 2025–2026 연구 동향 문서의 전체 내용을 웹 읽기 흐름으로 재구성한 것이다. 논문별 수치와 사례는 첨부 자료의 서술에 근거하며, 외부 문헌을 별도로 재검증해 수치를 추가하거나 교정하지 않았다.
2025년의 전환:
Model-centric AI에서 Science-centric AI로
The unit of progress shifts from a predictor to a scientific organization.
Scientific Multimodal Foundation Model은 단순한 VLM이 아니다
신약개발의 과학특화 멀티모달 파운데이션 모델은 molecule, protein, cell, tissue, disease를 서로 다른 표상으로 다루는 범용 사전학습 계층이다. 입력과 증거의 공간은 SMILES, molecular graph, 3D structure, protein sequence, DNA/RNA, omics, cell image, assay, text를 가로지른다.
첨부 자료는 multimodality를 세 층위로 나눈다. 첫째는 molecule–protein–3D–omics–image를 함께 표현하는 Representation Multimodality다. 둘째는 paper–assay–clinical–toxicity–omics처럼 서로 다른 증거를 통합하는 Evidence Multimodality다. 셋째는 search–code–docking–simulation–robotic experiment처럼 서로 다른 행동을 수행하는 Action Multimodality다.
세계를 표현한다
MAMMAL은 protein, small molecule, gene, transcriptomic information을 cross-modal하게 다루고, PhenoModel은 molecular structure와 cell morphology를 연결한다.
증거를 묶는다
문헌, assay, 임상, 독성, omics가 같은 결론을 지지하거나 반박하는 근거망으로 연결된다.
과학 행동을 선택한다
검색, 코드 실행, docking, simulation, 실험 장비 실행이 동일한 scientific workflow 안에 들어온다.
AI Co-Scientist는 연구자용 챗봇과 다르다
첨부 자료의 정의를 압축하면 AI Co-Scientist는 주어진 과학적 목표에 대해 문헌과 데이터를 탐색하고, 새로운 가설을 만들고, 경쟁 가설을 비판·순위화하고, 필요한 계산이나 실험을 설계·실행하며, 결과에 따라 기존 가설을 수정하거나 폐기할 수 있는 자율적 또는 반자율적 과학추론 시스템이다.
Google AI Co-Scientist는 Generate → Debate → Evolve 구조와 Generation, Reflection, Ranking, Evolution 같은 역할 분리를 보여 주었다. Robin은 literature → hypothesis → experimental proposal → wet-lab result → autonomous data analysis → revised hypothesis라는 연속적인 feedback loop를 실험생물학에 연결했다.
왜 2025년이 전환점인가
첨부 자료는 2023–2024년의 중심 질문을 “AI가 분자나 단백질을 얼마나 잘 예측하거나 생성하는가?”로 요약하고, 2025년 이후 질문이 “AI가 여러 모델과 데이터를 이용해 과학자의 연구과정 자체를 수행할 수 있는가?”로 바뀌었다고 본다.
| System / Model | Year | Source document’s key advance |
|---|---|---|
| AI Co-Scientist | 2025 | Generate–Debate–Evolve 기반 scientific hypothesis society. |
| TxGemma / Agentic-Tx | 2025 | therapeutic foundation model을 더 큰 agentic workflow의 도구로 사용. |
| BioDiscoveryAgent | ICLR 2025 | 실험 결과를 보고 다음 genetic perturbation을 선택. |
| PharmAgents | 2025 | target → lead → preclinical의 virtual-pharma multi-agent pipeline. |
| Robin | Nature 2026 | hypothesis–experiment–analysis–revision 폐루프의 초기 실증. |
| Biomni | Science 2026 | biomedical tools·DB·code·wet-lab를 통합하는 범용 agent. |
| DiscoVerse | 2026 | 장기 제약 R&D archive에 근거한 traceable co-scientist. |
| ChatInvent | 2026 | 실제 discovery pipeline의 multi-agent molecular design과 synthesis planning. |
| MAMMAL | 2026 | molecule–protein–gene 등 cross-modal biomedical foundation model. |
| PhenoModel | 2025/26 | molecular structure–cell morphology multimodal foundation model. |
이 흐름을 첨부 자료는 Model-centric AI → Agent-centric AI → Science-centric AI로 정리한다.
예측 문제를 넘어
Sequential Scientific Decision Problem으로
Drug discovery is not a single prediction; it is a sequence of decisions under uncertainty.
기존 AI 신약개발은 흔히 \((\text{protein},\text{ligand})\rightarrow\text{binding affinity}\) 또는 \(\text{molecule}\rightarrow\text{toxicity}\) 같은 정적 prediction을 풀었다. 실제 연구자의 질문은 훨씬 길다. disease mechanism을 이해하고, 아직 충분히 탐색되지 않은 target을 고르고, binding affinity·selectivity·ADMET·synthesizability를 함께 고려해 candidate를 고르며, 현재 uncertainty를 가장 잘 줄일 next assay까지 선택해야 한다.
즉 최적화 대상은 candidate 하나가 아니다. hypothesis \(h\), intervention \(x\), next experiment \(e\)의 묶음이다. scientific utility와 therapeutic utility를 키우되 experiment cost와 safety risk를 함께 고려해야 한다. 그리고 모든 결론에는 provenance와 uncertainty가 따라야 하며, 새로운 결과가 기존 믿음과 충돌하면 belief revision이 일어나야 한다.
신약개발은 원래부터 multimodal problem이다
disease mechanism에는 genomics와 transcriptomics가 필요하고, target에는 protein sequence와 3D structure가 필요하며, compound에는 graph·SMILES·conformation이 필요하다. biological effect는 Cell Painting, assay, pathology로 관찰되고, translational decision에는 toxicity, PK/PD, clinical evidence와 문헌이 들어간다. 최근 foundation model은 원래 존재하던 이 heterogeneity를 계산적으로 한 구조 안에 놓기 시작한 것이다.
신약개발은 원래부터 multi-agent problem이다
실제 제약 연구조직은 medicinal chemist, structural biologist, disease biologist, bioinformatician, toxicologist, DMPK scientist, statistician, clinician의 협업으로 움직인다. 첨부 자료는 그래서 monolithic agent보다 전문 역할이 분화된 scientific organization architecture가 자연스럽다고 본다. PharmAgents의 “virtual pharma”, ChatInvent의 multi-agent 확장이 이 관점을 구체화한다.
가장 비싼 자원은 계산이 아니라 잘못된 실험이다
prediction accuracy를 조금 더 높이는 것보다 “다음에 어떤 실험을 해야 가장 많은 정보를 얻는가?”가 더 중요한 경우가 있다. 첨부 자료는 향후 metric이 AUROC 하나보다 다음과 같은 방향에 가까워질 수 있다고 본다.
BioDiscoveryAgent가 genetic perturbation experiment selection을 직접 문제로 삼은 이유도 여기에 있다. 첨부 자료는 Claude 3.5 Sonnet 기반 시스템이 6개 dataset에서 기존 Bayesian-optimization baseline 대비 평균 21% 향상을 보고했다고 정리한다.
과학적 세계를 표현하고,
가설을 경쟁시키고, 증거로 묶는다
A Co-Scientist needs more than generation: it needs representation, instruments, evidence, falsification and laboratory feedback.
첨부 자료는 이 분야의 core concept를 여섯 가지 기능으로 정리한다. 이 여섯 기능이 따로 존재하는 것이 아니라 하나의 scientific loop 안에서 연결될 때 Co-Scientist가 된다.
Multimodal Scientific Representation
MAMMAL과 PhenoModel 같은 모델이 molecule, protein, gene, phenotype을 richer world representation으로 연결한다.
Hypothesis Society
Generator, Evidence Retriever, Skeptic, Experiment Designer, Statistician, Safety Critic, Judge가 서로 다른 역할을 맡는다.
Foundation Models as Scientific Instruments
Boltz-2, TxGemma, MAMMAL, docking, MD, FEP, ADMET, search를 문제에 따라 선택한다.
Evidence-Grounded Reasoning
scientific claim을 문헌·assay·clinical archive의 source-linked evidence와 연결한다.
Falsification, Not Just Generation
가설을 많이 만드는 능력보다 틀린 가설을 제거하는 Skeptic/Counter-Evidence 역할을 중시한다.
Lab-in-the-Loop
Hypothesis → Experiment → Observation → Revision을 반복해 실험결과가 다음 가설의 input이 되게 한다.
Evidence-grounding의 실제 사례: DiscoVerse
첨부 자료에 따르면 Roche의 DiscoVerse는 Preclinical, Clinical 등 역할이 다른 agent와 Supervisor를 사용하면서 제약회사 내부 자료를 source-linked evidence로 연결한다. 180개 molecule과 40년 이상의 R&D archive를 대상으로 평가했고, 7개 benchmark query에서 recall ≥ 0.99, precision 0.71–0.91을 보고했다.
이 수치의 의미는 단순 QA 성능보다 traceability에 있다. 과학적 conclusion이 어느 source에 기대고 있는지 보이지 않으면 Co-Scientist의 판단을 검증하거나 반박하기 어렵다.
Lab-in-the-loop의 실제 사례: Robin
첨부 자료는 Robin이 dry age-related macular degeneration 연구에서 RPE phagocytosis를 치료 전략으로 제안하고 ripasudil과 KL001을 실험적으로 확인한 뒤, RNA-seq 결과를 분석해 ABCA1이라는 후속 mechanistic hypothesis까지 생성했다고 정리한다. 결과가 report의 마지막 줄이 아니라 다음 hypothesis의 시작점이 된 것이다.
하나의 거대 모델보다
계층적 과학 시스템이 유망하다
Supervisor, specialist agents, scientific foundation models, tools and wet-lab form a hierarchy of scientific control.
첨부 문헌을 종합한 가장 유망한 구조는 monolithic super-model이 아니라 계층적 orchestration이다. Research Objective를 Supervisor가 분해하고, Hypothesis Agent·Evidence Agent·Skeptic Agent가 경쟁·협력한다. Experiment Designer와 Model/Tool Selection Agent가 다음 행동을 설계하고, MAMMAL·Boltz-2·TxGemma·PhenoModel 같은 전문 모델과 docking·MD·FEP·ADMET·omics·search를 호출한다. wet-lab 결과는 counter-evidence와 uncertainty로 다시 올라와 hypothesis revision을 일으킨다.
Method 1 — Generate–Debate–Evolve
Google AI Co-Scientist가 대표적이다. 여러 candidate hypothesis를 생성하고 pairwise comparison과 critique로 competition시킨 뒤 더 좋은 hypothesis를 진화시킨다. 핵심은 single-shot generation보다 selection pressure를 넣는 데 있다.
Method 2 — Hierarchical Supervisor Architecture
Supervisor가 연구문제를 decomposition하고 전문 agent에 할당한다. DiscoVerse와 enterprise multi-agent workflow에 특히 적합한 방식으로 첨부 자료는 설명한다.
Method 3 — Tool-Augmented ReAct
Reason → Tool Call → Observe → Revise를 반복한다. TxGemma/Agentic-Tx와 최근 drug-discovery agent 연구가 이 구조와 밀접하다고 자료는 설명한다.
Method 4 — Cross-modal Foundation Modeling
MAMMAL은 heterogeneous biomedical entity를 통합하고, PhenoModel은 molecule–phenotype image alignment를 수행한다. 첨부 자료는 이들을 agent가 호출할 수 있는 learned scientific simulator 혹은 representation instrument로 본다.
Method 5 — Retrieval + Provenance
단순 RAG에서 한 걸음 더 나아가 모든 scientific claim에 source를 연결한다. DiscoVerse는 이를 pharmaceutical archive 수준에서 구현한 사례로 제시된다.
Method 6 — Agentic Experimental Design
BioDiscoveryAgent는 이전 perturbation 결과를 다음 prompt에 넣고 next experiment를 선택한다. 첨부 자료는 이를 LLM agent를 acquisition function처럼 사용하는 중요한 방향으로 해석한다.
Method 7 — Lab-in-the-loop Hypothesis Revision
Robin의 경우 실험결과가 시스템의 final output이 아니라 다음 가설의 input이 된다. 이 차이가 “assistant”와 “closed-loop scientist”를 가르는 핵심 중 하나다.
무엇을 맞힐 것인가보다
무엇을 실험할 것인가
The frontier questions move from representation accuracy to causal choice, cross-model verification and prospective science.
10개의 핵심 Research Question
| RQ | Question |
|---|---|
| RQ1 · Multimodal Representation | 분자–단백질–세포–질병–환자 수준을 하나의 representation 또는 interoperable latent space로 연결할 수 있는가? |
| RQ2 · Hypothesis Generation | 기존 논문의 재조합이 아니라 novel, plausible, testable한 hypothesis를 만들 수 있는가? |
| RQ3 · Causal Discovery | correlation과 causal intervention을 구분해 druggable target을 선택할 수 있는가? |
| RQ4 · Agent Collaboration | debate, tournament, supervisor, swarm 중 어떤 협력 구조가 scientific validity를 가장 높이는가? |
| RQ5 · Experiment Selection | 한정된 비용에서 어떤 실험이 hypothesis space를 가장 빨리 줄이는가? |
| RQ6 · Cross-model Verification | Boltz-2, MAMMAL, TxGemma, docking, MD/FEP의 결과가 충돌할 때 어떻게 판단할 것인가? |
| RQ7 · Closed-loop Learning | negative result까지 long-term memory에 저장하고 이후 hypothesis generation을 바꿀 수 있는가? |
| RQ8 · Prospective Evaluation | QA benchmark가 아니라 실제 target/hit과 검증 실험 성공률로 평가할 수 있는가? |
| RQ9 · Accountability | 각 claim에 evidence, provenance, uncertainty, counter-evidence를 자동 부착할 수 있는가? |
| RQ10 · Human–AI Division | 어디까지 autonomous하게 하고 어느 지점에서 scientist approval을 요구할 것인가? |
첨부 자료는 특히 RQ5와 RQ7, 즉 experiment selection과 negative-result-aware closed-loop learning을 assistant와 co-scientist를 가르는 중요한 경계로 본다.
주요 응용 영역
Target Identification & Validation
omics, literature, KG, perturbation evidence를 합쳐 disease-driving target을 찾고 검증한다. AI Co-Scientist의 liver fibrosis 사례가 예로 제시된다.
Drug Repurposing
drug–disease–target–phenotype evidence를 재조합해 새로운 indication을 찾는다. AML과 dAMD 사례가 언급된다.
Hit Discovery / Selection
Target → Virtual Library → Affinity → Phenotype → ADMET → Hit을 여러 모델로 평가한다. PhenoModel은 phenotype-aware hit discovery 가능성을 보여 준다.
Lead Optimization
Affinity, Selectivity, ADMET, PK, Toxicity, Synthesizability를 동시에 최적화한다. PharmAgents는 초기 virtual-pharma pipeline을 제시한다.
Phenotypic Drug Discovery
PhenoModel/PhenoScreen은 osteosarcoma와 rhabdomyosarcoma cell line에서 phenotypically active compound 탐색 사례를 제시한다.
Genetic Perturbation
BioDiscoveryAgent가 어느 gene 또는 gene combination을 perturb할지 반복적으로 선택해 target biology와 validation을 연결한다.
Pharmaceutical Reverse Translation
clinical observation → mechanism → target → next design으로 되돌아가며, 실패한 R&D program을 negative scientific memory로 활용한다.
Real-world Molecular Design
ChatInvent는 actual discovery pipeline에서 molecular design과 synthesis planning을 지원하고 multi-agent architecture로 확장된다.
Multi-Agent라고 해서
증거가 여러 개가 되는 것은 아니다
Reliability depends on epistemic diversity, uncertainty propagation and prospective validation—not agent count.
현재 가장 큰 난제는 model size가 아니라 신뢰할 수 있는 과학적 폐루프다. 첨부 자료는 cross-modal alignment, missing modality, assay/domain shift, causal reasoning, hallucination, correlated agent failure, uncertainty propagation, long-horizon coherence, prospective validation, provenance/reproducibility, safety/dual use, human–AI governance를 주요 challenge로 든다.
| Challenge | Core issue |
|---|---|
| Cross-modal alignment | molecule, protein, image, omics는 semantic scale 자체가 다르다. |
| Incomplete multimodality | 모든 compound에 3D·omics·phenotype data가 존재하지 않는다. |
| Assay/domain shift | laboratory, protocol, cell line에 따라 data distribution이 바뀐다. |
| Causal reasoning | correlation에서 intervention target을 구별하기 어렵다. |
| Hallucination | 잘못된 reference나 mechanism이 scientific hypothesis처럼 보일 수 있다. |
| Correlated agent failure | 동일 foundation model을 쓰는 agent들이 같은 오류를 반복할 수 있다. |
| Uncertainty propagation | docking → ADMET → agent decision을 거치며 uncertainty가 누적된다. |
| Long-horizon coherence | 수십~수백 단계 연구에서 초기 assumption을 잃기 쉽다. |
| Prospective validation | retrospective benchmark가 실제 discovery success를 보장하지 않는다. |
| Provenance / Reproducibility | data·model·version·prompt가 conclusion을 만든 경로를 기록해야 한다. |
| Safety / Dual use | biological design 능력의 증가는 안전관리와 직접 연결된다. |
| Human–AI governance | AI 제안과 human approval의 경계를 정의해야 한다. |
같은 model family를 쓰는 다섯 agent는 같은 latent bias를 공유할 수 있다. 그래서 첨부 자료는 Agent Diversity보다 Epistemic Diversity를 요구한다. 서로 다른 model family, evidence source, physics-based simulator, symbolic rule, experimental data가 서로를 검증해야 한다.
Boltz-2의 독립 평가 사례는 이 교훈을 선명하게 만든다. 첨부 자료에 따르면 2026년 대규모 독립 평가에서 일부 affinity 영역은 weak-to-moderate correlation을 보였고, top-ranked ligand 수준에서 physics-based free-energy 계산과의 일치가 충분하지 않았다고 보고됐다. 따라서 어떤 scientific foundation model의 output도 “사실”이 아니라 evidence 중 하나로 다뤄야 한다.
여섯 개의 열린 문제
- 01FM–Agent coupling. MAMMAL·PhenoModel 같은 multimodal FM과 AI Co-Scientist·Robin·DiscoVerse 같은 agent architecture는 아직 대부분 “Agent → Scientific Model as Tool” 수준으로 느슨하게 연결돼 있다.
- 02Novelty vs truth. 새로운 가설은 쉽게 만들 수 있지만 \(\text{Novel}\not\Rightarrow\text{Correct}\)다. falsifiability, causal plausibility, experimental validation이 함께 평가돼야 한다.
- 03Debate is not independent verification. Generator와 Critic이 같은 foundation model이면 동일한 bias를 공유한다. LLM critic + causal model + KG reasoner + physics simulator + experimental critic처럼 epistemic mechanism을 달리해야 한다.
- 04Chain-level uncertainty. affinity와 ADMET가 모두 불확실한데 마지막 agent가 candidate를 단정하면 안 된다. claim-level uncertainty graph가 필요하다.
- 05Negative result memory. 실패한 실험을 기억하지 않으면 system은 같은 실패를 반복한다. 장기 R&D archive는 미래 Co-Scientist의 중요한 negative memory가 된다.
- 06Retrospective-to-prospective gap. random split보다 scaffold/temporal split, distribution shift, activity cliff, calibration을 보아야 하며, 최종 평가는 “AI 없이 할 때보다 얼마나 적은 실험으로 검증된 새 지식을 얻었는가”에 가까워야 한다.
다음 단계는 더 큰 LLM이 아니라
더 과학적인 의사결정 구조다
Federated scientific models, causal world models, falsification-first agents, value-of-information, epistemic memory and human-in-command.
1. One Giant Model → Federation of Scientific Foundation Models
첨부 자료는 모든 것을 하나의 model이 담당하는 구조보다 Orchestrator와 여러 Specialized FM이 협력하는 federation을 더 유망하게 본다. 한 연구문제에서 MAMMAL은 disease/molecule representation을, AlphaGenome-type model은 regulatory biology를, Boltz-type model은 protein–ligand structure를, PhenoModel은 cellular phenotype을, TxGemma는 therapeutic property reasoning을 맡을 수 있다.
따라서 핵심 능력은 “무엇을 예측하는가”뿐 아니라 어떤 모델을 어느 상황에서 얼마나 믿을 것인가를 판단하는 Meta-Scientific Reasoning이 된다.
2. Multimodal Model → Multiscale Causal World Model
현재 multimodal FM은 representation alignment에 강하다. 다음 단계는 Molecule → Protein → Pathway → Cell → Tissue → Patient의 multiscale causal relation을 다루는 biomedical world model이다. 그렇게 되면 “이 ligand가 target에 붙는가?”에서 “이 perturbation이 downstream pathway와 어떤 환자군의 efficacy/toxicity를 어떻게 바꾸는가?”로 질문을 확장할 수 있다.
3. Generative Agent → Falsification-First Co-Scientist
향후 가장 중요한 agent는 Generator보다 Falsifier일 수 있다. hypothesis \(H\)가 만들어지면 supporting evidence, contradicting evidence, hidden confounder, alternative explanation, decisive experiment를 별도 mechanism이 찾는 구조다.
4. Accuracy Optimization → Value-of-Information Optimization
가장 정확한 candidate를 고르는 것보다 “어떤 experiment가 uncertainty를 가장 많이 줄이는가”를 계산해야 한다. 첨부 자료는 Bayesian experimental design과 AI agent를 결합한 다음 목적을 제안한다.
5. Scientific Memory → Evidence / Provenance Graph
단순 vector memory는 “무엇을 기억하는가”를 저장한다. Co-Scientist는 “왜 그것을 믿는가”까지 저장해야 한다. 첨부 자료가 제안하는 Scientific Epistemic Memory는 hypothesis가 어떤 paper와 assay의 지지를 받고, 어떤 assay에 반박되며, 어떤 assumption에 의존하고, 어떤 model version이 예측했고, 어떤 experiment가 검증했으며, 어떤 uncertainty를 가지고 어떤 hypothesis로 수정됐는지를 graph로 남기는 방식이다.
| Epistemic edge | Example meaning |
|---|---|
| supports | Paper P31, Assay A12가 Hypothesis H17을 지지한다. |
| contradicts | Assay A19가 H17을 반박한다. |
| depends_on | H17이 Assumption C3에 의존한다. |
| predicted_by | Boltz-2 특정 version이 prediction을 만들었다. |
| tested_by | Experiment E8이 hypothesis를 직접 시험했다. |
| uncertainty | claim-level uncertainty를 저장한다. |
| revised_to | H17이 새로운 evidence 때문에 H24로 수정된다. |
6. Human-in-the-Loop → Human-in-Command
완전한 human removal보다 AI가 탐색하고 인간이 research objective, safety boundary, decisive experiment를 통제하는 구조가 현실적이라고 첨부 자료는 본다. governance는 automation의 반대말이 아니라 scientific accountability의 일부다.
7. Autonomous AI → Self-Driving Drug-Discovery Laboratory
궁극적인 loop는 Disease → Hypothesis → Target → Design → Make → Test → Analyze → Revise다. 2026년 Nature Chemical Biology 논평도 agentic AI와 laboratory automation의 결합을 self-driving laboratory와 adaptive DMTA cycle의 방향으로 논의한다고 첨부 자료는 설명한다.
현재의 성숙도: L0에서 L5까지
| Level | Form | Source document’s status |
|---|---|---|
| L0 | Predictor | 이미 성숙 |
| L1 | Scientific Assistant | 널리 가능 |
| L2 | Tool-using Agent | 빠르게 성숙 |
| L3 | Multi-Agent Co-Scientist | AI Co-Scientist, DiscoVerse, ChatInvent |
| L4 | Closed-loop Experimental Scientist | Robin, Biomni 등이 초기 실증 |
| L5 | Autonomous Drug-Discovery Scientist | 아직 미해결 |
첨부 자료가 가장 높은 연구가치를 두는 곳은 L3에서 L4/L5로 넘어가는 경계다. 단순히 LLM agent 하나를 더 만드는 것이 아니라, 여러 scientific foundation model을 causal/provenance-aware scientific memory 위에서 협력시키고, 서로 독립적인 epistemic mechanism을 가진 agent가 가설을 생성·반증하며, Value-of-Information에 따라 다음 계산과 실험을 선택하고, positive/negative result에 따라 belief를 수정하는 시스템이다.
연구 신규성의 교차점
첨부 자료는 특히 MAMMAL / PhenoModel / Boltz-2 / TxGemma 같은 과학특화 FM + Hyper-Relational Knowledge Graph + Causal Inference + Falsification Agent + Lab-in-the-Loop의 통합을 유망한 연구 공백으로 본다. 각 축은 이미 따로 발전하고 있지만, 이를 하나의 검증 가능한 scientific decision system으로 묶는 문제는 아직 열려 있다는 판단이다.