Static structure
holo structure 하나를 truth로 놓으면 conformational ensemble과 induced fit을 잃는다.
Spatially Grounded Multiscale World Models for Multi-Agent AI Co-Scientists in Drug Discovery
신약개발 AI가 단백질의 3차원 구조를 보는 것만으로 과학자가 되는 것은 아니다. 과학자는 구조를 본 다음, 그 구조가 움직이면 무엇이 달라지고, 약물을 넣으면 세포와 조직이 어떻게 변하며, 자기 예측이 틀렸을 때 어떤 세계관을 버려야 하는지를 묻는다.
2025년 이후 연구는 이 질문에 필요한 부품을 빠르게 갖추고 있다. FLOWR.ROOT·DrugBLIP·DiffSMol·Token-Mol 같은 모델은 molecular spatial intelligence를 넓히고, ProTDyn과 protein generative world model은 static structure를 conformational dynamics로 확장한다. VCWorld·AlphaCell·VCHarness는 intervention-conditioned cellular world model을 향하고, STORM·SEAL은 조직의 공간적 분자상태를 multimodal foundation model 안으로 가져온다. Co-Scientist·Robin·Virtual Lab은 이 모델들을 실제 hypothesis–experiment loop로 조직할 agentic scaffold를 제공한다.
2026년 8월 현재, 과학특화 멀티모달 파운데이션 모델·Spatial Intelligence·World Models·Multi-Agent AI Co-Scientist를 분자→단백질→세포→조직 전 스케일에서 하나의 end-to-end drug-discovery system으로 통합하고 prospective wet-lab까지 검증한 표준 연구는 아직 확립되지 않았다. 따라서 이 글의 통합 아키텍처는 여러 선행연구의 수렴에서 도출한 연구설계 제안이다.
그럼에도 방향은 뚜렷하다. Spatial Intelligence는 세계의 구조를 이해하고, World Model은 그 세계의 변화법칙을 모델링한다. 둘을 결합하면 AI Co-Scientist는 단지 “현재 구조가 어떠한가”를 답하는 데서 벗어나 “이 구조에 개입하면 다음 세계는 어떻게 달라질 것인가”를 가상실험할 수 있다.
Spatial Intelligence는 구조와 관계를, World Model은 action-conditioned transition을 담당한다. 신약개발에서는 두 능력을 3D/4D biology에 맞게 다시 정의해야 한다.
일반 AI에서 Spatial Intelligence는 객체의 위치, 방향, 거리, 관계, 관점 변화와 3차원 구조를 지각하고 추론하는 능력이다. 2026년 CVPR의 SenseNova-SI는 800만 규모 spatial data와 taxonomy를 이용해 multimodal foundation model의 공간능력을 체계적으로 강화했다.
신약개발에서 공간은 방 안의 위치가 아니라 원자 → 잔기 → binding pocket → protein complex → cell → tissue microenvironment로 이어지는 생물학적 공간이다. 따라서 Molecular/Biological Spatial Intelligence는 분자·단백질·세포·조직의 3D/4D 위치·방향·형상·접촉·변형을 표현하고, 이러한 관계가 결합·기능·질병·약물반응에 미치는 영향을 추론하는 능력으로 정의하는 편이 유용하다.
엄격한 의미의 world model은 현재 latent state \(z_t\)와 action \(a_t\)를 조건으로 다음 상태를 예측한다.
신약개발에서 action은 로봇의 이동명령이 아니라 ligand edit, mutation, dose, drug treatment, gene knockout, combination therapy, assay condition이 된다.
따라서 drug-discovery world model의 핵심은 예측 하나가 아니라 가상개입의 결과를 rollout할 수 있는 transition model이다.
| Concept | 핵심 질문 | Drug Discovery 예 |
|---|---|---|
| Spatial Intelligence | 어디에 있으며 어떻게 맞물리는가? | ligand pose, pocket geometry, cell neighbourhood |
| World Model | 무엇을 하면 다음에 어떻게 변하는가? | ligand edit 후 affinity, drug 후 cell state |
| Spatial World Model | 이 공간상태에 개입하면 구조와 기능이 어떻게 바뀌는가? | induced fit, conformational transition, spatial tumor response |
둘을 합치면 AI Co-Scientist에 필요한 3D/4D scientific imagination이 된다.
현재 biomedical AI에는 세 가지 큰 gap이 존재한다. 첫째, sequence representation과 실제 3D binding physics 사이의 Sequence–Structure Gap이다. 둘째, static structure와 conformational ensemble 사이의 Static Structure–Dynamic Biology Gap이다. 셋째, 좋은 binder와 실제 cellular/therapeutic response 사이의 Molecular–Cellular Gap이다.
미래 Co-Scientist의 문제는 이 사슬의 각 링크를 따로 최적화하는 것이 아니라, 개입에 따라 사슬 전체가 어떻게 변하는지를 모델링하고 어느 단계에서 prediction이 깨지는지를 찾아내는 것이다.
AI Co-Scientist, molecular spatial foundation models, protein dynamics, virtual cell, spatial biology가 동시에 가까워지고 있다.
2025년 Nature의 Virtual Lab은 LLM Principal Investigator와 specialist agents가 ESM, AlphaFold-Multimer, Rosetta를 활용해 92개의 nanobody를 설계하고 실험으로 검증했다. 2026년 Co-Scientist는 Generate–Critique–Rank–Evolve 구조를 drug repurposing, target discovery, AMR mechanism에 적용했다. Robin은 hypothesis generation과 experimental data analysis를 같은 multi-agent workflow 안에 연결했다.
이 변화의 핵심은 AI가 더 긴 답을 쓰게 된 것이 아니다. 여러 scientific models와 tools를 조합하고, 결과가 예상과 다르면 다음 hypothesis를 바꾸는 운영구조가 생겼다는 데 있다.
2026년 Nature Communications의 FLOWR.ROOT는 SE(3)-equivariant backbone 안에서 pocket-aware 3D ligand generation과 pIC50, pKi, pKd, pEC50 prediction, confidence estimation, scaffold hopping, fragment growing을 결합한다. DrugBLIP은 SE(3)-equivariant graph transformer로 protein–molecule 3D interaction을 학습한다. 2025년 DiffSMol은 3D ligand shape와 protein pocket guidance를 molecular generation에 직접 사용했고, Token-Mol은 2D/3D molecular information과 property를 discrete token으로 통합했다.
이 흐름은 drug foundation model이 언어모델의 문법을 넘어서 물리적 배치와 상호작용을 다루기 시작했음을 의미한다.
2025년 Generative World Models for Protein Folding Pathways는 generative AI로 protein folding/conformational transition을 모델링하고 equilibrium MD와 비교했다. 2026년 ICLR의 ProTDyn은 conformational ensemble sampling과 multi-timescale protein dynamics generation을 unified framework로 묶는다.
2026년 Biohub의 Language Modeling Materializes a World Model of Protein Biology는 ESMC representation을 기반으로 protein structure prediction, binder design, function/structure organization과 대규모 protein map을 연결한다. 다만 여기서 “world model”은 classical reinforcement learning의 action-conditioned transition model보다 넓은 representational/generative sense에 가깝다. 이 용어 차이를 구별해야 한다.
ICLR 2026의 VCWorld는 structured biological knowledge와 LLM reasoning을 결합해 perturbation-induced signaling cascade와 mechanistic hypothesis를 생성한다. AlphaCell은 full protein-coding transcriptome representation과 continuous state-transition modeling을 사용하는 generative virtual-cell world model을 제안한다. VCHarness는 AI coding agent와 multimodal biological FM으로 perturbation-response model 자체를 자동 설계·개선하는 방향을 보여준다.
조직 수준에서는 STORM이 18개 장기의 120만 spatial transcriptomic profile과 matched histology를 이용해 morphology, expression, spatial context를 통합하고, SEAL은 spatial transcriptomics의 localized molecular information을 pathology foundation model에 주입한다.
geometry, equivariance, conformational ensemble, latent state, intervention, transition, counterfactual rollout, uncertainty가 하나의 scientific world model을 구성한다.
| Core Concept | 의미 | Co-Scientist 역할 |
|---|---|---|
| 3D Geometry | atom/residue coordinates | pocket–ligand understanding |
| SE(3) Equivariance | 회전·이동에도 일관된 표현 | 3D molecular reasoning |
| Spatial Interaction Graph | contact, H-bond, π-stack 등 | binding mechanism |
| Conformational Ensemble | 여러 protein state | induced fit / dynamics |
| Latent World State | 압축된 biological state | simulation state |
| Action / Intervention | drug, mutation, dose 등 | counterfactual experiment |
| Transition Model | intervention 후 상태변화 | world dynamics |
| Counterfactual Rollout | “이것을 하면?” 가상실험 | candidate prioritization |
| Spatial Omics | 위치를 가진 molecular state | tissue context |
| Virtual Cell | perturbation-response simulator | MoA / resistance |
| Uncertainty | simulation confidence | abstention / experiment selection |
| Multi-Agent Planning | 여러 전문 model을 조합 | scientific decision |
scalar property라면 rigid transformation에 대해 다음과 같은 invariance가 바람직하다.
좌표를 출력하는 모델이라면
와 같은 equivariance가 필요하다. 그래서 FLOWR.ROOT, DrugBLIP, MolX 같은 3D drug models에서 SE(3)/E(3)-equivariant architecture가 반복해서 등장한다.
다만 reflection까지 동일하게 취급하는 E(3) invariance는 chirality에서 주의해야 한다. enantiomer는 mirror image이지만 pharmacological activity가 전혀 다를 수 있다. 중요한 것은 “공간적으로 invariant해야 한다”가 아니라 어떤 transformation에 invariant/equivariant해야 하는가를 분자물리학에 맞게 고르는 일이다.
연구자는 “이 compound의 affinity는 얼마인가?”보다 “이 methyl group을 빼면?”, “T790M이 C797S로 더 변하면?”, “A와 B를 같이 주면?”, “dose를 절반으로 줄이면?”, “hypoxic niche라면?”을 반복해서 묻는다.
이 질문은 static predictor보다 intervention-conditioned world model에 가깝다. 그리고 Co-Scientist의 가치는 수천 개 counterfactual을 cheap하게 rollout한 다음, 어느 가설을 실제 실험으로 보낼지를 결정하는 데 있다.
Spatial hallucination, model bias, long-horizon rollout error, cross-scale gap, hidden confounder를 다루지 않으면 world model은 정교한 자기확증 엔진이 될 수 있다.
holo structure 하나를 truth로 놓으면 conformational ensemble과 induced fit을 잃는다.
Ångström atom에서 millimeter tissue까지 여러 자릿수의 scale을 연결해야 한다.
drug action은 molecule·dose·route·duration·cell state·context가 묶인 hyper-relational intervention이다.
transition prediction이 좋아도 intervention mechanism을 인과적으로 증명하지는 않는다.
작은 one-step error가 resistance·toxicity 같은 긴 trajectory에서 누적된다.
sequence·3D·omics·image·text가 같은 biological state를 가리키는지 정렬해야 한다.
같은 biased model로 10만 번 가상실험해도 10만 번 틀릴 수 있다.
존재하지 않는 residue contact, 불가능한 distance, stereochemistry 오류가 자연어로 생성될 수 있다.
존재하지 않는 transition law 안에서 self-consistent simulation을 반복할 수 있다.
batch, donor, cell cycle, culture condition을 biological law로 오인할 수 있다.
new target + new scaffold + new mutation + new cell type 조합이 핵심 난제다.
공간적으로 가깝다는 사실이 causal relation을 의미하지 않는다.
가장 위험한 시스템은 incoherent한 AI가 아니라, 틀린 세계 안에서 매우 일관되게 reasoning하는 AI다. World Model Falsifier와 외부 실험이 필요한 이유가 여기에 있다.
연구주제로 가장 강한 조합은 RQ3 + RQ5 + RQ8 + RQ11 + RQ12다. 이는 3D 정확도 개선이 아니라 multiscale biological world model을 만들고, 가상실험하고, 실제 실험으로 반증·수정하는 문제이기 때문이다.
Molecular geometry, protein dynamics, cellular perturbation, spatial tissue를 각각 specialist world model로 두고 Co-Scientist가 counterfactual rollout과 model debate를 조직하는 구조가 현실적이다.
atomic state를 \(X=(V,E,\mathbf R)\)로 두고, atom/residue feature, chemical/intermolecular edge, 3D coordinate를 함께 인코딩한다.
FLOWR.ROOT, DrugBLIP, MolX가 이 방향의 선행근거다. Co-Scientist가 “para position에 chlorine을 추가하자”고 제안하면 Spatial Agent는 수정된 분자의 pose, clash, interaction, conformational strain을 다시 평가한다.
static structure 한 장이 아니라 conformational ensemble과 trajectory를 생성한다. ProTDyn과 generative protein world-model 연구가 이런 방향을 뒷받침한다. Co-Scientist는 “이 ligand가 closed conformation을 안정화시키는가?” 같은 동적 hypothesis를 검사할 수 있다.
기존 구조기반 생성은 종종 pocket을 rigid하게 고정한다. 더 자연스러운 formulation은 protein pocket과 ligand를 함께 변화시키는 것이다.
Apo2Mol은 apo pocket에서 ligand와 holo pocket conformation을 함께 생성하는 방향을 탐색한다. 이는 spatial intelligence가 spatial world model로 넘어가는 중요한 단계다.
VCWorld는 structured biological knowledge와 LLM reasoning으로 perturbation-induced signaling을 모델링하고, AlphaCell은 continuous state transition을 이용한다. 이를 drug Co-Scientist에 넣으면 Drug → Target binding → Signalling → Gene expression → Phenotype의 중간상태를 가상으로 검사할 수 있다.
각 cell \(i\)에 위치와 상태가 존재한다고 보면 tissue world state는 다음과 같이 생각할 수 있다.
STORM·SEAL 같은 연구가 histology와 spatial molecular state를 연결하는 representation을 발전시키고 있다. 다음 단계는 \(S_t+Drug\rightarrow S_{t+1}\)의 Virtual Tissue World Model이다.
한 거대한 모델이 atom에서 tissue까지 직접 처리하기보다 각 scale의 specialist state를 계층적으로 연결하는 편이 현실적이다.
Ligand Edit → Pocket Geometry Change → Binding/Dynamics → Signalling → Cell State → Tissue Response의 연쇄를 명시적으로 구성한다.
실제로 모든 compound를 합성할 수 없으므로 먼저 world model 안에서 가상실험을 수행한다.
예측효용이 높고 uncertainty가 의미 있게 큰 후보를 wet lab으로 보낸다.
world model 하나만 믿으면 self-confirmation이 생긴다. Structure Model은 high affinity를 예측하지만 Dynamics Model은 unstable complex, Cell World Model은 weak phenotype, Tissue World Model은 poor penetration을 예측할 수 있다.
Hit discovery, lead optimization, mutation-aware design, binder design, target validation, combination therapy, resistance, spatial precision oncology로 확장되지만 평가기준은 결국 prospective decision quality여야 한다.
pocket fit, interaction geometry, conformational strain, affinity, selectivity를 함께 평가한다. FLOWR.ROOT의 구조가 직접적인 선행사례다.
“methyl을 fluorine으로 바꾸자” 같은 edit를 pose → protein relaxation → affinity → cell response로 rollout한다.
kinase resistance mutation처럼 pocket geometry가 달라지는 경우 \(P_{WT}\neq P_{Mutant}\)이므로 static WT structure만으로 설계해서는 안 된다.
Virtual Lab은 ESM + structure tools + multi-agent reasoning이 실험 가능한 binder design으로 이어질 수 있음을 보여준다. Biohub의 ESM world-model 연구도 therapeutic targets에 대한 experimentally validated binder design을 보고했다.
cellular world model 안에서 \(do(Target=KO)\) 또는 \(do(Target=inhibited)\)를 실행해 disease phenotype의 변화를 가상으로 점검한다.
두 약물의 순서가 다를 때 \(T(T(z,A),B)\)와 \(T(T(z,B),A)\)가 달라질 수 있으므로 temporal action ordering까지 reasoning해야 한다.
Drug → Initial Response → Adaptive Signalling → Resistance라는 long-horizon trajectory를 모델링한다.
tumor, immune, stromal cell과 spatial niche를 구분하고 intervention 후 조직재편을 예측한다. STORM 같은 spatial multimodal FM이 이 방향의 foundation을 제공한다.
2026년 Nature Reviews Drug Discovery Perspective가 강조하듯 AI drug discovery는 benchmark improvement보다 실제 의사결정 개선을 입증해야 한다. Spatial World-Model Co-Scientist의 최종 metric도 simulation accuracy 자체가 아니라 prospective decision quality와 experimental success가 되어야 한다.
4D molecular intelligence, causal world models, epistemic memory, multi-world-model debate, active learning, self-correction을 다음 세대 구조로 제안한다.
다음 세대 molecular foundation model은 static coordinate \(X\)가 아니라 trajectory \(X(t)\)를 학습해야 한다. protein–ligand interaction을 한 장의 pose가 아닌 지속되는 상태전이로 이해하는 4D Molecular Foundation Model이 중요한 방향이다.
world model의 action을 단일 ligand가 아니라 mutation, dose, environment, disease context까지 포함하는 intervention tuple로 확장한다.
단순 transition \(p(z_{t+1}\mid z_t,a)\)를 넘어 intervention semantics를 명시해야 한다.
Spatial Representation + World Dynamics + Causal Intervention의 결합이 필요하다.
각 prediction에는 source, model version, assay context, supporting evidence, contradicting evidence, uncertainty가 붙어야 한다. Epistemic HRKG와 결합하면 world state를 다음처럼 표현할 수 있다.
여기서 \(Z_t\)는 neural spatial/world state, \(G_t\)는 epistemic HRKG, \(D_t\)는 dynamics model, \(U_t\)는 uncertainty다.
Protein Geometry World Model, Protein Dynamics World Model, Cellular World Model, Spatial Tissue World Model, Causal Model이 동일 intervention에 대한 미래를 각각 예측한다. Falsifier는 consensus를 만드는 대신 disagreement를 찾는다.
AI가 무엇을 예측할지뿐 아니라 “내가 지금 세계에 대해 가장 잘 모르고 있는 것이 무엇인가?”를 판단한다. uncertainty가 큰 영역을 실험으로 채우며 world model의 coverage를 넓힌다.
이것은 단순 fine-tuning과 다르다. 실험 실패가 들어오면 scientific belief와 transition model 자체를 수정해야 한다.
| Study | Year | 핵심 기여 | 본 주제에서의 의미 | Source |
|---|---|---|---|---|
| Virtual Lab | 2025 | multi-agent + protein tools + wet-lab nanobody design | spatial tool-using Co-Scientist | Nature |
| Co-Scientist | 2026 | generate–critique–rank–evolve + biomedical validation | agent backbone | Nature |
| Robin | 2026 | hypothesis–experiment–analysis–revision | closed-loop discovery | Nature |
| FLOWR.ROOT | 2026 | SE(3) pocket-conditioned 3D generation + affinity | molecular spatial FM | Nature Communications |
| DrugBLIP | 2026 | SE(3)-equivariant protein–molecule interaction | spatial interaction reasoning | Bioinformatics |
| MolX | 2026 preprint | geometric pocket–ligand foundation model | joint spatial representation | bioRxiv |
| DiffSMol | 2025 | shape/pocket-guided 3D molecular diffusion | geometry-conditioned design | Nature Machine Intelligence |
| Token-Mol | 2025 | tokenized 2D/3D drug design | discrete multimodal spatial representation | Nature Communications |
| ProTDyn | ICLR 2026 | protein ensemble + dynamics foundation model | dynamic protein world model | OpenReview |
| Generative Protein World Models | 2025 preprint | protein folding/pathway modeling | conformational world model | bioRxiv |
| ESM World Model of Protein Biology | 2026 preprint | protein sequence–structure–function/design space | broad protein world model | bioRxiv |
| VCWorld | ICLR 2026 | knowledge + LLM cellular perturbation simulation | virtual-cell world model | ICLR |
| AlphaCell | 2026 preprint | continuous cellular perturbation dynamics | cell-state transition model | bioRxiv |
| VCHarness | 2026 preprint | AI agent automatically constructs virtual-cell models | agent builds world model | bioRxiv |
| STORM | 2026 preprint | histology + spatial transcriptomics FM | tissue spatial intelligence | arXiv |
| SEAL | 2026 preprint | ST-guided pathology foundation model | morphomolecular spatial grounding | arXiv |
주의: 이 문헌들은 서로 다른 scale·task·dataset을 다루므로 성능 수치를 직접 비교해서는 안 된다. preprint는 peer-reviewed paper와 구분해 표기했다. ProTDyn의 서지정보는 앞선 검토에서 ICLR 2026 연구로 확인된 내용을 반영했으며, 본문에서는 방법론적 방향만 사용한다.
Can a spatially grounded, multiscale biological world model enable a multimodal multi-agent AI Co-Scientist to simulate molecular and cellular interventions, falsify competing drug mechanisms, and select experiments that improve real-world therapeutic discovery?
한국어로 풀면 다음과 같다. 분자·단백질·세포·조직의 공간상태를 연결한 다중스케일 biological world model을 구축하고, AI Co-Scientist가 그 안에서 ligand modification·mutation·drug treatment를 가상실험하며, 서로 경쟁하는 mechanism을 반증하고, 실제 wet-lab 실험으로 세계모델 자체를 지속적으로 수정했을 때 기존 static structure model 또는 LLM/RAG 기반 Co-Scientist보다 신약개발의 실험 성공률과 의사결정 품질을 높일 수 있는가.
역할을 간단히 나누면 더욱 분명하다. Multimodal Foundation Model은 세계를 읽고, Spatial Intelligence는 그 세계의 구조를 이해하며, World Model은 세계가 어떻게 변할지를 상상하고, Multi-Agent Co-Scientist는 여러 미래를 논쟁하며, 실험은 어떤 세계모델이 현실에 가까운지를 결정한다.
이 관점에서 World Model은 단순 simulator가 아니다. AI Co-Scientist가 가지고 있는 실행 가능한 과학적 세계관이다. Spatial Intelligence는 그 세계관을 실제 분자·세포·조직의 물리적 구조에 붙잡아 두는 grounding mechanism이다.