Hit discovery
wnovelty , waffinity ↑
Beyond IsoDDE — PRISM-DDE: OOD 일반화·동역학·물리 검증·생성설계를 통합하는 차세대 Drug Design Engine 연구 프레임워크
구조를 잘 맞히는 모델과 약을 설계하는 엔진은 같은 것이 아니다. 2026년의 연구 지형은 바로 이 간극 위에 서 있다.
2025–2026년 biomolecular foundation model 연구는 AlphaFold 3 이후의 구조예측 경쟁을 넘어섰다. protein–ligand co-folding, binding affinity, conformational sampling, controllable generation, test-time scaling, prospective validation을 하나의 drug-design problem으로 재정의하는 단계에 진입한 것이다. Isomorphic Labs의 IsoDDE는 이 변화를 가장 명시적으로 “Drug Design Engine”이라는 개념으로 제시했다. protein–ligand OOD generalisation, antibody–antigen interaction, pocket identification, binding affinity prediction을 하나의 계산적 시스템으로 묶었다. 그러나 상세 architecture와 전체 training/data recipe가 공개되지 않아 재현성과 독립 검증에는 제약이 따른다.
같은 시기 다른 연구들은 각기 다른 병목을 공격했다. Pearl은 SO(3)-equivariant diffusion과 대규모 synthetic data를, SeedFold는 Pairformer width scaling과 linear triangular attention을, Boltz-2는 structure–affinity joint modelling을, OpenDDE는 structural-token reasoning과 prediction–design unification을 제안했다. 그런데 Runs N’ Poses, FoldBench, prospective Mac1 evaluation, affinity data-leakage 연구는 한목소리로 경고한다. 높은 benchmark accuracy가 곧 새로운 chemical space에서의 causal molecular reasoning이나 prospective drug-discovery utility를 의미하지는 않는다는 것이다.
본 분석은 이를 바탕으로 차세대 연구 아키텍처 PRISM-DDE(Physics-grounded, Robust, Interventional, Sampling-aware, Multi-objective Drug Design Engine)를 제안한다. 핵심 명제는 단순하다.
Drug Design Engine은 오히려 다음의 결합으로 정의해야 한다.
한 시대의 기준점이 만들어지면 경쟁은 그 기준점 안에서 벌어진다. 문제는 기준점 자체가 이동했다는 사실을 언제 알아차리는가이다.
AlphaFold 3는 Pairformer와 atom-coordinate diffusion을 이용해 protein, nucleic acid, small molecule, ion, modified residue를 하나의 all-atom prediction framework에서 처리했다. biomolecular interaction prediction의 기준점은 그렇게 만들어졌다. Pairformer는 AlphaFold 2의 Evoformer보다 MSA processing을 줄였고, structure module은 raw atomic coordinates를 diffusion으로 직접 생성한다.
2025–2026년 연구의 변화는 한 줄의 흐름으로 요약된다.
IsoDDE는 이 흐름에서 중요한 전환점이다. 공식적으로 공개된 predictive core는 structure prediction, pocket identification, binding affinity를 포함하며, Isomorphic Labs는 이를 generative molecular design과 연결된 더 큰 시스템의 일부로 설명한다. 그러나 공개 기술보고서가 제시하는 것은 전체 generative system이 아니라 predictive capability의 일부다. 개념은 앞서갔지만 검증 가능한 실체는 아직 뒤에 있다.
순위표를 읽는 방식으로는 아무것도 알 수 없다. 각 모델이 무엇을 병목으로 지목했는지를 읽어야 지형이 보인다.
Table 1 — Architecture / Training / Data / Benchmark / Strength / Weakness
| Model | Architecture | Training strategy | 주요 데이터 · cutoff | 대표 Benchmark | 핵심 Strength | 핵심 Weakness |
|---|---|---|---|---|---|---|
| AlphaFold 32024 · baseline | MSA module + 48-block Pairformer + atom-coordinate diffusion + confidence module. 거의 모든 PDB molecular type을 unified tokenisation으로 처리 | multistage training, diffusion training, distillation / cross-distillation; crop-size fine-tuning | structural training cutoff 2021-09-30 | protein–ligand, protein–protein, antibody–antigen, protein–nucleic acid 및 PoseBusters 계열 | 범용 all-atom biomolecular modelling의 기준점. Pairformer + diffusion 패러다임 확립 | affinity를 직접적인 핵심 output으로 다루지 않으며 static structure 중심. 이후 연구에서 training-similarity 의존성, induced-fit / novel ligand 문제가 확인됨 |
| IsoDDE2026 | 상세 architecture는 비공개. 공개된 system-level capability는 structure prediction + pocket identification + affinity prediction이며 generative design capability와 통합된다고 설명 | 전체 training recipe와 post-training strategy 비공개. inference에서 multi-sampling / confidence ranking 활용 | 공개 보고서는 structure evaluation과 관련해 AF3와 동일한 2021-era cutoff 조건 등을 기술하지만 전체 data mixture는 비공개 | Runs N’ Poses, FoldBench, low-homology Ab–Ag, FEP+4, OpenFE, CASP16, pocket benchmark | 어려운 OOD protein–ligand 및 Ab–Ag에서 강한 보고 성능. affinity와 pocket을 DDE 개념 안으로 통합 | architecture / data mixture / weights / training procedure 비공개 → 독립적 ablation·재현·성능 귀속이 어려움. prospective drug-campaign 수준의 외부 검증이 공개 evidence보다 훨씬 더 필요함 |
| Boltz-22025 | Boltz-style Pairformer / cofolding trunk + denoising module + confidence + 별도 affinity module. predicted coordinates와 pair representation을 affinity reasoning에 활용 | structure / confidence / affinity objectives를 결합. experimental-method conditioning, distance constraint, multi-chain template 등 controllability 추가 | PDB 구조, distillation / MD 계열 데이터. affinity에는 PubChem, ChEMBL, BindingDB 등 대규모 assay data 활용 | FEP benchmark, OpenFE, CASP16, MF-PCBA 및 structural benchmarks | 공개적으로 사용 가능한 structure–affinity joint model. FEP보다 매우 낮은 계산비용을 목표로 하며 downstream fine-tuning 가능 | 이후 large-scale 평가에서 affinity ranking의 energetic resolution 한계가 관찰됨. affinity benchmark 자체의 leakage 문제도 제기됨. 원래 affinity training recipe의 완전한 재현성은 제한적 |
| Pearl2025 | invariant trunk + lightweight triangular operations + SO(3)-equivariant diffusion. generalized multi-chain / holo-like templates | synthetic-data scaling + curriculum. equivariant geometric inductive bias + controllable inference | PDB ≤2021-09-30. synthetic data도 같은 cutoff 이전 public structures에서 파생. scaling experiment는 910 proteins, 582,065 synthetic structures | Runs N’ Poses, PoseBusters, proprietary pocket-conditional benchmark | synthetic data와 equivariance의 결합으로 sample efficiency와 PL cofolding generalisation 강화. strict RMSD + physical-validity criterion에서 강점 | OOD 성능 저하가 여전히 존재하고 큰 induced-fit 같은 long-tail failure가 남음. 특히 best@k는 강하지만 confidence 기반 pose selection이 충분히 좋지 않음. affinity / dynamics는 핵심 범위 밖 |
| SeedFold2025 | AF3 계열 trunk를 width scaling. vanilla triangular attention 및 Linear Triangular Attention 변형. 512-width SeedFold와 384-width SeedFold-Linear | model / data / architecture 동시 scaling. large-scale distillation | 총 26.5M samples, experimental set 대비 약 147배 확장 | FoldBench | Pairformer width가 주요 capacity bottleneck임을 실증. triangular attention의 cubic bottleneck을 완화. protein-related task에서 강력 | 모델 variant별 task-specific 우열이 존재. structure prediction 중심이며 affinity, dynamics, generation, uncertainty calibration은 별도 문제로 남음 |
| OpenDDE2026 | 655M parameters, cz=384. 48 Pairformer blocks + 24 diffusion transformer blocks + structural-token refiner + shape-complementarity objectives + confidence / distance heads | warm-up + 다단계 curriculum. precision → breadth → precision. 후반부에는 prediction과 de novo design conditional training을 함께 수행 | cutoff 2021-09. weighted PDB, AFDB multimers, Teddymer, MGnify, Swiss-Prot, disordered data, SAbDab 등 | FoldBench, PXMeter-AB, FoldBench-AB, new 2026ARK-AB, test-time scaling | 공개 checkpoint / code / training details. atomic structural-token reasoning. prediction과 design을 한 diffusion formalism으로 통합하려는 명시적 시도 | 저자들도 현재 시스템을 완전한 DDE라기보다 folding-centred foundation으로 규정. affinity, conformational ensemble, active learning, experimental feedback은 향후 과제. 큰 oracle–ranking gap 역시 confidence / ranking 문제를 노출 |
주 — 표의 목적은 순위 산정이 아니다. 각 모델이 어떤 병목을 자기 문제로 삼았는지, 그리고 그 선택이 어떤 대가를 남겼는지를 나란히 놓는 데 있다.
이 표에서 가장 중요한 것은 “누가 1등인가”가 아니다. 각 모델이 서로 다른 병목을 공격하고 있다는 점이 중요하다.
따라서 Post-IsoDDE의 연구 질문은 다음과 같이 물어서는 안 된다.
물어야 할 것은 이것이다.
Runs N’ Poses는 training cutoff 이후의 2,600개 고해상도 protein–ligand system을 사용한다. 그리고 pocket과 ligand가 training data와 얼마나 유사한지에 따라 cofolding 성능이 크게 달라짐을 보여주었다. 신약설계에서 중요한 것은 평균 benchmark accuracy가 아니다. novel-target × novel-pocket × novel-ligand 영역의 worst-case 성능이다.
미래 DDE의 primary metric은 평균 성공률이 아니라 이 GOOD가 되어야 한다.
Pearl은 synthetic data를 추가할수록 fixed hold-out set 성능이 단조롭게 향상됨을 보여주며, SO(3)-equivariant diffusion을 통해 회전 대칭을 architecture 자체에 넣는다. 반면 SeedFold는 Pairformer의 depth보다 width가 더 중요한 scaling axis라고 분석하고, triangular attention 계산을 효율화한다. 두 연구는 같은 성능 향상을 서로 다른 원인으로 설명한다.
이 다섯 축은 따로 연구해야 한다. 하나의 ‘성능’이라는 말로 뭉뚱그리는 순간 무엇이 효과를 냈는지 알 수 없게 된다.
Boltz-2는 structure representation과 predicted geometry를 별도의 affinity module에 연결했다. 매우 중요한 방향 전환이다. 그러나 이후의 대규모 독립 평가에서는 affinity correlation이 target과 chemical regime에 따라 약해질 수 있고, top-ranked compounds에서는 physics-based free-energy estimate와의 일치가 충분하지 않은 사례가 보고되었다.
그리고 한 걸음 더 나아가면,
Pearl은 여러 sample 중 좋은 pose를 생성할 수 있음에도 confidence ranking이 이를 안정적으로 선택하지 못하는 문제를 명시한다. OpenDDE에서도 oracle selection과 model-ranking performance 사이에 큰 차이가 존재한다. 좋은 답을 만들어 놓고도 그것이 좋은 답인지 알아보지 못하는 상태다.
따라서 미래 시스템의 objective는 다음 하나로 끝나지 않는다.
여기에 반드시 다음이 더해져야 한다.
이는 uncertainty calibration과 test-time reasoning을 독립 연구축으로 취급해야 한다는 뜻이다.
네 개의 묶음으로 나눈다. 차세대 cofolding 기반모델, 일반화와 신뢰성, 예측에서 생성설계로, 정적 구조에서 앙상블과 동역학으로. 이 배열 자체가 하나의 논증이다.
Isomorphic Labs Team. Accurate Predictions of Novel Biomolecular Interactions with IsoDDE
2026 · Zenodo
IsoDDE의 직접적인 1차 자료다. protein–ligand generalisation, antibody–antigen, affinity, pocket identification을 하나의 Drug Design Engine의 predictive core로 제시한다. 동시에 architecture와 전체 training recipe가 공개되지 않았다는 점이 후속 연구의 중요한 reproducibility gap을 만든다.
Passaro et al. Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction
2025 · bioRxiv
all-atom cofolding과 binding-affinity prediction을 결합한 대표적인 공개 연구다. experimental-method conditioning, distance constraints, multi-chain templates 등을 추가하고 FEP 수준에 접근하는 affinity prediction을 목표로 했다. Post-IsoDDE 연구에서 “geometry → thermodynamics” 연결의 핵심 출발점이다.
Genesis Research Team et al. Pearl: A Foundation Model for Placing Every Atom in the Right Location
2025 · arXiv
SO(3)-equivariant diffusion, large-scale synthetic training data, generalized multi-chain templating을 결합한다. Runs N’ Poses / PoseBusters에서 stringent accuracy + physical-validity metric을 강조했으며 synthetic data scaling과 generalisation 사이의 관계를 실험적으로 분석했다.
Yi et al. SeedFold: Scaling Biomolecular Structure Prediction
2025 · arXiv
Pairformer width scaling, linear triangular attention, 26.5M-scale distilled dataset을 결합한다. “모델을 더 깊게” 하기보다 pair representation을 더 넓게 만드는 것이 효과적이라는 관찰은 차세대 DDE의 backbone scaling 전략에 직접적인 의미를 가진다.
Aureka AI OpenDDE Project. Folding, Reasoning, and Scaling with Open-source Drug Discovery Engine
2026 · arXiv
IsoDDE에 가장 직접적으로 대응하는 공개 연구 가운데 하나다. structural tokens, shape-complementarity, prediction / design unified diffusion, explicit training-compute scaling과 test-time scaling을 다룬다. 특히 oracle과 ranked performance의 차이는 차세대 시스템에서 uncertainty / ranking이 핵심 병목임을 보여준다.
ByteDance AML AI4Science Team et al. Protenix — Advancing Structure Prediction Through a Comprehensive AlphaFold3 Reproduction
2025 · bioRxiv
AF3 architecture를 높은 충실도로 재현·공개하여 all-atom biomolecular modelling의 실험 가능한 기반을 제공했다. 저자들도 memorisation 가능성을 제한점으로 명시하며, open reproducibility 관점에서 중요한 reference implementation 역할을 한다.
Protenix Team et al. Protenix-v1: Toward High-Accuracy Open-Source Biomolecular Structure Prediction
2026 · bioRxiv
동일 cutoff / model-scale / inference budget 조건에서 AF3와 정면 비교하고 inference-time scaling을 중요한 특성으로 분석한다. 또한 현실 응용을 위해 2025-06-30까지 데이터를 확장한 별도 variant를 제공함으로써 benchmark model과 deployment model의 분리를 보여준다.
Protenix Team et al. Protenix-v2: Broadening the Reach of Structure Prediction and Biomolecular Design
2026 · bioRxiv
structure prediction에서 biomolecular design으로 Protenix 계열을 확장한다. folding foundation model이 prediction endpoint에 머물지 않고 design prior가 되는 2026년의 큰 흐름을 보여주는 자료다.
IntFold Team. IntFold: A Controllable Foundation Model for General and Specialized Biomolecular Structure Prediction
2025 · arXiv
범용 구조예측 backbone 위에 allosteric-state control, user constraints, affinity adaptation과 prediction ranking을 추가하려는 연구다. “하나의 universal model + task adapter / conditioning” 전략은 제안 DDE의 modular specialization 설계에 중요한 근거가 된다.
Qiao et al. NeuralPLexer3: Accurate Biomolecular Complex Structure Prediction with Flow Models
NeurIPS 2025
diffusion 대신 physics-inspired conditional flow modelling을 활용하여 arbitrary biomolecular complex의 all-heavy-atom generation을 수행한다. DDE의 generative kernel을 반드시 diffusion으로 한정할 필요가 없으며 flow matching이 빠른 sampling의 대안이 될 수 있음을 보여준다.
Corley et al. Accelerating Biomolecular Modeling with AtomWorks and RF3
2025 · bioRxiv
AtomWorks라는 modular data framework와 RF3를 함께 공개하여 model architecture보다 training-data engineering과 reproducible preprocessing 자체가 frontier model 연구의 핵심임을 강조한다. chirality 처리 및 nucleic-acid distillation 개선도 중요하다.
Škrinjar et al. Evaluating Generalization in Protein–Ligand Cofolding Methods
2026 (preprint 2025) · Nature Structural & Molecular Biology
Runs N’ Poses의 2,600 post-cutoff complexes를 통해 ligand와 pocket의 training similarity가 cofolding accuracy에 미치는 영향을 정량화한다. “시간분할만으로 OOD benchmark가 되는가?”라는 질문을 제기한 핵심 논문이다.
Xu et al. Benchmarking All-Atom Biomolecular Structure Prediction with FoldBench
2025 · Nature Communications
protein monomer부터 protein–protein, antibody–antigen, protein–ligand, nucleic acid까지 low-homology system을 cross-domain으로 평가한다. 특정 모델이 모든 interaction class에 보편적으로 우월하지 않음을 보여주는 점이 중요하다.
Masters, Mahmoud & Lill. Investigating Whether Deep Learning Models for Co-Folding Learn the Physics of Protein–Ligand Interactions
2025 · Nature Communications
단순 pose accuracy가 아니라 ligand / protein perturbation에 모델이 물리적으로 합리적인 방향으로 반응하는지를 묻는다. Post-IsoDDE에서는 “정답 구조를 복원했는가”와 함께 interventional consistency를 평가해야 한다는 근거가 된다.
Large Scale Prospective Evaluation of Co-Folding across 557 Mac1–Ligand Complexes and Three Virtual Screens
2025 / 2026 · eLife
AF3, Chai-1, Boltz-2가 557개의 post-cutoff Mac1 ligand에 대해 절반 이상의 pose를 2 Å 이내로 재현했지만, specific conformational rearrangement와 hit-ranking에는 한계가 있었다. 특히 cofolding score와 physics-based docking score가 서로 보완적인 정보를 제공한다는 결과는 hybrid AI–physics architecture를 강하게 지지한다.
Comparative Assessment of the Utility of Co-Folding and Docking for Small-Molecule Drug Design
2025 · bioRxiv
Runs N’ Poses를 이용해 novel ligand와 novel pocket 영역에서 conventional physics-based docking이 cofolding보다 우수한 경우를 보고한다. benchmark protocol에 따른 caveat가 존재하지만 “AI가 docking을 완전히 대체한다”는 단순한 서사를 반박하며 adaptive hybrid inference의 필요성을 보여준다.
Influence of Molecular Representation and Charge on Protein–Ligand Structural Predictions by Popular Co-Folding Methods
2026 · bioRxiv
동일 molecule이라도 representation과 protonation / charge state의 차이가 cofolding 결과를 바꿀 수 있음을 분석한다. DDE input pipeline에서 protonation, tautomer, charge, stereoisomer uncertainty를 ‘전처리 문제’가 아니라 model uncertainty의 일부로 다뤄야 함을 의미한다.
Identifying and Addressing Systematic Data Leakage in Protein–Ligand Affinity Benchmarks
2026 · bioRxiv
affinity benchmark에서 ligand novelty와 assay / document leakage를 통제하지 않으면 높은 성능이 true interaction reasoning이 아니라 interpolation으로 발생할 수 있음을 분석하고 Novelty-Tiered Affinity Benchmark를 제안한다. 차세대 affinity 연구에서 거의 필수적인 evaluation principle이다.
Wan et al. On the Reliability of AI Methods in Drug Discovery: Evaluation of Boltz-2 for Structure and Binding Affinity Prediction
2026 · arXiv
3CLPro 16,780 compounds와 TNKS2 21,702 compounds를 이용해 Boltz-2를 large-scale로 검증한다. global affinity correlation이 제한적이고 top-ranked subset에서는 fine-grained physics evaluation과의 일치가 충분하지 않은 결과를 통해 screening과 lead optimization의 요구조건이 다름을 강조한다.
Affinity Fine-Tuning of Boltz-2: An Open Framework for Protein–Ligand Potency Prediction in Drug Discovery
2026 · bioRxiv
pretrained affinity model을 project-specific experimental assay data에 적응시키는 접근을 제시한다. foundation model을 고정된 oracle로 사용하기보다 campaign-specific continual adaptation을 수행해야 한다는 방향을 보여준다.
Cho et al. BoltzDesign1: Inverting All-Atom Structure Prediction Model for Generalized Biomolecular Binder Design
2025 · bioRxiv
Boltz-1의 Pairformer와 confidence output에 gradient를 역전파하여 sequence를 최적화한다. 별도 generative model을 만드는 대신 predictor 자체를 inverse-design objective로 뒤집을 수 있음을 보여준다는 점에서 중요하다.
Stark et al. BoltzGen: Toward Universal Binder Design
2025 / 2026 · bioRxiv
prediction과 design을 동일한 all-atom framework에서 통합하고 covalent bonds, binding sites, structural constraints 등을 design specification으로 제공한다. 여러 wet-lab campaign과 다양한 target에 대한 validation을 수행했다는 점에서 “universal binder design engine” 방향의 중요한 연구다.
Butcher et al. De Novo Design of All-Atom Biomolecular Interactions with RFdiffusion3
2025 · bioRxiv
protein만이 아니라 ligand, nucleic acid 등 non-protein atoms를 명시적으로 포함하여 atom-level constraints 아래 단백질을 생성한다. enzyme active-site geometry와 DNA / protein interaction처럼 정밀한 atomic constraint를 generative model에 직접 부여하는 방향을 제시한다.
Didi et al. Scaling Atomistic Protein Binder Design with Generative Pretraining and Test-Time Compute
2026 · arXiv
Proteína-Complexa는 large-scale synthetic Teddymer pretraining과 flow-based atomistic generation을 결합하고, best-of-N, beam search, Feynman–Kac steering, MCTS 등 test-time search를 binder design에 적용한다. 차세대 DDE에서 inference compute를 능동적으로 배분하는 search controller 설계의 직접적인 근거다.
Jing et al. Promera: A Unified Model for Biomolecular Structure Prediction, Filtering, and Design
2026 · bioRxiv
structure prediction, binder / non-binder filtering, controllable binder design을 하나로 묶는다. 특히 design candidate 생성 자체보다 filtering confidence를 별도 핵심 기능으로 본다는 점이 Post-IsoDDE architecture에 중요하다.
Kong et al. Programming Biomolecular Interactions with All-Atom Generative Model — AnewOmni
2026 · bioRxiv
5M 이상의 biomolecular complex를 이용한 atom-to-block latent representation과 programmable graph prompts를 제안하며 small molecule, peptide, nanobody를 하나의 generative framework에서 설계한다. molecular modality 간 interaction physics의 transfer 가능성을 연구한다는 점에서 “multi-modal therapeutic design engine”의 강력한 참고점이다.
Lewis et al. Scalable Emulation of Protein Equilibrium Ensembles with Generative Deep Learning — BioEmu
2025 · Science
200 ms 이상의 MD information, static structure와 stability data를 이용해 equilibrium conformational ensembles를 생성한다. cryptic pocket formation, local unfolding, domain rearrangement를 다룬다는 점에서 “single-structure DDE”에서 “ensemble-aware DDE”로 이동하기 위한 대표적인 연구다.
Wang et al. Learning the All-Atom Equilibrium Distribution of Biomolecular Interactions at Scale — AnewSampling
2026 · bioRxiv
15M 이상의 protein–ligand trajectory conformations를 활용해 all-atom equilibrium distribution을 학습하는 quotient-space generative framework를 제안한다. 특히 ligand torsion과 side-chain의 coupled dynamics를 정적 pose가 아닌 probability distribution으로 다룬다.
Feng et al. Physically Grounded Generative Modeling of All-Atom Biomolecular Dynamics — BioKinema
2026 · bioRxiv
Langevin-dynamics-inspired temporal attention과 hierarchical forecasting / interpolation으로 continuous-time all-atom trajectories를 생성한다. equilibrium ensemble을 넘어 induced fit, allosteric response, ligand unbinding 같은 kinetic pathways를 다루려 한다는 점에서 중요하다.
Liu et al. Reshaping Biomolecular Structure Prediction through Strategic Conformational Exploration with HelixFold-S1
2025 / 2026 · arXiv
무작위로 많은 conformer를 생성하는 대신 predicted inter-chain contact probability를 이용해 high-value conformational region에 inference compute를 집중한다. sampling budget을 ‘양’이 아니라 planning / search problem으로 재정의한다는 점에서 차세대 test-time reasoning 연구와 직접 연결된다.
공백은 성능이 모자란 자리가 아니다. 문제를 잘못 정의한 자리다. 여덟 개의 공백은 모두 ‘무엇을 측정하고 있는가’라는 하나의 질문으로 수렴한다.
대부분의 cofolding model은 본질적으로 하나 또는 몇 개의 candidate structure를 생성한다.
그러나 실제 interaction은 하나의 좌표가 아니라 분포다.
BioEmu, AnewSampling, BioKinema는 equilibrium state와 kinetics가 별개의 modelling problem임을 보여준다. 따라서 single pose를 맞히는 모델과 실제 drug action을 설명하는 모델 사이에는 dynamics gap이 존재한다.
현재 benchmark의 평균점수는 protein family, pocket, ligand scaffold가 training set과 유사한 경우에 의해 과도하게 영향을 받을 수 있다. Runs N’ Poses와 FoldBench는 이러한 generalisation 문제를 직접 드러낸다. 필요한 것은 하나의 숫자가 아니라 조건부 평가다.
이것이 novelty-conditioned evaluation이다.
단일 affinity score는 실제 binding thermodynamics 전체를 표현하지 못한다. 더 현실적인 표현은 다음과 같다.
여기에 더해 결합의 시간 구조까지 포함해야 한다.
Boltz-2 이후 affinity prediction이 크게 주목받았다. 그러나 independent evaluation과 leakage 연구는 high benchmark correlation을 그대로 medicinal-chemistry utility로 해석해서는 안 됨을 보여준다.
Pearl과 OpenDDE에서 공통적으로 관찰되는 현상은 다음 한 줄로 요약된다.
이는 모델이 정답에 가까운 구조를 생성할 능력은 있지만 그것이 정답인지 판별할 능력은 부족하다는 뜻이다. 따라서 candidate generation과 candidate verification을 동일 network의 confidence head 하나에 맡기는 것은 충분하지 않을 수 있다.
prospective Mac1 평가와 cofolding-vs-docking 연구는 AI와 classical docking이 서로 다른 오류를 범할 수 있음을 시사한다. 오류의 구조가 다르다면 둘은 경쟁 관계가 아니라 보완 관계다. 따라서 미래 architecture의 명제는 ‘AI냐 물리냐’가 아니다.
구조 데이터에서 correlation을 학습했다고 해서 interaction physics를 학습했다고 할 수는 없다. 다음 질문에 물리적으로 일관된 방향으로 답할 수 있어야 한다.
2025년 cofolding-physics 분석은 이러한 perturbational evaluation의 중요성을 제기한다. 따라서 Post-IsoDDE에서는 interventional molecular reasoning이 별도 연구목표가 되어야 한다.
affinity만 최적화하면 실제 drug candidate가 되지 않는다. 실제 설계는 벡터 값 문제다.
affinity, selectivity, kinetics, toxicity, solubility, permeability, metabolic stability, synthetic accessibility, novelty를 동시에 고려하는 Pareto problem이다.
회고적 평가의 구조는 이렇다.
실제 신약개발의 구조는 전혀 다르다.
따라서 최종적인 DDE benchmark는 benchmark dataset이 아니라 prospective Design–Make–Test–Learn loop여야 한다.
Physics-grounded, Robust, Interventional, Sampling-aware, Multi-objective Drug Design Engine. 다섯 글자는 장식이 아니라 앞의 여덟 공백에 대한 응답이다.
AI prediction을 docking, energy minimization, MD / FEP / ABFE 등과 선택적으로 결합한다.
time, target, pocket, ligand scaffold를 동시에 통제한 OOD learning과 calibrated uncertainty를 사용한다.
mutation, ligand edit, protonation / charge, conformational-state perturbation에 대한 causal consistency를 학습한다.
하나의 구조가 아니라 structure ensemble, equilibrium distribution, kinetics 및 adaptive test-time search를 다룬다.
affinity 하나가 아니라 selectivity, kinetics, ADMET, synthesizability, novelty까지 Pareto optimization한다.
기존 cofolding 문제는 짧게 쓸 수 있었다.
PRISM-DDE는 이를 확장한다. 표적 T, ligand L, 환경 E, experimental context C가 있을 때 다음을 학습한다.
inverse-design 문제는 다음으로 정의한다.
여기서 𝒞는 binding-site constraints, interaction motif, pharmacophore, covalent / non-covalent constraint, synthesis constraints, property constraints를 나타낸다.
개념적으로 다음 구조를 제안한다. 이 설계에서 가장 중요한 점은 모든 candidate에 FEP나 MD를 수행하지 않는다는 것이다. 계산은 균등하게 나누는 자원이 아니라 불확실한 곳에 몰아주는 자원이다.
candidate x에 대해 uncertainty U(x), novelty N(x), 서로 다른 model / physics predictor 사이의 disagreement D(x)를 계산한다.
g(x)=0이면 빠른 AI path만 사용한다. g(x)=1이면 위험도에 맞는 계산을 동적으로 호출한다.
따라서 목표는 단순 accuracy가 아니다.
이는 IsoDDE / Boltz-2의 빠른 AI prediction과 Mac1 · cofolding-vs-docking 연구에서 보이는 physics complementarity를 직접 연결한다.
PRISM-DDE는 하나의 structure를 ground truth로 취급하지 않는다.
affinity 역시 conformational ensemble을 통합하여 계산한다.
이 방식은 특히 다음 상황에서 유리할 것으로 가정한다.
BioEmu, AnewSampling, BioKinema가 각각 equilibrium ensemble과 kinetic trajectories를 별도의 학습대상으로 만들고 있다는 점이 이 설계의 근거다.
단순한 supervised structure loss 외에 intervention pair를 만든다. residue mutation 또는 ligand functional-group edit를 수행하는 방식이다.
그리고 예측 변화량이 실험값 또는 고충실도 physics reference와 일치하도록 학습한다.
즉 모델에 “정답 구조는 무엇인가”만 묻지 않는다.
이 질문을 학습시키는 것이 correlation-based cofolding에서 scientific molecular reasoning으로 이동하는 핵심 연구 novelty가 될 수 있다. perturbation에 대한 물리적 일관성을 평가해야 한다는 기존 연구가 이 방향의 필요성을 뒷받침한다.
현재 많은 모델은 confidence를 ranking에 사용한다. 그러나 Pearl · OpenDDE의 결과는 좋은 sample을 만들어 놓고도 제대로 선택하지 못할 수 있음을 보여준다. 따라서 PRISM-DDE는 uncertainty를 분해한다.
특히 representation uncertainty에는 다음을 포함한다.
molecular representation / charge 연구가 이러한 variation이 prediction에 영향을 줄 수 있음을 보여준다. 최종적으로는 abstention을 허용한다.
신약개발에서는 틀린 고신뢰 prediction보다 calibrated abstention이 훨씬 가치 있을 수 있다. 모르는 것을 모른다고 말하는 능력이 곧 성능이다.
candidate m에 대해 목적함수는 벡터다.
하나의 scalar reward로 모두 합치는 대신 전선을 유지한다.
필요하면 프로젝트 단계별 preference vector wt를 적용하여 candidate를 선택한다.
이 구조의 장점은 discovery stage에 따라 objective를 바꿀 수 있다는 것이다.
wnovelty , waffinity ↑
wselectivity , wADMET , wsynthesis ↑
제안은 검증 가능한 형태로 진술될 때만 연구가 된다. 일곱 개의 질문, 각각에 대응하는 가설, 그리고 그 가설을 반증할 수 있는 측정 규약이 필요하다.
protein sequence, pocket geometry와 ligand chemotype이 동시에 training distribution에서 멀어질 때 PRISM-DDE가 기존 cofolding systems보다 구조와 affinity의 정확도를 유지할 수 있는가?
H1novelty-aware training과 interventional consistency learning을 결합하면 Runs N’ Poses의 low-similarity tier에서 평균 accuracy뿐 아니라 worst-tier success rate가 향상될 것이다.
단일 predicted pose가 아닌 conformational ensemble을 이용하는 것이 binding-affinity와 selectivity ranking을 유의하게 개선하는가?
H2ΔĜensemble이 ΔĜsingle보다 congeneric-series ranking과 induced-fit target에서 유의하게 높은 Spearman / Pearson correlation을 보일 것이다.
모든 candidate에 고비용 physics를 적용하지 않고 uncertainty / OOD에 따라 선택적으로 physics를 호출해도 full-physics pipeline에 근접한 accuracy를 달성할 수 있는가?
H3adaptive gating은 full FEP / MD pipeline보다 훨씬 낮은 compute를 사용하면서 AI-only pipeline보다 높은 top-k enrichment와 lead-ranking accuracy를 달성할 것이다.
mutation, ligand edit, protonation 및 conformational perturbation을 이용한 interventional training이 모델의 genuine molecular reasoning을 향상시키는가?
H4intervention-trained model은 unseen mutation 및 matched molecular pair에서 ΔΔG prediction과 direction-of-effect accuracy를 개선할 것이다.
ensemble disagreement와 novelty-aware uncertainty를 결합하면 기존 confidence head보다 잘못된 구조 / affinity prediction을 더 안정적으로 식별할 수 있는가?
H5새 uncertainty layer는 더 낮은 ECE, 더 낮은 Brier score, 더 낮은 AURC, 더 높은 error-detection AUROC를 보일 것이다.
uniform sampling보다 adaptive test-time search가 동일 compute budget에서 더 높은 structural / design success를 달성하는가?
Proteína-Complexa와 HelixFold-S1은 이 연구 질문이 biomolecular modelling에서 실제로 유효한 방향임을 보여준다.
affinity-only generative model보다 Pareto-based design이 실험적으로 검증된 hit rate, selectivity, ADMET 및 synthetic success를 동시에 개선할 수 있는가?
이것이 최종적인 drug-design RQ다. 나머지 여섯 질문은 모두 이 질문에 답하기 위한 사전 조건이다.
그러나 random split은 사용하지 않는다. random split은 성능을 만들어내는 가장 손쉬운 방법이자 가장 무의미한 방법이다.
각 train / test pair에서 다음 네 조건을 동시에 측정한다.
추가로 affinity data에서는 동일 publication / document / assay family가 train과 test 양쪽에 들어가지 않도록 group splitting한다. 최근 affinity leakage 연구가 바로 이 종류의 protocol 필요성을 보여준다.
| Target | Ligand | 의미 | |
|---|---|---|---|
| Low | Low | Low | memorisation-friendly |
| Low | Low | High | scaffold hopping |
| Low | High | High | new pocket chemistry |
| High | High | Low | ligand transfer |
| High | High | High | true frontier |
핵심 benchmark는 마지막 영역이다. 이를 다음과 같이 정의할 수 있다.
전체 loss는 일곱 항의 합으로 구성한다.
OpenDDE의 shape-complementarity와 RF3 계열의 chirality / atomic conditioning 연구는 이런 objective의 중요성을 뒷받침한다.
대규모 구조 데이터로 sequence → all-atom geometry를 학습한다.
Pearl / SeedFold가 보여준 방향을 따라 다양한 synthetic structure와 distillation을 사용하되, train / test cutoff를 엄격히 유지한다.
affinity, ΔΔG, matched molecular pairs, mutation data를 학습한다.
MD / ensemble data를 활용해 p(X)와 p(Xt+Δt | Xt)를 학습한다.
mutation / chemical-edit pair를 이용해 causal direction consistency를 학습한다.
structure / interaction representation을 conditional design objective로 전환한다.
AI candidate 가운데 docking / MD / FEP / experimental evidence가 더 좋은 candidate에 preference를 부여한다. language model의 RLHF와 비슷한 개념이지만, 여기서 선호를 제공하는 주체는 사람이 아니다.
이 값들을 함께 보고한다. RMSD 하나만으로 평가하지 않는다.
similarity bin별 성공률 SR(b)를 계산하고 다음 두 지표를 추가한다.
이 두 metric을 통해 easy case가 평균을 끌어올리는 현상을 방지한다.
단순 Pearson r만 사용하지 않는다.
Pearson r · Spearman ρ · RMSE · MAE
pairwise ranking accuracy · congeneric-series Spearman · ΔΔG error
EF1% · BEDROC · PR-AUC
각 metric을 Nligand 및 Npocket tier별로 보고
이것이 최근 leakage analysis에 대한 직접적인 방법론적 대응이다.
static RMSD가 아니라 분포를 비교한다.
BioEmu · AnewSampling · BioKinema를 하나의 benchmark 축으로 연결하는 부분이다.
confidence는 correlation이 아니라 calibration으로 측정한다.
예를 들어 모델이 uncertainty 상위 20% prediction을 거부했을 때 위험이 얼마나 감소하는지를 평가한다.
매우 중요한 연구 문제다. generative model G가 candidate m을 만들고 동일 계열 structure predictor V가 이를 평가한다면 순환 편향이 생긴다.
BoltzDesign1 역시 동일 predictor를 design과 evaluation에 함께 활용할 때 발생할 수 있는 overfitting 문제를 제한점으로 논의한다. 자기가 만든 답을 자기가 채점하는 구조에서는 점수가 올라가도 그것이 무엇을 뜻하는지 알 수 없다.
따라서 PRISM-DDE에서는 최소한 세 단계의 독립적 평가를 사용한다.
| Ablation | 검증하려는 질문 |
|---|---|
| − Synthetic data | data scaling이 OOD에 실제 기여하는가? |
| − Equivariant blocks | geometric inductive bias의 효과는? |
| − Dynamics module | ensemble 정보가 affinity / selectivity를 개선하는가? |
| − Intervention loss | causal perturbation reasoning이 개선되는가? |
| − Physics gate | hybrid verification이 필요한가? |
| − OOD detector | uncertainty와 novelty를 분리할 필요가 있는가? |
| − Test-time search | 단순 sample 수 증가보다 search가 좋은가? |
| − Pareto layer | affinity-only design보다 실제 candidate quality가 좋아지는가? |
| − Active learning | 실험 feedback이 sample efficiency를 높이는가? |
AlphaFold 3 · IsoDDE(공개 결과와 비교 가능한 범위) · Boltz-2 · Pearl · SeedFold · OpenDDE · Protenix-v1/v2 · NeuralPLexer3
AutoDock Vina 계열 · strong flexible-docking baseline · MD / refinement · FEP / ABFE where feasible
MD ground truth · BioEmu · AnewSampling · BioKinema
BoltzDesign1 · BoltzGen · RFdiffusion3 · Proteína-Complexa · Promera · AnewOmni
벤치마크 하나를 더 추가하는 것으로는 이 논문의 주장을 증명할 수 없다. 증명은 실험실에서 이루어진다.
모델과 training data를 freeze한다.
low-similarity target을 선택한다. induced-fit kinase, cryptic / allosteric target, GPCR, antibody–antigen interface 등 서로 다른 어려움을 가진 3–5개 target class.
모델이 candidate를 생성하고 단계적으로 좁힌다.
최종 candidate는 model score를 공개하기 전에 synthesis / assay team에 blind transfer한다.
기존의 1차 지표는 RMSD 또는 Pearson r이었다. PRISM-DDE의 1차 지표는 다르다.
즉 cost-adjusted prospective discovery rate를 궁극의 지표로 제안한다. 예를 들어 다음과 같이 정의할 수 있다.
이렇게 하면 AI 연구의 metric을 실제 drug-discovery productivity와 직접 연결할 수 있다.
protein, pocket, ligand, time novelty를 동시에 통제하는 Novelty Cube와 Frontier-Generalisation metric을 제안한다.
static cofolding과 equilibrium / kinetic modelling을 동일 interaction representation에 결합한다.
mutation 및 chemical perturbation에 대한 consistency objective를 도입하여 correlation과 causal interaction reasoning을 구분한다.
uncertainty와 novelty에 따라 docking / FEP / MD compute를 동적으로 할당한다.
generator와 verifier를 분리하고 model–physics–experiment의 Triangulated Verification을 도입한다.
affinity-only optimisation을 넘어 selectivity, kinetics, ADMET, synthesis, novelty를 Pareto problem으로 정의한다.
최종 성공 기준을 retrospective score가 아니라 prospective experimental validation으로 설정한다.
PRISM-DDE: Physics-Grounded and OOD-Robust Drug Design through Interventional Reasoning, Ensemble Sampling, and Multi-Objective Generation
Beyond Static Co-Folding: PRISM-DDE for Physics-Grounded, Uncertainty-Calibrated and Closed-Loop Drug Design
Beyond IsoDDE: A Physics-Grounded Ensemble Reasoning Engine for Prospective Molecular Design
세 번째는 강한 제목이지만 실제 투고 논문에서 특정 기업 모델을 title에 넣으면 연구의 일반성이 좁아진다. 첫 번째가 가장 안정적이다.
그리고 바로 다음 문단으로 이어간다.
현재의 모델들은 대체로 이 질문에 답한다.
IsoDDE는 질문을 여기까지 확장했다.
Post-IsoDDE의 진짜 목표는 다음 질문이어야 한다.
이 차이는 매우 크다. 첫 번째는 prediction이다. 두 번째는 interaction modelling이다. 세 번째는 다른 종류의 일이다.
이 관점에서 보면 Post-IsoDDE Drug Design Engine의 경쟁력은 더 이상 “structure RMSD가 조금 더 낮은가”에 있지 않다. 진정한 경쟁력은 여섯 항의 곱으로 결정되어야 한다.
곱이라는 점이 중요하다. 어느 하나가 0이면 전체가 0이다. 그리고 이 기준을 채택하면 PRISM-DDE의 가장 중요한 연구 질문은 한 문장으로 압축된다.
이 질문이 구조예측 모델과 진정한 Drug Design Engine을 구분하는 경계다.