AI Research Notes · Drug DesignPost-IsoDDE Horizon Scan · 26 Aug 2026
Horizon Scan · AI Drug Design · August 2026

생성은 충분히 좋아졌다.
이제 병목은 ‘무엇을 믿고 고를 것인가’이다.

Post-IsoDDE: Selection, Ranking, and State Awareness as the New Bottlenecks in AI Drug Design

IsoDDE 이후의 경쟁을 단순한 구조예측 정확도 싸움으로 보면 중요한 변화를 놓친다. 2026년 8월의 최신 연구들은 서로 다른 과제에서 같은 경고를 보낸다. 좋은 구조와 분자를 후보 집합 안에 넣는 능력은 빠르게 향상되었지만, 그중 실제로 맞는 후보를 고르고, 올바른 단백질 상태를 선택하고, 다음 실험으로 보낼 결정을 내리는 능력은 아직 그 속도를 따라가지 못한다.

Post-IsoDDE decision bottleneckConceptual illustration. Multiple generated candidates converge on a narrow ranking gate, then pass through state validation before prospective experiment. GENERATESELECTSTATE GATEEXPERIMENT poses · pockets · molecules · antibodiesconfidence / rescoringensemble · physics · contextprospective readout RANKERoracle ≠ selectedcalibration matters conformationalstate space wet-lab truth THE EMERGING BOTTLENECK
Coverage: 17–26 Aug 2026Focus: IsoDDE · AF3 · Boltz-2 · generative designMode: evidence + analysis

Central thesis

이번 추적에서 가장 강한 신호는 “더 큰 cofolding model” 자체가 아니다. 항체, cyclic peptide, cryptic pocket, GPCR allostery, FEP, molecular generation을 가로질러 반복되는 패턴은 하나다. 생성 능력은 선택 능력보다 앞서 있다.

좋은 후보를 만들 수 있다는 사실은, 그 후보가 좋은지 알아볼 수 있다는 뜻이 아니다. 그리고 좋은 구조를 고를 수 있다는 사실도, 그것이 실제 약물개발의 다음 실험을 올바르게 선택한다는 뜻은 아니다.
\[\boxed{\text{Can generate}\;\not\Rightarrow\;\text{Can rank}\;\not\Rightarrow\;\text{Can predict affinity}\;\not\Rightarrow\;\text{Can choose the next experiment}}\]
Part I · Signal

이번 주 연구들이 가리키는 하나의 방향

각 연구는 서로 다른 modality와 benchmark를 다루지만, 실패 지점은 놀라울 만큼 비슷하다.

§1 · Evidence map

새 모델보다 중요한 것은 ‘어디에서 실패하는가’이다

연구평가 대상핵심 관찰드러난 병목
AIntibody511 antibodies · 29 organizations일부 <100 pM 항체 성공, 그러나 HCDR3 cluster ranking은 대부분 random보다 나쁨Prospective ranking
AF3/Boltz-2 + FEPcSrc · 133 ligands · >1,400 FEP runs구조의 출처보다 올바른 conformational state 선택이 중요State selection
Pair Representation Scaling86 two-state proteins재학습 없이 latent pair scaling으로 alternative state recovery 확장State coverage / test-time control
Cyclic peptide cofolding111 complexes · 22,200 poses좋은 pose는 자주 생성하지만 native confidence가 best pose를 충분히 못 고름Pose ranking
Cryptic pocket benchmarkCryptoBench · 178 structures후보는 이미 많이 제안되지만 top-k에 올리는 능력이 약함Pocket ranking
12-model generative benchmark176 protein–ligand systems모든 기준을 지배하는 단일 generator가 없음Stage-specific routing
GPCR allosteric screening4 Class-A GPCRsGaMD ensemble + conventional docking이 Boltz-2 단독보다 실용적Dynamics-aware hybridization
MASCOTMulti-objective lead optimization역할특화 agents가 chemical edit와 실험평가를 연결Decision orchestration

Analysis이 결과들을 한데 놓으면 Post-IsoDDE 연구의 중심축은 prediction accuracy만이 아니라 candidate selection fidelity로 이동한다. 구조·분자·pocket을 생성하는 foundation model이 강해질수록, 그 위에서 uncertainty를 보정하고 후보를 재랭킹하고 상태를 검증하는 층이 전체 성능을 지배하게 된다.

Part II · Prospective Reality Check

AIntibody: 구조를 잘 맞히는 것과 치료용 항체를 고르는 것은 다르다

CASP형 blind benchmark가 antibody discovery의 실제 의사결정으로 이동했다.

§2 · AIntibody

511개 항체를 실제로 합성하고 같은 조건에서 재었다

Source fact2026년 8월 19일 Nature Biotechnology에 출판된 AIntibody benchmark는 29개 기관이 제출한 511개 AI-designed 또는 AI-predicted antibody를 대상으로 affinity maturation, HCDR3 cluster 내 affinity ranking, out-of-library CDR design의 세 과제를 prospective·blinded 방식으로 평가했다. 설계된 sequence는 full IgG로 합성되고, SPR·KinExA 및 developability panel을 통해 동일 조건에서 측정되었다.

일부 팀은 developability를 유지하면서 100 pM 이하 affinity를 달성했다. 그러나 성공은 과제 전반으로 일반화되지 않았다. 특히 HCDR3 cluster에서 고-affinity clone을 고르는 문제는 한 모델을 제외하고 random clone picking보다 나빴다. 논문은 첫 challenge가 매우 잘 연구된 SARS-CoV-2 RBD와 풍부한 sequencing/affinity 정보를 제공했다는 점에서 현재 역량의 비교적 유리한 상한 조건으로 해석해야 한다고 명시한다.

항체 구조를 그럴듯하게 만드는 능력과, 실제로 합성할 다음 항체를 고르는 능력 사이에는 아직 큰 간극이 있다.

AnalysisIsoDDE가 antibody–antigen interface와 CDR-H3 modelling에서 강한 구조예측 결과를 제시했다면, AIntibody는 그 다음 질문을 던진다. “그 구조 정보를 사용해 어떤 sequence를 실제 실험으로 보낼 것인가?” 앞으로 antibody benchmark는 DockQ와 RMSD만이 아니라 prospective affinity ranking, developability, calibration, novelty distance를 함께 봐야 한다.

원문: Nature Biotechnology — A blinded, prospective benchmark of in silico antibody discovery...

Part III · State Awareness

정확한 좌표보다 더 어려운 문제: ‘어떤 상태의 단백질인가’

AF3와 Boltz-2는 단일 구조예측기를 넘어 conformational state를 탐색하는 도구로 재해석되고 있다.

§3 · Structure → FEP

ML 구조를 넣으면 FEP가 자동으로 좋아지는가?

Source fact8월 20일 Journal of Chemical Information and Modeling에 출판된 연구는 cSrc의 133개 congeneric ligand에 대해 experimental crystal structure, homology model, ML-predicted structure를 FEP 입력으로 비교했고, 1,400회가 넘는 계산을 수행했다. 핵심 메시지는 구조의 출처보다 micro/macro conformational state가 올바른가가 reliability에 큰 영향을 준다는 점이다.

Analysis이는 “AF3/Boltz-2 pose → FEP”를 단순 직렬 연결하는 pipeline에 경고를 준다. Drug Design Engine에는 구조 생성기와 affinity 계산 사이에 state validation이 필요하다.

\[\text{cofolded pose}\rightarrow\boxed{\text{state validation}}\rightarrow\text{protonation/solvation}\rightarrow\text{FEP or affinity model}\]

원문: JCIM — To ML-Predict or Not to ML-Predict

§4 · Pair Representation Scaling

재학습 없이 alternative state를 더 꺼내는 한 개의 스칼라

Source fact같은 날 정식 출판된 Pair Representation Scaling은 Pairformer 앞의 latent pair representation에 \(z'=(1+\beta)z\)를 적용한다. auxiliary model도, retraining도, 별도의 두 번째 forward pass도 필요하지 않는다. 86개 two-state target에서 AF3와 Boltz-2의 sampled ensemble을 넓혀 default inference가 놓친 alternative state recovery를 개선했으며, 효과는 AF3에서 특히 강했다.

\[z'=(1+\beta)z\]

AnalysisBioEmu처럼 ensemble 자체를 목적으로 학습한 생성모델과, 기존 cofolding model의 latent distribution을 test-time에 steering하는 접근이 새로운 비교축을 만든다. 향후 benchmark는 top-1 RMSD뿐 아니라 state coverage per unit compute를 재야 한다.

원문: JCIM — Biasing Conformational Sampling in AlphaFold 3 and Boltz-2...

Part IV · Ranking Crisis

후보는 있는데, 그 후보를 위로 올리지 못한다

Cyclic peptide와 cryptic pocket이라는 서로 다른 과제가 같은 실패 모드를 드러낸다.

§5 · Cyclic peptide–protein cofolding

Pose generation은 강하다. Confidence가 그 속도를 따라가지 못한다.

Source fact8월 21일 공개된 benchmark는 다섯 종류 cyclization chemistry를 포함한 111개 nonredundant complex에 대해 Boltz와 Protenix가 target당 100개의 pose를 생성하도록 하여 총 22,200개 pose를 평가했다. median top-pose DockQ는 약 0.89이고 96–98% target에서 Acceptable 이상의 구조를 생성했다. 그러나 native ranking score와 실제 pose quality의 Spearman 상관은 약 0.53–0.66이었고, 약 12% pose는 confidence가 높지만 구조 품질이 낮았다. 최고 품질 pose가 rank 1이 아닌 경우도 거의 모든 target에서 관찰되었다.

연구진은 interface hydrogen-bond density 같은 외부 구조 feature를 native score에 더한 gradient-boosted rescoring으로 ranking을 개선했다.

Generation capability > Selection capability여러 후보를 만들 수 있을 때, 실제 병목은 best candidate를 알아보는 데 생긴다.

원문 DOI: 10.64898/2026.08.20.746104

§6 · Cryptic pockets

“못 찾는다”보다 “잘못 순위를 매긴다”가 더 큰 문제일 수 있다

Source factCryptoBench test fold 178개 structure에서 fpocket은 qualifying candidate를 74.2% target에 제안했지만 top-5 recovery는 43.8%였다. P2Rank는 coverage 66.3%, top-5 recovery 63.5%, IF-SitePred는 70.8%와 61.8%였다. 네 detector의 candidate union은 92.1% coverage에 도달했다. 즉 candidate가 전혀 없는 target보다 candidate는 존재하지만 위로 올라오지 않는 target이 훨씬 많다.

Coverage

후보를 제안할 수 있는가

실제 cryptic site와 겹치는 proposal이 candidate pool 안에 존재하는 비율이다.

Conversion

사용자가 볼 만큼 위로 올리는가

이미 생성한 올바른 후보를 top-k budget 안에 surface하는 능력이다.

AnalysisIsoDDE의 ligand-free pocket identification을 평가할 때도 AUPRC 하나만으로는 부족하다. coverage, ranking conversion, candidate budget, ensemble cost를 분리해서 보는 편이 더 실전적이다.

원문: bioRxiv — Cryptic binding sites are detected but not ranked

Part V · Hybrid Design

단일 거대 모델보다 ‘상황에 맞게 조합하는 엔진’

Generative model과 physics, dynamics, synthesis-aware scoring을 단계별로 routing하는 방향이 더 강한 실험적 근거를 얻고 있다.

§7 · 12-model generative benchmark

모든 기준을 동시에 이기는 generator는 없었다

Source fact8월 17일 공개된 systematic benchmark는 12개 molecular generation/optimization method를 176개 curated protein–ligand system에서 비교했다. chemical validity, uniqueness, molecular/scaffold diversity, QED, synthetic accessibility, docking, physicochemical/ADMET property, computational resource requirement까지 함께 평가했다.

결과는 architecture별 trade-off를 보여준다. pocket-conditioned 3D model은 receptor geometry 활용에, flow-based model은 sampling efficiency에, reference-conditioned method는 analogue generation에, synthesis-aware method는 chemical feasibility에 강했다. 하지만 모든 기준에서 일관되게 우수한 단일 architecture는 없었다.

연구진은 MD-derived receptor ensemble, ensemble docking, protein–ligand interaction graph를 통합한 State-Aware Functional Classifier(SAFC)를 추가해 docking·drug-likeness·synthetic accessibility와 일부 상보적인 ranking signal을 제시했다.

Analysis‘Unified Drug Design Engine’은 반드시 monolithic neural network일 필요가 없다. 더 현실적인 통합은 서로 다른 강점을 가진 모델을 stage별로 선택하는 orchestration architecture일 수 있다.

원문: bioRxiv — Systematic Benchmarking of AI-Based Molecular Generation Models...

§8 · GPCR allostery

Boltz-2는 dynamics-aware docking을 아직 대체하지 못했다

Source fact네 개 Class-A GPCR의 allosteric modulator screening을 비교한 연구에서는 PDB structure와 GaMD ensemble을 Glide HTVS, AutoDock Vina, DOCK3.8, Boltz-2에 적용했다. GaMD ensemble은 네 target 모두에서 적어도 한 program의 early enrichment를 개선했다. Glide ensemble docking만이 네 target에서 일관되게 개선되었고, M2 receptor에서는 PDB structure 대비 active recovery가 거의 9배 향상되었다.

Boltz-2는 GaMD template 변화에 상대적으로 둔감했고 이 workload에서는 conventional docking보다 낮은 screening 성능을 보였다. 저자들은 Boltz-2 affinity signal을 physics/empirical docking의 대체라기보다 complementary signal로 해석했다.

\[\boxed{\text{MD ensemble}+\text{cofolding}+\text{physics docking}+\text{consensus ranking}}\]

Inference특히 cryptic/allosteric target에서는 “foundation model 하나가 flexibility를 내부적으로 해결한다”는 가정보다 explicit dynamics와 learned model을 결합하는 편이 당분간 더 안전한 전략일 수 있다.

원문 DOI: 10.64898/2026.08.12.744492

Part VI · Agentic Layer

Foundation model 위에 ‘의사결정하는 medicinal chemistry layer’가 올라오기 시작했다

생성·예측 모델의 출력물을 실제 multi-objective optimization과 실험 후보 선택으로 연결하는 문제다.

§9 · MASCOT

Multi-Agent Molecular Optimization이 실험 후보까지 이어졌다

Source factMASCOT은 chemically constrained graph-editing search 위에 세 역할을 둔다. trade-off agent는 potency·PK·safety 등 경쟁 objective의 우선순위를 조절하고, strategy agent는 molecular edit의 제안 방식을 바꾸며, reflection agent는 이전 의사결정의 교훈을 다음 search에 반영한다.

SARS-CoV-2 Mpro task에서 mean docking-score improvement가 가장 강한 baseline의 3.6배였고, remimazolam optimization에서는 RM-1을 우선순위화한 뒤 후속 derivative로 RM-7을 도출했다. preprint는 animal study에서 RM-7이 더 높은 potency, 빠른 functional recovery, 넓은 safety margin, flumazenil reversibility를 보였다고 보고한다.

\[\text{objective trade-off}\rightarrow\text{chemical edit}\rightarrow\text{evaluation}\rightarrow\text{reflection}\rightarrow\text{experimental candidate}\]

Analysis이 흐름은 IsoDDE/OpenDDE/Boltz류 foundation model을 ‘도구’로 보고, 그 위에 어떤 계산을 호출하고 어떤 후보를 다음 단계로 보낼지를 결정하는 agentic layer를 두는 설계를 현실적으로 만든다. 다만 현재 MASCOT은 preprint이므로 독립 재현과 전체 실험 프로토콜 검증이 필요하다.

원문 DOI: 10.64898/2026.08.17.745149

Part VII · Post-IsoDDE Agenda

차세대 Drug Design Engine은 무엇을 최적화해야 하는가

정답 후보를 생성하는 것보다, 그 정답을 식별하고 불확실하면 abstain하며 실험으로 보내는 능력이 핵심 연구대상이 된다.

§10 · A common architecture

Foundation model + state + physics + calibrated ranking + experiment

이번 문헌을 종합하면 다음 구조가 자연스럽게 도출된다.

\[\boxed{\text{Foundation Model}\rightarrow\text{Conformational Ensemble}\rightarrow\text{Physics/Context Gate}\rightarrow\text{Calibrated Ranker}\rightarrow\text{Multi-objective Agent}\rightarrow\text{Prospective Experiment}}\]
Layer 1

Generate broadly

Protein–ligand pose, pocket, antibody, molecular candidate를 충분히 넓게 생성한다.

Layer 2

Select honestly

confidence calibration, structural rescoring, state validation으로 false certainty를 줄인다.

Layer 3

Learn prospectively

wet-lab outcome을 다음 ranking·generation cycle의 evidence로 되돌린다.

§11 · Oracle–selected gap

Post-IsoDDE benchmark에 추가할 가장 단순하면서 강한 지표

한 모델이 \(K\)개의 후보를 생성한다고 하자. 그중 실제 quality가 가장 높은 후보를 oracle이라고 하고, 모델 confidence/ranker가 선택한 후보를 selected라고 하면 다음 차이를 측정할 수 있다.

\[\Delta_{\mathrm{select}}=Q(X_{\mathrm{oracle}})-Q(X_{\mathrm{selected}})\]

Analysis이 지표는 “좋은 답을 생성할 능력”과 “그 답을 알아볼 능력”을 분리한다. cyclic peptide에서는 이미 이 gap이 크게 드러났고, cryptic pocket에서도 coverage와 conversion의 분리가 같은 논리를 보여준다. antibody design에서는 oracle을 retrospective experimental best clone으로 정의해 prospective selection loss를 측정할 수 있다.

Research inference향후 IsoDDE, OpenDDE, Boltz, Pearl, SeedFold류 시스템을 비교할 때 top-1 accuracy와 함께 oracle-selected gap, calibration error, compute-normalized state coverage, abstention utility를 공동 지표로 두면 훨씬 강한 scientific benchmark가 된다.

§12 · Model status

이번 추적 기간에 확인된 공식 release 상태

2026년 8월 17–26일 범위에서 새 IsoDDE 본체 기술보고서나 새로운 OpenDDE/Pearl/SeedFold/BioEmu의 primary model release는 확인되지 않았다. 현재 공개 기준점은 아래와 같다.

IsoDDE · technical report · 2026-02-10OpenDDE Preview · 2026-07-03Pearl · public model announcement · 2025-10-28SeedFold · arXiv · 2025-12-30BioEmu · Science · 2025

IsoDDE는 공식적으로 protein–ligand generalisation, pocket identification, binding affinity, antibody–antigen modelling의 통합 predictive core를 강조한다. OpenDDE는 공개 checkpoint와 코드를 제공하는 all-atom cofolding foundation model로 대조군을 만든다. Pearl은 physics-generated synthetic data와 scaling을 핵심 전략으로 내세우고, SeedFold는 Pairformer width scaling·linear triangular attention·large-scale distillation로 model/data scale의 경로를 제시한다. BioEmu는 equilibrium ensemble generation이라는 별도의 축을 확립했다.

공식/주요 자료: IsoDDE · OpenDDE · Pearl · SeedFold · BioEmu

§13 · Bottom line

AI 신약설계의 다음 경쟁은 ‘생성’보다 ‘판단’이다

IsoDDE는 구조예측을 drug-design engine으로 확장하는 기준점을 만들었다. 하지만 2026년 8월의 문헌은 그 다음 병목을 선명하게 만든다. 모델이 더 많은 후보를 생성할수록, confidence calibration과 state selection, physics-based checking, cross-model consensus, prospective experimental feedback의 가치가 오히려 커진다.

다음 세대의 Drug Design Engine은 “정답을 한 번에 만들어내는 모델”보다 “후보를 넓게 만들고, 모르는 것을 구분하고, 다음 실험을 가장 잘 고르는 시스템”에 가까울 가능성이 높다.여러 최신 연구의 종합적 해석이며, 단일 논문이 직접 입증한 결론은 아니다.

References

Primary studies · August 2026
01
A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability
Nature Biotechnology · 19 Aug 2026

511 antibodies from 29 organizations; prospective, blinded experimental benchmark across affinity and developability.

https://www.nature.com/articles/s41587-026-03238-6
02
To ML-Predict or Not to ML-Predict: The Impact of Machine Learning-Predicted Protein Structures on FEP Accuracy and Data Augmentation
Journal of Chemical Information and Modeling · 20 Aug 2026

133 cSrc ligands and more than 1,400 FEP calculations; conformational-state choice remains critical.

https://pubs.acs.org/doi/10.1021/acs.jcim.6c01024
03
Biasing Conformational Sampling in AlphaFold 3 and Boltz-2 via Pair Representation Scaling
Journal of Chemical Information and Modeling · 20 Aug 2026

Inference-time pair-representation scaling across 86 two-state targets.

https://pubs.acs.org/doi/10.1021/acs.jcim.6c02094
04
Benchmarking confidence estimation and rescoring for cyclic peptide-protein complex predictions
bioRxiv preprint · 21 Aug 2026

111 complexes, 22,200 poses; ranking rather than generation emerges as the dominant failure mode.

https://doi.org/10.64898/2026.08.20.746104
05
Cryptic binding sites are detected but not ranked: coverage, conversion, and detector consensus
bioRxiv preprint · 17 Aug 2026

Separates candidate coverage from top-k conversion on 178 CryptoBench structures.

https://www.biorxiv.org/content/10.64898/2026.08.11.743381v1
06
Systematic Benchmarking of AI-Based Molecular Generation Models for Structure-Based Drug Design
bioRxiv preprint · 17 Aug 2026

12 generative/optimization methods over 176 protein–ligand systems; no single architecture dominates all criteria.

https://www.biorxiv.org/content/10.64898/2026.08.14.744939v1
07
Benchmarking Docking Protocols for GPCR Allosteric Modulators
bioRxiv preprint · Aug 2026

GaMD ensembles and conventional docking remain important for GPCR allosteric screening; Boltz-2 is complementary.

https://doi.org/10.64898/2026.08.12.744492
08
A multi-agent molecular optimization framework leads to a rapid-recovery intravenous anesthetic candidate with an improved safety margin
bioRxiv preprint · 20 Aug 2026

MASCOT links role-specialized agents, graph editing, medicinal chemistry objectives, and experimental pharmacology.

https://doi.org/10.64898/2026.08.17.745149
Model baselines · context
09
The Isomorphic Labs Drug Design Engine unlocks a new frontier beyond AlphaFold
Isomorphic Labs · 10 Feb 2026

Official IsoDDE overview covering protein–ligand generalisation, pocket identification, affinity prediction, and biologics modelling.

https://www.isomorphiclabs.com/articles/the-isomorphic-labs-drug-design-engine-unlocks-a-new-frontier
10
OpenDDE Preview
Aureka Research · Jul 2026

Open-source all-atom biomolecular foundation model and public cofolding checkpoint.

https://github.com/aurekaresearch/OpenDDE
11
Pearl: A Foundation Model for Placing Every Atom in the Right Location
Genesis · 2025

Physics-generated synthetic data, equivariant diffusion, and scaling as a route to protein–ligand generalisation.

https://arxiv.org/abs/2510.24670
12
SeedFold: Scaling Biomolecular Structure Prediction
arXiv · 30 Dec 2025

Pairformer width scaling, linear triangular attention, and large-scale distillation.

https://arxiv.org/abs/2512.24354
13
Scalable emulation of protein equilibrium ensembles with generative deep learning
Science · 2025 · BioEmu

Protein equilibrium ensemble generation as a distinct foundation-model objective.

https://doi.org/10.1126/science.adv9817