후보를 제안할 수 있는가
실제 cryptic site와 겹치는 proposal이 candidate pool 안에 존재하는 비율이다.
IsoDDE 이후의 경쟁을 단순한 구조예측 정확도 싸움으로 보면 중요한 변화를 놓친다. 2026년 8월의 최신 연구들은 서로 다른 과제에서 같은 경고를 보낸다. 좋은 구조와 분자를 후보 집합 안에 넣는 능력은 빠르게 향상되었지만, 그중 실제로 맞는 후보를 고르고, 올바른 단백질 상태를 선택하고, 다음 실험으로 보낼 결정을 내리는 능력은 아직 그 속도를 따라가지 못한다.
이번 추적에서 가장 강한 신호는 “더 큰 cofolding model” 자체가 아니다. 항체, cyclic peptide, cryptic pocket, GPCR allostery, FEP, molecular generation을 가로질러 반복되는 패턴은 하나다. 생성 능력은 선택 능력보다 앞서 있다.
각 연구는 서로 다른 modality와 benchmark를 다루지만, 실패 지점은 놀라울 만큼 비슷하다.
| 연구 | 평가 대상 | 핵심 관찰 | 드러난 병목 |
|---|---|---|---|
| AIntibody | 511 antibodies · 29 organizations | 일부 <100 pM 항체 성공, 그러나 HCDR3 cluster ranking은 대부분 random보다 나쁨 | Prospective ranking |
| AF3/Boltz-2 + FEP | cSrc · 133 ligands · >1,400 FEP runs | 구조의 출처보다 올바른 conformational state 선택이 중요 | State selection |
| Pair Representation Scaling | 86 two-state proteins | 재학습 없이 latent pair scaling으로 alternative state recovery 확장 | State coverage / test-time control |
| Cyclic peptide cofolding | 111 complexes · 22,200 poses | 좋은 pose는 자주 생성하지만 native confidence가 best pose를 충분히 못 고름 | Pose ranking |
| Cryptic pocket benchmark | CryptoBench · 178 structures | 후보는 이미 많이 제안되지만 top-k에 올리는 능력이 약함 | Pocket ranking |
| 12-model generative benchmark | 176 protein–ligand systems | 모든 기준을 지배하는 단일 generator가 없음 | Stage-specific routing |
| GPCR allosteric screening | 4 Class-A GPCRs | GaMD ensemble + conventional docking이 Boltz-2 단독보다 실용적 | Dynamics-aware hybridization |
| MASCOT | Multi-objective lead optimization | 역할특화 agents가 chemical edit와 실험평가를 연결 | Decision orchestration |
Analysis이 결과들을 한데 놓으면 Post-IsoDDE 연구의 중심축은 prediction accuracy만이 아니라 candidate selection fidelity로 이동한다. 구조·분자·pocket을 생성하는 foundation model이 강해질수록, 그 위에서 uncertainty를 보정하고 후보를 재랭킹하고 상태를 검증하는 층이 전체 성능을 지배하게 된다.
CASP형 blind benchmark가 antibody discovery의 실제 의사결정으로 이동했다.
Source fact2026년 8월 19일 Nature Biotechnology에 출판된 AIntibody benchmark는 29개 기관이 제출한 511개 AI-designed 또는 AI-predicted antibody를 대상으로 affinity maturation, HCDR3 cluster 내 affinity ranking, out-of-library CDR design의 세 과제를 prospective·blinded 방식으로 평가했다. 설계된 sequence는 full IgG로 합성되고, SPR·KinExA 및 developability panel을 통해 동일 조건에서 측정되었다.
일부 팀은 developability를 유지하면서 100 pM 이하 affinity를 달성했다. 그러나 성공은 과제 전반으로 일반화되지 않았다. 특히 HCDR3 cluster에서 고-affinity clone을 고르는 문제는 한 모델을 제외하고 random clone picking보다 나빴다. 논문은 첫 challenge가 매우 잘 연구된 SARS-CoV-2 RBD와 풍부한 sequencing/affinity 정보를 제공했다는 점에서 현재 역량의 비교적 유리한 상한 조건으로 해석해야 한다고 명시한다.
AnalysisIsoDDE가 antibody–antigen interface와 CDR-H3 modelling에서 강한 구조예측 결과를 제시했다면, AIntibody는 그 다음 질문을 던진다. “그 구조 정보를 사용해 어떤 sequence를 실제 실험으로 보낼 것인가?” 앞으로 antibody benchmark는 DockQ와 RMSD만이 아니라 prospective affinity ranking, developability, calibration, novelty distance를 함께 봐야 한다.
원문: Nature Biotechnology — A blinded, prospective benchmark of in silico antibody discovery...
AF3와 Boltz-2는 단일 구조예측기를 넘어 conformational state를 탐색하는 도구로 재해석되고 있다.
Source fact8월 20일 Journal of Chemical Information and Modeling에 출판된 연구는 cSrc의 133개 congeneric ligand에 대해 experimental crystal structure, homology model, ML-predicted structure를 FEP 입력으로 비교했고, 1,400회가 넘는 계산을 수행했다. 핵심 메시지는 구조의 출처보다 micro/macro conformational state가 올바른가가 reliability에 큰 영향을 준다는 점이다.
Analysis이는 “AF3/Boltz-2 pose → FEP”를 단순 직렬 연결하는 pipeline에 경고를 준다. Drug Design Engine에는 구조 생성기와 affinity 계산 사이에 state validation이 필요하다.
Source fact같은 날 정식 출판된 Pair Representation Scaling은 Pairformer 앞의 latent pair representation에 \(z'=(1+\beta)z\)를 적용한다. auxiliary model도, retraining도, 별도의 두 번째 forward pass도 필요하지 않는다. 86개 two-state target에서 AF3와 Boltz-2의 sampled ensemble을 넓혀 default inference가 놓친 alternative state recovery를 개선했으며, 효과는 AF3에서 특히 강했다.
AnalysisBioEmu처럼 ensemble 자체를 목적으로 학습한 생성모델과, 기존 cofolding model의 latent distribution을 test-time에 steering하는 접근이 새로운 비교축을 만든다. 향후 benchmark는 top-1 RMSD뿐 아니라 state coverage per unit compute를 재야 한다.
원문: JCIM — Biasing Conformational Sampling in AlphaFold 3 and Boltz-2...
Cyclic peptide와 cryptic pocket이라는 서로 다른 과제가 같은 실패 모드를 드러낸다.
Source fact8월 21일 공개된 benchmark는 다섯 종류 cyclization chemistry를 포함한 111개 nonredundant complex에 대해 Boltz와 Protenix가 target당 100개의 pose를 생성하도록 하여 총 22,200개 pose를 평가했다. median top-pose DockQ는 약 0.89이고 96–98% target에서 Acceptable 이상의 구조를 생성했다. 그러나 native ranking score와 실제 pose quality의 Spearman 상관은 약 0.53–0.66이었고, 약 12% pose는 confidence가 높지만 구조 품질이 낮았다. 최고 품질 pose가 rank 1이 아닌 경우도 거의 모든 target에서 관찰되었다.
연구진은 interface hydrogen-bond density 같은 외부 구조 feature를 native score에 더한 gradient-boosted rescoring으로 ranking을 개선했다.
원문 DOI: 10.64898/2026.08.20.746104
Source factCryptoBench test fold 178개 structure에서 fpocket은 qualifying candidate를 74.2% target에 제안했지만 top-5 recovery는 43.8%였다. P2Rank는 coverage 66.3%, top-5 recovery 63.5%, IF-SitePred는 70.8%와 61.8%였다. 네 detector의 candidate union은 92.1% coverage에 도달했다. 즉 candidate가 전혀 없는 target보다 candidate는 존재하지만 위로 올라오지 않는 target이 훨씬 많다.
실제 cryptic site와 겹치는 proposal이 candidate pool 안에 존재하는 비율이다.
이미 생성한 올바른 후보를 top-k budget 안에 surface하는 능력이다.
AnalysisIsoDDE의 ligand-free pocket identification을 평가할 때도 AUPRC 하나만으로는 부족하다. coverage, ranking conversion, candidate budget, ensemble cost를 분리해서 보는 편이 더 실전적이다.
원문: bioRxiv — Cryptic binding sites are detected but not ranked
Generative model과 physics, dynamics, synthesis-aware scoring을 단계별로 routing하는 방향이 더 강한 실험적 근거를 얻고 있다.
Source fact8월 17일 공개된 systematic benchmark는 12개 molecular generation/optimization method를 176개 curated protein–ligand system에서 비교했다. chemical validity, uniqueness, molecular/scaffold diversity, QED, synthetic accessibility, docking, physicochemical/ADMET property, computational resource requirement까지 함께 평가했다.
결과는 architecture별 trade-off를 보여준다. pocket-conditioned 3D model은 receptor geometry 활용에, flow-based model은 sampling efficiency에, reference-conditioned method는 analogue generation에, synthesis-aware method는 chemical feasibility에 강했다. 하지만 모든 기준에서 일관되게 우수한 단일 architecture는 없었다.
연구진은 MD-derived receptor ensemble, ensemble docking, protein–ligand interaction graph를 통합한 State-Aware Functional Classifier(SAFC)를 추가해 docking·drug-likeness·synthetic accessibility와 일부 상보적인 ranking signal을 제시했다.
Analysis‘Unified Drug Design Engine’은 반드시 monolithic neural network일 필요가 없다. 더 현실적인 통합은 서로 다른 강점을 가진 모델을 stage별로 선택하는 orchestration architecture일 수 있다.
원문: bioRxiv — Systematic Benchmarking of AI-Based Molecular Generation Models...
Source fact네 개 Class-A GPCR의 allosteric modulator screening을 비교한 연구에서는 PDB structure와 GaMD ensemble을 Glide HTVS, AutoDock Vina, DOCK3.8, Boltz-2에 적용했다. GaMD ensemble은 네 target 모두에서 적어도 한 program의 early enrichment를 개선했다. Glide ensemble docking만이 네 target에서 일관되게 개선되었고, M2 receptor에서는 PDB structure 대비 active recovery가 거의 9배 향상되었다.
Boltz-2는 GaMD template 변화에 상대적으로 둔감했고 이 workload에서는 conventional docking보다 낮은 screening 성능을 보였다. 저자들은 Boltz-2 affinity signal을 physics/empirical docking의 대체라기보다 complementary signal로 해석했다.
Inference특히 cryptic/allosteric target에서는 “foundation model 하나가 flexibility를 내부적으로 해결한다”는 가정보다 explicit dynamics와 learned model을 결합하는 편이 당분간 더 안전한 전략일 수 있다.
원문 DOI: 10.64898/2026.08.12.744492
생성·예측 모델의 출력물을 실제 multi-objective optimization과 실험 후보 선택으로 연결하는 문제다.
Source factMASCOT은 chemically constrained graph-editing search 위에 세 역할을 둔다. trade-off agent는 potency·PK·safety 등 경쟁 objective의 우선순위를 조절하고, strategy agent는 molecular edit의 제안 방식을 바꾸며, reflection agent는 이전 의사결정의 교훈을 다음 search에 반영한다.
SARS-CoV-2 Mpro task에서 mean docking-score improvement가 가장 강한 baseline의 3.6배였고, remimazolam optimization에서는 RM-1을 우선순위화한 뒤 후속 derivative로 RM-7을 도출했다. preprint는 animal study에서 RM-7이 더 높은 potency, 빠른 functional recovery, 넓은 safety margin, flumazenil reversibility를 보였다고 보고한다.
Analysis이 흐름은 IsoDDE/OpenDDE/Boltz류 foundation model을 ‘도구’로 보고, 그 위에 어떤 계산을 호출하고 어떤 후보를 다음 단계로 보낼지를 결정하는 agentic layer를 두는 설계를 현실적으로 만든다. 다만 현재 MASCOT은 preprint이므로 독립 재현과 전체 실험 프로토콜 검증이 필요하다.
원문 DOI: 10.64898/2026.08.17.745149
정답 후보를 생성하는 것보다, 그 정답을 식별하고 불확실하면 abstain하며 실험으로 보내는 능력이 핵심 연구대상이 된다.
이번 문헌을 종합하면 다음 구조가 자연스럽게 도출된다.
Protein–ligand pose, pocket, antibody, molecular candidate를 충분히 넓게 생성한다.
confidence calibration, structural rescoring, state validation으로 false certainty를 줄인다.
wet-lab outcome을 다음 ranking·generation cycle의 evidence로 되돌린다.
한 모델이 \(K\)개의 후보를 생성한다고 하자. 그중 실제 quality가 가장 높은 후보를 oracle이라고 하고, 모델 confidence/ranker가 선택한 후보를 selected라고 하면 다음 차이를 측정할 수 있다.
Analysis이 지표는 “좋은 답을 생성할 능력”과 “그 답을 알아볼 능력”을 분리한다. cyclic peptide에서는 이미 이 gap이 크게 드러났고, cryptic pocket에서도 coverage와 conversion의 분리가 같은 논리를 보여준다. antibody design에서는 oracle을 retrospective experimental best clone으로 정의해 prospective selection loss를 측정할 수 있다.
Research inference향후 IsoDDE, OpenDDE, Boltz, Pearl, SeedFold류 시스템을 비교할 때 top-1 accuracy와 함께 oracle-selected gap, calibration error, compute-normalized state coverage, abstention utility를 공동 지표로 두면 훨씬 강한 scientific benchmark가 된다.
2026년 8월 17–26일 범위에서 새 IsoDDE 본체 기술보고서나 새로운 OpenDDE/Pearl/SeedFold/BioEmu의 primary model release는 확인되지 않았다. 현재 공개 기준점은 아래와 같다.
IsoDDE는 공식적으로 protein–ligand generalisation, pocket identification, binding affinity, antibody–antigen modelling의 통합 predictive core를 강조한다. OpenDDE는 공개 checkpoint와 코드를 제공하는 all-atom cofolding foundation model로 대조군을 만든다. Pearl은 physics-generated synthetic data와 scaling을 핵심 전략으로 내세우고, SeedFold는 Pairformer width scaling·linear triangular attention·large-scale distillation로 model/data scale의 경로를 제시한다. BioEmu는 equilibrium ensemble generation이라는 별도의 축을 확립했다.
IsoDDE는 구조예측을 drug-design engine으로 확장하는 기준점을 만들었다. 하지만 2026년 8월의 문헌은 그 다음 병목을 선명하게 만든다. 모델이 더 많은 후보를 생성할수록, confidence calibration과 state selection, physics-based checking, cross-model consensus, prospective experimental feedback의 가치가 오히려 커진다.
511 antibodies from 29 organizations; prospective, blinded experimental benchmark across affinity and developability.
https://www.nature.com/articles/s41587-026-03238-6133 cSrc ligands and more than 1,400 FEP calculations; conformational-state choice remains critical.
https://pubs.acs.org/doi/10.1021/acs.jcim.6c01024Inference-time pair-representation scaling across 86 two-state targets.
https://pubs.acs.org/doi/10.1021/acs.jcim.6c02094111 complexes, 22,200 poses; ranking rather than generation emerges as the dominant failure mode.
https://doi.org/10.64898/2026.08.20.746104Separates candidate coverage from top-k conversion on 178 CryptoBench structures.
https://www.biorxiv.org/content/10.64898/2026.08.11.743381v112 generative/optimization methods over 176 protein–ligand systems; no single architecture dominates all criteria.
https://www.biorxiv.org/content/10.64898/2026.08.14.744939v1GaMD ensembles and conventional docking remain important for GPCR allosteric screening; Boltz-2 is complementary.
https://doi.org/10.64898/2026.08.12.744492MASCOT links role-specialized agents, graph editing, medicinal chemistry objectives, and experimental pharmacology.
https://doi.org/10.64898/2026.08.17.745149Official IsoDDE overview covering protein–ligand generalisation, pocket identification, affinity prediction, and biologics modelling.
https://www.isomorphiclabs.com/articles/the-isomorphic-labs-drug-design-engine-unlocks-a-new-frontierOpen-source all-atom biomolecular foundation model and public cofolding checkpoint.
https://github.com/aurekaresearch/OpenDDEPhysics-generated synthetic data, equivariant diffusion, and scaling as a route to protein–ligand generalisation.
https://arxiv.org/abs/2510.24670Pairformer width scaling, linear triangular attention, and large-scale distillation.
https://arxiv.org/abs/2512.24354Protein equilibrium ensemble generation as a distinct foundation-model objective.
https://doi.org/10.1126/science.adv9817