01 · Evidence boundary
먼저 세워야 할 전제: 공개된 것은 Drug Design Engine의 전부가 아니라 predictive core의 일부다
IsoDDE를 가장 직접적으로 설명하는 공개 문헌은 Isomorphic Labs가 2026년 2월 10일 공개한 기술보고서 Accurate Predictions of Novel Biomolecular Interactions with IsoDDE다. 이 보고서는 IsoDDE를 단일 구조예측 모델이 아니라 structure prediction, pocket identification, binding affinity prediction을 포함하는 통합 Drug Design Engine으로 규정한다.
그러나 공개 범위에는 중요한 한계가 있다. 상세 architecture, 전체 training recipe, data mixture, post-training 방식, weights, code가 공개되어 있지 않다. 2026년 7월 OpenDDE 논문은 바로 이 비공개성을 독립 검증과 재현성의 핵심 제약으로 지적한다. 따라서 이 글은 “IsoDDE 내부 구현을 복원한다”는 글이 아니다.
대신 네 종류의 증거를 연결한다. ① IsoDDE 공식 보고서, ② IsoDDE가 직접 활용하는 2025–2026 benchmark 연구, ③ Boltz-2·Pearl·SeedFold·BioEmu처럼 같은 문제를 푸는 동시대 연구, ④ IsoDDE에 대한 직접적 후속 응답인 OpenDDE다.
02 · Definition
IsoDDE란 무엇인가: “분자 카메라”에서 “분자 연구실”로
IsoDDE(Isomorphic Labs Drug Design Engine)는 새로운 생물학적 표적과 새로운 화학공간에서도 분자 간 상호작용의 구조, 결합 위치, 결합 강도를 높은 정확도로 예측하고, 이를 생성적 분자설계로 연결하려는 통합 계산 신약설계 시스템이다.
공개된 predictive core는 크게 세 축으로 보인다. Structure Prediction은 원자들이 어디에 위치하는지를, Pocket Identification은 ligand가 결합할 만한 위치가 어디인지를, Binding Affinity는 실제로 얼마나 강하게 결합하는지를 다룬다. Isomorphic Labs는 이 predictive fidelity를 방대한 molecular design space를 탐색하는 generative capability와 통합한다고 설명하지만, 공개 technical report는 전체 생성 시스템보다 predictive core에 초점을 맞춘다.
Structure Prediction
복합체의 all-atom 3차원 구조를 예측한다. “원자들이 어디에 있는가?”라는 질문이다.
Pocket Identification
ligand-binding 가능성이 높은 residue와 pocket 위치를 찾는다. “어디에 붙을 수 있는가?”라는 질문이다.
Binding Affinity
구조가 존재한다는 사실을 넘어 interaction strength를 추정한다. “얼마나 강하게 붙는가?”라는 질문이다.
Generative Design
원하는 interaction을 만들기 위해 새로운 분자를 탐색한다. 공개 보고서가 세부 구현을 공개한 영역은 아니다.
03 · Problem definition
구조, pocket, affinity, confidence를 하나의 decision-support problem으로 묶는다
표적 생체분자를 \(T\), ligand 또는 biomolecular partner를 \(L\), 복합체의 3차원 좌표를 \(X\), binding pocket을 \(P\), affinity를 \(A\), confidence를 \(C\)라고 하자. IsoDDE의 predictive problem은 개념적으로 다음처럼 쓸 수 있다.
구조예측은 다음 조건부 분포를 근사하는 문제다.
pocket identification은 residue 수준의 binding probability로 볼 수 있다.
affinity prediction은 예측 구조와 molecular identity를 받아 interaction strength를 추정한다.
여기서 \(A\)는 \(K_D\), \(\Delta G\), \(pEC_{50}\) 등과 연결될 수 있다. 그리고 여러 생성 결과 중 어느 구조를 믿을지 결정하는 confidence estimation이 별도 하위 문제로 들어간다.
Runs N’ Poses 평가에서 IsoDDE는 25개의 sample을 생성하고, confidence가 가장 높은 결과를 선택하는 방식으로 평가되었다. 이 때문에 inference-time sampling과 confidence ranking은 부가 기능이 아니라 실제 benchmark score를 결정하는 구성 요소다.
진짜 질문은 accuracy보다 generalisation이다
2025년에 제안된 Runs N’ Poses benchmark를 2026년 Nature Structural & Molecular Biology 논문이 분석한 결과, training cutoff 이후 공개된 2,600개 protein–ligand system에서 기존 cofolding 모델들의 ligand pose 성능이 상당 부분 memorization에 의존하는 현상을 보고한다. 신약개발에서 가치 있는 문제는 바로 아직 보지 못한 target, pocket, scaffold, ligand이므로 이 현상은 근본적이다.
04 · Core concepts
IsoDDE를 읽을 때 놓치면 안 되는 다섯 개념
4.1 All-atom cofolding
protein과 ligand를 독립적으로 만든 뒤 docking하는 대신 상호작용하며 형성되는 구조를 함께 예측한다. ligand가 들어오며 protein conformation이 바뀌는 induced fit이나 apo 상태에서는 닫혀 있던 pocket이 열리는 상황에서 특히 중요하다.
4.2 Generalisation over memorisation
training set과 유사한 pocket·ligand에서 잘 맞히는 것보다 similarity가 낮아졌을 때 성능을 유지하는지가 핵심이다.
IsoDDE 보고서가 제시한 Runs N’ Poses의 가장 어려운 0–20 similarity bin에서는 IsoDDE 성공률이 50%, AF3는 약 23% 수준으로 보고된다. 보고서가 “more than doubles AF3”라고 표현하는 근거다.
4.3 Induced fit와 cryptic pocket
protein은 고정된 자물쇠가 아니다. ligand 접근에 따라 side chain, helix, loop, domain이 움직이고 apo structure에 없는 pocket이 열릴 수 있다. IsoDDE가 강조하는 8EA6, 8E23 사례는 이런 OOD induced-fit·cryptic-pocket event를 겨냥한다.
4.4 Structure와 affinity는 별개의 문제다
정확한 pose를 맞혔다고 potency ranking도 자동으로 정확해지는 것은 아니다. IsoDDE는 affinity를 독립 capability로 평가한다. 공개 benchmark에서 평균 Pearson \(r\)은 FEP+4 계열에서 IsoDDE 0.85 대 FEP+ 0.78, OpenFE benchmark에서 0.73 대 0.72로 보고된다.
4.5 Test-time sampling과 confidence
앞으로의 경쟁은 한 번의 forward pass 정확도뿐 아니라 추가 test-time compute를 얼마나 효과적으로 정확도로 바꾸는지의 문제다. OpenDDE 역시 sampling budget과 confidence를 독립적인 연구축으로 다룬다.
05 · Introduction
왜 지금 IsoDDE인가: “Can we predict structure?” 다음의 질문
AlphaFold 2와 AlphaFold 3 이후 생명과학 AI는 3차원 구조 자체를 machine-learning object로 다룰 수 있다는 사실을 확인했다. 그러나 drug discovery에서 필요한 질문은 구조 하나에서 끝나지 않는다.
어떻게 생겼는가?
붙는가?
어떻게 변하는가?
붙는가?
설계해야 하는가?
FoldBench는 1,522개의 low-homology biological assembly를 평가하면서 antibody–antigen, allosteric protein–ligand, nucleic-acid system이 여전히 어렵다고 보여준다. protein–ligand 성능의 제한 인자 중 하나가 protein folding 자체보다 ligand similarity일 수 있다는 결과도 제시한다.
2025–2026년 연구의 중심은 이 질문으로 이동한다. IsoDDE는 이 전환점에 놓여 있다.
06 · Motivation
신약개발은 본질적으로 OOD이며, AI와 physics 사이의 계산 간극을 줄여야 한다
첫째, 신약개발은 OOD 문제다
기존 데이터와 거의 같은 ligand만 찾는다면 새로운 medicine을 설계하는 의미가 제한된다. Runs N’ Poses가 pose memorization을 정면으로 문제 삼은 이유다.
둘째, experimental structure만으로 탐색 공간을 감당하기 어렵다
새로운 protein state, cryptic pocket, antibody interface마다 구조를 실험적으로 규명한 뒤 candidate를 탐색하면 공간이 너무 크다. IsoDDE는 experimental-grade에 가까운 predictive fidelity를 계산적으로 확보하는 것을 generative design의 선결조건으로 둔다.
셋째, physics-based accuracy와 AI scalability 사이에 간극이 있다
FEP 같은 free-energy method는 중요한 기준이지만 계산비용과 system-specific preparation이 크다. 반대로 fast ML은 빠르지만 OOD generalisation과 affinity reliability가 충분하지 않을 수 있다. IsoDDE가 affinity 성능을 FEP 계열과 직접 비교하는 이유다.
넷째, prediction과 design을 하나의 representation 안에서 묶으려는 흐름이 있다
OpenDDE는 structure prediction과 de novo design을 conditional diffusion formulation으로 통합하려 하고, Pearl은 synthetic data·SO(3)-equivariant diffusion·controllable inference를 결합한다. 이 분야는 folding benchmark를 넘어 foundation model for molecular design 문제로 이동한다.
07 · Challenges
핵심 난제는 일곱 가지다
① 데이터 누출과 memorisation
PDB 기반 학습에서는 sequence, pocket geometry, ligand scaffold가 train/test 사이에 미묘하게 겹칠 수 있다. 그래서 단순 time split만으로는 충분하지 않을 수 있다. FoldBench는 low-homology filtering을 하고 Runs N’ Poses는 protein pocket similarity와 ligand shape similarity를 동시에 본다.
② Protein은 정적 구조가 아니다
drug binding은 induced fit, loop movement, side-chain rearrangement, cryptic pocket opening을 포함한다. 실제 대상은 단일 structure가 아니라 conformational ensemble일 수 있다.
BioEmu가 equilibrium ensemble generation을 별도의 foundation-model 문제로 다루는 이유가 여기에 있다.
③ Affinity는 geometry보다 어렵다
electrostatics, solvent, entropy, protonation, conformational rearrangement는 단순 좌표 정확도로 환원되지 않는다. pose prediction 90%와 medicinal-chemistry potency ranking 정확도는 서로 다른 주장이다.
④ Antibody–antigen은 매우 어려운 OOD interface다
CDR loop, 특히 CDR-H3의 구조적 다양성과 epitope 다양성이 크다. IsoDDE의 별도 334-complex held-out evaluation에서 high-fidelity \(\mathrm{DockQ}>0.8\) 예측 비율은 39%로 보고된다. AF3와 Boltz-2 대비 개선이지만, 동시에 61%가 high-fidelity 기준을 넘지 못했다는 뜻이기도 하다.
⑤ Prediction confidence가 실제 uncertainty인가?
높은 confidence가 familiarity score에 불과한지, 실제 오류확률과 calibration되는지는 별개의 문제다. sampling-based selection이 강해질수록 uncertainty quality가 더 중요해진다.
⑥ 물리적으로 그럴듯함과 화학적으로 맞음
낮은 RMSD라도 stereochemistry, clash, bond geometry, local packing이 잘못될 수 있다. OpenDDE가 shape-complementarity loss, clash-aware contact quality, local-frame/torsion supervision을 도입한 배경이다.
⑦ 재현성
현재 IsoDDE의 가장 큰 학술적 제약이다. OpenDDE 연구진은 full training recipe, data mixture, inference procedure, post-training strategy, engineering optimization을 알 수 없어 observed gain을 architecture, scale, data, inference 중 어디에 귀속할지 판단하기 어렵다고 명시한다.
08 · Research questions
IsoDDE 이후의 연구를 여덟 개 질문으로 재구성하면
protein, pocket, ligand가 모두 training distribution에서 멀어져도 정확한 interaction geometry를 예측할 수 있는가?
training structure를 재조합하는가, 아니면 transferable molecular interaction rules를 학습하는가?
AI가 FEP 수준의 affinity ranking을 훨씬 낮은 계산비용으로 안정적으로 제공할 수 있는가?
ligand 정보 없이 sequence 또는 apo representation만으로 unknown/cryptic druggable site를 발견할 수 있는가?
antibody–antigen처럼 매우 variable한 interface에서도 reliable atomic modelling이 가능한가?
confidence가 실제 오류확률을 설명하며, sampling budget을 늘리면 성능이 체계적으로 개선되는가?
“이 구조가 무엇인가?”와 “어떤 구조를 만들어야 하는가?”를 동일 foundation model에서 풀 수 있는가?
benchmark 향상이 prospective drug-discovery campaign에서도 재현되며 외부 연구자가 독립 검증할 수 있는가?
09 · Methods A
IsoDDE에서 공개적으로 확인되는 방법론
1. All-atom biomolecular structure modelling
protein–ligand, protein–protein, antibody–antigen 등 여러 interface를 통합적으로 취급한다. 공식 보고서는 AF3 계열의 pair representation, triangular operations, generative diffusion이라는 연구 계보를 설명하지만 IsoDDE 자체의 상세 architecture는 공개하지 않는다.
2. OOD-centred benchmarking
Runs N’ Poses similarity bins, FoldBench low-homology split, training-cutoff control을 통해 단순 benchmark accuracy보다 generalisation을 앞세운다.
3. Multiple-sample generation + confidence ranking
Runs N’ Poses에서는 25 sample 중 confidence가 가장 높은 결과를 채택한다. antibody 분석에서는 sampling을 크게 늘렸을 때 성능이 개선되는 현상도 살핀다.
4. Independent binding-affinity capability
structure confidence의 proxy가 아니라 experimentally measured affinity와 correlation을 직접 평가한다. FEP+, OpenFE, CASP16 benchmark가 포함된다.
5. Residue-level pocket probability
protein surface의 residue별 ligand-binding probability를 추정해 알려진 site뿐 아니라 새로운/cryptic pocket discovery로 문제를 확장한다.
10 · Methods B
IsoDDE 주변 2025–2026 방법론 프런티어
Pearl (2025)
대규모 synthetic data, SO(3)-equivariant diffusion, controllable inference, generalized templating을 결합한다. 데이터 부족과 rotational symmetry를 동시에 공략한다.
SeedFold (2025)
Pairformer width scaling, linear triangular attention, 대규모 distillation을 결합한다. 성능 향상을 model scale × efficient attention × data scale의 곱으로 보는 시각이 강하다.
OpenDDE (2026)
atomic-level reasoning, shape-complementarity objective, local all-atom refinement, unified prediction/design conditional diffusion을 제안한다. surface orientation, spacing, clash까지 loss에 반영하며 coordinate regression에서 interaction geometry reasoning으로 이동한다.
OpenDDE scaling 분석에서는 training compute가 늘수록 antibody–antigen 성능이 좋아지는 경향이 있지만 이득은 점차 감소한다. 연구진 스스로 scale만으로 설명되지 않으며 architecture, data quality, training strategy가 중요하다고 해석한다.
11 · Applications
IsoDDE 계열 시스템이 특히 유용할 수 있는 다섯 영역
1. First-in-class target drug discovery
known ligand와 구조가 거의 없는 신규 target일수록 claimed generalisation의 가치가 커진다. 공식 보고서도 first-in-class target과 새로운 modulatory mechanism을 중요한 목표로 든다.
2. Cryptic/allosteric pocket discovery
2026년 Nature에는 Cereblon(CRBN)의 새로운 allosteric site가 실험적으로 보고되었다. IsoDDE 보고서는 이를 retrospective test로 사용해 ligand identity 없이도 기존 pocket과 새로운 cryptic site에 pocket signal을 만들었다고 보고한다. ligand를 준 cofold에서는 lenalidomide와 SB-405483 pose에 각각 약 0.12 Å, 0.33 Å RMSD를 보고한다.
3. Hit-to-lead / lead optimisation
candidate potency ranking을 빠르게 수행할 수 있다면 medicinal chemist의 synthesis 우선순위 결정에 직접 연결될 수 있다. affinity benchmark는 바로 이 decision problem을 겨냥한다.
4. Antibody and biologics design
antibody–antigen interface와 CDR-H3 modelling이 좋아지면 antibody optimization, epitope targeting, 향후 generative antibody design으로 이어질 수 있다.
5. 새로운 mechanism-of-action 탐색
cryptic/allosteric site가 예측 가능해지면 orthosteric inhibitor뿐 아니라 molecular glue, PPI modulation 등 새로운 modulation 방식으로 탐색 공간이 확장될 수 있다.
12 · Open problems
Benchmark victory와 drug-discovery victory는 같지 않다
12.1 Prospective validation
retrospective benchmark는 정답이 이미 존재한다. 실제 연구에서는 prediction 후 합성·실험을 하고 나서야 정답을 안다.
따라서 prospective blind loop가 궁극적 기준이 되어야 한다.
12.2 Conformational ensemble
protein–ligand interaction은 단일 structure가 아니라 energy landscape 문제다.
metastable state, cryptic state, transition, induced-fit ensemble까지 모델링해야 한다. OpenDDE도 conformational ensemble modelling을 future extension으로 둔다.
12.3 Affinity에서 kinetics로
다음 질문은 “얼마나 세게 붙는가?”에서 “어떤 경로로 붙고, 얼마나 오래 머물며, 어떤 conformational state를 선택하는가?”로 갈 가능성이 크다.
12.4 Selectivity
한 target에 강하게 붙는 것만으로 약이 되지 않는다.
binding prediction은 selectivity landscape prediction으로 확장되어야 한다.
12.5 Synthetic accessibility와 ADMET
IsoDDE 보고서가 공개적으로 보여주는 핵심은 structure, pocket, affinity다. 실제 molecule design은 efficacy, selectivity, toxicity, solubility, permeability, metabolic stability, synthesizability를 동시에 만족해야 한다. predictive core를 곧바로 완전한 autonomous drug designer와 동일시하면 안 된다.
12.6 Calibration과 abstention
실전에서는 “모르겠을 때 모른다고 말하는 능력”이 중요하다. OOD uncertainty가 높을 때 실험 검증을 우선하게 만드는 calibrated uncertainty, selective prediction, abstention이 핵심 연구문제다.
12.7 재현성과 scientific auditability
현재 가장 명백한 open problem이다. full recipe가 공개되지 않았기 때문에 성능 gain을 data scale, architecture, post-training, inference-time sampling으로 독립 분해하기 어렵다.
13 · Future directions
IsoDDE 이후 Drug Design Engine은 어디로 가는가
방향 1. Molecular Camera → Molecular Simulator
가장 그럴듯한 구조 한 장이 아니라 conformational ensemble과 energy landscape를 생성해야 한다.
사진을 찍는 모델에서 분자의 움직임을 시뮬레이션하는 모델로의 전환이다.
방향 2. Molecular Simulator → Molecular Engineer
forward prediction을 inverse design으로 뒤집는다.
“이 ligand가 붙는가?”에서 “이 target의 이 pocket에 원하는 affinity와 selectivity를 가지도록 어떤 molecule을 만들어야 하는가?”로 질문이 바뀐다. OpenDDE의 unified conditional diffusion은 이 방향의 초기 형태다.
방향 3. 단일 optimum → Pareto drug design
실제 medicine은 하나의 objective로 최적화할 수 없다. Drug Design Engine은 결국 상충 제약 사이 Pareto frontier를 탐색하는 scientific decision engine으로 가야 한다.
방향 4. Offline AI → Design–Make–Test–Learn
실험실은 validation facility가 아니라 AI가 다음 학습 데이터를 능동적으로 획득하는 sensor가 된다. OpenDDE도 molecular design, affinity, active learning, experimental feedback를 mature DDE의 확장 영역으로 둔다.
방향 5. Bigger models → Better scientific scaling
SeedFold와 OpenDDE는 scale이 유효할 수 있음을 보여주지만 gain은 sublinear하며 scale만으로 모든 차이가 설명되지 않는다.
차세대 경쟁은 parameter count 하나보다 이 다섯 축의 조합이 될 가능성이 크다.
14 · Big picture
핵심적으로 어떻게 이해하면 좋은가
IsoDDE는 “단백질 구조를 예측하는 AI”라는 패러다임을 “새로운 생물학적 공간에서 구조·pocket·affinity를 함께 추론하여 실제 분자설계 의사결정을 지원하는 AI”로 확장하려는 시도다.
Prediction
Prediction
Simulation
Design
Drug Discovery
이 관점에서 IsoDDE의 가장 중요한 혁신 후보는 특정 benchmark 숫자 하나가 아니다. 정확한 structure-prediction model과 drug-design engine 사이의 경계를 허물려 한다는 점이다.
동시에 가장 큰 미해결 질문도 분명하다.
향후 2–3년간 IsoDDE뿐 아니라 Boltz, OpenDDE, Pearl, Protenix, SeedFold 계열을 평가할 때 이 질문이 가장 중요한 과학적 기준이 될 가능성이 높다. 이는 2025–2026 문헌을 종합한 연구 전망이다.
15 · Sources
2025년 이후 핵심 문헌과 직접 출처
아래 목록은 첨부 연구 메모가 사용한 주요 2025–2026 공개 자료다. IsoDDE 자체의 직접 자료와 benchmark, 경쟁/후속 시스템, 관련 생물물리 연구를 구분해 읽는 것이 중요하다.
- Isomorphic Labs Team. Accurate Predictions of Novel Biomolecular Interactions with IsoDDE (2026). Technical report · Official introduction.
- Škrinjar et al. Evaluating generalization in protein–ligand cofolding methods. Nature Structural & Molecular Biology, 2026 (preprint 2025). Nature.
- Xu et al. Benchmarking all-atom biomolecular structure prediction with FoldBench. Nature Communications, 2025/2026. Nature Communications.
- Passaro et al. Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction (2025). bioRxiv.
- Hitawala & Gray. What does AlphaFold3 learn about antibody and nanobody docking, and what remains unsolved? (2025). DOI.
- Gilson et al. Assessment of pharmaceutical protein–ligand pose and affinity predictions in CASP16 (2025). DOI.
- Genesis Research. Pearl: A Foundation Model for Placing Every Atom in the Right Location (2025). arXiv.
- ByteDance Seed. SeedFold: Scaling Biomolecular Structure Prediction (2025). arXiv.
- Corley et al. Accelerating Biomolecular Modeling with AtomWorks and RF3 (2025). bioRxiv.
- Lewis et al. Scalable emulation of protein equilibrium ensembles with generative deep learning. Science, 2025 — BioEmu. DOI · Preprint.
- Dippon et al. Identification of an allosteric site on the E3 ligase adapter cereblon. Nature, 2026. Nature.
- Aureka AI. Folding, Reasoning, and Scaling with Open-source Drug Discovery Engine (OpenDDE) (2026). arXiv · HTML.
- Stark et al. BoltzGen: Toward Universal Binder Design (2025). bioRxiv.
Next-step note from the source: 후속 확장으로 IsoDDE–AlphaFold3–Boltz-2–Pearl–SeedFold–OpenDDE의 Architecture/Training/Data/Benchmark/Strength/Weakness 비교, 2025–2026 관련 20–30편 annotated bibliography, 이를 바탕으로 한 “Post-IsoDDE Drug Design Engine”의 Research Gap·RQ·Methodology 설계가 가능하다.