2026년 9월 10일 확인에서 추가할 만큼 의미 있는 변화는 하나다. 핵심은 새로운 LLM 멀티에이전트가 아니라, Virtual Cell이 정적 single-cell 표현에서 실제 약물 perturbation의 시간적 반응을 학습하고 환자·오가노이드 의사결정으로 전이되는 operational scientific foundation model로 이동했다는 점이다.
Nature에 출판된 ProteinTalks는 18개 breast-cancer cell line, 63개 FDA 승인 항암제, 59개 drug combination에서 수집한 3,800만 개 이상의 temporal protein-abundance measurement를 이용해 drug perturbation 이후 proteome의 시간적 변화를 학습한다. 같은 pretrained representation을 drug efficacy, synergy, resistance, patient stratification, organoid drug prioritization에 재사용한다. 여기서 중요한 변화는 “이 세포가 무엇인가?”보다 “이 세포에 이 약을 가했을 때 앞으로 어떻게 변하는가?”를 묻는다는 데 있다.
이 글은 첨부된 AI Co-Scientis-Trends-0910.md 전체를 1차 출처로 재구성한다. 첨부 자료는 9월 9–10일 기준으로 새로운 Drug-Discovery Multi-Agent/Agentic RAG, prospective autonomous wet-lab closed loop, 대형 신규 multimodal biomolecular foundation model의 동급 변화는 확인하지 못했다고 명시한다. 아래의 통합 Co-Scientist 구조는 ProteinTalks 자체의 기능이 아니라 자료가 제안한 후속 연구방향이다.
Training Systems
18breast-cancer cell line 수.
Perturbations
63 + 59FDA 승인 항암제 63개와 drug combination 59개.
Temporal Measurements
>38M여러 시간점에 걸친 protein-abundance measurement.
Unseen Drugs
81 / ~88%training에 없던 81개 약물의 cellular response prediction에서 Nature 해설이 전한 약 88% 정확도.
Virtual Cell의 경쟁축이 embedding에서 intervention trajectory로 옮겨간다
세포 상태를 압축하는 표현학습만으로는 약물이 만든 미래 상태를 설명하기 어렵다. 시간축과 실제 의사결정이 들어오면서 문제정의가 바뀐다.
이번 변화는 모델 크기보다 scientific utility에 관한 변화다
이번 최신 확인에서 의미 있는 변화는 ProteinTalks의 Nature version of record 출판이다. 2025년 bioRxiv에서 공개됐던 모델이 2026년 9월 9일 peer-reviewed 논문으로 출판되며, 대규모 동적 proteomics에서 학습한 representation을 실제 drug-discovery task와 patient-derived system까지 연결했다.
평가대상이 더 큰 embedding을 만드는 능력에서 실제 perturbation 이후의 biological trajectory와 translational decision을 지원하는 능력으로 이동한다는 점이 중요하다.
“세포를 무엇으로 표현할까?”라는 질문의 한계
현재 Virtual Cell 연구의 상당 부분은 scRNA-seq를 중심으로 Cell state → latent representation을 만들거나, perturbation 이후 하나의 endpoint state를 예측한다. 이런 접근은 표현력은 높일 수 있지만 실제 약물작용이 시간에 따라 어떻게 전개되는지를 충분히 다루지 못한다.
신약개발에서 phenotype은 한 시점의 사진이 아니다. pathway activation, compensation, resistance, cell death 같은 과정은 서로 다른 시간축에서 나타난다. 따라서 temporal trajectory는 부가적인 feature가 아니라 intervention의 기전을 해석하는 핵심 상태변수에 가깝다.
3,800만 개 이상의 시간분해 proteomics가 dynamical latent space를 만든다
약물과 조합을 처리한 세포가 시간에 따라 어떤 proteome 상태로 이동하는지를 학습한다.
18개 cell line, 63개 drug, 59개 combination, 여러 시간점
연구진은 18개 breast-cancer cell line에 63개 FDA 승인 항암제와 59개 drug combination을 처리하고 여러 시간점에서 측정하여 3,800만 개 이상의 temporal protein-abundance measurement를 구축했다.
ProteinTalks는 이 데이터로 단순한 정적 cell embedding이 아니라 drug perturbation에 따라 시간이 흐르면서 proteome이 어떻게 변하는지를 표현하는 transferable dynamical latent representation을 사전학습한다.
하나의 pretrained representation을 여러 drug-discovery decision에 재사용한다
Drug Efficacy
개별 약물의 cellular response와 효능 예측에 사용한다.
Drug Synergy
조합 효과 예측과 새로운 drug combination 탐색에 사용한다.
Resistance
drug-resistance와 관련된 protein 탐색에 사용한다.
Translational Use
patient response stratification과 patient-derived organoid drug prioritization에 사용한다.
5,585 protein groups와 training에 없던 81개 약물
첨부 자료가 인용한 Nature 해설에 따르면 ProteinTalks는 5,585개 protein group의 시간적 변화를 학습했고, training에 포함되지 않은 81개 약물의 cellular response prediction에서 약 88% 정확도를 보였다.
또한 치료 전 501명의 triple-negative breast cancer 환자 biopsy에서 측정한 3,651개 단백질을 이용한 분석에서 실제 치료 결과와 연결되는 예측력을 보였으며, 연구진은 cell line에서 patient-derived organoid와 clinical biopsy로 모델을 전이했다.
5,585 protein groups, 81개 unseen drug, 약 88% 정확도, 501명 TNBC biopsy, 3,651개 단백질 수치는 첨부 자료가 Nature의 독립 해설에서 인용한 값이다.
현재 상태의 분류가 아니라 미래 molecular trajectory의 예측이다
ProteinTalks의 가장 중요한 차별점은 time을 명시적으로 모델의 입력과 출력 해석에 넣는 데 있다.
“6시간·24시간·48시간 뒤 무엇이 달라지는가?”를 묻는다
기존 framing이 Cell state → latent representation 또는 Perturbation → endpoint state였다면, ProteinTalks는 그 사이의 시간적 진화과정을 직접 학습한다. Nature 논문도 기존 Virtual Cell의 주요 한계로 large-scale time-resolved perturbation data의 부족을 명시한다.
“이 세포가 무엇인가?”라는 정적 질문이 “이 세포에 약물을 가하면 언제 어떤 생물학적 상태로 이동하는가?”라는 동적 질문으로 바뀐다.
반응의 순서는 resistance와 combination 설계를 바꾼다
약물반응에서 특정 protein network가 먼저 변하고 다른 pathway가 뒤이어 보상적으로 활성화된다면, 같은 endpoint라도 mechanism 해석은 달라진다. 시간축은 efficacy뿐 아니라 resistance mechanism과 combination timing을 해석할 때도 중요한 단서가 된다.
첨부 자료의 연구기회를 확장하면, temporal representation은 향후 sequential combination therapy나 adaptive dosing hypothesis를 평가하는 world-model state로 사용할 수 있다. 다만 ProteinTalks 자체가 이런 causal intervention policy를 자동 최적화했다고 의미하지는 않는다.
foundation model이 embedding benchmark를 떠나 실제 연구결정에 들어간다
operational scientific foundation model의 의미는 representation 하나가 여러 downstream scientific decision에 재사용된다는 데 있다.
Hit/Lead에서 precision therapy까지 이어지는 사용면
ProteinTalks의 operationality는 foundation model을 embedding benchmark에만 두지 않고 동일 representation을 drug efficacy, synergy, resistance, patient stratification, organoid prioritization으로 재사용한다는 데 있다.
transfer의 방향이 translational medicine 쪽으로 뻗는다
cell line에서 배운 동적 proteomic representation이 patient-derived organoid와 clinical biopsy 분석으로 이어진다는 점은 중요한 방향성이다. 이는 Virtual Cell의 가치를 “세포 시뮬레이션이 얼마나 정교한가”보다 실제 임상·중개 의사결정에 얼마나 전이되는가로 평가할 근거를 제공한다.
property predictor보다 더 흥미로운 이유는 “언제 무엇이 변하는가”를 묻게 해준다는 데 있다
AI Co-Scientist는 static Drug–Target–Disease KG만으로는 intervention 후 상태변화를 충분히 표현하기 어렵다.
Candidate A 뒤의 protein-network trajectory에서 다음 combination으로
AI Co-Scientist가 “Candidate A를 처리하면 어떤 protein network가 언제 변하는가?”를 질문한다고 가정할 수 있다. ProteinTalks가 predicted trajectory를 제공하고, agent가 그 변화에서 resistance mechanism을 추론하고, 문헌과 KG를 검색해 alternative combination을 찾은 뒤 organoid validation으로 연결하는 식이다.
Drug–Target–Disease KG에서 intervention trajectory graph로
첨부 자료는 AI Co-Scientist의 내부 세계모델도 정적인 Drug–Target–Disease KG에서 intervention과 시간의존 molecular state를 직접 표현하는 dynamic causal/epistemic model로 확장할 필요가 있다고 본다.
이 구조에서 KG는 단순한 entity-relation 저장소가 아니라 “어떤 intervention이 어떤 조건과 시간축에서 어떤 molecular transition을 유발했는가”를 기록하는 provenance-rich state transition substrate가 된다.
깊은 single modality의 성공이 곧 multimodal causal world model은 아니다
이번 결과는 오히려 다음 연구공백을 더 선명하게 만든다.
강점은 perturbation proteomics의 깊이에 있다
ProteinTalks 자체는 multimodal foundation model이 아니다. 현재 강점은 perturbation proteomics라는 매우 깊은 단일 modality에 있다. 첨부 자료는 이를 BioMatrix 같은 multimodal biomolecular foundation model과 상보적으로 본다. ProteinTalks가 temporal dynamics를 깊게 모델링한다면, multimodal FM은 sequence, structure, text, omics 사이의 범위를 넓힌다.
실제 약물반응은 proteome 하나만으로 결정되지 않는다
가장 큰 공백은 이 모든 modality와 context를 함께 받아 intervention 이후 trajectory를 생성하는 Multimodal Perturbation World Model이다.
“A를 투여했다”에서 “A 대신 B였다면?”으로
ProteinTalks의 관측 기반 예측에서 더 나아가려면 “A 대신 B를 투여했다면?”, “A+B를 이 순서로 투여했다면?”, “특정 resistance protein을 동시에 억제했다면?”을 causal하게 비교해야 한다.
예측과 counterfactual causal reasoning은 같은 능력이 아니다. observed perturbation pattern을 잘 맞추는 모델이 개입의 인과효과를 자동으로 식별했다고 볼 수는 없다.
주어진 perturbation을 예측하는 것에서 다음 perturbation을 스스로 선택하는 것으로
현재 모델은 주어진 perturbation의 결과를 예측하는 방향이 강하다. AI Co-Scientist가 되려면 uncertainty를 보고 가장 정보가치가 높은 perturbation, dose, time point, combination을 스스로 선택해야 한다.
가장 유망한 다음 구조는 예측, 검색, 반사실, 실험선택, 검증, 업데이트를 한 루프로 묶는 것이다
Virtual Cell이 실험 전에 여러 intervention을 시험하는 biological world model이 되려면 prospective evidence loop가 필요하다.
Multimodal Scientific FM에서 wet-lab belief update까지
첨부 자료가 제안하는 통합 구조는 Multimodal Scientific FM → Dynamic Virtual Cell → Agentic RAG/HRKG → Counterfactual Hypotheses → Value-of-Information Experiment Selection → Organoid/Wet Lab → Temporal Observation → Belief/Model Update이다.
디지털 세포 시뮬레이터가 아니라 실험정책을 지원하는 scientific world model
이 구조가 구현되면 Virtual Cell은 단순한 “디지털 세포 시뮬레이터”를 넘어 AI Co-Scientist가 실험 전에 여러 intervention을 시험하고, 그중 가장 판별력 높은 실험을 고르는 biological world model 역할을 할 수 있다.
ProteinTalks가 보여주는 핵심 변화는 더 큰 cell embedding이 아니다. 약물을 가했을 때 시간에 따라 실제 생물학이 어떻게 변할지를 예측하고, 그 representation을 환자·오가노이드 수준의 의사결정으로 전이할 수 있는가가 새로운 경쟁축이다.
Operational virtual cell as a scientific world model이번 업데이트를 여덟 문장으로 압축하면
01 · One major update
9월 9일 Nature 출판된 ProteinTalks가 이번 추적에서 추가할 만큼 의미 있는 신규 변화다.
02 · Dynamics over static state
Virtual Cell의 질문이 cell identity 표현에서 perturbation 뒤 future proteome trajectory 예측으로 이동한다.
03 · Large temporal proteomics
18개 cell line, 63개 drug, 59개 combination, 3,800만 개 이상의 temporal measurement를 활용한다.
04 · Translational transfer
representation이 efficacy, synergy, resistance, patient stratification, organoid prioritization에 재사용된다.
05 · Patient-level signal
Nature 해설은 501명 TNBC biopsy의 3,651개 단백질 분석과 실제 치료 결과 연결을 전한다.
06 · Not multimodal yet
ProteinTalks는 깊은 perturbation proteomics 모델이지 multimodal biomolecular FM은 아니다.
07 · Three major gaps
multimodal perturbation world model, counterfactual reasoning, agentic experiment selection이 핵심 공백이다.
08 · Next Co-Scientist
dynamic virtual cell + Agentic RAG/HRKG + counterfactuals + VoI + prospective wet-lab update가 가장 직접적인 통합 연구방향이다.
References
첨부 연구동향 문서가 직접 제시한 공식 논문, 해설, 코드, 비교 연구 링크를 유지한다.
ProteinTalks의 peer-reviewed version of record. 대규모 dynamic proteomics, transferable dynamical representation, drug-discovery 및 patient-derived applications의 핵심 출처.
5,585 protein groups, unseen 81개 약물 약 88% 정확도, 501명 TNBC biopsy와 3,651개 단백질 관련 수치를 첨부 자료가 인용한 해설.
첨부 자료가 제시한 공개 코드 저장소.
ProteinTalks의 deep temporal single-modality modeling과 multimodal biomolecular foundation model의 상보성을 설명할 때 첨부 자료가 비교 대상으로 제시한 연구.