AI Research Blog · AI Co-Scientist × Drug Discovery2026-09-21
Research Watch · Four Structural Changes

논문을 읽는 AI에서, 실험하고 수정하는 연구조직으로

Epistemic Closed-Loop Drug Discovery Co-Scientist: Autonomous Labs, Virtual Biotech, Executable Papers, and Specialized Biology Models

SPECIALIZEDSCIENTIFIC FMEXECUTABLEPAPER / SKILLMULTI-AGENTORGANIZATIONEVIDENCE /TOOL EXECUTIONPHYSICALEXPERIMENTMEASUREMENT+ PROVENANCENEXTDECISIONsuccess and failure revise the next hypothesis, tool, and experimentCONCEPTUAL SYNTHESIS · NOT A MEASURED CHART

이번 주의 변화는 하나의 더 큰 모델이 아니다. AI Co-Scientist의 기본 단위가 ‘답변 생성기’에서 실행 가능한 논문, 전문 연구조직, 자동화 wet-lab, 측정과 다음 의사결정을 연결하는 폐루프 시스템으로 이동하고 있다는 점이다.

신규 핵심은 네 건이다. Andromeda 2는 실제 formulation batch를 반복 실행한다. Virtual Biotech는 수만 개 agent로 therapeutic R&D 조직을 모사한다. Paper2Agent는 논문을 MCP 기반 executable scientific agent로 바꾼다. LongevityBench/Longevity Claw는 domain-specialized multi-omics model과 scientific tools를 결합한다. 반면 BioMatrix/MIMIC급의 새로운 범용 biomolecular multimodal foundation model은 이번 주에 추가로 확인되지 않았다.

Part I · Andromeda 2

Agent가 제약 formulation 실험을 실제로 반복 실행한다

Andromeda 2의 핵심은 recommendation이 아니라 experimental evidence와 다음 batch를 물리적으로 연결하는 successive closed loop다.

Source factCraig et al., “Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory”(arXiv:2609.19099)는 2026년 9월 16일 공개됐다. 시스템은 structured in-house experimental evidence를 읽고 computational tool과 experimental tool을 호출해 다음 formulation batch를 설계하고, 자동화 실험실에서 실행한 뒤 측정 결과를 이용해 이어지는 batch를 다시 설계한다.

EXPERIMENTAL EVIDENCEAGENT REASONINGTOOL EXECUTIONWET-LAB MEASUREMENTNEXT BATCH

Paclitaxel SEDDS

50%

high-performance formulation hit rate

Andromeda 1

17%

probabilistic optimization baseline

Wet-lab DoE

2%

comparison baseline

Evidence Ablation

−34%

structured in-house evidence 제거 시 평균 AUC 감소

네 가지 target product profile 목표를 모두 만족한 formulation 수는 방법별로 각각 12개, 6개, 0개였다. 원 자료가 강조하는 차별점은 LLM → recommendation → human experiment가 아니라 실험 결과가 다음 실험전략을 직접 바꾸는 물리적 반복 루프다.

Boundary현재 범위는 SEDDS formulation과 paclitaxel 중심이다. 이를 곧바로 target discovery부터 clinical translation까지 포괄하는 end-to-end drug-discovery Co-Scientist로 확대해 해석해서는 안 된다.

Drug-development stage

Lead Optimization 이후 Formulation Development / Preclinical Development에 특히 직접적이다.

Open problem

agent의 experimental strategy가 epistemic uncertainty와 Value-of-Information을 명시적으로 최적화하는 과학적 실험설계인지가 아직 불분명하다.

Research direction. assay/formulation provenance + uncertainty + negative result + cost + expected information gain을 공유 memory에 보존하고 Bayesian/causal planner가 다음 batch를 선택하는 Evidence-Grounded Self-Driving Formulation Co-Scientist가 자연스러운 후속 구조다.

arXiv · Andromeda 2

Part II · Virtual Biotech

수만 개 Agent가 신약개발 ‘조직’ 전체를 소프트웨어로 옮긴다

중요한 것은 agent 수 자체보다 biotech R&D의 division, evidence scale, review structure를 하나의 시스템 아키텍처로 재현한 데 있다.

Source factZhang, Eckmann, Miao, Mahon & Zou의 “The Virtual Biotech: A multi-agent AI framework for therapeutic discovery and development”는 2026년 9월 17일 Science version of record로 출판됐다(DOI 10.1126/science.aeg6779). 시스템은 Chief Scientific Officer agent 아래 target discovery, safety, modality selection, clinical development 등의 전문 부서를 둔다.

Specialist agents

37,000+

전문 역할을 가진 대규모 agent population

Clinical trials

55,984

분석된 clinical trial 수

Market reach

+48%

cell-type specificity가 높은 target 약물의 시장 도달확률 연관성

Adverse events

−32%

같은 분석에서 보고된 adverse-event 연관성

시스템은 genetic, single-cell, spatial, clinical evidence를 통합해 lung cancer의 B7-H3 antibody–drug conjugate 전략을 제안했고, 종료된 ulcerative-colitis trial의 실패 원인도 분석했다.

VIRTUAL CSOBIOLOGICAL / CLINICAL / SAFETY / MODALITY DIVISIONSDOMAIN TOOLS + DATASCIENTIFIC REVIEWERITERATIVE REFINEMENT

InterpretationVirtual Biotech가 새 ADC molecule을 만들어 직접 임상시험에서 성공시킨 것은 아니다. 시스템은 2025년 1월 이전 정보만 사용해 B7-H3 ADC 전략을 제안했고, 이후 다른 제약사가 독립적으로 같은 modality/target 전략을 개발해 임상 성과를 냈다. 따라서 prospective wet-lab validation이 아니라 temporal holdout에 가까운 독립적 수렴 증거로 보는 것이 정확하다. 연구진도 새로운 후보의 실제 실험 검증이 다음 단계라고 명시한다.

신약개발 범위는 Target Identification에서 Clinical Translation까지 넓다. 이 시스템은 후보 생성 AI라기보다 R&D portfolio decision AI에 가깝다. 어떤 target을 선택하고 어떤 modality를 쓰며 어떤 실패 trial을 재해석할 것인가를 한 조직 구조에서 다룬다.

Research gap. 37,000개 agent가 곧 37,000개의 독립적인 과학적 근거를 뜻하지는 않는다. 동일 논문·database·clinical trial을 여러 agent가 반복 사용할 수 있으므로 evidence ancestry/dependence를 모델링해야 한다. 또한 최종 hypothesis가 wet-lab으로 넘어가고 그 결과가 다시 조직의 belief를 수정하는 반복 DMTA loop는 아직 없다.

Science DOI · Virtual Biotech project · Stanford Medicine

Part III · Paper2Agent

RAG가 논문을 ‘검색’하는 데서 논문의 방법을 ‘실행’하는 데로 이동한다

Paper2Agent는 manuscript와 supplement를 읽는 것을 넘어 codebase와 dataset을 재현 가능한 환경과 테스트된 MCP tool로 변환한다.

Source factMiao et al., “Reimagining research papers as interactive and reliable AI agents”는 2026년 9월 16일 Nature version of record로 출판됐다. Paper2Agent는 manuscript, supplementary material, dataset, codebase를 분석해 MCP(Model Context Protocol) server를 만들고, 논문의 핵심 방법을 executable tool, resource, workflow prompt로 변환한다.

별도의 environment agent, extraction agent, testing agent가 실행환경을 만들고 방법을 추출하고 코드를 재현·검증한다. 실패한 tool은 수정하거나 제외한다. 따라서 Agentic RAG의 기본 단위가 다음처럼 변한다.

PAPERREPRODUCIBLE ENVIRONMENTEXECUTABLE METHODTESTED MCPPAPER AGENT

Computational biology

100

자동 처리된 paper

Agentized papers

74

성공적으로 agent화된 computational-biology paper

Tutorial benchmarks

300

agentized paper 검증용 benchmark

Out-of-scope rejection

100%

잘못 연결한 adversarial paper–question test

추가로 26개의 data/discovery paper, 10개의 비생물 계산논문을 처리했고, 26개 discovery paper에는 main text와 supplement를 함께 사용해야 하는 100개 synthesis task가 적용됐다.

AlphaGenome 논문을 Paper2Agent로 변환한 사례에서는 variant scoring, gene expression, chromatin accessibility 등 여러 기능을 실제 호출할 수 있는 agent가 만들어졌다. 여러 paper MCP를 동시에 하나의 agent에 연결할 수도 있어 논문마다 하나의 executable scientific expert가 존재하고 서로 협업하는 구조로 확장될 수 있다.

신약개발에서는 target validation, single-cell analysis, genomics, PK/PD, biomarker analysis에 직접적인 의미가 있다. 문헌에서 “무엇을 주장했는가?”를 검색하는 수준을 넘어 “이 논문의 공개된 분석방법을 현재 compound/omics dataset에 실제로 실행하라”까지 수행할 수 있기 때문이다.

Epistemic gap. “논문의 방법을 재현할 수 있다”와 “그 방법이 현재 연구문제에 과학적으로 적절하다”는 다른 문제다. 저자들도 최종 연구방향 선택과 evidence 평가는 인간 책임임을 명시한다. 다음 단계는 각 paper-agent에 validity domain, dataset shift, assay condition, uncertainty, external replication, retraction/correction status를 붙이는 Epistemic Paper Agent다.

Nature · Paper2Agent

Part IV · LongevityBench / Longevity Claw

더 큰 범용모델보다 domain-specialized multi-omics model + agent가 강할 수 있다

Cell 연구는 규모 경쟁만으로 생물학적 Co-Scientist를 설계하는 가정을 흔들고, 작은 전문모델과 도구 결합의 실용적 가능성을 보여준다.

Source factZhavoronkov et al., “An open benchmark and language models for AI in aging biology”는 2026년 9월 17일 Cell 189, 5980–5994.e8에 출판됐다(DOI 10.1016/j.cell.2026.08.026). LongevityBench는 clinical, genetics, epigenomics, transcriptomics, proteomics의 17개 task로 구성된다.

평가에서는 하나의 frontier model이 모든 biological modality에서 우세하지 않았고, omics 기반 biological-age prediction이 특히 어려웠다. 연구진은 0.6B–9B 규모의 다섯 Longevity-LLM을 domain-specific clinical/multi-omics data로 fine-tuning했으며, 작은 모델들이 훨씬 큰 범용 frontier system과 맞먹거나 더 나은 결과를 냈다. 공식 leaderboard에서는 L-Qwen3.5-9B가 aggregate rank 1위를 기록한다.

Longevity Claw: 전문모델을 도구와 연구 workflow에 연결한다

Longevity Claw는 전문모델을 gene-set enrichment, aging-clock calculation, population profiling, evidence retrieval/synthesis, target evaluation tool과 결합한 agentic research platform이다. 14개 hallmarks of aging에 대해 multi-step workflow를 수행해 328개 후보 유전자를 제안했고, 독립적으로 알려진 aging target reference set에 대해 최대 5.6배 enrichment가 보고됐다.

DOMAIN-SPECIALIZED MULTI-OMICS MODELSCIENTIFIC TOOLSEVIDENCE RETRIEVALMULTI-STEP TARGET PRIORITIZATION

중요한 차별점은 단순 “생물학용 LLM”이 아니라 이 전체 open pipeline을 제공한 데 있다. 이는 “가장 큰 general LLM을 Co-Scientist의 brain으로 쓰는 것이 항상 최선인가?”라는 가정에 도전한다. domain-specific data와 작은 전문모델을 잘 설계하면 local/offline scientific agent에서도 경쟁력 있는 biological reasoning이 가능할 수 있다는 신호다.

주요 적용 단계는 Target Identification / Target Prioritization이다. 노화 분야의 구조를 fibrosis, oncology, immunology로 옮겨 disease-specific model + omics + RAG + analysis tools를 구성할 수 있다.

Research gap. reference-set enrichment는 유망하지만 “이미 알려진 생물학과 잘 맞는다”와 “새 치료표적이 실제로 작동한다”는 다르다. 결정적 후속 연구는 target hypothesis → perturbation experiment → negative/positive result → model/KG update를 연결하는 것이다.

Cell paper · LongevityBench · Insilico Medicine

Part V · Infrastructure Signal

Frontier-model 회사가 scientific software를 넘어 laboratory hardware까지 내려온다

논문 밖의 산업 변화도 Co-Scientist 아키텍처의 물리적 실행 계층이 빠르게 현실화되고 있음을 보여준다.

Source factAnthropic은 9월 18일 Reuters 인터뷰를 통해 실제 wet-biology lab 운영을 확인했다. life-sciences 책임자는 자체 시설과 외부 파트너를 통해 실제 생물학 실험을 수행하고 있으며, Claude가 robotic laboratory execution을 더 많이 담당하도록 하는 방향을 설명했다. 회사는 해당 시설이 drug discovery만을 위한 것은 아니라고 명확히 했다.

또한 9월 17일 Life Sciences Verification Program을 공개해 drug discovery, research biology, clinical development 등의 고급 생명과학 작업에 Mythos/Opus/Sonnet 모델 접근을 제공하기 시작했다. 8월 공개된 Model Hardware Standard는 microscope, liquid handler, robotic arm 등을 AI agent가 공통 interface로 제어하도록 설계됐다.

LLMSCIENTIFIC SOFTWARELAB HARDWAREPHYSICAL EXPERIMENT

Analysis이는 독립된 한 논문의 성능 결과가 아니라 산업·연구 인프라의 방향 신호다. frontier-model 회사가 모델 API와 research software를 넘어 physical experiment execution 계층까지 수직적으로 연결하려는 움직임으로 읽을 수 있다.

Reuters report via The Star · Anthropic Life Sciences Verification Program

Part VI · Synthesis

가장 큰 공백은 ‘Epistemic Closed-Loop Drug-Discovery Co-Scientist’다

네 연구는 서로 다른 층을 채우지만, 근거의 독립성·정보가치·negative result·belief revision을 하나의 반복 시스템에 묶은 사례는 아직 비어 있다.

Research signalWhat becomes executableStrongest evidenceRemaining gap
Andromeda 2Formulation experiment loopsuccessive automated wet-lab batches; hit-rate comparison; evidence ablationexplicit uncertainty / Value-of-Information planning
Virtual BiotechBiotech organizational workflow37,000+ agents; 55,984 trials; cross-scale evidence integrationevidence ancestry/dependence + iterative wet-lab belief update
Paper2AgentPublished scientific methodtested MCP tools; 74 agentized computational-biology papers; adversarial rejectionvalidity domain + shift + replication + correction state
Longevity ClawDomain-specialized model/tool workflow17-task benchmark; 328 target candidates; up to 5.6× enrichmentprospective wet-lab validation + causal discrimination
Anthropic lab stackModel-to-hardware execution layerwet-lab operation confirmed; verification program; hardware interface directionpublic quantitative validation of autonomous scientific discovery

이번 주 네 연구를 하나의 흐름으로 묶으면 AI Co-Scientist architecture는 다음 형태로 수렴한다.

SPECIALIZED SCIENTIFIC FMEXECUTABLE PAPER / SKILLMULTI-AGENT ORGANIZATIONEVIDENCE / TOOL EXECUTIONPHYSICAL EXPERIMENTMEASUREMENTNEXT DECISION

특히 세 변화가 중요하다. 첫째, retrieval이 executable knowledge로 변한다(Paper2Agent). 둘째, agent가 개별 assistant에서 실제 R&D 조직 구조로 확장된다(Virtual Biotech). 셋째, formulation처럼 실제 약물개발의 물리적 단계가 closed loop에 들어오기 시작한다(Andromeda 2).

Still missing이번 검색에서는 multimodal biomolecular FM + evidence-lineage-aware Agentic RAG + specialist agents + executable literature/skills + Value-of-Information experiment planning + automated wet lab + negative-result memory + belief revision을 하나의 시스템에서 반복적으로 검증한 연구는 확인되지 않았다.

가장 강한 연구기회는 더 많은 agent를 배치하는 데 있지 않다. 각 가설의 근거·독립성·조건·불확실성·반대증거를 추적하고, 정보가치가 가장 높은 다음 실험을 선택하며, 실제 실패와 성공이 다음 hypothesis와 model selection을 바꾸는 Epistemic Closed-Loop Drug-Discovery Co-Scientist에 있다.Synthesis grounded in the attached research watch

Epistemic memory

provenance, evidence ancestry, uncertainty, negative result, validity domain, retraction/correction state를 first-class state로 유지한다.

Decision rule

다음 실험은 단순 expected performance가 아니라 cost와 expected information gain, falsifiability, safety를 함께 고려해 선택한다.

References

Source Set

01
Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory
arXiv:2609.19099 · 2026-09-16

arxiv.org/abs/2609.19099

02
The Virtual Biotech: A multi-agent AI framework for therapeutic discovery and development
Science · 2026-09-17 · DOI 10.1126/science.aeg6779

doi.org/10.1126/science.aeg6779 · virtualbiotech.ai

03
Stanford Medicine — Virtual biotech company puts thousands of AI scientist agents to work on drug discovery
Stanford Medicine · 2026-09

Stanford Medicine

04
Virtual Biotech official project page
Project resource

virtualbiotech.ai

05
Reimagining research papers as interactive and reliable AI agents
Nature · 2026-09-16

nature.com/articles/s41586-026-11044-y

06
An open benchmark and language models for AI in aging biology
Cell · 2026-09-17 · DOI 10.1016/j.cell.2026.08.026

Cell · LongevityBench

07
Insilico Medicine opens AI longevity discovery toolkit
Longevity Claw resource

insilico.com

08
Anthropic quietly sets up biology lab as it ramps AI drug program
Reuters report via The Star · 2026-09-18

The Star

09
Introducing the Life Sciences Verification Program
Anthropic · 2026-09-17

anthropic.com/news/life-sciences-verification-program