AI Co-Scientist × Drug DiscoveryGolden Circle · 06 Sep 2026
Golden Circle/Why → How → What/AI Co-Scientist

과학적으로 올바른 방법을 알고,
물리적으로 가능한 실험만 실행하는 AI

From Scientific Skills to Verified Lab Workflows: A Golden-Circle Architecture for Drug-Discovery Co-Scientists

Golden Circle of procedural and physical verification Why identifies the gap between knowing, doing and doing correctly; How combines scientific skills with formal laboratory-state verification; What builds a closed-loop drug-discovery Co-Scientist. WHYcorrect ≠ executableHOWskills + semanticsWHATverified closed loop WHYRAG knows facts. Tools can act.Neither guarantees the right procedure or a valid experiment. HOWSelect a versioned skill, then verify state transitions.Scientific Agent Skills + Computable Laboratory WHATEvidence → Skill → Verified Workflow → Evidence UpdateProcedurally and Physically Verifiable Co-Scientist
Editorial Abstract

2026년 9월 6일의 의미 있는 변화는 새로운 거대 drug-discovery foundation model이나 prospective wet-lab Co-Scientist가 나온 데 있지 않다. AI Co-Scientist stack에서 “과학적 절차를 올바르게 선택하는 층”과 “그 절차를 현재 물리 실험실 상태에서 검증하는 층”이 명시적인 연구대상으로 분리되기 시작했다.

Scientific Agent Skills는 지식과 절차를 분리해 “어떻게 올바르게 수행해야 하는가”를 재사용 가능한 versioned skill로 만든다. Computable laboratory 연구는 실험실 자체를 typed state space로 표현하고 실행 전에 precondition과 constraint를 검증한다. 두 흐름을 연결하면 AI Co-Scientist의 기본 stack은 LLM → RAG → Tool → Result에서 Evidence → Procedure/Skill → Planner → Formal Workflow → Scientific Tool/Robot → Verified State Transition → Evidence Update로 이동한다.

WHY

사실을 아는 것과 과학적으로 올바른 절차를 고르는 것, 그리고 그 절차를 물리적으로 실행할 수 있는 것은 서로 다른 문제다.

HOW

Versioned Scientific Skill을 evidence와 context로 선택하고, formal laboratory state와 constraint verification을 거쳐 실행한다.

WHAT

Evidence와 실패를 다시 provenance-aware graph로 환류하는 Procedurally and Physically Verifiable Drug-Discovery Co-Scientist를 만든다.

Evidence boundary

첨부 Research Watch는 2026년 9월 6일 현재 이 전체 루프를 실제 small-molecule drug discovery에서 prospective하게 검증한 새 연구는 확인하지 못했다고 명시한다. 따라서 아래 통합 아키텍처는 두 신규 연구가 제공한 procedural layer와 physical-execution verification layer를 연결한 후속 연구 제안이지, 이미 end-to-end로 실증된 단일 시스템이 아니다.

Part I · WHY

왜 RAG와 Tool만으로는 과학적 Co-Scientist가 되기 어려운가

Co-Scientist의 실패는 정보를 못 찾거나 API를 못 호출해서만 생기지 않는다. 잘못된 절차를 선택하거나, 올바른 절차라도 현재 실험실 상태에서 실행 불가능한 경우가 있다.

§1 · The missing layers

“무엇을 아는가”, “어떻게 해야 하는가”, “지금 실행 가능한가”는 다른 질문이다

기존의 단순한 agentic stack은 흔히 LLM → RAG → Tool → Result로 설명된다. 하지만 과학 연구에서는 이 구조가 두 종류의 지식을 한데 섞는다. 하나는 fact knowledge이고, 다른 하나는 procedural knowledge다. 여기에 실제 물리 환경의 execution validity가 별도로 존재한다.

RAG
Know

무엇이 알려져 있는가. 예: 단백질, assay, 문헌, KG의 사실과 evidence.

Skill
Do Right

어떻게 과학적으로 올바르게 수행해야 하는가. 예: QC, 통계 절차, reporting rule.

Lab Semantics
Can Execute

현재 laboratory state에서 그 workflow가 실제로 허용되고 실행 가능한가.

잘 실행되는 코드와 과학적으로 방어 가능한 분석은 같지 않고, syntactically valid한 robot command와 scientifically/physically valid한 experiment도 같지 않다. 이번 두 연구는 바로 이 두 간극을 각각 별도의 문제로 만든다.

§2 · Drug-discovery stakes

신약개발에서는 절차 오류와 상태 오류가 결과 자체를 무효화한다

Multiple-testing correction, batch effect, coordinate convention, confounder 처리 같은 procedural error는 코드가 정상 종료돼도 분석 결론을 무효화할 수 있다. 반대로 후보 24종을 dose-response assay로 비교하라는 계획이 과학적으로 타당하더라도 plate 상태, reagent availability, sample identity, concentration, instrument capability, remaining volume, protocol dependency가 맞지 않으면 실제 실험으로 실행할 수 없다.

따라서 신약개발 Co-Scientist가 필요한 이유는 단순 자동화가 아니다. 과학적 절차 타당성과 물리적 실행 가능성을 서로 다른 검증 계층으로 분리하고, 둘을 evidence와 연결해야 하기 때문이다.

Part II · HOW / Scientific Skills

과학적 절차를 versioned skill로 만들어 필요할 때 불러온다

Scientific Agent Skills는 RAG가 사실을 검색하듯, AI Scientist가 분야별 절차 지식을 검색·재사용할 수 있도록 procedural layer를 분리한다.

§3 · arXiv:2609.00065

Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents

Kassis 외의 연구로, 공개본은 2026년 9월 4일로 표시되어 있다. 연구진은 genomics, cheminformatics, medical imaging, study design, scientific communication 등을 포함하는 163개의 scientific procedure skill을 공개했다.

핵심은 knowledge와 procedure를 분리하는 것이다. 일반 RAG가 “이 단백질에 대해 무엇이 알려져 있는가?”를 검색한다면, Scientific Agent Skills는 “RNA-seq differential expression에서 어떤 statistical procedure를 써야 하는가?”, “어떤 identifier namespace가 authoritative한가?”, “어떤 caveat와 reporting rule을 적용해야 하는가?”처럼 과학자가 작업을 올바르게 수행하기 위해 알아야 하는 절차 지식을 versioned instruction으로 저장한다.

§4 · Skill package

SKILL.md, reference, executable script를 단계적으로 공개한다

각 skill은 최소한 SKILL.md를 가지며 필요하면 reference document와 executable script를 포함한다. 중요한 설계는 모든 상세 procedure를 항상 context에 밀어 넣지 않는다는 점이다.

Standing Contextskill name + description
Selection현재 task와 context에 맞는 skill 선택
Instruction LoadSKILL.md를 단계적으로 읽음
Reference / Script필요할 때만 추가 자원 로드

progressive disclosure 구조에서 전체 library의 standing description은 200k-token context의 약 7.1%이며, 대부분의 문서는 실제 선택될 때까지 로드되지 않는다.

§5 · Layer separation

RAG, Tool, Skill의 역할을 명확히 나눈다

LayerQuestionDrug-discovery example
Evidence RAG무엇을 알고 있는가?ChEMBL, PubMed, KG에서 evidence 검색
Scientific Skill어떻게 올바르게 수행해야 하는가?assay QC, DTA analysis, molecular filtering, statistical validation
Tool무엇을 실행할 수 있는가?RDKit, docking, Boltz, PK/PD simulator
Agent어떤 evidence·skill·tool을 사용할 것인가?현재 연구목표와 experimental context에 따라 조합 결정

이 구분은 화학 계산이나 분석 규칙까지 LLM의 암묵적 지식에 맡기지 말아야 한다는 방향과 맞닿아 있다. 절차를 외부화하면 어떤 규칙과 caveat가 적용됐는지 추적하고 versioning할 수 있다.

§6 · Limitation and gap

163개 skill이 있다는 사실과, 올바른 skill을 고른다는 것은 다르다

논문은 resource paper이며 163개 skill 전체에 대한 task-level effectiveness를 검증하지 않았고, agent가 올바른 skill을 선택하는 비율도 측정하지 않았다고 명시한다. 따라서 skill library를 추가하면 AI Scientist가 자동으로 더 정확해진다고 결론 내릴 수 없다.

가장 중요한 연구공백은 Evidence-Aware Scientific Skill Selection이다. 동일한 procedure도 assay type, target family, 데이터 규모, species, experimental design에 따라 적절성이 달라지므로 skill 자체에 applicability 조건이 필요하다.

Skill + Applicability Condition + Provenance + Version + Evidence + Expected Error

이 typed skill representation을 Hyper-Relational KG로 표현하고 Agentic RAG가 현재 연구 context에 맞는 skill을 검색한다면, 사실 검색과 절차 선택을 하나의 evidence-aware reasoning 문제로 연결할 수 있다.

Part III · HOW / Computable Laboratory

실험실 자체를 계산 가능한 state space로 만들고 dispatch 전에 검증한다

Scientific Skill이 “어떻게 해야 하는가”를 말한다면, computable laboratory는 “지금 이 상태에서 실제로 할 수 있는가”를 판단한다.

§7 · arXiv:2609.03621

A computable representation of the physical laboratory enables verifiable workflows

Xiaobo Li 외의 연구로 2026년 9월 3일 공개됐다. 핵심 기여는 laboratory automation을 LLM → API/tool call → robot으로 단순화하지 않고, 물리 실험실 자체를 컴퓨터가 reasoning할 수 있는 computable state space로 표현한 데 있다.

Typed research objects

sample, reagent, plate, instrument 등 연구 객체를 타입이 있는 상태로 표현한다.

Capability-bound operations

각 operation을 실제 장비 capability와 결합해 가능한 행동을 제한한다.

Compositional workflow algebra

dependency, decision, iteration, concurrency를 포함한 workflow를 조합 가능한 프로그램으로 표현한다.

Workflow는 시간에 따라 변화하는 laboratory state에 대한 프로그램이다. 실행 전에는 stateful simulation으로 object transformation을 추적하고 operation precondition과 laboratory constraint를 검사한다. 이 표현은 modular agentic robotic laboratory의 executable Function Skills와 연결된다.

§8 · Execution semantics

Robot API 앞에 formal verification layer를 삽입한다

이 연구의 차별점은 hardware interface가 아니라 “이 행동이 현재 실험실 상태에서 수행 가능한가?”를 형식적으로 판단하는 execution semantics에 있다.

Scientific Intentagent가 원하는 실험 목표
Formal Workflowtyped operation program
State / Constraint Verificationprecondition + lab constraints
Executable Skillrobotic laboratory Function Skill

따라서 syntactically valid한 robot command와 scientifically/physically valid한 experiment를 구분할 수 있는 기반이 생긴다.

§9 · Closed-loop DMTA

“무슨 실험을 할까?”에서 “지금 합법적으로 실행 가능한 실험은 무엇인가?”로

Closed-loop DMTA에서 Co-Scientist가 후보 24종의 dose-response assay를 결정했다고 하자. 실제 실행은 plate 상태, reagent availability, sample identity, concentration, instrument capability, remaining volume, protocol dependency가 모두 일치해야 한다.

Candidate Selection24 compounds
Assay Plandose-response design
Formal Lab Stateplate · reagent · sample · instrument
Precondition Checkvolume · capability · dependency
Robotic Executionverified dispatch
Observationassay result
Updated Lab Statestate transition

이 구조는 AI Co-Scientist를 “실험 아이디어를 생성하는 시스템”에서 현재 laboratory state에서 어떤 실험이 허용되고 실행 가능한지를 reasoning하는 시스템으로 확장한다.

§10 · Research gap

Physical Laboratory State를 Epistemic Laboratory State로 확장한다

논문의 핵심 검증은 workflow/state representation과 dispatch 전 verification이다. 이것이 실제 drug-discovery DMTA에서 장기간 prospective하게 실패율, 재현성, 비용을 얼마나 개선하는지는 아직 검증되지 않았다.

첨부 Research Watch가 제안하는 가장 가치 높은 확장은 laboratory state에 물리 상태뿐 아니라 다음 정보를 함께 넣는 Epistemic Laboratory State다.

compound lot / assay version / cell passage / instrument calibration / model prediction / uncertainty / provenance / negative result

이렇게 하면 Agentic RAG의 evidence state와 robotic laboratory의 physical state를 하나의 모델에서 연결할 수 있다.

Part IV · WHAT

무엇을 만들 것인가: Procedurally and Physically Verifiable Drug-Discovery Co-Scientist

Scientific Skills의 procedural correctness와 computable laboratory의 physical validity를 하나의 evidence-driven closed loop로 결합한다.

§11 · New software stack

LLM → RAG → Tool의 세 단계를 일곱 단계의 검증가능한 연구 stack으로 확장한다

이번 두 연구를 함께 보면 Co-Scientist software stack의 분화가 명확해진다.

Evidenceliterature · KG · assays · omics
Procedure / Skillversioned scientific know-how
Plannercontext-aware orchestration
Formal Workflowtyped operations
Scientific Tool / RobotRDKit · docking · Boltz · lab
Verified State Transitionpreconditions satisfied
Evidence Updateresult · failure · provenance

“과학적으로 올바른 방법을 아는 것”과 “그 방법을 물리적으로 안전하게 실행하는 것”은 서로 다른 연구 문제이며, 좋은 Co-Scientist는 둘 다 해결해야 한다. Scientific Agent Skills가 전자를 다루고 computable laboratory가 후자를 다룬다.

§12 · Integrated architecture

Evidence-Aware Skill Selection과 Epistemic Laboratory State를 연결한다

통합 시스템은 multimodal scientific foundation model과 Agentic RAG가 evidence를 수집한 뒤, 현재 experimental context에 맞는 versioned scientific skill을 선택한다. 선택된 procedure는 Planner가 formal workflow로 변환하고, typed laboratory-state model이 실행 전 검증한다. 실행 결과와 실패는 provenance-aware evidence graph로 돌아가 다음 의사결정을 수정한다.

ModulePrimary responsibilityTyped state / metadata
Evidence RAG / HRKG사실·근거 검색source, time, assay, provenance, contradiction
Skill Selector현재 context에 맞는 procedure 선택applicability, version, evidence, expected error
Plannerevidence와 skill을 실행 계획으로 조합goal, dependency, decision, iteration
Lab-State Verifierworkflow가 현재 상태에서 실행 가능한지 검증object type, capability, volume, calibration, precondition
Tool / Robot검증된 scientific operation 수행executable Function Skill
Evidence Updater결과·실패·불확실성을 다음 reasoning state로 환류observation, negative result, uncertainty, provenance
§13 · Research thesis

Procedurally and Physically Verifiable Drug-Discovery Co-Scientist

Multimodal scientific foundation model과 Agentic RAG가 evidence를 수집하고, 현재 experimental context에 맞는 versioned scientific skill을 선택하며, 계획된 실험을 typed laboratory-state model로 검증한 뒤 실행하고, 결과와 실패를 provenance-aware evidence graph로 돌려보내 다음 의사결정을 수정하는 시스템.

Research direction synthesized in the attached Research Watch

이 구조의 핵심은 tool-use를 늘리는 것이 아니다. 무엇을 해야 하는지 결정하는 절차적 지식과, 그 절차가 지금 실행 가능한지를 판정하는 물리적 의미론을 모두 외부화하고 검증 가능하게 만드는 것이다.

§14 · What remains unproven

End-to-end prospective validation은 아직 빈칸이다

2026년 9월 6일 현재 이번 검색에서는 Evidence → Skill → Verified Physical Workflow → Wet-Lab Result → Belief Revision 전체 루프를 실제 small-molecule drug discovery에서 prospective하게 검증한 새 연구는 확인되지 않았다.

Scientific Agent Skills는 163개 skill의 task-level effectiveness와 agent skill-selection accuracy를 아직 전면 검증하지 않았다. Computable laboratory 연구 역시 실제 장기간 drug-discovery DMTA에서 실패율, 재현성, 비용의 개선을 prospective하게 입증한 것은 아니다.

따라서 현재의 가장 정확한 결론은 다음과 같다. 완전한 Co-Scientist가 완성된 것이 아니라, 이를 가능하게 하는 procedural layer와 physical-execution verification layer가 각각 독립적인 연구대상으로 등장했다.

References

Sources in the Research Watch

아래 링크는 첨부 파일에 포함된 공식 논문·저장소만 사용했다.

01
A computable representation of the physical laboratory enables verifiable workflows
arXiv:2609.03621 · 03 Sep 2026

Typed research objects, capability-bound operations, compositional workflow algebra와 stateful simulation을 통해 laboratory workflow를 dispatch 전에 검증한다.

02
Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents
arXiv:2609.00065 · 04 Sep 2026

과학자가 작업을 올바르게 수행하기 위한 procedural knowledge를 163개의 versioned scientific skill로 외부화한다.

03
Scientific Agent Skills — Official Repository
GitHub · K-Dense-AI

Scientific Agent Skills의 공개 skill library 저장소.