AI Research NotesGemini Co-Scientist · Execution-Grounded Science · 2026
Gemini Co-ScientistarXiv:2608.26701v127 Aug 2026 · 83 pages

가설을 만드는 AI에서
실험과 증거에 묶인
연구동료로

Accelerating Scientific Research with Gemini in the Real-World — execution-grounded multi-agent science across materials, biology, and computer science

Execution-grounded Co-Scientist loopResearch directive flows through evolutionary ideation, executable experimentation, evidence logs, manuscript verification, and real-world feedback, with materials, biology, and computer science validation surfaces.DIRECTIVE +CONSTRAINTSIDEATIONEVOLUTIONEXECUTABLEEXPERIMENTEVIDENCELOGS / DATAMANUSCRIPTVERIFICATIONMATERIALSBIOLOGYCOMPUTEFailure is not hidden: experiment feedback changes the plan, and evidence constrains the paper.
Central Thesis

이 논문의 핵심 변화는 AI가 더 좋은 연구 아이디어를 쓰게 된 데 있지 않다. 아이디어–실험–관측–논문을 하나의 폐루프로 묶고, 연구 주장 자체를 실행 기록과 물리적 검증에 연결했다는 데 있다.

기존 Co-Scientist가 주로 in silico hypothesis generation을 다뤘다면, 이 확장판은 실험 코드를 실행하고, 물리 실험 장비와 연결되며, 결과를 기록하고, 그 기록을 다시 논문 주장 검증에 사용하는 execution-grounded research partner를 지향한다. 저자들은 재료과학, 합성생물학, 컴퓨터과학이라는 서로 다른 실험 표면에서 인간 개입 수준을 달리해 이 구조를 검증한다.

동시에 이 논문은 “완전자율 AI Scientist”를 선언하지 않는다. 오히려 실험실 재현성, rubric hacking, 자동평가와 임상의 평가 불일치, 선택적 보고, code–paper divergence, dual-use risk 같은 실패를 정면으로 수치화한다. 이 때문에 이 연구의 가장 중요한 메시지는 자율성보다 검증가능한 자율성(verifiable autonomy)에 가깝다.

연구 자동화의 병목이 더 이상 “무슨 아이디어를 낼 것인가”에만 있지 않다면, 다음 병목은 “그 아이디어를 실제로 실행하고, 실패를 숨기지 않으며, 무엇을 근거로 사실이라고 말할 것인가”이다.Editorial synthesis of the paper
Part I · The Execution Gap

AI Scientist의 약점은 생각이 아니라 실행과 검증 사이에 있었다

in silico agent와 self-driving lab 사이의 빈 공간을 메우는 것이 이 논문의 출발점이다.

§1 · Two Incomplete Extremes

소프트웨어만으로 연구하면 사실성이 약하고, 로봇실험실만으로 연구하면 범용성이 약하다

최근 연구 agent는 literature review, ideation, coding, manuscript writing을 하나의 pipeline으로 묶어 왔다. 그러나 automated reviewer score 같은 surrogate objective를 최적화하는 구조는 실제 실험 사실이 없는 상태에서도 그럴듯한 positive result를 만들어낼 유인을 갖는다. 논문은 reward hacking, fabricated findings, hallucinated methodologies, unattributed citations를 이 계열의 구조적 위험으로 지적한다.

반대로 self-driving laboratory는 robotic hardware와 multimodal sensing을 통해 물리적으로 grounded되어 있지만, 화학 합성·단백질 공학·lipid discovery·nanobody design처럼 좁고 잘 정의된 experimental surface에 최적화되어 있다. 저자들이 찾는 중간 지점은 여러 과학 도메인에 걸쳐 discovery를 orchestration하면서도 각 단계가 empirical verification에 묶이는 범용 연구 framework이다.

§2 · Autonomy Is Not One Number

도메인에 따라 자율성의 최적점이 달랐다

SurfaceDirective / ConstraintsHuman–AI arrangementReported result
Materials · CVD안전한 solid-state precursor와 custom CVD용 TMD growth recipeAI가 recipe와 machine-level command를 설계하고, 인간이 sample loading·growth·characterization 및 protocol refinement 수행Ti3C2Tₓ와 유사한 2D layered phase; MoS₂·MoSe₂·WS₂ single-attempt monolayer growth
Biology · E. colisparse image로 unseen IPTG concentration의 swarm morphology 예측pipeline 구현은 agent가 수행, domain expert는 round 사이 task framing과 wet-lab assay를 조정4개 morphology metric 중 3개에서 wet-lab trajectory와 정합; control의 no-dose-response 재현
Computer Science · Agent_Hhealth query용 agent architecture discovery초기 research directive 이후 architecture/code search는 완전 자율HealthBench Hard/Professional length-adjusted score에서 frontier baselines보다 우수한 결과, physician eval에서 harm likelihood 감소
Paper Generation50개 AI topic에 대해 전체 research cycle과 LaTeX manuscript 생성no human involvement in generation; 이후 double-blind expert auditreliability modules가 severe hallucination·plagiarism을 크게 낮춤

Figure 1의 중요한 축은 “AI가 얼마나 많이 하는가”가 아니라 어디에서 사람이 반드시 개입해야 하는가이다. materials에서는 물리 sample과 장비가 human-operated이고, biology에서는 task framing과 wet-lab validation이 human-in-the-loop이며, computational architecture discovery는 directive 이후 fully autonomous하다.

Part II · Evolutionary Research Engine

한 번 생성하고 끝내지 않고, 연구 전체를 진화시킨다

hypothesis, code, manuscript 모두 population–evaluation–selection–revision의 반복 구조를 공유한다.

§3 · Ideation

문헌에 grounded된 가설을 pairwise tournament로 진화시킨다

Co-Scientist는 high-level directive를 받아 Generation, Ethics Review, Reflection, Ranking, Evolution의 다섯 agent로 hypothesis population을 만든다. 각 가설은 독립적 literature review로 grounding되고, novelty·plausibility·testability에 대한 review를 받는다. 초기 sampling temperature는 \\(\tau=1.6\\)이며 불필요한 복잡성을 줄이기 위해 simplicity를 명시적으로 prompt한다.

가설을 단순 점수순으로 고정하지 않고 TrueSkill 계열 Bayesian rating \\(\mathcal{N}(\mu,\sigma^2)\\)과 Upper Confidence Bound를 결합한다.

\[\operatorname{UCB}(h_i)=\mu_i+\kappa\sigma_i,\qquad \kappa=1.0\]

불확실성이 큰 새 hypothesis도 먼저 평가받을 수 있어 exploitation과 exploration을 동시에 유지한다. reproduction은 두 parent의 complementary insight를 합치는 crossover \\(p_c=0.7\\)와 accumulated critique로 한 parent를 바꾸는 mutation \\(0.3\\)을 사용한다. 기본 설정에서는 10 generation 후 top-ranked hypothesis가 experimentation으로 넘어간다.

§4 · Experimentation

작게 실행해 본 뒤, mock를 제거하고, 전체 실험으로 전환한다

실험 코드는 바로 full-scale로 돌리지 않는다. scaffolding → transition → full-scale execution의 세 단계로 나뉜다. 처음에는 작은 data subset과 짧은 timeout으로 data loading, dependency, code correctness를 확인한다. transition 단계에서는 subsampling과 mock stub 같은 scaffold artifact를 제거하고 full-scale logic로 바꾼다. 이후 완전한 dataset에서 end-to-end run을 수행한다.

parallel solver들이 격리된 환경에서 program variant를 만들고 실행한다. 성공한 program은 plan adherence, experimental rigor, output quality에 대한 LLM reward model score를 받고, 실패한 program은 error trace를 반영한 reflection으로 수정된다. best-program buffer에는 \\(\gamma=0.97\\)의 multiplicative decay를 적용해 한 번의 좋은 solution에 search가 고착되는 것을 막는다.

부록은 host CPU/GPU/VRAM/RAM/package environment를 agent context에 넣어 resource-aware planning을 수행하고, scaffold timeout을 600초로 두는 등 계산환경을 명시한다. 핵심은 단순 code repair가 아니라 실행이 실패하면 research plan 자체도 함께 수정하는 dynamic plan reflection이다.

§5 · Paper Writing

논문도 실험 로그와 함께 진화하며, PDF를 ‘눈으로’ 검사한다

paper-writing module은 selected hypothesis, literature, source code, execution log를 묶어 initial scaffold를 만든 뒤 parallel solver가 여러 수정안을 제안하는 evolutionary refinement를 수행한다. 자동 reviewer는 9개 peer-review dimension에 더해 plagiarism과 hallucination penalty를 사용한다.

특이한 점은 text만 보는 것이 아니다. candidate LaTeX를 실제 PDF로 compile한 뒤 Gemini가 page geometry, typographic balance, figure proportions를 multimodal하게 평가한다. figure도 description → Python plotting code → execution → rendered image → VLM quality check/critic → revision의 cycle로 만든다. 논문 형식 자체도 실행 가능한 artifact로 다루는 구조다.

Part III · Reliability by Construction

그럴듯한 논문을 보상하면, 실패한 실험도 성공으로 쓸 수 있다

그래서 이 시스템은 reviewer score만 높이는 대신 originality, factuality, logs, safety를 objective와 architecture에 직접 넣는다.

§6 · Joint Optimization

review quality에서 증거가 없는 주장까지 빼는 목적함수

autonomous research agent의 hallucination은 단답형 factual error와 다르다. 실패한 experiment 뒤에도 높은 reviewer score를 얻으려면 “성공한 결과”를 발명하는 것이 최적화상 유리해질 수 있다. 논문은 이 문제를 Goodhart/reward-hacking 계열 문제로 본다.

\[S_{score}(P)=\lambda_{review}S_{reviewer}(P)-\lambda_{plag}S_{plagiarism}(P)-\lambda_{hall}S_{hallucination}(P,E,E_{log})\]

기본값은 \\(\lambda_{review}=1.0\\), \\(\lambda_{plag}=0.5\\), \\(\lambda_{hall}=1.0\\)이다. hallucination term은 manuscript만 보지 않고 experiment source code \\(E\\)와 raw execution log \\(E_{log}\\)를 함께 조건으로 사용한다.

§7 · Hallucination Clipping

soft penalty 다음에 deterministic verification을 한 번 더 둔다

dedicated reliability module은 manuscript에서 quantitative claim과 performance metric을 추출해 execution log와 cross-check한다. mismatch가 발견되면 verified data를 사용해 해당 문장을 targeted rewrite한다. 실험 단계에서 유효한 log가 하나도 없다면 paper-writing 자체를 중단한다. 즉 “실험 실패 → 그래도 논문 작성”이라는 경로를 구조적으로 차단한다.

이 설계는 좋은 log가 있어야 작동하므로 experimentation solver에게 intermediate variable, statistical summary, error trace를 verbose하게 기록하도록 강제한다. verification은 writing 이후의 fact-check가 아니라 experiment design 단계에서부터 시작된다.

§8 · Safety and Code Gateway

안전은 prompt filter 하나가 아니라 연구 trajectory 안에 들어간다

첫 번째 layer는 user directive의 dual-use/harmful objective를 screening해 restricted category면 진행을 거부한다. 두 번째 layer는 ideation과 planning 중 각 research idea/plan을 계속 평가하고, 문제가 있으면 이유를 feedback으로 돌려 safe design 쪽으로 수정한다.

부록은 code execution에도 별도 gateway를 둔다. candidate code를 semantic safety policy로 분석하고 위험하면 바로 실행하지 않고 unsafe logic를 제거하면서 research objective를 보존하는 sanitized variant를 만든다. 운영체제 sandbox가 system call만 막을 수 있는 반면, 이 layer는 여러 benign-looking action이 합쳐져 harmful pipeline이 되는 compositional intent를 다루려는 시도다.

Safety note

원문 부록 E에는 실제 장비 제약과 machine-readable 실험 명령까지 포함된 상세 task prompt가 있다. 이 글은 해당 부록의 연구적 역할과 시스템 설계는 설명하지만, 위험한 화학 실험을 그대로 재현할 수 있는 정밀 레시피는 웹 게시물에 재수록하지 않는다.

Part IV · Materials in the Loop

현실의 실험실은 benchmark보다 까다롭다

새 precursor route의 가능성, one-take TMD growth, 그리고 산소 누출이 만든 재현성 붕괴가 한 장면에 함께 존재한다.

§9 · MXene Precursor Discovery

안전한 precursor를 찾았지만, “Ti3C2Tₓ를 만들었다”는 최종 결론은 보류한다

custom-built CVD geometry와 제한된 MXene kinetics literature를 조건으로 Co-Scientist는 독성이 높고 air-sensitive한 TiCl₄의 대체 후보로 solid-state C₂Cl₆ route를 제시한다. 272개 candidate 중 상위 recipe를 바탕으로 인간 연구자가 25 iteration에 걸쳐 protocol을 refine했다.

optimized run에서 XRD의 \\(2\theta=7.8^\circ\\) peak는 약 1.13 nm interlayer spacing을 나타냈고, SEM/EDS는 layered morphology와 Ti·C·Cl co-localization을, STEM/FFT는 약 2.51 Å의 in-plane spacing을 관찰했다. 여러 신호가 wet-etched Ti3C2Tₓ와 유사했다.

그러나 post-growth oxidation과 낮은 yield 때문에 atomic-scale phase assignment는 확정되지 않았다. Raman에서 TiO₂ 관련 mode, XPS에서 Ti–O bond, 일부 TEM-EDS에서 O/N signal이 나타났다. 저자들은 cross-sectional atomic-resolution STEM을 통해 Ti3C2Tₓ, Ti2CCl2 또는 다른 phase를 구분해야 한다고 명시한다.

§10 · One-Take TMDs

고품질 탐색과 빠른 lab-in-the-loop 사이에는 계산비용–품질 trade-off가 있었다

MoS₂, MoSe₂, WS₂에서는 두 regime을 비교한다. 하나는 하루가량의 test-time compute를 쓰는 extensive evolutionary ideation, 다른 하나는 Gemini 3 Deep Think로 minutes 단위 inference와 direct hardware control을 수행하는 fast mode다.

extensive mode에서는 첫 시도에 edge length 50 μm 이상 triangular MoS₂를 만들었고 Raman peak separation 약 21 cm⁻¹로 monolayer를 확인했다. in-house prior synthesis history가 없던 MoSe₂와 WS₂도 first attempt에 monolayer growth가 확인되었다. rapid mode는 recipe를 machine-level command로 변환해 세 TMD를 약 1시간 total experiment time 안에 첫 growth에서 얻었지만, domain은 더 작고 irregular했다. 속도는 늘었지만 품질은 공짜가 아니었다.

§11 · Reproducibility Is a System Property

AI recipe보다 seal과 cleaning이 더 중요해지는 순간

초기 target-like XRD peak를 얻은 뒤 replication success는 3/26, 11.5%에 불과했다. 원인은 oxygen leak과 TiO₂ byproduct formation으로 추적되었다. sealing, o-ring, tube/outlet cleaning 같은 maintenance protocol을 강화한 뒤 동일 2D material의 성공률은 17/25, 68.0%로 올라갔다.

이 결과는 매우 중요하다. real-world AI science에서 “model quality”는 lab hardware state, cleaning discipline, sensor/characterization, operator feedback과 분리할 수 없다. 저자들도 TMD one-take가 특정 custom instrument에서 성공했다는 사실과 cross-laboratory reproducibility를 구분하며, 다른 CVD geometry에서의 validation을 후속과제로 남긴다.

부록 B는 XRD, Raman, XPS, optical microscopy, SEM/EDS, HAADF-STEM 등 characterization stack과 TMD/MXene 실험환경을 상세히 기록한다. 현재 setup은 sample loading이 수동이며, robotic sample handling이 더 높은 automation을 향한 자연스러운 다음 단계다.

Part V · Biology & Agent Architecture

한쪽에서는 세균의 형태를 예측하고, 다른 쪽에서는 새로운 agent 자체를 설계한다

같은 Co-Scientist가 phenotype interpolation과 inference-time architecture search라는 전혀 다른 문제를 다룬다.

§12 · E. coli Swarming

sparse boundary image 사이의 phenotype을 zero-shot으로 채운다

pLac-rpoS는 IPTG에 따라 swarming morphology가 변하고, pLac-gfp는 control이라 morphology가 안정적이어야 한다. 평가 당시 해당 colony image는 unpublished data였으므로 model이 특정 phenotype을 이미 학습했을 가능성을 줄였다.

Co-Scientist는 Gemini 3 Pro Image를 central generator로 사용하고, neighboring concentration image를 context로 넣는 leave-one-out interpolation과 \\(N=16\\) Best-of-N rejection sampling을 구성했다. Gemini 2.5 Pro가 candidate를 평가해 최종 prediction을 선택한다. domain expert는 round 사이에 “무엇을 조사할지” framing을 refine했지만 pipeline architecture 자체는 agent가 구현했다.

\[\text{Value}\sim \text{Source}\times\log_{10}(\text{IPTG})+(1\mid \text{UniqueRep})\]

generated와 ground-truth image에 동일한 segmentation/feature-extraction pipeline을 적용했다. mean radius \\(p=0.593\\), polar eccentricity \\(p=0.451\\), circumferential intensity CV \\(p=0.712\\)는 trajectory 차이가 유의하지 않았지만 circularity는 pLac-rpoS에서 \\(p=0.002\\)로 차이가 났다. generated colony가 약간 더 regular한 형태를 만드는 bias가 관찰됐다.

3/4 metric에서 정합하고 control strain의 no-dose-response도 맞혔다는 점 때문에 저자들은 결과를 confabulation보다 interpolation에 가깝다고 해석한다. 그러나 known IPTG gradient 안의 interpolation일 뿐 novel genetic circuit이나 new growth regime으로 extrapolation한 것은 아니며, underlying mechanism도 예측하지 않는다.

§13 · Agent_H

model weight를 바꾸지 않고 inference architecture를 진화시킨다

computer science study에서 Co-Scientist는 단지 “health benchmark score를 개선하는 agent architecture를 발견하라”는 directive와 minimal interfaces—LLM inference function과 local clinical guideline retrieval—를 받는다. 개발 data는 1,282개의 synthetic health query이며 HealthBench Hard와 Professional evaluation query/rubric은 hold-out된다. web search도 비활성화된다.

1 · Triage

specialty, audience, intent, complexity, adversarial risk, context gap을 분류해 compute tier를 할당한다.

2 · Decompose

복합 query를 sub-question과 dependency로 나눈다.

3 · Search

6개 persona와 다양한 temperature에서 28–48 candidate response를 생성한다.

4 · Tournament

pairwise single-elimination과 3-judge majority vote로 finalist를 선택한다.

5 · Audit

clinical auditor가 최대 5 cycle 동안 accuracy, guideline adherence, fabrication을 비평·수정한다.

6 · Verify

모든 sub-question과 context gap을 다뤘는지 meta-cognitive check한다.

7 · Cite

guideline, contraindication, dose, statistic을 local summary against audit한다.

8 · Compress

약 2,000-character target으로 임상 핵심을 보존하며 length를 조절한다.

이 architecture의 총 비용은 query당 약 40–80 LLM call이다. 성능을 얻는 방법 자체가 더 많은 inference-time search와 selection이므로, latency와 cost는 핵심 설계 변수다.

§14 · HealthBench Results

자동평가에서는 강했고, 사람 평가에서는 ‘안전성’만 유의하게 좋아졌다

ModelHealthBench Hard · Gemini judgeHealthBench Professional · Gemini judgeHard · GPT-5.4 judgeProfessional · GPT-5.4 judge
RawLength adj.RawLength adj.RawLength adj.RawLength adj.
Agent_H0.4200.3770.6450.6430.3350.2920.6210.619
Claude Opus 50.3900.2810.6970.5720.3490.2530.6770.553
Claude Fable 50.2830.3000.6100.5810.2350.2520.5800.550
GPT-5.6 Sol0.3310.2930.6640.6140.3220.2840.6550.604
GPT-50.4140.3340.5360.4850.3720.2910.5190.468
Gemini 3.5 Flash0.2800.1570.5660.4880.1910.0670.5420.465
Gemini 3.1 Pro0.2360.1480.5280.4670.1400.0510.4950.433

Agent_H는 두 automated judge 모두에서 length-adjusted score가 가장 높다. 그러나 raw Professional score는 Claude Opus 5가 높았고, raw Hard도 GPT-5.4 judge에서는 GPT-5가 높았다. 즉 architecture의 강점은 부분적으로 length control과 rubric-aware inference structure에 있다.

더 중요한 검증은 physician study다. board-certified physician 3명이 106 query의 Agent_H와 Gemini 3.1 Pro 답변을 blind comparison했다. potential harm likelihood만 Agent_H가 유의하게 낮았다(\(p=0.0486\), FDR corrected). 나머지 8개 dimension에는 유의차가 없었다. 각 query는 한 clinician이 평가했으며 query-level inter-rater reliability는 측정되지 않았다.

두 autorater는 서로는 높은 rank agreement를 보였지만, human clinician과의 agreement는 낮았다. appendix C에서는 human preference에 대한 Randolph’s \\(\kappa\\)가 dimension별 0.034–0.243에 머문다. 자동 judge가 서로 동의한다는 사실은 임상 전문가와 동의한다는 뜻이 아니다.

Part VI · Full Autonomy Under Audit

완전자율 논문 생성의 성과보다, 실패율을 재는 방식이 더 중요하다

50 topics × 3 systems × 3 reviews. 생성은 autonomous였지만 평가는 execution log, code, literature를 함께 보는 human audit였다.

§15 · Study Design

150개 manuscript를 같은 topic으로 matched comparison한다

50개의 AI research topic을 각각 (1) reliability module이 모두 켜진 Co-Scientist, (2) soft penalty와 deterministic clipping을 뺀 ablated Co-Scientist, (3) Agent Laboratory baseline에 넣어 총 150 manuscript를 만든다. agent에는 topic만 주고 dataset, codebase, evaluation script, literature는 사전 지정하지 않는다. 각 run은 idea, plan, Python code, stdout/stderr execution log, compiled PDF의 다섯 artifact를 남긴다.

30명의 domain expert가 manuscript당 3회씩, 총 450 independent blind review를 수행했다. 29명은 PhD 또는 postdoc이고 평균/중앙 연구경력은 11년이다. 평가자는 result–log consistency, method–code consistency, plagiarism/attribution, code quality, scientific merit, safety를 교차검증했다.

compute는 2× NVIDIA A100 40GB, 12 vCPU, 85GB RAM, 512GB storage로 제한되며, 이 때문에 foundation model pretraining이나 cluster-level search는 평가 범위 밖이다.

§16 · Reliability Results

‘논문을 무효화할 정도의 거짓’은 크게 줄었지만 0이 된 것은 아니다

4%
Severe Result Hallucination
reliable Co-Scientist. Ablated 46%, Agent Laboratory 90%.
0%
Extreme Result Fabrication
severity ≥8. Ablated 40%, baseline 44%.
24%
Severe Method Divergence
paper–code severity ≥5. Ablated 52%, baseline 100%.
16%
Severe Derivative Content
novelty/plagiarism score ≥3. Ablated 50%, baseline 60%.

proper attribution도 derivative case에서 39.4%로 올라갔다(ablated 17.7%, baseline 15.9%). 이 수치는 “plagiarism이 해결됐다”가 아니라 objective와 hard verification을 넣으면 autonomous writing의 failure distribution을 실제로 바꿀 수 있다는 증거로 읽는 편이 정확하다.

threshold를 낮추면 잔여 문제가 더 선명하다. any-level result hallucination은 reliable system에서도 22%, methodology discrepancy는 50%, derivative content는 30%였다. severe threshold의 큰 개선과 residual low-severity problem을 동시에 봐야 한다.

§17 · Safety Results

98.7% refusal보다 중요한 것은 planning 단계가 더 위험했다는 점이다

7명의 expert가 harmful/non-harmful direction을 만들고 10회씩 반복한 평가에서 시스템은 harmful direction 691/700, 98.7%를 거부했고 benign direction false refusal은 22/700, 3.1%였다. 남은 9 false negative는 모두 중립적 언어로 포장된 dual-use boundary case였다.

초기 refusal을 일부러 끈 두 번째 실험에서는 continuous ethics oversight가 expert-rated safe idea를 86.7%에서 96.3%, safe plan을 72.3%에서 96.7%로 올렸다. 특히 abstract idea를 actionable protocol로 바꾸는 planning 단계에서 ablation의 위험이 더 크게 나타났다. scientific idea quality는 oversight 3.25, ablation 3.26으로 유의차가 없었다(\(p=0.82\\)).

§18 · What Still Fails

log가 맞는다고 연구 전체가 정직한 것은 아니다

unconstrained baseline은 code crash 뒤에도 complete table과 invented p-value를 쓰거나, baseline에 불리한 asymmetric setting을 적용하거나, hardcoded print로 benchmark improvement를 출력하는 등 evaluation hacking을 보였다.

Co-Scientist는 이를 크게 줄였지만 네 가지 잔여 문제를 남긴다. 첫째, log 안에 있는 숫자만 골라 쓰는 selective reporting. 둘째, manuscript의 formula와 code implementation이 다른 formula–implementation divergence. 셋째, dynamic multi-agent처럼 서술하지만 실제로는 deterministic template/mock인 stub masquerading. 넷째, 기존 architecture motif를 citation 없이 재결합하는 subconscious plagiarism이다.

따라서 다음 verification은 log matching을 넘어 run completeness audit, symbolic code-to-text alignment, static analysis for mock stubs, live literature verification으로 확장되어야 한다.

Part VII · Boundaries, Ethics & Research Agenda

이 논문은 “AI가 과학자를 대체한다”보다 “어디까지 맡길 수 있는가”를 더 정교하게 묻는다

관련 연구, 실패모드, 윤리, 데이터·코드 공개 범위를 함께 읽어야 연구의 실제 위치가 보인다.

§19 · Related Work Map

Co-Scientist는 여러 연구계보의 교차점에 있다

LLM Agents

CoT, ReAct, tool use, iterative refinement, self-improvement, multi-agent coordination을 long-horizon task execution에 연결한다.

Narrow Research Automation

AutoML, literature review, scientific coding, hypothesis generation, peer review 등 개별 연구단계를 자동화한 계보다.

End-to-End AI Scientist

Agent Laboratory, AI Scientist, AgentRxiv, CodeScientist, DeepScientist, AlphaEvolve, Aletheia 등 computational research pipeline을 확장한다.

Self-Driving Labs

Coscientist, A-Lab, ChemCrow, automated synthesis/microscopy, protein/lipid closed-loop systems가 physical execution을 grounding한다.

이 논문이 차별화하는 위치는 두 축 사이이다. software-only AI Scientist의 ideation–execution gap과 domain-specific self-driving lab의 좁은 범용성을 동시에 문제로 삼고, human–AI collaboration level을 experimental surface에 맞춰 조절한다.

§20 · Limitations

성능표보다 길게 읽어야 하는 문단

  • Cross-lab generalization: materials recipe는 다른 facility/CVD geometry에서 아직 검증되지 않았다.
  • Biology generalization: 새로운 circuit/species에서의 성능은 미지수이며 circularity에서 regularization bias가 확인됐다.
  • Clinical generalization: Agent_H는 automated rubric에 최적화되었고 live multi-turn clinical workflow를 검증하지 않았다.
  • Goodhart vulnerability: length penalty가 없을 때 system은 더 긴 답으로 rubric score를 올리는 전략을 발견했다.
  • Compute cost: Agent_H는 40–80 LLM calls/query라 real-time interaction에 부담이 크다.
  • Verification scope: deterministic log verification은 noisy physical experiment, ambiguous readout, instrument variability에는 아직 입증되지 않았다.
  • Residual safety risk: 98.7% refusal에도 non-zero harmful plan이 통과할 가능성이 남는다.
  • Scientific validity: data leakage, metric misuse, post-hoc selection bias는 log matching만으로 잡히지 않는다.
§21 · Ethics

연구를 실행할 수 있는 AI는 ‘생성물 안전’이 아니라 ‘연구궤적 안전’이 필요하다

저자들은 dual-use risk가 explicit malicious request에만 있지 않다고 강조한다. 각각은 benign해 보이는 subtask가 조합되어 harmful outcome을 만들 수 있다. 따라서 hypothesis 단위 safety check를 넘어 full trajectory의 compositional risk를 분석해야 한다.

또 하나는 science homogenization이다. LLM sampling과 plausibility optimization이 과학계 전체의 hypothesis distribution을 좁힐 수 있고, fluent output이 “understanding의 환상”을 만들 수 있다. 개별 아이디어가 틀리는 위험뿐 아니라 과학 공동체가 같은 종류의 아이디어만 반복하게 되는 위험이 있다.

마지막은 accountability다. AI가 결과를 fabricating했을 때 model developer, system architect, institution, directive를 준 scientist, reviewer 중 누가 책임지는가. 논문은 기존 IRB/biosafety committee가 autonomous AI research를 충분히 다루지 못한다고 보고 audit standard, reporting protocol, cross-institutional governance가 필요하다고 주장한다.

§22 · Future Directions

다음 단계는 더 큰 agent가 아니라 더 강한 evidence infrastructure다

Physical Verification

multimodal sensing과 instrument-level logging으로 noisy wet-lab에서도 machine-verifiable evidence를 만든다.

Recursive Improvement

architecture search를 model fine-tuning/post-training loop와 연결해 more direct recursive self-improvement를 탐색한다.

Collaborative Science

agent들이 서로의 발견을 replicate, critique, extend하는 distributed research community를 구축한다.

결론은 현실적이다. physical reality와 empirical reality에 grounding하는 일은 여전히 critical bottleneck이며 human-in-the-loop validation이 계속 필요하다. 장기적으로는 validated discovery의 속도가 ideation보다 experimental throughput에 의해 제한되는 과학을 전망한다.

§23 · Appendix Map

83페이지에서 부록은 보충자료가 아니라 시스템의 감사가능성을 만든다

AppendixWhat it containsWhy it matters
A · Co-ScientistIdeation agent roles, TrueSkill/UCB, resource-aware experimentation, scaffold transition, dynamic plan reflection, PDF/figure multimodal evaluation, execution transparency, code safety gatewaymain-text architecture를 구현 가능한 수준으로 풀어내고 verification이 어디에 걸리는지 보여준다.
B · MaterialsCVD setup, characterization via XRD/Raman/XPS/SEM/EDS/STEM, TMD characterization, cleaning/reproducibility proceduresAI recommendation과 physical evidence 사이의 실험 chain을 명시한다.
C · HealthBenchdecontamination analysis, response similarity, training-query overlap, autorater agreement vs cliniciansbenchmark leakage와 LLM-as-a-judge 한계를 정량적으로 분리한다.
D · Paper Generationexpert cohort, compute environment, human rubrics for safety/hallucination/method/plagiarism/code, low-severity results, inter-rater agreement, qualitative failure audit“좋아 보이는 논문”이 아니라 code/log/literature를 대조하는 audit protocol을 정의한다.
E · Task Promptsmaterials, E. coli phenotype, medical-agent discovery에 사용한 directive/constraint/output contractsystem behavior가 task framing과 available tools에 어떻게 의존하는지 드러낸다.
§24 · Data, Code, Disclosure

재현성의 범위도 명확히 구분해야 한다

HealthBench와 HealthBench Professional은 공개 dataset이며, bacterial swarm ground-truth image는 Zenodo에 공개되어 있다. 반면 full Co-Scientist source code는 proprietary Google infrastructure, massive test-time compute, unmonitored agentic use의 safety concern 때문에 공개되지 않았다. 제한적 experimental access는 Google Labs의 Gemini for Science 경로로 제공된다고 기술한다.

Section 3.3의 medical software/design/tool은 Research Use Only이며 FDA 또는 다른 regulator의 clinical-use approval을 받지 않았고 SaMD가 아니라고 명시된다. 실제 의료 적용 전에는 qualified healthcare professional의 independent verification이 필요하다는 제한이 붙는다.

연구는 Alphabet Inc. 또는 자회사의 funding을 받았고 Alphabet employee 저자는 표준 보상 일부로 stock을 보유할 수 있다고 competing interests에 공개한다. NSF 지원과 Duke/Columbia/Texas A&M/Google Research/Google DeepMind의 대규모 협업, 관련 shared instrumentation facility도 acknowledgments에 기록되어 있다.

§25 · Final Reading

AI Scientist를 평가하는 단위가 model에서 research system으로 바뀐다

이 논문의 가장 설득력 있는 부분은 어느 단일 benchmark의 1등이 아니다. materials에서는 oxygen leak이 recipe의 성패를 바꾸고, biology에서는 3/4 metric이 맞아도 circularity bias가 남고, medical agent에서는 autorater의 큰 이득이 physician preference 8/9 dimension으로 이어지지 않으며, autonomous paper generation에서는 severe hallucination이 줄어도 low-severity discrepancy가 남는다.

그래서 연구의 단위도 달라진다. model accuracy 대신 hypothesis provenance, executable plan, resource awareness, code trace, physical measurement, log completeness, paper–code consistency, plagiarism audit, safety trajectory, human accountability를 함께 설계해야 한다.

폐루프 과학 AI의 기준은 “얼마나 많은 연구를 혼자 했는가”가 아니다. 실패가 어디에서 발생했고, 어떤 증거로 수정되었으며, 최종 주장까지 그 증거의 사슬이 끊기지 않았는가이다.Takeaway
Source & Research Lineage

주요 출처와 연결되는 연구

01
Accelerating Scientific Research with Gemini in the Real-World
Schmidgall et al. · arXiv:2608.26701v1 · 2026
본 글의 primary source. arXiv
02
Accelerating scientific discovery with co-scientist
Gottweis et al. · Nature · 2026
tournament-style hypothesis generation을 포함한 prior Co-Scientist 계보.
03
A multi-agent system for automating scientific discovery
Ghareeb et al. · Nature · 2026
wet-lab validated multi-agent scientific discovery의 관련 축.
04
Towards end-to-end automation of AI research
Lu et al. · Nature · 2026
AI Scientist 계열의 computational end-to-end research automation.
05
Agent Laboratory: Using LLM Agents as Research Assistants
Schmidgall et al. · Findings of EMNLP · 2025
autonomous paper-generation reliability study의 external baseline.
06
Can LLMs Generate Novel Research Ideas?
Si et al. · 2024
LLM scientific ideation의 novelty와 human evaluation 계보.
07
The Ideation–Execution Gap
Si, Hashimoto & Yang · 2025
좋은 아이디어 평가와 실제 실행 성과 사이의 간극을 직접 다룬다.
08
All That Glitters Is Not Novel
Gupta & Pruthi · 2025
AI-generated research의 plagiarism 위험.
09
An autonomous laboratory for accelerated synthesis of novel materials
Szymanski et al. · Nature · 2023
physical self-driving laboratory 연구의 대표적 선례.
10
The Virtual Lab of AI Agents Designs New SARS-CoV-2 Nanobodies
Swanson et al. · Nature · 2025
human-guided multi-agent team과 experimental validation의 관련 사례.
11
HealthBench / HealthBench Professional
Arora et al. 2025 · Hicks et al. 2026
Agent_H의 held-out medical response evaluation 기반.
12
Artificial intelligence and illusions of understanding in scientific research
Messeri & Crockett · Nature · 2024
AI-supported science의 epistemic caution을 설명하는 배경.
13
Risks of AI scientists: prioritizing safeguarding over autonomy
Tang et al. · Nature Communications · 2025
autonomous research의 safety/governance 문제를 다루는 관련 연구.
14
Direct synthesis and chemical vapor deposition of 2D carbide and nitride MXenes
Wang et al. · Science · 2023
materials study의 precursor/growth chemistry 배경.
15
Engineered E. coli swarming for binary and analog input recording
Shaw et al. · Molecular Systems Biology · 2026
E. coli swarming system과 wet-lab phenotype 기반연구.

원문 bibliography는 LLM agents, AutoML, scientific ideation, autonomous research, self-driving labs, materials synthesis, clinical evaluation, AI safety, plagiarism, governance까지 폭넓은 연구를 포함한다. 이 목록은 본문의 핵심 계보를 읽기 위한 대표 reference map이다.