AI scientist가 가설을 쓰고 코드를 실행하고 논문을 편집한다고 해서 과학을 다루는 것은 아니다. 연구의 출발점은 작업 목록이 아니라 관찰이다. 현미경 영상의 국소 형태, 세 채널 파형의 동시성, 시간에 따라 휘는 궤적, 3-D point cloud의 주축처럼 요약하기 전에만 존재하는 관계가 질문의 재료이기 때문이다.
OmniScientist는 기존 AI-scientist pipeline이 workflow는 거의 끝까지 자동화하면서도 text·code·label·precomputed scalar에 의존해 raw scientific evidence를 잃는 문제를 겨냥한다. perception layer와 ideation·experiment·writeup 3개 agent를 결합하고, stage boundary는 deterministic Python predicate로 검증한다.
workflow가 완성돼도 관찰 세계가 잘리면 질문은 빈곤해진다
문제는 모든 것을 token으로 바꿀 수 있느냐가 아니라, 바꾼 뒤 과학적으로 결정적인 관계가 살아 있느냐이다.
caption과 scalar feature가 지워 버리는 것
이미지·micrograph·spectrum·waveform·audio·video·3-D structure·trajectory는 공간적 인접성, 시간 순서, channel 간 동조, procedural sequence를 갖는다. caption은 local morphology를 놓칠 수 있고 unordered feature vector는 temporal order를 지우며 몇 개 scalar는 cross-channel inconsistency를 숨길 수 있다. 기존 agent가 인간이 고른 representation을 먼저 받으면 anomaly, hypothesis, claim space도 그 interface에 묶인다.
representation이 아니라 필요한 reasoning으로 분류한다
images, video, micrographs, radar, astronomy/remote sensing, plots, audio, 3-D.
documents, formulae, variables, rules, sequences, knowledge graphs, logical/causal relations.
tables, measurements, distributions, curves, correlation, significance, regression.
experimental steps, code execution, agent traces, simulations, protocols, dynamic evolution.
‘범용’이라는 말을 5개 discipline family와 36개 데이터셋으로 시험한다
task는 dataset·subject·target property·raw artifact만 주고 method는 agent가 선택한다.
12개의 equation에서 5,088,434-edge graph까지
PHYSICAL — NFFA-EUROPE 2,655 image; RRUFF 2,000 spectrum; UCI superconductor 21,263 table; PubChem 30 image; Feynman 12 formula. EARTH & SPACE — EuroSAT 5,000 image; Galaxy Zoo 1,000 image; GZ DECaLS 210 image; GWOSC 1,500 signal; STEAD 1,500 signal; WHOI-Plankton 3,000 image; Digital Rocks 375 3-D; SEVIR 384 video; IBTrACS 400 trajectory. LIFE & MEDICAL — Kather CRC 5,000 image; Chest X-ray 3,000 image; MedMNIST CT 1,496 3-D; CinC 2016 2,000 audio; Sleep-EDF 1,520 signal; Cell Tracking Challenge 280 video; DNA(H3) 10,000 sequence. AGRI & ECO — PlantVillage 3,002 image; Indian Pines 2,000 spectrum; CalMS21 2,500 image; Bird Audio Detection 2,000 audio; Watkins MMSD 1,697 audio; Pheno4D 223 3-D; Caltech Fish 120 video; White-stork GPS 49 trajectory. ENGINEERING & INFO — MCB 1,500 3-D; SemanticKITTI 1,200 3-D; comma2k19 3,000 image; MIMII 1,784 audio; NGSIM 209 trajectory; PDEBench 1,500 field; ogbl-biokg 5,088,434 edges graph.
36개 중 28개가 perceptual evidence다. 나머지 8개 symbolic·quantitative-statistical·procedural case는 visual inspection이 덜 중요한 breadth control이다. 새 discipline은 core engine을 바꾸지 않고 specification file 하나를 추가하는 방식으로 확장한다.
raw cue가 research question으로 변하는 16개 사례
Radiology는 patchy opacity를 보고 d=1.25의 patchiness와 held-out AUC 0.85 vs raw-pixel 0.63을 검증한다. Galaxy cross-survey는 같은 galaxy의 survey depth 차이를 보고 83.8% vs 81.0%, McNemar p=0.63, κ=0.75를 얻는다. Raman은 dominant band swap에 따라 rank order 0.77 vs 0.23(n=218)을 보고, seismology는 noise-labelled trace의 onset을 보고 21.7%(163/750), CI [18.8,24.9]의 coherent transient를 찾는다.
Ecoacoustics는 masking에 따라 detector AUC 0.73→0.58, meteorology는 cell merger가 +9.65 VIL-unit/5-min(p=3.2×10^-14), fisheries는 site가 range-density slope variance 81%(H=84.8,p<0.0001)를 설명한다고 보고한다. Plant 3-D에서는 maize가 동시에 2개 leaf를 시작하지 않지만 tomato는 94% event에서 그러고, CAD 1,500개는 11 morphotype(AMI=0.31)을 이룬다. Cyclone high-curvature period는 24h 뒤 slower intensification(p=2.4×10^-4), material family leave-out RMSE는 random k-fold보다 3.1–7.0× 높다. Symbolic regression은 Cramér–Rao slope -1.002 vs -1, R²=0.998, genomics는 9.5–11bp periodicity AUROC 0.55 vs shuffled 0.50, knowledge graph는 disease protein의 function diversity excess를 held-out edge에서도 확인한다.
agent는 자유롭게 탐색하지만 stage boundary는 코드가 결정한다
모든 것을 그림으로 바꾸지 않는다
signal·audio·video·3-D·trajectory에는 native numeric reader와 visual reader가 함께 있다. FFT peak, trend, duration, PCA axis 같은 native property를 먼저 읽고 spatial pattern이 중요할 때 rendering을 호출한다. visual inspection은 budget으로 제한한다.
세 agent가 하는 일
materials inventory → OpenAlex/Crossref prior-art search → ≥5 candidate → novelty risk/feasibility → falsifiable proposal. selected idea를 겨냥한 검색을 포함해 최소 3 focused search가 필요하다.
controlled run_python에서 design·execute·debug를 반복한다. primary, baseline, ablation, mechanism, breakdown, sensitivity, grouping-aware evaluation을 요구한다.
machine learning, biomedical, earth/space, physics, chemistry 5개 structural specification을 사용한다. 각 section은 experiment record의 필요한 slice만 받는다.
null/collapse 또는 exit-gate failure 때 outer pipeline이 다시 ideation으로 돌아갈 수 있으며 최대 2 fallback을 허용한다.
autonomy에는 explicit boundary가 있다
Ideation 최대 24 step·8 literature query, visual budget min(24,max(8,2g)); Experiment 최대 50 step·8 visual inspection, run_python 150초 timeout; generation call당 8,000 token이다. 완전한 run은 JSON record, Markdown summary, replayable trace, compiled PDF를 남긴다.
가장 중요한 장치는 agent가 아니라 ‘아직 끝낼 수 없다’고 말하는 predicate다
model self-critique가 아니라 Python predicate다
Idea check는 question·hypothesis·protocol·falsification, candidate breadth, prior art, fully computational feasibility, minimal claim, novelty evidence, claim scope, effective sample, leakage, visual audit를 검사한다. absolute novelty 표현 “first/never explored”도 거부한다.
Rigour check는 real execution, real data load, raw perception usage, key-number stdout provenance, ≥4 analysis battery, multiple-comparison correction, circularity, lead selection/significance, demotion을 본다. Claim check는 draft의 모든 number를 grounded record와 대조하고 polish가 number·citation·claim·model name을 바꾸면 revision을 통째로 되돌린다.
anti-fabrication과 anti-HARKing
if no successful run: reject for every reported number n: if n not in real stdout: reject if dataset not loaded from disk: reject if >=2 reported p-values: correct over ALL tests, including demoted ones if headline not in supported analyses: reject keep unsupported non-headline analyses in trace
한 stage는 Model(history,tools) → tool execution → finalize → deterministic gate 순으로 반복한다. gate가 실패하면 이유를 새 observation으로 append하고 계속하며 budget을 다 쓰면 outer pipeline이 실패 또는 backtrack한다.
가장 흔한 오류는 fabrication이 아니라 result selection이었다
36개 primary run에서 finalize attempt 115개가 거부됐다. Idea check 28회/23 run: schema 12, effective sample 4, novelty evidence 3, visual audit 3, claim scope 3, novelty language 2, breadth 1. Rigour check 87회/32 run: non-significant analysis demotion 51회/26 run, missing verdict 21, bad lead 9, no key numbers 3, provenance 1, no raw perception 1, multiple-test undercount 1이다.
36편을 만들었다는 사실보다 무엇을 고정하고 무엇을 바꿨는지가 중요하다
reasoning만 swap하고 perception은 고정한다
Reasoning backbone은 Sonnet 5(primary), GPT-5.6, GLM-5.2, Kimi K2.7, Qwen3.5 9B/27B/122B, Gemma-4 26B/31B다. perception model은 모든 run에서 Sonnet 5로 고정한다. judge는 deepseek-v4-flash와 gemini-2.5-flash-lite이며 novelty, soundness, clarity, significance, reproducibility, multimodal grounding, factual accuracy 7차원을 0–10으로 평가한다.
primary manuscripts completed
Sonnet full-suite mean overall score
paired perception head-to-head wins
reported run cost range
quality와 coverage
| Backbone | Overall | Factual | Cases→Completed | Mean composite |
|---|---|---|---|---|
| Sonnet 5 | 6.3 | 7.7 | 36→36 | 6.5 |
| GPT-5.6 | 5.6 | 7.7 | 10→9 | 5.7 |
| GLM-5.2 | 6.5 | 7.5 | 18→17 | 6.7 |
| Kimi K2.7 | 6.2 | 8.0 | 9→6 | 6.5 |
| Qwen3.5 122B/27B/9B | 5.1/5.1/4.0 | 6.5/6.4/4.8 | 34→30 / 36→32 / 32→18 | 5.4/5.3/4.1 |
| Gemma-4 31B/26B | 4.8/4.2 | 6.5/5.1 | 36→32 / 34→25 | 5.0/4.3 |
Table 3 Overall 6.3과 Table 4 mean composite 6.5는 다른 metric이다. 7차원 mean과 judge의 independent overall score를 섞으면 안 된다. dimension ordering은 backbone 사이 median pairwise Spearman 0.82이고 factual accuracy가 일관되게 가장 높은 축이다.
modality보다 backbone 차이가 더 크다
Sonnet modality composite는 image 6.4, signal 6.1, audio 7.1, video 6.4, 3-D 7.0, trajectory 6.3이고 discipline mean은 earth 6.5, life 6.7, agriculture 6.8, engineering 6.6, physical 5.8이다. case-level permutation test에서 cross-domain variation이 유의하지 않았다고 보고한다. 36-case all-discipline mean은 6.3/7.0/7.0/6.3/6.1/5.1/7.7/6.3이고 factual accuracy ≥7.0인 case는 30/36이다.
Judge validation은 Krippendorff α=0.66(사전 target >0.6), self-preference bias 0, verbosity bias ρ=0.16이다. 다만 이는 model-judge evaluation이며 human peer-review panel과 동일하지 않다.
raw evidence를 scalar feature로 바꾸면 연구 trajectory가 바뀐다
5개 blind pair에서 full perception이 85%의 comparison을 이긴다. 가장 큰 main gain은 MM grounding +2.8, significance +1.8이며 factual accuracy는 같은 provenance check 때문에 거의 동일하다. Appendix의 5 blind + 1 cardiology vision-off macro Δ는 novelty +1.14, soundness +0.75, clarity +0.72, significance +1.31, reproducibility +1.03, MM grounding +1.53, factual +0.72, overall +1.39다.
component ablation에서 prior-art search 제거는 composite 6.9→5.7로 가장 크게 떨어뜨리고 single-pass agentic loop도 큰 손실을 낸다. novelty check 제거는 novelty score를 1점 낮춘다.
score보다 중요한 것은 질문이 어떻게 달라졌는가이다
supported, mixed, refuted를 모두 남긴다
Supported: radiology patchiness, pathology compositional sub-clusters, cardiology recording-protocol confound(AUC 0.60→0.35 out-of-source), ecoacoustic recording-set degradation, unsupervised CAD morphotypes, Cramér–Rao exponent variance, disease-protein function diversity. Mixed: galaxy morphology 83.8 vs 81.0(비유의), remote-sensing color shortcut 76.2 vs 83.2, seismology 21.7% transient + instrument hypothesis refuted, species-specific leaf-initiation pattern, material leave-family-out RMSE 3.1–7×. Refuted: rejection-driven repair가 omission에 지배된다는 CS/ML trace hypothesis.
STEAD noise label을 감사하다
약 1,500개 3-component seismogram의 절반이 earthquake, 절반이 noise다. agent는 noise label에서 onset-and-decay envelope를 보고 “noise-labelled trace 중 coherent transient는 몇 개인가”를 묻는다. STA/LTA, amplitude, rectilinearity, planarity, cross-channel coincidence를 결합하고 3 label-agnostic surrogate null의 99th percentile을 threshold로 쓴다.
결과는 163/750=21.7%, 95% CI [18.8,24.9]. coincidence를 제거하면 2.0%로 amplitude-only baseline과 같아진다. null FAR 1.07%, sensitivity 19–25%, 417-station cluster bootstrap CI [17.9,25.8]%, BH/HH/HN prevalence 32.8/29.1/2.4%다. 사전 instrument-type hypothesis는 refuted되어 headline에서 내려간다.
pneumonia를 ‘밝기’가 아니라 ‘공간적 얼룩짐’으로 측정하다
pediatric chest radiograph를 직접 보고 dense patch와 clear region이 섞인 mottled texture를 관찰한 뒤 sliding-window local Shannon-entropy map의 spatial dispersion을 patchiness로 정의한다. effect size는 all d=1.25, development 1.26, held-out 1.29; Mann-Whitney p<0.0001, mean entropy와 독립 p=1.7×10^-5다.
held-out AUC는 raw pixel 0.634, mean entropy 0.840, patchiness 0.847, mean+patchiness 0.851. window 8/16/32 px의 d=1.63/1.50/1.36, ROI sweep d=1.18–1.35로 robustness를 확인한다.
범용성의 근거와 한계는 pipeline의 세부 규칙에 함께 적혀 있다
Ideation은 search/perception-heavy, Experiment는 code-heavy
36 trace에서 ideation 평균 8.8 step, 19.9 tool call, 8.7 search, 8.4 perception, 4.5 visual call, code 0이다. experiment는 36.0 step, 37.2 tool call, 31.8 run_python, 1.0 perception이다. verdict는 supported 16(44%), mixed 17(47%), refuted 2(6%), no verdict 1(3%). manuscript 36/36, rigour-exit experiment 35/36, full judge scoring 36/36이다. manuscript analysis 265(7.4/run), demoted trace 67(1.9/run)이다.
337 calls 중 173개가 visual channel
visual: look_at_image 49, signal 47, 3d 40, table 22, audio 17, video 13, trajectory 7. native: analyze_signal 40, audio 39, 3d 21, read_trace 20, trajectory 19, video 3. 모든 tool이 최소 한 번 사용됐다. modality를 무조건 image로 만들지 않는다는 설계 원칙이 실제 trace에도 나타난다.
새 분야는 method code가 아니라 specification으로 들어온다
Appendix E의 seismology spec은 role, subject(60-second E/N/Z, 100Hz), property, open request, member path와 dataset metadata만 담고 detector/hypothesis는 적지 않는다. Appendix F의 writeup skeleton은 ML: Introduction→Related Work→Method→Experiments→Conclusion→Limitations(abstract 150–220), Biomedical: Introduction→Results→Discussion→Methods(150–200), Earth/Space: Introduction→Data→Methods→Results→Discussion→Conclusions, Physics: Introduction→Theory and Methods→Results→Discussion→Conclusion, Chemistry: Introduction→Experimental Section→Results and Discussion→Conclusions(후자의 세 형식 150–250)이다.
prompt의 요구를 gate predicate로 다시 구현한다
Ideation prompt는 inspect materials → 3–6 focused literature search → observation에서 question → ≥5 projects → computational feasibility/power → novel+feasible selection → falsifiable proposal 순서를 요구한다. novelty는 “appears under-explored”처럼 제한하고 claim scope, statistical power, no-leakage를 명시한다.
Experiment prompt는 primary/baseline/ablation/mechanism/breakdown/sensitivity/group-disjoint evaluation을 요구하고 “every number must come from real run_python output”을 iron rule로 둔다. multiple comparisons는 실제 실행한 모든 test를 세고, weak result는 demote하며, nothing clears the bar이면 mixed/null을 정직하게 보고하도록 한다.
completion, review score, scientific truth는 같은 말이 아니다
Appendix cost는 Sonnet $2.63/paper·29min, GPT-5.6 $4.34·12min, local Qwen3.5-27B $0.06·30min, Gemma-4-31B $0.03·14min이며 experiment가 비용 대부분을 차지한다. 논문은 computationally executable project만 허용하고 physical experiment proposal은 ideation gate가 거부한다.
AI scientist의 다음 병목은 reasoning depth만이 아니라 epistemic interface일 수 있다
사람이 만든 label과 feature table만 받는 agent는 그 representation이 허용한 가설 공간 안에서만 똑똑해질 수 있다. OmniScientist의 기술적 핵심은 하나의 거대한 model보다 raw observation channel, 역할별 agent, replayable execution record, finalize를 거부하는 predicate, section별 evidence slicing이라는 작은 계약들의 결합에 있다.