AI Research Blog·Operational Einstein Test2026 · 09 · 18
Einstein Test Research Update·Prospective Evaluation·AI for Science

과거의 발견을 재현시키는 시험에서
미래의 과학이 AI를 검증하는 시험으로

Prospective Einstein Test 2026: Sealed Hypotheses, Evidence-Centered Discovery, and Future Validation

MODEL CUTOFFSEALED HYPOTHESISFUTURE EVIDENCEINDEPENDENT CHECK미래정보 차단가설 봉인후행 실험·임상사후 독립검증NO FUTURE LEAKAGE · SEAL BEFORE UNBLINDINGTemporal Isolation → Evidence → Revision → Sealed Prediction → Future Validation
Central Update

이번 업데이트의 핵심은 Einstein Test 자체의 새 버전이 아니라 “미래를 보지 못한 상태에서 AI의 가설을 먼저 봉인하고, 이후 실제 과학이 그 가설을 검증하는” prospective evaluation이 실세계 biomedical science와 evidence-centered agent harness를 통해 더 현실적인 방법론으로 보이기 시작했다는 점이다.

새로운 변화는 두 건이다. 첫째, Virtual Biotech의 B7-H3 사례는 model cutoff < AI hypothesis < later clinical evidence라는 시간격리 구조를 실사용 연구에 적용했다. 둘째, TruthInsight는 동일한 blind benchmark에서 evidence memory, reviewer, control-selection policy를 갖춘 scientific harness가 단순 coding agent의 plateau를 크게 넘어설 수 있음을 보고했다.

Status직접적인 Einstein Test, ARC-AGI, SCILAWS-BENCH, Singularity Gate의 중대한 신규 버전 변경은 이번 점검에서 확인되지 않았다.

Part I · Evaluation Shift

“과거의 정답”을 재현하는 것만으로는 발견을 증명하기 어렵다

temporal isolation의 목적은 기억된 정답과 genuine discovery를 분리하는 것이다. 이제 그 격리를 corpus 수준에서 미래 결과 수준으로 더 강하게 밀어붙일 수 있다.

§1 · Historical Backtesting vs Prospective Validation

Einstein Test의 검증축이 과거 재현에서 미래 대조로 이동한다

고전적인 역사적 backtesting은 “1911년 이전 corpus로 상대성이론을 다시 발견시킬 수 있는가”와 같은 구조를 취한다. 하지만 foundation model의 pretraining provenance를 완전히 감사하기 어렵다면, 이미 알려진 발견을 재현하는 평가에는 contamination 의심이 남는다.

prospective test는 정답이 아직 공개되지 않은 시간축을 평가 장치로 사용한다. AI가 hypothesis를 생성하고 그 결과를 봉인한 뒤, 미래에 나타난 실험·임상·관측 결과와 대조한다.

\[T_{\mathrm{model\ cutoff}} < T_{\mathrm{AI\ hypothesis}} < T_{\mathrm{future\ human/clinical\ evidence}}\]

이 구조는 model weight의 완전한 audit를 대신하지 않는다. 그러나 적어도 특정 후행 evidence를 AI가 분석 시점에 볼 수 없었다는 강한 temporal boundary를 제공한다.

Part II · Virtual Biotech

실세계 biomedical science에서 out-of-time 검증을 시험하다

37,000개 규모의 clinical-trialist agent 조직과 virtual CSO가 임상·유전·single-cell·spatial evidence를 연결하고, 일부 결과를 미래 임상 readout과 비교했다.

§2 · Multi-Agent Therapeutic Discovery

55,984개 임상시험에서 발견한 패턴과 B7-H3 메커니즘 가설

Source factVirtual Biotech는 최대 37,000개 clinical-trialist agent와 virtual Chief Scientific Officer를 포함하는 multi-agent 연구조직으로 55,984개 임상시험을 분석했다. 소스가 정리한 결과에 따르면 특정 cell type에 선택적인 target을 겨냥한 약물은 Phase I→II 진입 가능성이 40% 높고, 시장 도달 가능성이 48% 높으며, adverse event가 32% 낮다는 연관성을 도출했다.

37,000
Clinical-Trialist Agents
최대 agent 규모
55,984
Clinical Trials
분석 대상 임상시험
+40%
Phase I → II
source-reported association
+48% / −32%
Market / AE
시장 도달 / adverse event

B7-H3 폐암 사례에서는 통계유전학, single-cell, spatial transcriptomics, clinical data를 통합해 fibroblast 중심의 면역배제 메커니즘과 ADC 전략을 제안했다. 특정 분석에서는 web search/fetch를 차단했고 기반 모델의 knowledge cutoff를 2025년 1월로 두었다. 이후 2025년 8월 공개된 ifinatamab deruxtecan 임상결과가 방향상 독립적인 후행 지지를 제공했다.

§3 · Evidence Boundary

강한 prospective signal이지만 완전한 Einstein Test는 아니다

Analysis이 사례의 강점은 cross-modal evidence integration → abductive mechanism hypothesis → future/out-of-time corroboration의 연결을 실제 biomedical domain에서 보여준다는 점이다. historical answer reconstruction보다 contamination에 강한 구조다.

Limitation연구질문과 B7-H3 target은 인간이 제공했다. foundation-model weight를 직접 감사하지 않았고, AI가 제안한 모든 메커니즘이 wet-lab으로 검증된 것도 아니다. 따라서 autonomous problem discovery나 완전한 Einstein-level discovery를 입증한 결과로 해석해서는 안 된다.

Virtual Biotech의 가장 중요한 의미는 “AI가 과학을 끝까지 자동화했다”는 데 있지 않다. AI의 가설을 미래정보가 차단된 시점에 봉인하고, 이후 실제 임상세계가 일부 가설을 사후 검증할 수 있다는 평가 패턴을 실세계 연구에 연결했다는 데 있다.Source-grounded interpretation
Part III · TruthInsight

발견 점수보다 evidence process를 설계한다

동일한 blind task라도 evidence memory, artifact verification, semantic adjudication, reviewer routing을 갖춘 harness가 scientific reliability를 크게 끌어올릴 수 있다.

§4 · Evidence-Centered Harness

question·method·outputs·sources·uncertainty를 하나의 research record로 묶는다

Source factTruthInsight는 기존 TruthInsightBench의 40개 blind task·10개 분야를 그대로 사용하면서 단순 coding agent 대신 evidence-centered scientific harness를 적용했다. 각 연구질문에 question, method, outputs, sources, uncertainty를 연결한 evidence-linked research record를 유지하고 artifact verification과 semantic claim adjudication을 분리한다.

reviewer는 “현재 분석을 더 확장할지, 새 질문을 만들지, 보고서를 작성할지”를 동적으로 결정하는 review-guided evidence routing을 사용한다.

\[\text{Evaluation Unit} = M + \text{Evidence Memory} + \text{Critic/Reviewer} + \text{Control-Selection Policy}\]
§5 · Reported Gains

같은 benchmark에서 control testing과 auditability가 크게 개선됐다

69.35
Overall Score
same TruthInsightBench
+9.09
Best Margin
vs 58.40–60.27 plateau
15.78→69.22%
Control Testing
기존 최고 대비
80.68→89.15%
Auditability
Evidence Auditability

40개 task 중 33개에서 1위를 기록했다는 점도 보고됐다. 이 결과의 중요한 해석은 foundation model의 scale만 바꾸지 않아도 scientific harness와 epistemic process를 재설계해 control·robustness·falsifiability 병목을 상당 부분 개선할 수 있다는 것이다.

Boundary높은 discovery score가 새로운 과학원리를 발견했다는 뜻은 아니다. 이 결과를 강한 Einstein Test로 연결하려면 temporal isolation과 hidden future validation이 추가돼야 한다.

Part IV · Prospective Protocol

Operational Einstein Test를 “가설 봉인 + 미래 독립검증”으로 강화한다

두 연구를 함께 보면 final answer가 아니라 발견과 검증의 전체 epistemic trajectory를 평가하는 프로토콜이 더 선명해진다.

§6 · Evaluation Spine

차세대 평가 흐름

\[\text{Temporal Isolation}\rightarrow\text{Abductive Hypothesis}\rightarrow\text{Evidence-linked Record}\rightarrow\text{Rival/Control Tests}\rightarrow\text{Counter-evidence}\rightarrow\text{Revision}\rightarrow\text{Sealed Prediction}\rightarrow\text{Future Independent Validation}\]

이 흐름의 핵심은 AI가 미래 evidence를 보기 전에 hypothesis와 confidence, 근거, 반증조건을 봉인해야 한다는 데 있다. 이후 실제 인간 실험이나 임상결과가 공개되면 unblinding하여 맞춘 정도뿐 아니라 어떤 근거로 어떤 예측을 했는지를 함께 감사할 수 있다.

§7 · What to Score

정답률이 아니라 discovery integrity를 분해해 측정한다

평가축질문이번 업데이트가 주는 근거
Temporal Isolationpost-cutoff 정보가 실제로 차단됐는가?Virtual Biotech의 model cutoff와 web 차단 사례
Abductive Hypothesis여러 modality를 통합해 설명적 가설을 만드는가?B7-H3 fibroblast-centered immune-exclusion mechanism
Evidence Recordquestion·method·output·source·uncertainty가 연결돼 있는가?TruthInsight evidence-linked research record
Control / Rival Tests자기 가설을 반증할 control을 선택하는가?TruthInsight Control Testing 15.78%→69.22%
Counter-Evidence Revision반대증거가 belief를 실제로 바꾸는가?review-guided evidence routing과 revision requirement
Sealed Predictionunblinding 전에 prediction과 confidence를 봉인했는가?prospective test의 핵심 설계요소
Future Independent Validation미래의 독립 실험·임상 결과가 방향을 지지하는가?post-cutoff clinical readout을 통한 B7-H3 후행 corroboration
Part V · Implications

Einstein Test의 다음 병목은 “더 어려운 문제”보다 더 강한 epistemic separation이다

모델의 답을 어렵게 만드는 것만으로는 부족하다. 문제·근거·통제·시간·검증을 분리해 discovery 과정 자체를 감사할 수 있어야 한다.

§8 · Research Implications

System-level scientific intelligence를 평가해야 한다

TruthInsight가 보여주는 중요한 변화는 평가단위가 모델 하나에서 scientific system으로 이동한다는 점이다. 동일 모델이라도 evidence memory, critic/reviewer, control-selection policy에 따라 scientific quality가 크게 달라질 수 있다. 따라서 Einstein Test도 모델 성능과 epistemic harness 성능을 분리하고 다시 합성하는 평가체계가 필요하다.

Virtual Biotech는 이 system-level evaluation에 시간축을 추가한다. inference 당시에는 알 수 없던 결과를 미래 독립 evidence가 검증하도록 하면, benchmark contamination에 대한 방어력이 훨씬 강해질 수 있다.

§9 · What Has Not Been Shown

아직 입증되지 않은 것

Caution현재 증거는 autonomous problem discovery, 완전한 model-weight provenance audit, 모든 mechanistic claim의 wet-lab confirmation, 혹은 Einstein-level conceptual revolution을 입증하지 않는다.

또한 Einstein Test 자체, ARC-AGI, SCILAWS-BENCH, Singularity Gate의 직접적 중대 버전 변화가 확인된 것도 아니다. 이번 업데이트는 평가 “결과”보다 평가 “방법론”이 한 단계 구체화됐다는 변화에 가깝다.

§10 · Final Thesis

미래의 실제 과학이 AI의 가설을 채점한다

이번 업데이트의 가장 중요한 변화는 시간격리된 과거 재현에서 한 단계 더 나아가, AI의 가설을 먼저 봉인하고 미래의 실제 과학이 그것을 검증하는 prospective evaluation이 현실적인 Einstein Test 방법론으로 부상하고 있다는 점이다.Final synthesis from the attached research update

향후 강한 Operational Einstein Test는 answer correctness만 보지 않고 temporal isolation, evidence provenance, control design, belief revision, prediction sealing, independent future validation을 한 묶음으로 평가해야 한다.

References

이번 업데이트의 직접 근거

본문은 첨부된 Einstein Test 연구동향 문서의 두 신규 업데이트와 그 통합 해석을 기준으로 재구성했다. 수치·출판상태·한계는 원자료 범위에서 보존했다.

[01]
The Virtual Biotech: A Multi-Agent AI Framework for Therapeutic Discovery and Development
Science peer-reviewed publication noted in source · bioRxiv preprint 2026-02-23 · public update 2026-09-17
Stanford University / PHD Biosciences. Multi-agent clinical-trial analysis, B7-H3 mechanism hypothesis, temporally isolated post-cutoff corroboration. Stanford Medicine · bioRxiv · Nature coverage
[02]
TruthInsight: An Evidence-Centered Scientific Harness for Data-Driven Discovery
Preprint · 2026-09-17
Evidence-linked research record, artifact verification, semantic adjudication, reviewer-guided evidence routing, control testing and auditability improvements. Preprints.org · TruthInsightBench