AI Research NotesEinstein Test · Discovery Trajectory · Temporal Backtesting
Einstein Test / AGI Scientific-Discovery Research Watch · 02 Sep 2026

정답을 맞혔는지가 아니라
어떻게 발견했고, 어디서 틀렸고,
왜 이론을 바꿨는지를 평가한다

From cognitive traces and autonomy gradients to temporal backtesting and falsifiable belief revision.

QUESTIONDISCOVERYABDUCTIONMECHANISTICTHEORYDISCRIMINATINGEXPERIMENTFUTUREPREDICTIONBELIEFREVISIONTEMPORAL CONTAMINATION AUDITAUTONOMY WITHDRAWALPROCESS TRACE EVALUATIONfuture knowledge sealed · guidance removed · failures and revisions scored explicitly
Central Shift

이번 흐름의 가장 중요한 변화는 AGI/AI Scientist 평가의 단위가 “정답”에서 “발견 과정”으로 이동하고 있다는 점이다. COGTRL은 scientific cognitive trace를 학습대상으로 만들고, CEDAR-GRPO는 abduction process를 reward로 만들며, Model Discovery Agent는 이론 실패–새 이론–판별실험–posterior update를 closed loop로 묶는다.

ASI-Bench는 인간의 연구방법 지침을 단계적으로 제거해 autonomy 자체를 측정하고, Historical Backtesting과 Singularity Gate는 시간을 이용해 contamination을 구조적으로 차단한다. SEE는 그 이전 단계인 evidence-bounded inference조차 아직 충분하지 않다는 신호를 준다.

다음 Einstein Test는 “상대성이론을 맞혔는가?”가 아니라 “질문을 스스로 찾았는가, 경쟁가설을 버렸는가, 판별실험을 설계했는가, 미래증거 앞에서 믿음을 수정했는가?”를 추적해야 한다.
Part I · Eight Signals

2026년 8월의 여덟 신호가 하나의 평가철학으로 수렴한다

§1 · Research Map
DateResearch / BenchmarkCore ContributionEinstein Test Relevance
2026-08-31COGTRL제약 검토→실패 대안→수정→다음 단계 선택의 cognitive trace를 trajectory-level RL로 학습. 3B에서 method quality 평균 +7.85.최종 theory match가 아니라 discovery trajectory를 평가해야 한다는 직접적 구현 근거.
2026-08-18ASI-Bench11개 과학분야 60 project-level task. full guidance 50.91 → method-only 29.10 → autonomous method selection 26.62.인간 지침을 제거하며 autonomy gradient를 측정.
2026-08-17Historical Backtestinghistorical cutoff에서 질문 생성 후 future corpus로 답변·부분답변·독립제기·반증 평가. 4 cutoffs, 798 questions, 200-question prospective set.future 자체를 ground truth로 쓰는 contamination-resistant Einstein Test 기반.
2026-08-14CEDAR-GRPOevidence coverage와 evidence→explanation 방향성을 reward. unseen tasks 평균 +7.4, correctness-only GRPO 대비 +2.7.abduction을 학습 가능한 process capability로 operationalize.
2026-08-14LLMs Don't Pay for the Jumpembodiment보다 epistemic error에 실제 cost를 부과해 belief revision을 강제하는 coupling이 핵심일 수 있다는 반론.반증된 가설을 유지하면 손실을 받고 revision해야 하는 test structure의 근거.
2026-08-10Model Discovery AgentLLM proposer + Bayesian SMC/SBI verifier + Value of Information experiment selector. M-open에서 새 model structure 생성.현재 이론 실패→새 이론→판별실험→posterior update 루프의 알고리즘 사례.
2026-08-07Science Edge Evaluation실제 peer-reviewed multimodal evidence. 19 MLLM 최고 48.7%, tool-using visual agent 52.7%.theory novelty 이전에 evidence-bounded inference를 독립 metric으로 요구.
2026-07-15 v1.1Singularity Gatetraining cutoff 이후 discovery만 사용. GPT-5.6 Sol cutoff 2026-02-16 이전 public trace item 제거 후 재평가.prospective contamination-resistant gate. 단 problem selection과 long-horizon revision은 미측정.
Part II · Process-Aware Learning

발견의 실패·수정·backtracking을 학습 object로 만든다

§2 · COGTRL

COGTRL의 핵심은 과학자의 실제 연구과정에서 나타나는 제약 검토, 실패한 대안, 수정, 다음 연구단계 선택을 scientific step과 함께 trajectory-level reinforcement learning으로 학습한다는 점이다. 작은 모델이 훨씬 큰 모델에 접근할 수 있다는 결과보다 더 중요한 의미는 “어떤 reasoning path가 좋은 과학적 방법인가”를 보상할 수 있게 되었다는 데 있다.

§3 · CEDAR-GRPO

CEDAR-GRPO는 final correctness만 보상하지 않고 evidence coverage와 evidence→explanation directionality를 reward로 사용한다. 대안 탐색, 경쟁가설 제거, backtracking, uncertainty marking이 함께 개선된다.

Hypothesis Generation

새 설명후보를 만든다.

Rival Elimination

evidence로 경쟁가설을 제거한다.

Backtracking

실패한 경로에서 되돌아간다.

Uncertainty

증거가 부족한 지점을 남긴다.

이 네 요소는 Einstein-level abductive jump를 철학적 개념에서 독립적인 평가축으로 옮길 근거가 된다.

Part III · Autonomy & Temporal Isolation

얼마나 잘 푸는가보다, 지침을 제거했을 때 얼마나 남는가

§4 · ASI-Bench
50.91Full Guidance
29.10Method-Only
26.62Autonomous Selection
60Project-Level Tasks

ASI-Bench는 같은 task에서 problem/method/experiment guidance를 단계적으로 제거한다. 이 구조는 “정답을 아는 AI”와 “스스로 연구하는 AI”를 분리하는 autonomy-gradient evaluator다.

§5 · Historical Backtesting

이번 추적에서 Operational Einstein Test에 가장 직접적으로 도움이 되는 신규 연구다. 과거 시점 t에서 corpus를 동결하고 AI가 질문을 생성한 다음, 미래 corpus를 시간적으로 격리하여 질문이 답변·부분답변·독립제기·반증되었는지 측정한다.

Retrospective

2010–2024의 네 cutoff, 798개 질문

Prospective

2026-08-17 동결 200개 질문을 2027–2030 문헌으로 평가

Generalization

1911→1915 한 사례를 여러 historical cutoff와 질문발견으로 확장

핵심은 아직 존재하지 않는 future evidence 자체가 contamination-resistant evaluator가 된다는 점이다.

§6 · Singularity Gate

Singularity Gate는 각 모델 training cutoff 이후 처음 공개된 paradigm-changing discovery를 대상으로 한다. v1.1에서는 GPT-5.6 Sol cutoff인 2026-02-16 이전 public trace가 있는 item을 제거했다. 첨부 메모 시점에서는 어떤 평가모델도 하나의 발견을 완전히 재현하지 못했다.

다만 problem selection이나 장기간의 hypothesis revision은 측정하지 않는다. 따라서 AGI의 충분조건보다는 contamination-resistant prospective discovery gate에 가깝다.

Part IV · Mechanistic Discovery Loop

틀린 이론을 버리고 가장 판별력 있는 실험을 고르는가

§7 · Model Discovery Agent

Model Discovery Agent는 LLM을 mechanistic model proposer로, Bayesian SMC/SBI를 verifier로, Value of Information을 experiment selector로 결합한다. 현재 hypothesis class에 진실이 없는 M-open 상황에서 prediction failure를 감지해 새 model structure를 만들고 이를 구분하는 실험을 설계한다.

Current Theory Failure → New Model Proposal → Discriminating Experiment → Posterior Update

이는 Einstein Test의 핵심 루프를 실제 알고리즘으로 구현한 사례에 가깝다.

§8 · LLMs Don't Pay for the Jump

이 연구는 embodiment가 Einstein식 abduction의 필수조건이라는 주장에 대해, 더 근본적인 요소가 epistemic error에 실제 비용을 부과해 belief revision을 강제하는 coupling일 수 있다고 반론한다.

틀린 가설을 계속 유지해도 비용이 없다면, simulator나 world model을 붙였다는 사실만으로 scientific revision이 일어난다고 볼 수 없다.

따라서 Einstein Test는 반증된 theory를 유지할 때 loss를 주고, 실제 revision이 일어나야 다음 단계로 넘어가도록 설계할 수 있다.

Part V · Evidence-Bounded Inference

새 이론을 만들기 전에, 증거보다 더 많은 말을 하지 않는가

§9 · Science Edge Evaluation

SEE는 화학·생물·재료과학의 실제 peer-reviewed 실험과 multimodal evidence를 사용한다. 19개 MLLM 중 최고 정확도는 48.7%, tool-using visual agent도 52.7%다. 더 많은 정보가 반드시 더 신뢰할 수 있는 과학적 추론으로 이어지지 않는다는 결과가 중요하다.

따라서 future Einstein Test는 theory novelty만이 아니라 “주어진 실험증거를 넘어 과도한 주장을 하지 않는가”를 별도 metric으로 포함해야 한다.

§10 · Three Validity Controls
Temporal Contamination Audit
미래 지식이 들어오지 않았는가
Autonomy Withdrawal
인간 지침을 줄여도 연구가 유지되는가
Process Trace Evaluation
실패·수정·backtracking이 타당한가

이 세 제약을 동시에 적용하면 memorization, guided solving, lucky final answer를 각각 분리할 수 있다.

Part VI · Next Operational Einstein Test

여섯 단계 discovery trace에 temporal audit·autonomy withdrawal·process scoring을 겹친다

§11 · Six Stages
1 · QuestionQuestion Discovery
2 · AbductionNew explanatory hypothesis
3 · TheoryMechanistic formalization
4 · ExperimentDiscriminating intervention
5 · PredictionFuture / withheld consequence
6 · RevisionBelief update after counter-evidence

첨부 메모의 최종 제안은 이 여섯 단계 각각에 temporal contamination audit + autonomy withdrawal + process trace evaluation을 함께 적용하는 것이다.

§12 · Why Historical Backtesting Matters Most

“1911년 이전 데이터로 1915년 General Relativity를 재발견하라”는 유명한 단일 사례 대신, 여러 historical cutoff에서 먼저 질문 자체의 발견능력을 평가할 수 있다. 이후 abduction, mechanistic theory, discriminating experiment, future prediction, belief revision을 단계적으로 시간격리하면 훨씬 강한 falsifiable AGI scientific-discovery framework가 된다.

§13 · Evaluation Matrix
StageWhat to MeasureAutonomy WithdrawalTemporal HoldoutTrace Evidence
Question Discoveryscientific importance / anomaly selectionproblem hint 제거future relevance선택 이유
Abductionnovel explanationmethod hint 제거future theory hiddenrival generation/elimination
Mechanistic Theoryformal structuretemplate 제거later formalism hiddenassumption/constraint trace
Discriminating Experimentinformation gainexperiment prescription 제거result sealedVoI rationale
Future Predictionnovel consequencetarget 미제공future corpus holdoutconfidence/uncertainty
Belief Revisionfalsification responserevision hint 제거counter-evidence delayedrejected hypothesis + new belief
§14 · Evidence Boundary

How to read this research watch

COGTRL, ASI-Bench, Historical Backtesting, CEDAR-GRPO, Model Discovery Agent, SEE, Singularity Gate의 방법과 수치는 첨부 메모가 정리한 source-derived 내용이다. 여섯 단계 discovery trace와 세 validity control을 하나의 Operational Einstein Test로 결합한 구조는 이 흐름을 종합해 첨부 메모가 제안한 방향이다.

LLMs Don't Pay for the Jump는 embodiment 논쟁에 대한 반론이며 “embodiment가 불필요하다”는 확정적 실증결론으로 읽어서는 안 된다.

§15 · Final Synthesis
다음 Einstein Test는 천재적 정답 하나를 기다리는 시험이 아니라, 미래지식을 봉인하고, 인간지침을 단계적으로 제거하고, 질문선택부터 belief revision까지의 discovery trace를 검증하는 평가과학이 되어야 한다.
Primary Sources

이번 업데이트의 연구·벤치마크

01
COGTRL: Training LLMs for Scientific Discovery Assistance using Cognitive Traces via Reinforcement Learning
31 Aug 2026 · arXiv:2608.30109
arXiv
02
ASI-Bench: At the Dawn of Artificial Superintelligence
18 Aug 2026 · arXiv:2608.17271
arXiv
03
Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot
17 Aug 2026 · arXiv:2608.16795
arXiv
04
CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs
14 Aug 2026 · arXiv:2608.14791
arXiv
05
LLMs Don't Pay for the Jump
14 Aug 2026 · arXiv:2608.14397
arXiv
06
Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
10 Aug 2026 · arXiv:2608.09696
arXiv
07
Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery
07 Aug 2026 · arXiv:2608.06931
arXiv
08
The Singularity Gate
v1.1 · prospective paradigm-discovery benchmark
Project site

본 게시물은 첨부된 e-test-trends-0902.md의 전체 내용을 구조화해 재작성했다. 수치·날짜·연구별 주장과 한계는 첨부 메모의 범위를 보존했으며 외부 지식으로 조용히 수정하거나 보강하지 않았다.