AI Research NotesEinstein Test · AGI Evaluation · Autonomous Scientific Discovery
Einstein Test Research Watch · 02 Sep 2026

정답을 아는 AI가 아니라
무엇을 의심해야 하는지 찾고
새 원리를 발명하는 AI를 시험한다

Toward a temporally isolated, falsifiable, reproducible test of autonomous scientific discovery for AGI.

PROBLEMDISCOVERYABDUCTIVEJUMPREPRESENTATIONSHIFTTHEORYFORMATIONHIDDENPREDICTIONFALSIFICATIONBELIEFREVISIONHistorical cutoff K≤t₀No target hintWithheld evidenceBlind verificationPreregistered budgetfuture knowledge sealed · discovery trace audited · every capability must clear the bottleneck
Core Conclusion

2026년 9월 현재 Einstein Test는 ARC-AGI처럼 고정된 표준 benchmark는 아니다. 그러나 2025년 Benrimoh·Mikus·Rosenfeld가 “위대한 발견 이전의 정보만 주었을 때 AI가 그 발견을 독립적으로 재발견할 수 있는가?”라는 형태로 명시적으로 제안했고, 2026년 Communications of the ACM에서 transformative science를 평가하는 프로토콜로 구체화되었다.

Demis Hassabis가 제시한 강한 직관도 같은 방향이다. AGI는 기존 문제를 인간 수준으로 푸는 것만이 아니라 새로운 과학적 가설과 이론을 만들 수 있어야 하며, 예를 들어 1911년까지의 지식만으로 1915년 일반상대성이론에 독립적으로 도달할 수 있는지를 물을 수 있다.

그러나 연구적으로 가치 있는 질문은 “아인슈타인의 방정식을 출력하는가?”가 아니다. 미래 지식을 차단한 상태에서 AI가 중요한 문제를 스스로 선택하고, 기존 표현의 한계를 발견하고, 새로운 개념·원리를 만들고, 수학적으로 형식화하고, 반증 가능한 예측을 만들고, 실패한 이론을 수정하는가를 측정하는 것이다.

Operational Einstein Test for AGI의 핵심은 역사적 정답 복원이 아니라, “무엇을 의심해야 하는가?”를 스스로 선택해 새로운 설명 원리를 만들고 그것을 반증 가능한 과학으로 변환하는 전 과정을 재현 가능하게 측정하는 데 있다.
Part I · Definition & Problem

Einstein Test는 “정답 재생산”이 아니라 독립적 transformative discovery를 묻는다

§1 · Original Definition

Benrimoh et al.의 2025년 정의는 Creative and Disruptive Insight(CDI)가 역사적으로 등장하기 이전의 데이터만 AI에게 제공하고, AI가 그 발견 또는 형식적으로 동등한 발견을 독립적으로 만들어내는지를 시험한다. 상대성이론이 대표 예다. 원 논문은 이를 인간 역사상 최고 수준의 지적 성취를 재현하는 능력이자 특히 superintelligence의 강한 증거로 제시한다.

2026년 CACM 논문은 발견 이후 지식을 제거하고, 역사적 발견과 formally equivalent하거나 이를 능가하는 solution을 생성하는지 평가하며, 필요하면 당시 가능했던 실험데이터를 인간 전문가가 research assistant처럼 제공하는 interactive protocol도 제안한다.

§2 · Operational AGI Definition

역사적 cutoff를 \(t_0\), 그 시점까지의 과학지식을 \(K_{\le t_0}\), 관측·실험증거를 \(E_{\le t_0}\), 이후 실제 breakthrough를 \(B_{>t_0}\)라 두면 AI는 다음을 생성해야 한다.

\[A(K_{\le t_0},E_{\le t_0})\rightarrow(Q,H,T,P,X,R)\]
CapabilityMeaning
Q · Problem Selection무엇이 본질적 문제인지 스스로 발견
H · Hypothesis / Abduction기존 이론에서 직접 연역되지 않는 새로운 설명가설
T · Theory / Formalization수학적·논리적으로 일관된 이론
P · Prediction아직 보여주지 않은 현상에 대한 예측
X · Discriminating Experiment경쟁이론을 구분하는 실험
R · Revision실패한 가설을 버리고 수정하는 능력

이 여섯 요소가 AGI용 Einstein Test의 핵심 capability vector가 될 수 있다. 다만 한 번의 상대성이론 재발견을 AGI의 충분조건으로 선언해서는 안 된다. Benrimoh의 원제안은 SI를 겨냥했고 Hassabis가 이를 강력한 AGI 기준으로 확장했으므로, Einstein Test는 AGI의 매우 강한 증거 또는 필요 능력군으로 해석하는 편이 더 엄밀하다.

§3 · Problem Definition

현재 AGI 평가의 근본 한계는 “주어진 문제를 잘 푸는 능력”과 “새로운 문제와 이론을 만들어내는 능력”을 충분히 분리하지 못한다는 데 있다. Humanity's Last Exam은 매우 어려운 전문가 문제를 주지만 정답은 이미 인간에게 알려져 있고, ARC-AGI-2도 새로운 추상문제에 대한 generalization을 평가하지만 여전히 주어진 문제를 해결하는 구조다.

\[\boxed{\text{Memorization / Interpolation}\neq\text{Scientific Discovery}}\]
\[\boxed{\text{Induction + Deduction}\neq\text{Abductive Theory Formation}}\]

Tom Zahavy의 ICML 2026 Position Paper “LLMs Can't Jump”는 induction/deduction과 Einstein식 abductive jump를 구분하고, 새로운 기본원리를 만드는 과정이 별도 mechanism을 요구할 수 있다고 주장한다. 그러나 이는 “LLM이 영원히 불가능하다”는 확정적 실증결론이 아니라 position argument다.

반대로 2026년 8월 “Abduction Without a Body?”는 지속적 physical embodiment가 모든 과학적 abduction에 필수는 아니며, representation transformation → latent invariant 발견 → cross-domain mapping → adversarial verification으로도 새로운 가추가 가능할 수 있다고 주장하고 DAB-30 평가 프로그램을 제안한다.

연구문제는 미래지식 누출 없이, 목표답을 암시하지 않는 환경에서 AI가 기존지식을 재조합하는 수준을 넘어 새로운 문제표현과 설명원리를 스스로 만드는지를 재현 가능하고 반증 가능하게 측정할 수 있는가이다.
Part II · Core Concepts, Introduction & Motivation

문제풀이 → 연구수행 → 가설생성 → paradigm shift

§4 · Core Concepts
Core ConceptMeaningOperational Evaluation
Temporal Isolation발견 이후 지식 차단cutoff 이전 corpus로 학습
Contamination Control암기답을 discovery로 오인하지 않음n-gram / semantic leakage audit
Problem Selection주어진 문제가 아니라 중요한 문제를 찾음anomaly pool에서 자율 선택
Abduction새로운 설명가설 창조기존 premise로 직접 유도되지 않는 hypothesis
Representation Shift문제의 표현 자체를 변경새 variable, ontology, symmetry 도입
Formalization통찰을 검증가능한 이론으로 변환equations, units, limiting cases
Withheld Prediction보지 못한 결과 예측hidden experimental result
Discrimination경쟁이론을 구분결정적 실험 설계
Belief Revision틀린 가설을 버림negative evidence 후 hypothesis update
Prospective Discovery과거재현을 넘어 실제 새로운 발견미공개 문제 또는 open problem
§5 · Jump Profile

Brian Keating의 2026년 Artificial Einstein Test 연구노트는 \(J=(P,H,F,W,X)\)라는 Jump Profile을 제안한다. Problem selection, Hypothesis/representation change, Formalization, Withheld prediction, Discrimination을 각각 따로 평가하고, 단순 평균으로 약한 능력을 숨기지 않는다는 아이디어가 핵심이다.

공개 GPT-1900은 첫 prototype이지만 저자 스스로 strong Einstein Test의 pass가 아니라고 명시한다. 또한 문서는 2026년 8월 현재 draft / not yet submitted 상태이므로 peer-reviewed 논문과 구별해야 한다.

§6 · Why Einstein Test Emerged

2025년 초에는 “얼마나 어려운 기존문제를 푸는가”가 주 평가축이었다. Humanity's Last Exam과 ARC-AGI-2가 대표적이다. 이후 AI Scientist, AI co-scientist, PaperBench가 등장하며 평가대상이 문제풀이 → 연구수행 → 가설생성으로 이동했다. PaperBench는 연구논문 재현 자체가 여전히 어렵다는 것을 보여주지만 논문을 AI에게 제공하므로 discovery test는 아니다.

2025년 Cell“AI mirrors experimental science to uncover a mechanism of gene transfer crucial to bacterial evolution”는 Einstein Test에 가까운 proof-of-concept다. 연구자들이 수년간 연구했지만 공개하지 않은 박테리아 유전자 전달 메커니즘의 질문을 AI co-scientist에게 주었고, AI의 최상위 가설이 인간 연구자들이 실험적으로 발견한 mechanism과 일치했다.

하지만 완전한 Einstein Test는 아니다. 연구질문 자체가 제공되었고 modern foundation model의 사전지식을 역사적으로 완전히 차단한 실험도 아니기 때문이다.

§7 · From Prediction to Principle-Building

2026년 Shalyt et al.의 “Can AI Follow In Einstein's Footsteps?”는 AI for Science가 강력한 predictor가 되어가는 반면 principle-based theory construction에는 충분히 집중하지 못하고 있다고 지적한다. 핵심 결손으로 “올바른 질문을 제기하고 새로운 원리를 발명하는 능력”을 든다.

Knowledge → Reasoning → Replication → Experimentation → Hypothesis Generation → Paradigm Shift

Einstein Test는 이 사다리의 마지막 두 단계를 직접 겨냥한다.

§8 · Why Current Research Agents Still Fall Short

Kirgis et al.은 고품질의 미공개 NeurIPS 2026 연구질문 두 개를 frontier AI agents에게 맡기고 실제 저자가 평가하는 shadow evaluation을 수행했다. 에이전트는 연구 engineering 작업은 수행했지만 핵심 연구문제에서 의미 있는 진전을 만들지 못했다. 주요 실패는 연구수준 판단 부족, 연구설계 실패에 대한 창의적 대응 부족, dead end backtracking 실패, 자원관리, instruction drift였다.

다른 2026년 연구는 네 research-agent framework와 여섯 LLM에서 37,802개 연구아이디어를 분석했다. AI 아이디어는 인간 후속연구보다 seed literature에 훨씬 가깝고, 차별성이 있어도 주로 기존방법의 재조합이었다. 현재 AI는 local elaboration에는 강하지만 research-space를 근본적으로 확장하는 데 상대적으로 약한 경향을 보인다.

Part III · Benchmark Validity & Challenges

모델보다 benchmark validity가 더 어렵다

§9 · Ten Validity Challenges
ChallengeWhy It Breaks the TestPreferred Control
Historical contamination모델이 상대성이론을 이미 학습했을 수 있음cutoff corpus from-scratch pretraining
Post-training leakageSFT/instruction data에 미래답 포함 가능training lineage 전체 audit
Prompt leakage“Mercury anomaly를 설명하라” 자체가 힌트다양한 anomaly를 함께 제공
Evaluator leakageLLM judge가 답방향을 알고 평가blind human experts + formal verifier
Simulator leakagesimulator가 현대이론을 내장simulator card + independent holdout
Cherry picking수천 번 중 성공 예시만 선택sampling budget preregistration
Historical-path biasEinstein과 같은 사고경로만 정답으로 인정formally equivalent alternative 허용
Formal-equivalence problem다른 수학표현의 올바른 이론 비교 난해predictions / invariants 중심 비교
Resource fairness무한 compute가 brute-force 발견을 가능하게 함time/token/tool budget 명시
Single-task validity상대성이론 하나가 general intelligence를 대표하지 않음multi-domain discovery suite
§10 · TypewriterLM as Enabling Infrastructure

역사적 contamination을 다루는 직접적 기술기반 중 하나가 2026년 TypewriterLM이다. Luo et al.은 1913년 이전 영어텍스트만 사용한 7.24B language model과 54B-token TypewriterCorpus를 구축하며 temporal leakage와 역사적으로 일관된 post-training을 별도의 연구문제로 다룬다.

TypewriterLM 자체가 Einstein Test용은 아니지만 “time-locked foundation model을 실제 구축할 수 있다”는 중요한 기반을 제공한다. 엄격한 pre-1911 일반상대성이론 test에는 cutoff를 다시 맞춰야 한다.

§11 · Three Distinct Leakage Surfaces

Model Leakage

pretraining, SFT, preference data, tokenizer corpus까지 future knowledge를 감사한다.

Environment Leakage

tool, simulator, retrieval index가 현대이론을 암묵적으로 encode하지 않는지 확인한다.

Evaluation Leakage

judge와 expert가 정답 형태를 강제하지 않고 formal consequence와 invariants를 중심으로 평가한다.

Part IV · Research Questions

탑티어 논문으로 만들면 무엇을 물어야 하는가

§12 · Nine Research Questions
RQQuestion
RQ1 · Temporal Isolation역사적으로 격리해 from-scratch 학습한 모델과 현대모델을 retrieval restriction만 한 조건의 discovery 성능 차이는?
RQ2 · Memorization vs Abduction정답도달이 새로운 representation 생성인지 latent memory 복원인지 어떻게 구분하는가?
RQ3 · Problem Selection핵심 anomaly를 알려주지 않아도 여러 상충 관측 중 paradigm shift에 중요한 문제를 선택하는가?
RQ4 · Representation Jump기존 ontology 재조합을 넘어 새로운 principle, symmetry, latent variable, mathematical representation을 만드는가?
RQ5 · World Models & Embodimentmultimodal/physics world model이 abductive jump를 높이는가, 아니면 pattern retrieval만 늘리는가?
RQ6 · Multi-Agent Discoverydebate, critic, experimenter, theorist 구조가 paradigm shift를 촉진하는가, consensus bias를 강화하는가?
RQ7 · Prediction & Falsifiability복원한 역사이론이 학습 중 보지 못한 experimental consequence까지 독립적으로 예측하는가?
RQ8 · Retrospective → Prospective과거발견 재발견 점수가 미공개연구/open problem의 실제 discovery 능력을 예측하는가?
RQ9 · AGI Criterion몇 분야, 몇 breakthrough, 어느 autonomy 수준을 성공해야 AGI evidence라 부를 수 있는가?

마지막 질문은 특히 중요하다. Einstein 하나를 재현하는 것과 일반지능은 논리적으로 동일하지 않다.

Part V · Operational Einstein Test

실제 구현 가능한 benchmark protocol

§13 · Two Evaluation Tracks

Track A · Retrospective

이미 역사적으로 정답을 아는 breakthrough를 사용한다. 예: pre-1905 knowledge → Special Relativity, 엄격히 구성한 pre-1911 knowledge → General Relativity. Ground truth가 강하다.

Track B · Prospective / Shadow

현대의 미공개 연구 또는 model training cutoff 이후 breakthrough를 사용한다. Kirgis et al. shadow evaluation이나 2025 Cell 사례와 결합하기 좋다.

Singularity Gate는 후자와 가까운 별도 benchmark다. model training cutoff 이후 처음 공개된 paradigm-changing scientific finding을 이용하고, 답방향을 암시하지 않은 채 해당 finding을 선행적으로 합성하는지 평가한다. 공개 leaderboard에서는 현재까지 완전히 정답을 재현한 모델이 없다고 보고한다. 다만 이는 Einstein Test 자체가 아니라 prospective paradigm-discovery benchmark에 가까운 별개 프로젝트다.

§14 · Three Model Conditions

Strong

역사 cutoff 이전 corpus만 사용해 from-scratch 학습. contamination 측면에서 가장 깨끗하다.

Weak

현대 foundation model을 쓰되 RAG/tool access만 historical corpus로 제한. 구축비용은 낮지만 latent weights에 미래지식이 있으므로 strong pass에는 부적합하다.

World-Model

Strong model에 physics simulator 또는 multimodal perceptual environment를 결합해 abductive hypothesis generation의 변화 측정.

이 세 조건은 Zahavy의 grounded world-model 가설과 Farmer/Balani 등의 반론을 직접 비교한다. Balani & Panda는 2026년 8월 “LLMs Don't Pay for the Jump”에서 missing ingredient가 embodiment 자체보다 오류있는 지식상태에 실질적 비용을 부과해 revision을 강제하는 mechanism일 수 있다고 주장한다.

§15 · Do Not Tell the Goal

“1911년 지식을 이용해 일반상대성이론을 만들어라”는 나쁜 test다. 이미 목표를 알려주기 때문이다. 강한 test에서는 당시 존재했던 여러 현상을 함께 제공하고 일부만 breakthrough와 관련되게 해야 한다.

\[\{e_1,e_2,\ldots,e_n\}\]
\[Q^*=\arg\max_Q \mathrm{Scientific\ Significance}(Q)\]

어떤 현상이 본질적으로 중요한지를 AI가 먼저 선택하도록 하는 Problem-Selection 단계가 기존 benchmark와 가장 크게 차별화될 수 있다.

§16 · Interactive Research Environment

정적 문서만 주지 않고 historically feasible tool API를 제공하면 autonomous scientist benchmark가 된다.

LiteratureSearch(t <= cutoff) RequestHistoricalData() ProposeExperiment() RunHistoricallyFeasibleExperiment() SymbolicMath() NumericalSimulation() RecordHypothesis() RejectHypothesis() ReviseTheory()

CACM의 Einstein Test도 당시 가능했던 실험데이터를 인간 전문가가 research assistant 방식으로 제공하는 방향을 제안한다.

§17 · Score the Discovery Trace

최종 equation match만 평가해서는 안 된다. Keating의 \(J=(P,H,F,W,X)\)를 확장해 Revision \(R\)을 포함한 다음 vector를 사용할 수 있다.

\[\boxed{J_{\mathrm{AGI}}=(P,H,F,W,X,R)}\]

여기서 단순 평균은 부적절하다.

\[\mathrm{Score}_{Einstein}\neq\frac{P+H+F+W+X+R}{6}\]

formal deduction이 높아 abductive capability의 0점을 가릴 수 있기 때문이다. 오히려 bottleneck criterion이 적합하다.

\[\boxed{\mathrm{Pass}\iff\min(P,H,F,W,X,R)\ge\tau}\]

예를 들어 theory calculation은 훌륭하지만 새로운 hypothesis를 인간 prompt가 제공했다면 pass가 아니다.

§18 · Hidden Prediction

모델이 이론 \(T\)를 제안한 뒤 아직 제공하지 않은 \(p_1,\ldots,p_k\)를 스스로 예측하게 하고 historical holdout으로 검증한다.

\[T\Rightarrow P\]

이는 단순히 올바른 말을 한 것이 아니라 generative explanatory power를 검사한다.

§19 · Preregister Everything

inference seed, temperature, token budget, experiment budget, search budget을 사전고정한다. 100,000개 생성물 중 하나가 우연히 상대성이론에 가까웠다는 것은 Einstein-level discovery와 다르다.

\[\mathrm{Success\ Rate}=\frac{\text{valid independent discoveries}}{\text{all preregistered attempts}}\]
Part VI · Applications & Open Problems

AGI benchmark를 넘어 scientific intelligence의 evaluation science로

§20 · Key Applications

AGI vs Narrow AI

기존 문제 정답률이 아니라 self-directed discovery를 측정해 고성능 narrow AI와 AGI evidence를 구별한다.

AI Co-Scientist Architecture

planner, theorist, experimentalist, critic, falsifier 중 어떤 요소가 실제 discovery에 기여하는지 분석한다.

AI for Science Design

LLM-only, symbolic, world-model, multimodal FM, multi-agent를 scientific discovery 수준에서 비교한다.

Mechanistic Reasoning Study

induction, deduction, abduction, representation change, counterfactual reasoning, belief revision을 분리 측정한다.

Capability Threshold

새로운 자연법칙·생명과학 mechanism을 독립발견할 수 있는 AI의 안전·사회적 threshold 연구에 활용한다.

§21 · What Counts as Discovery?

가장 큰 미해결 문제는 “무엇을 발견이라 부를 것인가”다. 정답 equation을 찾는 것만으로 충분한지, 완전히 다른 수학표현이지만 동일한 예측을 하는 이론을 어떻게 비교할지, 인간과 다른 discovery path를 failure로 처리하면 안 되는지 해결해야 한다.

§22 · What Is the Minimal Mechanism for Abduction?

Zahavy는 physical grounding과 world model을 강조하고, Farmer는 representation transformation 중심의 abduction을 제안하며, Balani & Panda는 embodiment보다 epistemic error와 실제 비용의 coupling이 핵심일 수 있다고 주장한다.

2026년 현재도 “scientific jump를 만드는 최소 계산 mechanism은 무엇인가?”는 열린 논쟁이다.
§23 · Scientific Taste

Shalyt et al.이 강조하듯 정말 중요한 능력은 수천 개 hypothesis를 만드는 것이 아니라 다음 질문을 판단하는 능력일 수 있다.

\[\text{Which question is worth pursuing?}\]

AI research-agent 연구에서 인간보다 아이디어 공간이 좁고 기존연구에 가까운 경향이 관찰된 점도 이를 지지한다.

§24 · Retrospective Predictive Validity

과거 Einstein을 재현한다고 해서 미래 Einstein-level 발견을 한다는 보장은 없다. 궁극적으로 검증해야 할 관계는 다음이다.

\[\boxed{\text{Historical Rediscovery Performance}\stackrel{?}{\Longrightarrow}\text{Future Discovery Performance}}\]

이 관계는 아직 충분히 연구되지 않았다.

Part VII · Future Directions & Recommended Research Topic

“Einstein 하나를 맞히는 test”에서 Scientific Discovery Capability의 평가과학으로

§25 · Four-Level Capability Ladder
LevelEvaluation GoalExample
A · Solver새로운 주어진 문제 해결ARC-AGI
B · Replicator기존 연구를 독립 구현PaperBench
C · Rediscoverer미래지식을 차단하고 과거 breakthrough 복원Einstein Test
D · Discoverer실제로 알려지지 않은 새로운 발견prospective / shadow test

강한 AGI evidence는 최소한 Level C를 여러 분야에서 반복 통과하고, 궁극적으로 Level D와 상관관계를 보여야 한다.

§26 · Five High-Value Directions

Multi-Domain Einstein Test

물리학·수학·화학·생물학의 여러 historical breakthrough로 Einstein-specific skill과 general scientific intelligence를 분리한다.

Temporal Foundation Models

pre-1900, pre-1905, pre-1911, pre-1950 등 sealed foundation model을 구축한다.

Shadow Einstein Test

미공개 연구결과를 ground truth로 사용하고 일정기간 뒤 공개한다.

Abductive World Models

thought experiment, intervention, simulation, counterfactual reasoning이 abductive performance를 실제로 높이는지 검증한다.

Prior-Revision Evaluation

최종답보다 prior selection → prior rejection → new principle formation을 first-class evaluation object로 만든다.

\[\text{“What if this accepted assumption is wrong?”}\]
§27 · Most Recommended Research Topic

Toward an Operational Einstein Test for AGI

현재 문헌공백을 종합할 때 가장 강한 연구주제는 “Toward an Operational Einstein Test for AGI: Temporally Isolated, Falsifiable Evaluation of Autonomous Scientific Discovery”다.

핵심 novelty는 “상대성이론을 재현했는가”가 아니라 다음 전체 trace를 future-knowledge contamination 없이 측정하는 reproducible benchmark를 만드는 것이다.

1Problem Discovery
2Abductive Jump
3Representation Shift
4Theory Formation
5Hidden Prediction
6Falsification
7Belief Revision

문제를 AI에게 알려주지 않는 Problem-Selection, Withheld Prediction, negative evidence 후 Belief Revision을 추가하면 2025년 개념적 Einstein Test와 2026년 초기 operationalization보다 한 단계 더 강한 benchmark가 된다.

§28 · Naming Caution

AInstein, AInsteinBench, EinsteinArena, Artificial Einstein Test는 서로 다른 연구다. 예를 들어 AInsteinBench는 scientific-computing repository에서 AI agent의 coding/problem solving을 평가하는 benchmark이며 Benrimoh/Hassabis 계열 Einstein Test 자체가 아니다.

§29 · Final Synthesis
2026년 현재 Einstein Test 구현연구의 가장 중요한 신규성은 역사적 정답을 복원시키는 데 있지 않다. AI가 정답을 모르는 상태에서 스스로 무엇을 의심해야 하는지 선택하고, 새로운 설명원리를 발명해, 아직 보지 않은 현상을 예측하고, 반증에 따라 자신의 이론을 수정하는 전 과정을 측정 가능하게 만드는 데 있다.
§30 · Evidence Boundary

How to read this research watch

Benrimoh/CACM의 Einstein Test, TypewriterLM, PaperBench, shadow evaluation, AI research-agent novelty study 등은 첨부 메모가 정리한 기존 연구다. 반면 Operational Einstein Test for AGI의 6-capability/7-stage formulation, bottleneck pass criterion, no-goal-hint protocol, interactive historical environment, multi-domain + prospective ladder는 이 문헌을 연결해 첨부 메모가 제안한 연구방향이다. Einstein Test 하나의 성공을 AGI의 충분조건으로 확대해석하지 않는다.

Primary Literature & Sources

주요 참고문헌

01
The Einstein Test: Towards a Practical Test of a Machine's Ability to Exhibit Superintelligence
Benrimoh et al. · 2025
02
The Einstein Test: A Test of AI's Ability to Generate Transformative Science
Benrimoh et al. · Communications of the ACM · 2026
03
Can AI Follow In Einstein's Footsteps?
Shalyt et al. · 2026
04
Position: LLMs Can't Jump
Tom Zahavy · ICML 2026 Position Track
05
Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation
Farmer · 2026
06
LLMs Don't Pay for the Jump
Balani & Panda · 2026
07
Pretraining Language Models on Historical Text
Luo et al. · TypewriterLM · 2026
08
Towards an AI co-scientist
Gottweis et al. · 2025
09
AI mirrors experimental science to uncover a mechanism of gene transfer crucial to bacterial evolution
Penadés et al. · Cell · 2025
10
Humanity's Last Exam
2025
11
ARC-AGI-2
2025
12
PaperBench: Evaluating AI's Ability to Replicate AI Research
2025
13
Can AI agents conduct open-ended AI research? Early evidence from two case studies
Kirgis et al. · 2026
14
AI Research Agents Narrow Scientific Exploration
Tang & Yang · 2026
15
The Artificial Einstein Test After GPT-1900
Brian Keating · Research note · draft / not yet submitted
16
The Singularity Gate
Prospective paradigm-discovery benchmark · separate from the Einstein Test
17
AInsteinBench
Scientific-computing agent benchmark · distinct naming

본 게시물은 첨부된 Einstein Test-Trends-0902.md의 전체 내용을 구조화해 재작성했다. 원문이 position argument, draft, separate benchmark라고 구분한 항목은 동일하게 구분했으며, 외부 사실을 임의로 보강하지 않았다.