AI Research Blog·AI Risk · RSI · Verification-Governed R&D2026 · 09 · 19
AI Risk Research Update·Recursive Self-Improvement·Scalable Oversight

AI 연구 자동화보다
검증·중단·권한통제가 더 빨리 확장될 수 있는가

Verification-Governed AI R&D: Frontier-Lab Automation, Formal Control, and Independent Evaluation

AI R&D AUTOMATIONCHEAPER ITERATIONEXTERNAL CONTROLINDEPENDENT AUDITAL4 비중 상승harness 비용 절감proof · closed worldembedded evaluatorCAPABILITY / ITERATION VELOCITYVERIFICATION / OVERSIGHT CAPACITY핵심 위험은 자기개선 가능성만이 아니라 두 속도의 격차다
Central Update

이번 업데이트에서 중요한 것은 intelligence explosion의 직접 증거가 아니다. AI가 실제 frontier-lab R&D를 더 많이 맡고, auto-research loop의 비용은 줄어드는 동시에, 검증·권한통제·제3자 평가를 외부 구조로 강화하려는 움직임이 함께 나타났다는 점이다.

반대로 RSI→intelligence explosion, alignment faking이나 scheming→현실의 loss of control, shutdown evasion/self-replication→실제 자율적 권력 확보를 새롭게 입증하는 독립 재현 또는 동료심사 결과는 이번 점검에서 확인되지 않았다.

\[\text{AI-R\&D Automation}\uparrow + \text{Iteration Cost}\downarrow \quad \text{vs.} \quad \text{Verification / Oversight / Independent Audit}\uparrow\]

따라서 문제 정의는 “AI가 자기개선을 할 수 있는가?”에서 capability-iteration 속도와 verification capacity의 성장속도가 얼마나 벌어지는가로 확장된다.

Part I · Frontier-Lab Automation

AI 주도 R&D를 실제 조직 안에서 계량하기 시작했다

Anthropic Institute의 측정은 RSI 논의를 benchmark에서 실제 frontier-lab 업무로 옮기지만, 동시에 AL5 완전자율 R&D가 아직 관찰되지 않았다는 반대 증거도 제공한다.

§1 · R&D Automation Index

AL4는 약 26%, AL5는 0

Source factAnthropic은 AI R&D 자동화 수준을 AL0–AL5로 분류했다. AL4는 사람이 고수준 목표를 제시하면 AI가 대부분의 과정을 주도하는 수준이고, AL5는 완전자율 수행이다. 2026년 8월 자체 측정에서 AL4 수준 AI-led R&D는 약 26%, 90% 이상은 최소한 인간과 적극적으로 협업하는 단계였으나 AL5는 관찰되지 않았다.

26%
AL4 AI-Led R&D
고수준 목표 이후 대부분의 과정을 AI가 주도
90%+
Human–AI Collaboration
최소한 적극적 협업 단계
0
AL5 Observed
측정 업무에서 완전자율 R&D 없음
59% / 97%
Label Agreement
정확 일치 / ±1단계 허용

Analysis이는 AI가 차세대 AI 개발 과정에 투입되는 feedback loop가 현실 운영에 들어왔다는 신호다. 그러나 현재를 closed-loop RSI로 부를 직접 근거는 아니다. 인간의 목표설정, 검토, 인프라, 데이터, 평가가 여전히 핵심 병목이다.

§2 · Oversight at Scale

약 3만 agent와 10억 건 이상의 의사결정을 감시한다

~30,000
Concurrent Agents
가장 많이 사용되는 내부 플랫폼
1B+
Decisions / Month
online monitor 분석량
0.002%
Online Blocks
실제 misbehavior prevalence가 아닌 경보/차단 비율
100k → ~50
Offline Review Funnel
주간 transcript flag → 인간 고우선 검토

일주일 compute-allocation snapshot에서는 AI R&D compute의 약 6%, AI가 수행하는 AI R&D compute의 약 12%가 safety 관련으로 분류됐다. 그러나 이 수치는 compute intensity와 분류기준 차이가 있어 전체 자원 중 safety 비중으로 곧바로 해석할 수 없다.

Limitation기업 자체 측정이고 일부 분류에는 Claude가 사용된다. 하나의 내부 플랫폼을 중심으로 측정했으며 보고서 자체도 제3자 검증 필요성을 인정한다.

Part II · Efficient Auto-Research

자동 연구 루프의 마찰비용을 줄이는 연구가 본격화된다

SoL-Pi는 foundation-model RSI가 아니라 agent harness 탐색이지만, 한 번의 auto-R&D cycle에 드는 token/API/context 비용을 낮추면 동일 예산에서 더 많은 반복이 가능해진다.

§3 · SoL-Pi

152개 방향, 535개 환경, 3,000회 이상의 run

Source factSoL-Pi는 AI optimizer가 약 152개의 개선 방향을 제안하고, 535개 executable environment에서 3,000회가 넘는 run과 6만 회 이상의 agent–environment interaction을 수행해 재사용 가능한 harness mechanism을 찾는다. capability metric은 optimizer와 분리해 고정하고 held-out EdgeBench는 탐색 피드백에 사용하지 않았다.

Action Fusion

여러 action을 결합해 반복 호출과 오버헤드를 줄이는 방향이다.

Online Context Compact

장기 실행의 context traffic을 줄이는 harness mechanism이다.

ObservationPack

관측 정보를 실행 친화적인 형태로 압축·정리한다.

Evidence-Preserving Reducer

효율화 과정에서도 근거가 사라지지 않도록 줄이는 방향이다.

−49%
Token Traffic
EdgeBench efficiency 설정, 기존 Pi 대비
−33.2%
Cost
평균 점수 93.7% 유지
44.8→47.2
Performance Setting
동시에 token 6.1% 감소
60k+
Interactions
agent–environment interaction 규모

Counter-evidenceTerminal-Bench에서는 기존 Pi가 18개 문제를 해결한 반면 SoL-Pi는 15개였고, IMO 2026 Lean에서는 SoL-Pi 3/6, Codex 5/6이었다. 효율적 loop가 항상 더 높은 capability를 뜻하지 않는다.

Risk implication의미는 intelligence explosion이 아니라 auto-R&D의 현실적 마찰비용이 낮아질 수 있다는 것이다. 같은 compute budget에서 iteration 수가 늘면 장기적으로 capability feedback 속도를 높일 가능성이 있다.

Part III · Formal Control

안전판정을 모델의 판단 밖에 있는 proof obligation으로 옮긴다

MAGS는 self-modifying agent의 결과를 더 좋은 LLM judge로 평가하는 대신, 사람이 승인한 frozen safety specification과 deterministic verifier를 실행 경계로 사용한다.

§4 · MAGS

Intent → Formal Specification → Artifact → Machine Proof → Execution

Source factMAGS는 agent가 생성한 프로그램을 Dafny 표현으로 바꾸고 verifier feedback으로 수정한 뒤, 검증이 완료된 코드만 실행 가능한 형태로 만든다.

\[\text{Intent}\rightarrow\text{Formal Specification}\rightarrow\text{Agent Artifact}\rightarrow\text{Machine Proof}\rightarrow\text{Execution}\]
100
CUDA Kernels
formal verification task
100
Terminal Scripts
formal verification task
20
Robotic-Arm Tasks
program-controlled embodied task
220 / 220
Task Coverage
각 task에서 적어도 하나의 frozen safety spec 보장 프로그램

Analysis중요한 변화는 일부 안전조건을 agent가 바꿀 수 없는 외부 실행 경계와 증명 의무로 옮긴 데 있다. 자기수정 권한과 실제 실행 권한을 분리하는 control primitive로 볼 수 있다.

Formal verification은 “명시된 specification을 만족한다”는 것을 보장한다. specification에 없는 위험이나 잘못 기술된 요구사항까지 보장하는 것은 아니다.Specification completeness boundary

따라서 “100% safety”가 아니다. auto-formalization이 실제 의도를 완전히 포착하지 못하면 proof가 있어도 기능 의미에서 실패할 수 있다.

Part IV · Action Boundary

권한검사 전에 먼저 “유효한 행동인가?”를 확인해야 한다

Closed-World Resolution 연구는 tool-use safety가 permission gate만으로는 충분하지 않으며, registry membership과 signature validation이 더 앞선 계층에 있어야 함을 보여준다.

§5 · Resolution Rung

Tool call을 permission보다 먼저 해석한다

일반적인 permission gate는 “요청된 tool을 실행해도 되는가?”를 검사한다. 그러나 모델이 registry에 존재하지 않는 tool 이름을 생성하면 authorization 이전부터 의미가 정의되지 않는다.

\[\text{Tool Request}\rightarrow\text{Closed-World Registry Membership}\rightarrow\text{Signature Validation}\rightarrow\text{Permission / Causal Gate}\rightarrow\text{Execution}\]

Source fact10개 hosted model을 두 invocation surface에서 평가해 322개의 genuine tool hallucination을 관찰했다. fabricated-tool call은 structured interface보다 unconstrained raw JSON에서 훨씬 많았고, 675B급 모델도 7–8B 모델보다 본질적으로 더 안전하지 않았다.

322
Tool Hallucinations
10 hosted models · 2 invocation surfaces
154
Merged MCP Failures
multi-server namespace collision/shadowing 환경
675B ≠ safer
Scale Result
7–8B 대비 본질적 안전 우위 없음
Closed World
Control Principle
행동 공간 membership을 먼저 검증

Boundary이는 scheming이나 goal misgeneralization의 증거가 아니다. 모델이 의도적으로 permission system을 우회했다는 주장도 아니다.

Analysis오히려 control architecture가 잘못된 추상화 계층에서 판단하면 평범한 model error도 통제 경계를 무력화할 수 있다는 시스템 안전 문제다.

Part V · Independent Evaluation

제3자 평가를 frontier-lab 내부 절차에 상시 삽입하려는 거버넌스

embedded evaluation은 공개 모델을 사후 시험하는 외부 감사보다 내부 모델·프로세스에 더 깊게 접근하는 상시 평가 구조를 목표로 하지만, 아직 효과성에 대한 empirical result는 없다.

§6 · Anthropic–Accenture / Faculty

자기평가 문제에 대한 제도적 응답

Source factAnthropic은 Accenture 및 Faculty와 외부 evaluator를 frontier-model 개발 절차 안에 상시 배치하는 embedded evaluation 모델을 추진한다고 발표했다. 목표는 외부 evaluator가 내부 모델·프로세스에 직원과 유사한 수준의 접근권을 받아 evaluation, red teaming, alignment 및 safeguard 검증을 수행하는 구조다.

Anthropic과 Accenture는 관련 역량 구축에 향후 5년간 각각 최소 10억 달러 규모의 투자를 예상한다고 밝혔다.

Analysis앞선 R&D Automation Index가 인정한 핵심 약점—AI 기업이 자기 모델로 자기 자동화 수준과 안전성을 평가하는 문제—에 대한 제도적 대응으로 읽을 수 있다.

Access

공개 모델 사후평가가 아니라 내부 모델·프로세스에 더 깊은 접근을 목표로 한다.

Evaluation

red teaming, alignment, safeguard validation을 개발 lifecycle 안에 삽입한다.

Independence

실제 조직·재정적 독립성과 결과 공개권이 얼마나 보장되는지는 아직 불분명하다.

Limitation아직 연구결과가 아니다. 독립성, 부정적 결과 공개권, access rights, 이해상충 관리 세부가 완전히 공개되지 않았다. 따라서 제3자 검증 문제가 해결됐다고 볼 수 없다.

Part VI · Evidence Verdict & Research Agenda

관찰된 것은 자동화와 control primitive의 동시 발전이다

현재 증거는 인간 없는 완전한 RSI를 보여주지 않는다. 오히려 capability iteration을 빠르게 만드는 요소와 verification·oversight를 확장해야 할 이유가 함께 강해지고 있다.

§7 · Observed vs Not Observed

증거의 경계를 보수적으로 유지한다

Observed

실제 frontier-lab R&D에서 AI 주도 비중이 계량될 정도로 커졌고, auto-research harness의 token/API/context overhead를 줄이는 방법과 formal verifier·closed-world resolver 같은 control primitive가 발전하고 있다.

Not Observed

AI가 사람 없이 연구목표를 정하고, 연구 알고리즘·foundation model·compute·evaluator·safety system까지 자체 수정하며 여러 세대에 걸쳐 가속되는 완전한 RSI loop는 확인되지 않았다.

이번 점검에서 alignment faking, scheming, sandbagging, shutdown evasion, self-replication, explicit power-seeking에 관해서도 기존 결론을 바꿀 새로운 독립 재현이나 반박은 확인되지 않았다.

§8 · What to Update in the Research Map

문제 정의는 “자기개선 가능성”에서 “속도 격차”로 이동한다

항목갱신새로 반영할 핵심
Definition높음AI R&D automation을 AL0–AL5처럼 단계화하고 bounded auto-R&D와 full autonomous RSI를 명시적으로 분리
Problem Definition매우 높음핵심 문제를 자기개선 가능 여부에서 capability-iteration 속도와 verification 속도의 격차로 확장
Core Concepts높음R&D Automation Index, immutable verifier, closed-world action resolution, embedded evaluator
Introduction높음frontier-lab에서 AL4 AI-led R&D 약 26%, AL5는 관찰되지 않았다는 operational evidence
Motivation & Background중간auto-research의 token/API-cost 병목이 감소 중이라는 증거
Challenges매우 높음self-evaluation bias, monitor false negatives, specification incompleteness, dynamic tool namespaces
Research Questions매우 높음AL4→AL5 전환 병목과 verification capacity가 R&D automation 속도를 따라가는지 평가
Approaches매우 높음frozen external evaluation, formal proof obligations, closed-world resolver, embedded third-party audit
Key Applications높음autonomous coding/R&D platform, tool-using agents, scientific agents
Open Problems매우 높음auto-R&D 효율 향상이 실제 recursive capability acceleration로 전환되는 임계조건
Future Directions매우 높음Verification-Governed AI R&D + Independent Evaluation + Immutable Action Boundary
§9 · Final Synthesis

Verification-Governed AI R&D

이번 브리핑의 핵심은 AI 연구 자동화가 실제 조직 내부에서 정량화할 만큼 높아졌다는 사실과, 그럼에도 인간 없는 RSI가 아직 관찰되지 않았다는 사실을 동시에 놓치지 않는 데 있다.

앞으로의 핵심 연구문제는 AI가 얼마나 빨리 개선되는가만이 아니다. 그보다 빠르게 독립적 검증·중단·권한통제를 확장할 수 있는가가 더 직접적인 안전 병목이 된다.Verification-governed AI R&D

Research agendaAL4→AL5 전환의 병목, capability–safety joint optimization, specification gaming, verifier isolation, dynamic namespace governance, embedded-evaluator independence를 한 프레임에서 연결하는 연구가 필요하다.

References

이번 업데이트의 직접 근거

본문은 첨부된 AI 위험·RSI 연구 추적 브리핑이 정리한 공식 보고서·preprint·거버넌스 발표의 수치와 한계를 기준으로 재구성했다.

[01]
Measurements for understanding the pace of AI development inside frontier labs
Anthropic Institute · 2026-09-17 · Official methodology report
AI R&D automation AL0–AL5, AL4 비중, monitoring scale, safety compute snapshot, self-measurement limitations. Official report
[02]
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
arXiv:2609.20519 · 2026-09-17 · Preprint
대규모 harness auto-research, token/cost efficiency, transferable mechanisms, capability trade-offs. arXiv
[03]
MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs
arXiv:2609.19391 · 2026-09-16 · Preprint
frozen safety specification, Dafny-based formalization, verifier feedback, proof-carrying execution boundary. arXiv
[04]
Closed-World Resolution Against Tool Hallucination in LLM Agents
arXiv:2609.19425 · 2026-09-16 · Preprint
registry membership, signature validation, tool hallucination, namespace collision/shadowing. arXiv
[05]
Partnering with Accenture on embedded evaluation
Anthropic · 2026-09-18 · Governance announcement
frontier-model development lifecycle에 외부 evaluator를 상시 삽입하려는 평가·거버넌스 파트너십. Official announcement