AI Research NotesVision-First AGI · World Models · Embodiment · Spatial Intelligence
Visual General IntelligenceWhite Paper · arXiv:2608.25924v126 Aug 2026 · CVPR 2026 VGI Workshop

언어보다 먼저,
세계는 보였고
지능은 그 구조를 배웠다

Visual General Intelligence: A White Paper — What evidence would convince us that intelligence has begun to arise from vision?

VISUAL EXPERIENCEPREDICTIMAGINERECONSTRUCTPERSISTENTWORLD MODELACTIVE LOOKACTION / TESTVGIGENERALISESee → infer structure → imagine and intervene → revise what is known
Central Question

이 백서는 AGI를 언어의 연장선으로만 보지 않는다. 질문은 더 오래된 감각으로 돌아간다. 이미지, 비디오, 기하, 움직임 같은 시각경험 자체에서 일반화 가능한 지능이 생겨날 수 있는가?

저자들은 하나의 VGI 정의, 하나의 모델, 하나의 벤치마크를 선언하지 않는다. computer vision이 AGI 시대에 다시 물어야 할 원칙을 여러 관점에서 펼친다. 시각은 언어를 거부하는 대안이 아니라, 언어에 결합되기 전에도 세계의 구조를 배우고 예측·상상·복원·행동할 수 있는 grounded 경로일 수 있다는 연구의제다.

Authors & Affiliations

Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian, Shangzhe Wu, Oishi Deb, Ryousuke Yamada, Christian Rupprecht, Jianyuan Wang, Kohsuke Ide, Koichi Namekata, Xianzheng Ma, Yiming Chen, Robert Geirhos, Aditi Raghunathan, Yuki M. Asano, Deva Ramanan, David Fouhey, Andrew J. Davison, Yilun Du, Jiajun Wu, Zhuang Liu. 소속은 AIST, Oxford VGG, CADDi, OpenAI, Cambridge, University of Technology Nuremberg, University of Tsukuba, Google DeepMind, Carnegie Mellon University, New York University, Imperial College London, Harvard, Stanford, Princeton이다.

VGI의 핵심은 더 강한 image classifier나 LLM에 붙은 visual encoder가 아니다. 시각경험에서 세계의 구조를 발견하고, 그 지식을 유지·수정하며, 가능한 미래를 예측하고, 보이지 않는 구조를 복원하고, 행동으로 지식을 시험하는 능력이다.White-paper-wide synthesis
Part I · Why Vision?

지능을 언어로 환원할 이유는 없다

시각은 단순 센서가 아니라 세계구조를 압축해 배우는 오래된 지능기능일 수 있다.

§1 · Evolutionary Premise

시각은 언어보다 훨씬 오래되었다

백서는 초기 AI와 cybernetics의 지능관—환경신호에서 표현을 얻고 규칙성을 포착하며, 미지 상황에 일반화하고, 예측이나 행동으로 연결하는 능력—에서 출발한다. camera-type eye와 compound eye 같은 고등 시각계는 Cambrian 시기까지 거슬러 올라가지만 인간의 고등언어와 문자는 생명사 전체에서 매우 최근이다. 저자들은 이 대비를 근거로 vision이 단순 입력 modality보다 세계구조를 포착하는 foundational intelligent function일 가능성을 제기한다.

§2 · The LLM Lesson, Reopened

GPT의 교훈을 vision에 그대로 복사하지는 않는다

언어에서는 web-scale text, 단순한 autoregressive objective, model/data scaling이 in-context learning과 few-shot adaptation 같은 폭넓은 능력으로 이어졌다. Sutton의 Bitter Lesson 역시 computation·data·learning·search와 함께 확장되는 일반방법의 힘을 상기시킨다.

하지만 백서는 “video에 next-token prediction을 그대로 적용하면 AGI”라고 결론내리지 않는다. single-modality scaling, video·geometry·spacetime 통합, generation·restoration·reconstruction의 결합 등 여러 가능성을 열어 둔다. 핵심은 VFM의 크기 자체가 아니라 어떤 학습전략·데이터구조·모델구조·평가가 unknown task와 environment로 일반화를 만드는가이다.

§3 · Vision and Language

VGI는 language rejection이 아니다

CLIP, Flamingo, BLIP/BLIP-2, LLaVA 같은 VLM/MLLM은 vision-language 연결의 강력한 축이다. 다만 image feature, image-text pair, instruction tuning, LLM reasoning이 얽혀 있어 어떤 능력이 visual experience에서 나오고 어떤 능력이 language model에서 나오는지 분리하기 어렵다. 그래서 논문은 vision-only, vision-first, language-mediated, multimodal 관점을 함께 유지한다.

Part II · Scale, Video & Creativity

생성은 인식보다 더 많은 세계를 설명해야 한다

Robert Geirhos와 Aditi Raghunathan은 generation을 generality와 creativity의 시험대로 본다.

§4 · Video Models as VFMs

Geirhos: video generation은 시각과제의 공통 인터페이스가 될 수 있다

Geirhos는 visual intelligence를 새 task마다 별도 training 없이 perception부터 reasoning까지 다양한 visual task를 풀 수 있는 능력으로 본다. 그의 주장은 large-scale training + generative objective + video에 집중된다. 분류는 texture나 background shortcut으로도 성공할 수 있지만 generation은 object, background, shape, texture, light, shadow를 함께 맞춰야 하므로 더 어려운 학습신호라는 논리다.

특히 video generation은 still image를 특수한 경우로 포함하는 일반적 visual framework다. 백서는 Veo 3 같은 generative video model이 task-specific training 없이 image-to-video generation만으로 edge detection, segmentation, keypoint localization, super-resolution, editing, style transfer, maze solving, graph traversal까지 보인 연구를 근거로 video model을 ‘visual foundation model 1.0’으로 보는 관점을 소개한다. 이는 white paper의 한 perspective이지 VGI 달성의 합의된 결론은 아니다.

§5 · Creativity as a Test

Raghunathan: 한 개의 정답보다 여러 개의 구조적으로 다른 답

open-ended visual intelligence는 정확성만으로 충분하지 않다. 기능제약이 있는 object design, robot plan, scientific diagram, physical scene의 여러 미래를 상상하려면 output이 coherent하면서 반복생성 사이에 structurally diverse하고 training experience와 비교해 original해야 한다.

논문은 combinational creativity와 exploratory creativity를 구분한다. 전자는 익숙한 요소 사이의 새 연결을, 후자는 규칙과 제약 아래 새로운 pattern·protein·mechanism·3D structure·action plan을 만드는 능력을 뜻한다. 평가축은 coherence, structural diversity, originality, utility다. controlled task에서는 teacherless/multi-token 또는 diffusion 계열이 conventional next-token learning보다 다양하고 독창적인 solution을 낸 결과도 인용한다. seed-conditioning으로 high-level possibility를 먼저 선택해 local randomness가 아니라 전체 plan의 다양성을 만들자는 아이디어도 제시된다.

Part III · Learning Through a Lifetime

한 번 학습하고 얼어붙는 모델은 열린 세계의 지능이 아니다

Yuki M. Asano와 Deva Ramanan et al.은 continual experience, multimodal raw signals, efficiency를 강조한다.

§6 · Learning from One Datum

Asano: visual lifetime 전체를 학습과정으로 만든다

현재 강한 visual model도 large curated dataset으로 한 번 training된 뒤 사실상 fixed model로 배치된다. Asano는 VGI라면 labels 없이, past data를 반복 reshuffle하지 않고, future task를 미리 알지 못해도 visual experience가 흘러가는 순서 그대로 계속 학습해야 한다고 본다.

그가 그리는 artificial visual brain은 perceptual system, spatial/dynamic world model, multi-timescale memory, task/action interface를 가진다. 이 모듈은 rigid symbolic component일 필요는 없지만 역할·update rule·timescale이 달라야 한다. depth, motion, persistence, containment, contact, affordance는 이름 붙기 전에도 존재하므로 language는 grounded world model을 query·steer·teach하는 interface가 될 수 있지만 sole organizing principle이 되어서는 안 된다는 주장이다.

§7 · Multimodal, Generative, Efficient

Ramanan et al.: vision alone가 아니라 multi-sensory world modeling

embodied intelligence에는 vision뿐 아니라 audio, tactile, proprioception이 중요하다. 현재 많은 multimodal model은 사실상 single-modal pretraining 뒤 adaptor로 modality를 붙인다. 이 관점은 raw multi-modal stream 자체를 공동 pretraining하고, GelSight 같은 near-field vision을 통해 tactile 구조까지 연결하는 방향을 제안한다.

동시에 generation을 self-supervised learning의 자연스러운 연장으로 보고 image/video뿐 아니라 multi-sensory signal generation을 탐색한다. 하지만 대규모 video generation의 비용 때문에 efficiency는 부차적 engineering 문제가 아니라 VGI 연구를 누가 수행할 수 있는지를 결정하는 연구조건이다. recurrence, 3D, multi-scale processing, stateful compact memory, 낮은 resolution/frame-rate의 long-horizon modeling이 후보로 제시된다.

또 하나의 도발적 주장은 data scale보다 data diversity가 핵심일 수 있다는 것이다. 반복적인 백만 장보다 다양한 경험이 중요하며, 올바른 data curation과 learner-driven collection을 위해 RL을 활용할 수 있다는 관점이다.

Part IV · Discovery & Spatial AI

과학에서는 ground truth가 없고, 공간에서는 기억이 지능이 된다

David Fouhey와 Andrew J. Davison은 VGI를 real-world validation과 persistent spatial state의 문제로 바꾼다.

§8 · Visual Intelligence and Discovery

Fouhey: 과학용 vision은 leaderboard보다 validation이 어렵다

과학발견에서는 더 많은 data를 수집할 수 없는 문제가 많다. 1970년의 specimen을 다시 채집할 수 없고, 태양물리는 제한된 각도와 시간의 photon만 관측하며, 진화생태학은 수천만 년의 과정을 다룬다. 따라서 일반 CV처럼 data volume을 knob처럼 올리는 전략이 맞지 않을 수 있다.

더 큰 문제는 ground truth 부재다. 새로운 measurement를 만들었을 때 boldface할 leaderboard number가 없고, 기존 물리법칙·다른 instrument·과거 data와 reconcile하는 긴 validation이 필요하다. multimodality는 입력채널 확대가 아니라 서로 다른 instrument와 equation을 연결해 reality check를 만드는 수단이다.

scientific sensor에는 unit와 systematics가 있다. model은 진짜 physical signal뿐 아니라 instrument artifact도 충실히 학습할 수 있으므로 real signal과 non-physical signal을 분리해야 한다. simulation도 완전한 답이 아니다. frontier science에서 새로운 관측이 필요한 이유 자체가 현실모델이 불완전하기 때문이다. Fouhey의 결론은 vision component가 discovery pipeline 전체의 일부이며 scientist가 실제로 묻는 것은 “AI data에 systematics가 있는가, 오류를 거의 모두 잡을 수 있는가”라는 점이다.

§9 · Spatial AI

Davison: world representation은 persistent하지만 adaptable해야 한다

Spatial AI는 주변환경의 rich하지만 efficient한 representation을 지속적으로 만들고, 이를 통해 사람처럼 일반적 interaction을 수행하는 능력으로 정의된다. SLAM은 partial observation을 누적해 long-horizon behavior를 가능하게 한 대표적 선례다. sparse map에서 dense map, semantic label, object-level scene graph, physics/dynamics property로 representation이 확장되어 왔다.

Davison은 end-to-end neural인지 handcrafted인지보다 전체 시스템의 storage and computational structure가 중요하다고 본다. future hardware는 fine-grained parallel cores, distributed local memory, sparse communication, asynchronous/event-driven operation, low-bit 혹은 analog representation으로 갈 가능성을 논의한다. Spatial AI의 graph structure를 hardware locality와 맞추고 loopy graph message passing—예컨대 belief propagation—을 활용하는 방향도 제안한다.

Part V · Embodiment & Physical Structure

보기만 하는 지능에서 개입하고 틀렸음을 배우는 지능으로

Yilun Du는 action-feedback loop를, Shangzhe Wu와 Jiajun Wu는 실행·검증 가능한 physical structure를 강조한다.

§10 · Visual Intelligence Is Embodied Intelligence

Du: action은 output이면서 새로운 evidence를 얻는 방법이다

VGI의 핵심 시험은 robot이 unfamiliar physical world에서 perceive, reason, act reliably할 수 있는가이다. embodiment는 downstream application이 아니다. viewpoint를 바꾸고 object에 개입하면 passive observation으로 알 수 없던 property를 드러내고 prediction을 시험하며 world model을 수정할 evidence를 얻는다.

이를 위해 generative world model, persistent scene representation, action-grounded planning, active visual learning, continual adaptation이 하나의 closed loop로 연결된다. high-level visual trajectory를 generative model이 제안하고 low-level controller가 continuous action으로 실행한 뒤 predicted outcome과 observed outcome을 비교하면 execution 자체가 plan과 world model의 시험이 된다.

SILVR는 agent interaction trajectory로 video-based robotic planner를 개선하고, World Action Verifier는 prediction failure를 targeted data collection과 world-model refinement에 이용한다. 새 경험을 흡수하면서 catastrophic forgetting을 막기 위한 wake–sleep 형태의 online experience / periodic consolidation 분리도 제안된다. 행동은 지각의 결과가 아니라 어떤 경험을 학습할지 선택하는 data-acquisition policy가 된다.

§11 · Seeing the Physical World via Code

Wu & Wu: photorealism은 physical understanding의 증거가 아니다

image는 physical phenomenon의 measurement다. underlying scene에는 entity, part hierarchy, geometry, material, mass, friction, stiffness, pose, lighting, camera, relation, dynamics가 존재한다. 보는 것은 pixel statistics를 흉내 내는 것보다 이 latent structure를 역으로 복원하는 일로 볼 수 있다.

convincing falling cup video를 만들었다고 mass·contact·friction을 표현했다고 말할 수는 없다. intervention, editing, action을 요구할 때 그 차이가 드러난다. 그래서 일부 structure는 explicit learning target이 되어야 할 수 있다.

Continuous

learned embedding은 expressive하고 optimization이 쉽지만 읽기·수정·검증이 어렵다.

Discrete / Symbolic

program과 named relation은 읽기·수정·검증이 쉽지만 vocabulary 밖으로 확장되기 어렵다.

논문은 두 표현을 결합하는 hybrid representation을 강조한다. Scene Language는 hierarchical program + semantic words + visual embedding을 결합하고, neuro-symbolic concept는 “left of”, “put down” 같은 개념을 symbolic program과 perception/action에 grounded된 neural network로 공동 표현한다.

더 흥미로운 연결은 coding agent다. physical structure가 code로 표현되면 program은 compile·execute·render·simulate되고 image나 motion과 비교해 수정될 수 있다. 이미 articulated 3D asset과 robot manipulation code에서 이런 loop가 나타난다. 다만 현재 coding agent가 language prior와 인간이 제공한 abstraction에 크게 의존하며, 관측만으로 abstraction 자체를 발견할 수 있는지는 열린 문제다.

Part VI · Vision-Native Intelligence

언어의 token을 vision에 억지로 찾기보다, 무엇을 볼지 스스로 배워야 한다

Zhuang Liu와 Kataoka et al.은 vision-native abstraction과 prediction–imagination–reconstruction의 합류를 제시한다.

§12 · When, Where, How Much to Look

Liu: pixels도 patches도 intelligence의 자연단위가 아니다

language는 word–phrase–sentence–document라는 인간이 만든 symbolic compression 위에서 학습한다. vision은 raw experience에 가깝고 pixel은 너무 낮은 수준이며 patch는 engineering convenience일 뿐이다. scene–object–part–relation 같은 hierarchy가 중요해 보이지만 language token에 대응하는 합의된 visual vocabulary는 없다.

모든 region을 최고해상도로 표현하면 compute가 낭비되고 너무 coarse하면 중요한 detail을 놓친다. useful visual unit는 task-dependent하다. 따라서 VGI의 중요한 시험은 주어진 image를 잘 푸는가가 아니라 언제 볼지, 어디를 볼지, 얼마나 자세히 볼지 스스로 결정하는가이다.

Liu는 scaling을 부정하지 않지만 raw scale과 visual diversity를 구분한다. dataset이 크더라도 repetitive·narrow·biased할 수 있다. 또한 현재 interface는 사람이 이미지를 capture/upload하고 중요성을 설명해야 하지만 robot·laboratory assistant·medical system은 world state를 스스로 perceive하고 change를 detect하며 feedback으로 action을 교정해야 한다.

§13 · Prediction + Imagination + Reconstruction

Kataoka et al.: 세 learning objective의 convergence

이 관점에서 visual intelligence는 partial visual observation에서 usable world structure를 배우고 hidden structure를 이해하며, 다음을 예측하고, 보이지 않는 것을 복원하고, 가능한 alternative를 상상하고, 그 지식을 새로운 visual task에 일반화하는 능력이다. human vision이 upper bound일 이유도 없다. 인간보다 더 넓고 정밀한 visual information에 노출된 machine은 human-centered perception을 넘어설 가능성이 있다.

Sequential Prediction

what comes next. coherent future prediction을 위해 object, geometry, motion, occlusion, physical constraint, camera movement, agent behavior, temporal causality, uncertainty를 학습한다.

Open-Ended Generation

what could exist. alternative future와 counterfactual을 상상하고 internal simulator처럼 consequence를 시험할 수 있다.

Reconstruction

what is hidden. inverse rendering, depth, multi-view, 3D/4D reconstruction을 통해 incomplete observation에서 hidden cause와 physical consistency를 회복한다.

이 세 objective가 VFM에서 visual intelligence, 잠재적으로 general intelligence로 가는 상보적 경로라는 것이 저자들의 central hypothesis다. true visual in-context learning도 중요한 가능성이다. 몇 개 visual example만 보고 새 task structure와 output format을 추론해 parameter update 없이 해결하는 능력이다. 그러나 world model만으로 intelligence가 완성되는 것은 아니다. 새 goal 아래 relevant abstraction을 선택하고 counterfactual을 만들고 new observation에 적응하며 행동해야 한다.

Part VII · What Would Count as Evidence?

VGI의 승부는 더 큰 모델이 아니라 무엇을 ‘지능의 증거’로 인정할지에 있다

백서의 plurality는 약점이 아니라 연구공간을 좁히지 않기 위한 의도적 결론이다.

§14 · Ten Perspectives at a Glance

Table 1을 다시 읽으면 하나의 system diagram이 보인다

ContributorCentral position
Robert GeirhosGenerative video models will become visual foundation models.
Aditi RaghunathanCreativity—coherence, structural diversity, originality, utility—should be a central VGI test.
Yuki M. AsanoVGI should continuously and autonomously learn from visual experience rather than remain fixed after pre-training.
Deva Ramanan et al.Visual intelligence should be multimodal, generative, computationally efficient, and include touch/audio/proprioception.
David FouheyScientific visual intelligence must work with limited data and validate against physical knowledge, instruments, units, and systematics.
Andrew J. DavisonSpatial AI needs persistent world representations integrating state estimation, learning, real-time computation, and efficient hardware.
Yilun DuVisual intelligence is embodied intelligence linking perception, world models, action, active exploration, and continual adaptation.
Shangzhe Wu & Jiajun WuVGI should recover compositional physical structure supporting editing, simulation, verification, and action.
Zhuang LiuVGI should be vision-native and learn when to look, where to look, and how much detail is needed.
Hirokatsu Kataoka et al.Prediction, imagination, and structural understanding may jointly provide a path from visual experience to general intelligence.
§15 · Shared Research Agenda

서로 다른 관점은 결국 같은 폐루프의 다른 부분이다

summary는 방향들을 scaling/generative modeling, creativity·prediction·imagination·reconstruction, continual/lifelong learning, persistent 3D/spatial representation, embodiment·active perception·action, multimodal learning·efficiency, scientific discovery·unfamiliar-environment generalization, compositional physical structure, language-centered가 아닌 vision-native intelligence로 묶는다.

공통 premise는 visual experience로부터 world structure를 discover and verify하고, knowledge를 유지·수정하며, possible future를 예측하고, alternative를 상상하고, unseen structure를 복원하고, unfamiliar situation에 behavior를 적응하는 것이다. description이나 reproduction뿐 아니라 representation이 edit·simulate되고 reliable action에 쓰일 수 있어야 한다.

§16 · Benchmark Redesign

static accuracy를 넘어 무엇을 측정할 것인가

Transfer
Unknown Tasks
새 task에 task-specific retraining 없이 일반화하는가.
Adapt
Continual
새 경험을 흡수하면서 이전지식을 유지하는가.
Seek
Active Look
불확실성을 줄이기 위해 무엇을 더 볼지 선택하는가.
Remember
Persistent State
world knowledge를 장기 유지·수정하는가.
Create
Creativity
coherent·diverse·original alternative를 만드는가.
Ground
Physics
appearance가 아니라 intervention 가능한 structure를 아는가.
Act
Reliability
unfamiliar setting에서 perception을 action으로 연결하는가.
Efficient
Compute
long-horizon real-time setting에서 지속 가능한가.

논문은 단일 benchmark가 이 모든 차원을 측정할 가능성이 낮다고 본다. isolated static tasks만으로 VGI를 판별하기 어렵다는 것이 더 중요한 결론이다.

§17 · Epistemic Boundary

이것은 VGI가 이미 달성되었다는 선언이 아니다

현재 system은 중요한 fragment를 보여준다. generative video model은 cross-task capability를 보이고, reconstruction model은 geometric/physical structure를 더 많이 회복하며, spatial system은 persistent representation을 유지하고, creative model은 alternative를 만들며, embodied agent는 interaction을 통해 perception과 behavior를 개선한다.

그러나 이 능력들은 여전히 fragmented, computationally expensive, temporally limited, data/training-condition dependent하다. 따라서 더 가까운 milestone은 “VGI가 이미 AGI의 길인가”를 묻는 것보다 visual experience 자체에서 intelligence emergence를 관찰할 수 있는지 검증하는 것이다.

설득력 있는 증거는 고정 visual task의 점수상승을 넘어 transferable capability, persistent and revisable world knowledge, creative and counterfactual reasoning, active information seeking, continual adaptation, spatial/physical consistency, reliable action을 함께 보여줘야 한다.Conclusion-level synthesis
§18 · Program Context

이 백서는 CVPR 2026 VGI Workshop에서 시작된 research agenda다

AIST 연구진의 작업은 일본 METI와 NEDO가 주도하는 FRONTia—AI robot과 Physical AI를 위한 domestic multimodal foundation model 개발 프로그램—아래 수행되었고 JST ASPIRE(JPMJAP2518)의 지원도 받았다. workshop에는 Matt Deitke, Alexei A. Efros, Kristen Grauman, Alex Kendall, Andrea Vedaldi 등의 초청발표가 논의에 기여했다.

하나의 architecture를 확정하는 논문이 아니라, computer vision community가 AGI 시대에 무엇을 측정하고 어떤 evidence를 요구할지 연구공간을 넓히는 position map으로 읽는 것이 정확하다.

Source & Reference Map

원문과 핵심 선행연구

01
Visual General Intelligence: A White Paper
Kataoka et al. · arXiv:2608.25924v1 · 26 Aug 2026
VGI의 10개 perspective를 모은 primary source. arXiv
02
CVPR 2026 Workshop on Visual General Intelligence
Workshop · 2026
본 white paper의 토론이 시작된 workshop. Workshop site
03
The Bitter Lesson
Richard Sutton · 2019
scale with computation, data, learning, search에 관한 배경.
04
Language Models are Few-Shot Learners
Brown et al. · NeurIPS 2020
large-scale autoregressive learning과 in-context/few-shot capability의 언어모델 선례.
05
Learning Transferable Visual Models from Natural Language Supervision
Radford et al. · ICML 2021
CLIP과 image-text contrastive learning.
06
Video Models are Zero-Shot Learners and Reasoners
Wiedemer et al. · 2025
generative video model의 cross-task zero-shot visual capability.
07
FutureMapping / FutureMapping 2
Davison 2018 · Davison & Ortiz 2019
Spatial AI의 computational structure와 graph/belief-propagation 관점.
08
MASt3R-SLAM
Murai et al. · CVPR 2025
real-time dense SLAM과 3D reconstruction prior.
09
3D-MEM: 3D Scene Memory for Embodied Exploration and Reasoning
Yang et al. · CVPR 2025
persistent scene-level 3D memory.
10
World Action Verifier
Liu et al. · 2026
prediction failure를 targeted data collection과 world-model refinement에 이용.
11
Self-Improving Loops for Visual Robotic Planning
Luo et al. · ICLR 2026
agent interaction trajectory를 활용한 visual planner improvement.
12
The Neuro-Symbolic Concept Learner
Mao et al. · ICLR 2019
perception과 symbolic concept grounding을 결합한 기반연구.
13
Building Intelligent Agents with Neuro-Symbolic Concepts
Mao, Tenenbaum & Wu · CACM 2026
perception-action grounded neuro-symbolic concepts.
14
The Scene Language
Zhang et al. · CVPR 2025
programs, words, embeddings를 결합한 hybrid scene representation.
15
ArtiCraft
Zhou et al. · 2026
agentic articulated 3D asset generation.

원문 bibliography는 115개 항목으로 Turing·Wiener·Marr부터 VLM, video reasoning, SLAM, embodied world models, scientific vision, dataset bias, neuro-symbolic concepts, 3D/4D reconstruction과 robot planning까지 폭넓게 연결한다. 위 목록은 본문 논지를 읽기 위한 대표 reference map이며 전체 인용목록은 원문에 보존되어 있다.