Continuous
learned embedding은 expressive하고 optimization이 쉽지만 읽기·수정·검증이 어렵다.
Visual General Intelligence: A White Paper — What evidence would convince us that intelligence has begun to arise from vision?
이 백서는 AGI를 언어의 연장선으로만 보지 않는다. 질문은 더 오래된 감각으로 돌아간다. 이미지, 비디오, 기하, 움직임 같은 시각경험 자체에서 일반화 가능한 지능이 생겨날 수 있는가?
저자들은 하나의 VGI 정의, 하나의 모델, 하나의 벤치마크를 선언하지 않는다. computer vision이 AGI 시대에 다시 물어야 할 원칙을 여러 관점에서 펼친다. 시각은 언어를 거부하는 대안이 아니라, 언어에 결합되기 전에도 세계의 구조를 배우고 예측·상상·복원·행동할 수 있는 grounded 경로일 수 있다는 연구의제다.
Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian, Shangzhe Wu, Oishi Deb, Ryousuke Yamada, Christian Rupprecht, Jianyuan Wang, Kohsuke Ide, Koichi Namekata, Xianzheng Ma, Yiming Chen, Robert Geirhos, Aditi Raghunathan, Yuki M. Asano, Deva Ramanan, David Fouhey, Andrew J. Davison, Yilun Du, Jiajun Wu, Zhuang Liu. 소속은 AIST, Oxford VGG, CADDi, OpenAI, Cambridge, University of Technology Nuremberg, University of Tsukuba, Google DeepMind, Carnegie Mellon University, New York University, Imperial College London, Harvard, Stanford, Princeton이다.
시각은 단순 센서가 아니라 세계구조를 압축해 배우는 오래된 지능기능일 수 있다.
백서는 초기 AI와 cybernetics의 지능관—환경신호에서 표현을 얻고 규칙성을 포착하며, 미지 상황에 일반화하고, 예측이나 행동으로 연결하는 능력—에서 출발한다. camera-type eye와 compound eye 같은 고등 시각계는 Cambrian 시기까지 거슬러 올라가지만 인간의 고등언어와 문자는 생명사 전체에서 매우 최근이다. 저자들은 이 대비를 근거로 vision이 단순 입력 modality보다 세계구조를 포착하는 foundational intelligent function일 가능성을 제기한다.
언어에서는 web-scale text, 단순한 autoregressive objective, model/data scaling이 in-context learning과 few-shot adaptation 같은 폭넓은 능력으로 이어졌다. Sutton의 Bitter Lesson 역시 computation·data·learning·search와 함께 확장되는 일반방법의 힘을 상기시킨다.
하지만 백서는 “video에 next-token prediction을 그대로 적용하면 AGI”라고 결론내리지 않는다. single-modality scaling, video·geometry·spacetime 통합, generation·restoration·reconstruction의 결합 등 여러 가능성을 열어 둔다. 핵심은 VFM의 크기 자체가 아니라 어떤 학습전략·데이터구조·모델구조·평가가 unknown task와 environment로 일반화를 만드는가이다.
CLIP, Flamingo, BLIP/BLIP-2, LLaVA 같은 VLM/MLLM은 vision-language 연결의 강력한 축이다. 다만 image feature, image-text pair, instruction tuning, LLM reasoning이 얽혀 있어 어떤 능력이 visual experience에서 나오고 어떤 능력이 language model에서 나오는지 분리하기 어렵다. 그래서 논문은 vision-only, vision-first, language-mediated, multimodal 관점을 함께 유지한다.
Robert Geirhos와 Aditi Raghunathan은 generation을 generality와 creativity의 시험대로 본다.
Geirhos는 visual intelligence를 새 task마다 별도 training 없이 perception부터 reasoning까지 다양한 visual task를 풀 수 있는 능력으로 본다. 그의 주장은 large-scale training + generative objective + video에 집중된다. 분류는 texture나 background shortcut으로도 성공할 수 있지만 generation은 object, background, shape, texture, light, shadow를 함께 맞춰야 하므로 더 어려운 학습신호라는 논리다.
특히 video generation은 still image를 특수한 경우로 포함하는 일반적 visual framework다. 백서는 Veo 3 같은 generative video model이 task-specific training 없이 image-to-video generation만으로 edge detection, segmentation, keypoint localization, super-resolution, editing, style transfer, maze solving, graph traversal까지 보인 연구를 근거로 video model을 ‘visual foundation model 1.0’으로 보는 관점을 소개한다. 이는 white paper의 한 perspective이지 VGI 달성의 합의된 결론은 아니다.
open-ended visual intelligence는 정확성만으로 충분하지 않다. 기능제약이 있는 object design, robot plan, scientific diagram, physical scene의 여러 미래를 상상하려면 output이 coherent하면서 반복생성 사이에 structurally diverse하고 training experience와 비교해 original해야 한다.
논문은 combinational creativity와 exploratory creativity를 구분한다. 전자는 익숙한 요소 사이의 새 연결을, 후자는 규칙과 제약 아래 새로운 pattern·protein·mechanism·3D structure·action plan을 만드는 능력을 뜻한다. 평가축은 coherence, structural diversity, originality, utility다. controlled task에서는 teacherless/multi-token 또는 diffusion 계열이 conventional next-token learning보다 다양하고 독창적인 solution을 낸 결과도 인용한다. seed-conditioning으로 high-level possibility를 먼저 선택해 local randomness가 아니라 전체 plan의 다양성을 만들자는 아이디어도 제시된다.
Yuki M. Asano와 Deva Ramanan et al.은 continual experience, multimodal raw signals, efficiency를 강조한다.
현재 강한 visual model도 large curated dataset으로 한 번 training된 뒤 사실상 fixed model로 배치된다. Asano는 VGI라면 labels 없이, past data를 반복 reshuffle하지 않고, future task를 미리 알지 못해도 visual experience가 흘러가는 순서 그대로 계속 학습해야 한다고 본다.
그가 그리는 artificial visual brain은 perceptual system, spatial/dynamic world model, multi-timescale memory, task/action interface를 가진다. 이 모듈은 rigid symbolic component일 필요는 없지만 역할·update rule·timescale이 달라야 한다. depth, motion, persistence, containment, contact, affordance는 이름 붙기 전에도 존재하므로 language는 grounded world model을 query·steer·teach하는 interface가 될 수 있지만 sole organizing principle이 되어서는 안 된다는 주장이다.
embodied intelligence에는 vision뿐 아니라 audio, tactile, proprioception이 중요하다. 현재 많은 multimodal model은 사실상 single-modal pretraining 뒤 adaptor로 modality를 붙인다. 이 관점은 raw multi-modal stream 자체를 공동 pretraining하고, GelSight 같은 near-field vision을 통해 tactile 구조까지 연결하는 방향을 제안한다.
동시에 generation을 self-supervised learning의 자연스러운 연장으로 보고 image/video뿐 아니라 multi-sensory signal generation을 탐색한다. 하지만 대규모 video generation의 비용 때문에 efficiency는 부차적 engineering 문제가 아니라 VGI 연구를 누가 수행할 수 있는지를 결정하는 연구조건이다. recurrence, 3D, multi-scale processing, stateful compact memory, 낮은 resolution/frame-rate의 long-horizon modeling이 후보로 제시된다.
또 하나의 도발적 주장은 data scale보다 data diversity가 핵심일 수 있다는 것이다. 반복적인 백만 장보다 다양한 경험이 중요하며, 올바른 data curation과 learner-driven collection을 위해 RL을 활용할 수 있다는 관점이다.
David Fouhey와 Andrew J. Davison은 VGI를 real-world validation과 persistent spatial state의 문제로 바꾼다.
과학발견에서는 더 많은 data를 수집할 수 없는 문제가 많다. 1970년의 specimen을 다시 채집할 수 없고, 태양물리는 제한된 각도와 시간의 photon만 관측하며, 진화생태학은 수천만 년의 과정을 다룬다. 따라서 일반 CV처럼 data volume을 knob처럼 올리는 전략이 맞지 않을 수 있다.
더 큰 문제는 ground truth 부재다. 새로운 measurement를 만들었을 때 boldface할 leaderboard number가 없고, 기존 물리법칙·다른 instrument·과거 data와 reconcile하는 긴 validation이 필요하다. multimodality는 입력채널 확대가 아니라 서로 다른 instrument와 equation을 연결해 reality check를 만드는 수단이다.
scientific sensor에는 unit와 systematics가 있다. model은 진짜 physical signal뿐 아니라 instrument artifact도 충실히 학습할 수 있으므로 real signal과 non-physical signal을 분리해야 한다. simulation도 완전한 답이 아니다. frontier science에서 새로운 관측이 필요한 이유 자체가 현실모델이 불완전하기 때문이다. Fouhey의 결론은 vision component가 discovery pipeline 전체의 일부이며 scientist가 실제로 묻는 것은 “AI data에 systematics가 있는가, 오류를 거의 모두 잡을 수 있는가”라는 점이다.
Spatial AI는 주변환경의 rich하지만 efficient한 representation을 지속적으로 만들고, 이를 통해 사람처럼 일반적 interaction을 수행하는 능력으로 정의된다. SLAM은 partial observation을 누적해 long-horizon behavior를 가능하게 한 대표적 선례다. sparse map에서 dense map, semantic label, object-level scene graph, physics/dynamics property로 representation이 확장되어 왔다.
Davison은 end-to-end neural인지 handcrafted인지보다 전체 시스템의 storage and computational structure가 중요하다고 본다. future hardware는 fine-grained parallel cores, distributed local memory, sparse communication, asynchronous/event-driven operation, low-bit 혹은 analog representation으로 갈 가능성을 논의한다. Spatial AI의 graph structure를 hardware locality와 맞추고 loopy graph message passing—예컨대 belief propagation—을 활용하는 방향도 제안한다.
Yilun Du는 action-feedback loop를, Shangzhe Wu와 Jiajun Wu는 실행·검증 가능한 physical structure를 강조한다.
VGI의 핵심 시험은 robot이 unfamiliar physical world에서 perceive, reason, act reliably할 수 있는가이다. embodiment는 downstream application이 아니다. viewpoint를 바꾸고 object에 개입하면 passive observation으로 알 수 없던 property를 드러내고 prediction을 시험하며 world model을 수정할 evidence를 얻는다.
이를 위해 generative world model, persistent scene representation, action-grounded planning, active visual learning, continual adaptation이 하나의 closed loop로 연결된다. high-level visual trajectory를 generative model이 제안하고 low-level controller가 continuous action으로 실행한 뒤 predicted outcome과 observed outcome을 비교하면 execution 자체가 plan과 world model의 시험이 된다.
SILVR는 agent interaction trajectory로 video-based robotic planner를 개선하고, World Action Verifier는 prediction failure를 targeted data collection과 world-model refinement에 이용한다. 새 경험을 흡수하면서 catastrophic forgetting을 막기 위한 wake–sleep 형태의 online experience / periodic consolidation 분리도 제안된다. 행동은 지각의 결과가 아니라 어떤 경험을 학습할지 선택하는 data-acquisition policy가 된다.
image는 physical phenomenon의 measurement다. underlying scene에는 entity, part hierarchy, geometry, material, mass, friction, stiffness, pose, lighting, camera, relation, dynamics가 존재한다. 보는 것은 pixel statistics를 흉내 내는 것보다 이 latent structure를 역으로 복원하는 일로 볼 수 있다.
convincing falling cup video를 만들었다고 mass·contact·friction을 표현했다고 말할 수는 없다. intervention, editing, action을 요구할 때 그 차이가 드러난다. 그래서 일부 structure는 explicit learning target이 되어야 할 수 있다.
learned embedding은 expressive하고 optimization이 쉽지만 읽기·수정·검증이 어렵다.
program과 named relation은 읽기·수정·검증이 쉽지만 vocabulary 밖으로 확장되기 어렵다.
논문은 두 표현을 결합하는 hybrid representation을 강조한다. Scene Language는 hierarchical program + semantic words + visual embedding을 결합하고, neuro-symbolic concept는 “left of”, “put down” 같은 개념을 symbolic program과 perception/action에 grounded된 neural network로 공동 표현한다.
더 흥미로운 연결은 coding agent다. physical structure가 code로 표현되면 program은 compile·execute·render·simulate되고 image나 motion과 비교해 수정될 수 있다. 이미 articulated 3D asset과 robot manipulation code에서 이런 loop가 나타난다. 다만 현재 coding agent가 language prior와 인간이 제공한 abstraction에 크게 의존하며, 관측만으로 abstraction 자체를 발견할 수 있는지는 열린 문제다.
Zhuang Liu와 Kataoka et al.은 vision-native abstraction과 prediction–imagination–reconstruction의 합류를 제시한다.
language는 word–phrase–sentence–document라는 인간이 만든 symbolic compression 위에서 학습한다. vision은 raw experience에 가깝고 pixel은 너무 낮은 수준이며 patch는 engineering convenience일 뿐이다. scene–object–part–relation 같은 hierarchy가 중요해 보이지만 language token에 대응하는 합의된 visual vocabulary는 없다.
모든 region을 최고해상도로 표현하면 compute가 낭비되고 너무 coarse하면 중요한 detail을 놓친다. useful visual unit는 task-dependent하다. 따라서 VGI의 중요한 시험은 주어진 image를 잘 푸는가가 아니라 언제 볼지, 어디를 볼지, 얼마나 자세히 볼지 스스로 결정하는가이다.
Liu는 scaling을 부정하지 않지만 raw scale과 visual diversity를 구분한다. dataset이 크더라도 repetitive·narrow·biased할 수 있다. 또한 현재 interface는 사람이 이미지를 capture/upload하고 중요성을 설명해야 하지만 robot·laboratory assistant·medical system은 world state를 스스로 perceive하고 change를 detect하며 feedback으로 action을 교정해야 한다.
이 관점에서 visual intelligence는 partial visual observation에서 usable world structure를 배우고 hidden structure를 이해하며, 다음을 예측하고, 보이지 않는 것을 복원하고, 가능한 alternative를 상상하고, 그 지식을 새로운 visual task에 일반화하는 능력이다. human vision이 upper bound일 이유도 없다. 인간보다 더 넓고 정밀한 visual information에 노출된 machine은 human-centered perception을 넘어설 가능성이 있다.
what comes next. coherent future prediction을 위해 object, geometry, motion, occlusion, physical constraint, camera movement, agent behavior, temporal causality, uncertainty를 학습한다.
what could exist. alternative future와 counterfactual을 상상하고 internal simulator처럼 consequence를 시험할 수 있다.
what is hidden. inverse rendering, depth, multi-view, 3D/4D reconstruction을 통해 incomplete observation에서 hidden cause와 physical consistency를 회복한다.
이 세 objective가 VFM에서 visual intelligence, 잠재적으로 general intelligence로 가는 상보적 경로라는 것이 저자들의 central hypothesis다. true visual in-context learning도 중요한 가능성이다. 몇 개 visual example만 보고 새 task structure와 output format을 추론해 parameter update 없이 해결하는 능력이다. 그러나 world model만으로 intelligence가 완성되는 것은 아니다. 새 goal 아래 relevant abstraction을 선택하고 counterfactual을 만들고 new observation에 적응하며 행동해야 한다.
백서의 plurality는 약점이 아니라 연구공간을 좁히지 않기 위한 의도적 결론이다.
| Contributor | Central position |
|---|---|
| Robert Geirhos | Generative video models will become visual foundation models. |
| Aditi Raghunathan | Creativity—coherence, structural diversity, originality, utility—should be a central VGI test. |
| Yuki M. Asano | VGI should continuously and autonomously learn from visual experience rather than remain fixed after pre-training. |
| Deva Ramanan et al. | Visual intelligence should be multimodal, generative, computationally efficient, and include touch/audio/proprioception. |
| David Fouhey | Scientific visual intelligence must work with limited data and validate against physical knowledge, instruments, units, and systematics. |
| Andrew J. Davison | Spatial AI needs persistent world representations integrating state estimation, learning, real-time computation, and efficient hardware. |
| Yilun Du | Visual intelligence is embodied intelligence linking perception, world models, action, active exploration, and continual adaptation. |
| Shangzhe Wu & Jiajun Wu | VGI should recover compositional physical structure supporting editing, simulation, verification, and action. |
| Zhuang Liu | VGI should be vision-native and learn when to look, where to look, and how much detail is needed. |
| Hirokatsu Kataoka et al. | Prediction, imagination, and structural understanding may jointly provide a path from visual experience to general intelligence. |
summary는 방향들을 scaling/generative modeling, creativity·prediction·imagination·reconstruction, continual/lifelong learning, persistent 3D/spatial representation, embodiment·active perception·action, multimodal learning·efficiency, scientific discovery·unfamiliar-environment generalization, compositional physical structure, language-centered가 아닌 vision-native intelligence로 묶는다.
공통 premise는 visual experience로부터 world structure를 discover and verify하고, knowledge를 유지·수정하며, possible future를 예측하고, alternative를 상상하고, unseen structure를 복원하고, unfamiliar situation에 behavior를 적응하는 것이다. description이나 reproduction뿐 아니라 representation이 edit·simulate되고 reliable action에 쓰일 수 있어야 한다.
논문은 단일 benchmark가 이 모든 차원을 측정할 가능성이 낮다고 본다. isolated static tasks만으로 VGI를 판별하기 어렵다는 것이 더 중요한 결론이다.
현재 system은 중요한 fragment를 보여준다. generative video model은 cross-task capability를 보이고, reconstruction model은 geometric/physical structure를 더 많이 회복하며, spatial system은 persistent representation을 유지하고, creative model은 alternative를 만들며, embodied agent는 interaction을 통해 perception과 behavior를 개선한다.
그러나 이 능력들은 여전히 fragmented, computationally expensive, temporally limited, data/training-condition dependent하다. 따라서 더 가까운 milestone은 “VGI가 이미 AGI의 길인가”를 묻는 것보다 visual experience 자체에서 intelligence emergence를 관찰할 수 있는지 검증하는 것이다.
AIST 연구진의 작업은 일본 METI와 NEDO가 주도하는 FRONTia—AI robot과 Physical AI를 위한 domestic multimodal foundation model 개발 프로그램—아래 수행되었고 JST ASPIRE(JPMJAP2518)의 지원도 받았다. workshop에는 Matt Deitke, Alexei A. Efros, Kristen Grauman, Alex Kendall, Andrea Vedaldi 등의 초청발표가 논의에 기여했다.
하나의 architecture를 확정하는 논문이 아니라, computer vision community가 AGI 시대에 무엇을 측정하고 어떤 evidence를 요구할지 연구공간을 넓히는 position map으로 읽는 것이 정확하다.
원문 bibliography는 115개 항목으로 Turing·Wiener·Marr부터 VLM, video reasoning, SLAM, embodied world models, scientific vision, dataset bias, neuro-symbolic concepts, 3D/4D reconstruction과 robot planning까지 폭넓게 연결한다. 위 목록은 본문 논지를 읽기 위한 대표 reference map이며 전체 인용목록은 원문에 보존되어 있다.