SUNGSOO'S AI BLOGPHYSICAL AI · HUMAN–ROBOT COLLABORATION
RESEARCH SYNTHESIS · 2025–2026 · 2026.09.23

서로를 이해하는 로봇,
함께 판단하는 협업지능

Human–Physical AI Collaborative Intelligence: Shared World Models and Embodied XAI

공유 세계모델, 인간 의도 추론, 공동 행동, 상호 적응과 설명가능성을 하나의 폐루프로 연결한다. 연구의 중심은 인간과 로봇이 서로의 판단을 이해하고 수정할 수 있는 공동지능이다.

인간과 Physical AI의 폐루프 협업지능인간 상태와 로봇 상태가 공유 세계모델에 연결되고, 공동 계획과 물리 행동의 결과가 설명과 반응을 거쳐 모델로 되돌아오는 개념도이다. 전 과정에 안전과 불확실성 감독이 적용된다.인간 상태로봇 상태공유 세계모델설명 · 인간 반응물리적 협업의도 · 주의 · 피로 · 선호능력 · 관측 · 불확실성공동 목표 · 역할 · 미래 예측XAI / DECISION PROVENANCEJOINT VLA / SHARED CONTROL공동 계획 · 협상역할과 제어권 조정관측 · 반응 · 수정의 반복전 과정: 안전 · 불확실성 감독 + 인간의 이해와 개입 가능성
첨부 연구동향을 종합한 개념 구조 · 표준 아키텍처나 전체 시스템의 실증 완료를 의미하지 않는다.
EDITORIAL ABSTRACT

이 글은 「Embodied XAI 기반 인간–로봇 상호작용 기술」에 관한 첨부 보고서의 논의를 Physical AI 기반 Human-AI Interaction의 변화피지컬 AI 협업지능(Physical AI Collaborative Intelligence) 관점으로 통합한다. 분석 범위는 2025년 1월부터 2026년 9월 23일까지이다.

문헌은 2025–2026년 peer-reviewed 논문을 우선했고, 최신성이 중요한 VLA·World Model 영역은 arXiv/기업 연구보고서를 보조적으로 사용했다. WEF, Capgemini, Deloitte 자료는 학술논문의 대체물이 아니라 산업·배포 동향을 확인하는 보조 근거로 구분했다.

핵심적으로 2025년 이후의 변화는 다음 한 문장으로 요약할 수 있다.

Physical AI 기반 HRI의 연구 중심이 “인간이 지시하고 로봇이 수행하는 Human→Robot 구조”에서 “인간과 로봇이 서로의 상태·의도·능력·불확실성을 추론하고, 미래를 공동 예측하고, 역할과 행동을 실시간 조정하며, 그 이유까지 서로 이해할 수 있는 Human↔Robot 공동지능”을 지향하는 방향으로 연구가 확장되고 있다.

이는 원자료가 제안한 Human Response Model, Robot Decision Trace, 불확실성 추적, HIL Simulation, 이종 로봇 Runtime, Benchmark를 단순한 XAI 요소기술이 아니라 Physical AI 협업지능의 핵심 기반계층으로 다시 해석할 수 있음을 의미한다.

해석 기준 · 논문이 보고한 결과, 산업·정책 관점, 이 글의 연구 제안을 구분한다. ‘연구안’은 첨부 Markdown에 기술된 내용을 뜻하며 별도 원제안서를 열람했다는 의미가 아니다. 구성요소의 합과 방향 화살표로 표현한 식은 설계를 설명하는 개념식이다. 예시 대화는 실제 시스템 발화나 측정 결과가 아니다.

전체 목차 · 13개 주제
  1. 1. Definition
  2. 2. Problem Definition
  3. 3. Core Concepts
  4. 4. Introduction — 왜 지금 협업지능인가?
  5. 5. Motivation and Background
  6. 6. Challenges
  7. 7. Research Questions
  8. 8. Approaches / Methods
  9. 9. Key Applications
  10. 10. Open Problems
  11. 11. Future Directions
  12. 12. 전체 연구방향을 하나의 구조로 통합하면
  13. 13. 가장 중요한 연구적 결론
PART 01 / HUMAN–PHYSICAL AI

협업지능의 정의와 문제

§01Definition

1.1 Physical AI

2026년 최신 world-model survey는 Physical AI를 현실세계의 동역학, 부분 관찰성, 센서 불확실성과 행동의 물리적 결과를 고려하면서 closed loop로 perception–prediction–planning–action을 수행하는 embodied system으로 본다. World Model은 이 과정에서 환경 dynamics와 action consequences를 학습해 미래를 예측하는 내부 모델로 기능한다. (Springer [1])

Deloitte 역시 산업적 정의에서는 Physical AI를 단순한 pre-programmed robot이 아니라 실시간으로 물리세계를 인식·이해·추론하고 행동하며 경험에 따라 적응하는 시스템으로 정의하고 있다. 이는 학술적 정의와 방향상 일치한다. (Deloitte [2])

따라서,

\[\text{Physical AI} = \text{Perception} +\text{World Modeling} +\text{Reasoning} +\text{Planning} +\text{Physical Action} +\text{Closed-loop Learning}\]

으로 요약할 수 있다.

1.2 Human–Robot Interaction에서 Human–Robot Collaboration으로

HRI는 사람과 로봇 사이의 모든 상호작용을 포괄하지만, HRC(Human–Robot Collaboration)는 그중에서도 인간과 로봇이 공유된 목표(shared goal)를 달성하기 위해 실시간으로 서로 행동을 조정하는 경우를 의미한다. 2026년 cognitive HRC systematic review는 이를 인간의 유연한 인지능력과 로봇의 정확성·반복성을 결합하는 partnership으로 설명한다. (Springer [3])

2026년 Embodied AI–HRC 논의에서는 한 단계 더 나아가 기존의

Instruction → Execution

구조에서

Perceive each other → Predict each other → Adapt → Coordinate → Learn together

구조로 패러다임이 이동한다고 정리한다. (OAEPress [4])

1.3 Physical AI 협업지능의 정의

“Physical AI Collaborative Intelligence” 자체가 아직 하나의 완전히 표준화된 학술 용어는 아니다. 그러나 Collaborative Intelligence, Cognitive HRC, Embodied HRC, Human-AI Teaming, Cooperative Intelligence 연구를 통합하면 다음과 같이 연구용 operational definition을 제시할 수 있다.

피지컬 AI 협업지능(Physical AI Collaborative Intelligence)이란 인간과 하나 이상의 Physical AI가 물리적·사회적 환경을 공동으로 인식하고, 서로의 의도·상태·능력·불확실성을 지속적으로 추론하며, 공유된 세계모델과 목표를 바탕으로 미래 행동을 상호 예측하고 역할·계획·제어권을 동적으로 조정하여, 개별 인간 또는 개별 로봇보다 높은 안전성·적응성·성과를 목표로 하는 폐루프 공동지능이다.

2026년 collaborative intelligence survey도 협업지능을 H2M, M2M, M2H learning으로 나누면서 인간과 로봇의 perception, decision-making, knowledge transfer, mission adaptation이 서로 연결되어야 함을 강조한다. (Frontiers [5])

따라서,

\[CI_{Physical} = PI_H + PI_R + MI_{shared}+CA+MA+XA\]

로 볼 수 있다.

  • \(PI_H\): Human Intelligence
  • \(PI_R\): Robot/Physical AI Intelligence
  • \(MI_{shared}\): Shared/Mutual Intelligence
  • \(CA\): Collaborative Action
  • \(MA\): Mutual Adaptation
  • \(XA\): Explainable/Accountable Interaction

위 덧셈식은 구성요소를 정리한 개념식이며, 서로 다른 지능을 공통 척도로 측정해 합산한 실증 법칙이 아니다.

중요한 것은 Human Intelligence + Robot Intelligence의 단순 합이 아니라 상호작용 과정에서 생성되는 Shared Intelligence이다.

1.4 Embodied XAI의 위치

이 구조에서 Embodied XAI가 중요해진다.

원자료에서 설명한 연구안은 robot perception–decision–plan–action 과정에서 실제 행동에 영향을 준 evidence, constraint, alternative와 uncertainty를 추적해 Robot Decision Provenance를 만들 것을 제안한다.

Physical AI 협업지능 관점에서는 Embodied XAI를 다음과 같이 재정의할 수 있다.

Embodied XAI는 Physical AI의 “설명 모듈”이 아니라 인간과 로봇이 shared mental model을 형성·수정하기 위한 epistemic communication layer이다.

즉,

Robot knows → Robot explains → Human understands → Human responds → Robot updates

의 루프를 만들어 주는 기술이다.

§02Problem Definition

Physical AI 협업지능의 근본 연구문제는 단순히

“로봇이 사람을 잘 인식하는가?”

가 아니다.

보다 정확히는,

불완전한 센싱과 불확실한 인간 행동이 존재하는 동적 물리환경에서 인간과 로봇이 서로의 의도와 능력을 추정하면서 shared goal을 유지하고, future state를 공동으로 예측하며, 안전하게 역할과 행동을 조정할 수 있는가?

이다.

2026년 pHHI review는 현재 기술을 humanoid modeling/control, human intent estimation, computational human model의 세 축으로 분류하면서 개별 축의 발전에 비해 이들을 통합하는 연구가 부족하다고 지적한다. (Springer [6])

이를 세분화하면 다음 여덟 가지 gap이 나타납니다.

문제기존 접근협업지능에서 필요한 변화
Perception Gap사람의 위치 검출상태·시선·의도·피로·주의 추론
Intent Gap현재 행동 분류미래 intention/trajectory prediction
World-model Gaprobot-centric worldHuman–Robot shared world model
Planning Gaprobot optimal planningjoint human–robot planning
Control Gaprobot authorityadaptive shared control
Communication Gapcommand/responseimplicit+explicit multimodal communication
Trust Gaptrust 증가calibrated reliance
Explainability Gap결과 설명decision provenance + uncertainty + alternatives
PART 02 / HUMAN–PHYSICAL AI

함께 이해하고 행동하는 핵심 개념

§03Core Concepts

3.1 Human State and Intent Modeling

협업의 출발점은 human sensing이다.

최근 연구는 단순 skeleton tracking을 넘어,

  • pose
  • hand trajectory
  • gaze
  • voice
  • gesture
  • tactile interaction
  • workload
  • fatigue
  • intervention history

를 이용해 what is the human doing?에서 what will the human do next?로 이동하고 있다.

2025년 Proactive robot task sequencing 연구는 human hand trajectory를 예측해 robot task sequence 자체를 미리 변경함으로써 reactive planning에서 proactive collaboration으로 이동했다. (ScienceDirect [7])

2026년 multi-worker anticipation 연구는 예를 들어 3초의 짧은 관찰로 장기 행동을 예측하되 인간 행동의 multimodal uncertainty까지 모델링한다. (ScienceDirect [8])

3.2 Physical Interaction and Tactile Intelligence

Physical AI 협업에는 vision만으로 부족한다.

사람과 로봇이 함께 물체를 들거나 접촉하는 순간에는

\[\text{Vision}+\text{Force}+\text{Touch}+\text{Proprioception}\]

이 필요한다.

2025년 tactile reflex 연구는 artificial skin에서 얻은 force를 이용해 사람과 접촉했을 때 robot이 reflex-like response를 발생시키도록 했다. (ScienceDirect [9])

더 나아가 Tactile-VLA(2025)는 vision–language–action에 tactile feedback을 직접 결합하여 contact-rich manipulation에서 force-aware reasoning을 가능하게 한다. (arXiv [10])

즉 tactile은 단순 “충돌 감지 센서”에서

협업 의도와 물리 상태를 전달하는 communication modality

로 발전하고 있다.

3.3 VLM → VLA → Collaborative VLA

2025년 VLA survey는 VLA가 vision, language, action을 하나의 policy로 통합해 task, object, environment, embodiment 간 generalization을 추구한다고 정리한다. (IEEExplore [11])

Gemini Robotics는 open-vocabulary instruction과 visual observation을 이용해 실제 robot action을 생성하며, 새로운 object/environment/embodiment로의 generalization을 강조했다. (arXiv [12])

π0.5 역시 heterogeneous robot data, language, web knowledge, high-level subtask를 co-training해 새로운 가정환경에서 장시간 manipulation generalization을 시연했다. (Physical Intelligence [13])

하지만 협업지능 관점에서는 일반 VLA보다 한 단계 더 필요한다.

기존:

\[(o_t,\ language)\rightarrow a^R_t\]

협업 VLA:

\[(o_t,H_t,G_{shared},A^H_{past}) \rightarrow (a^R_t,\hat a^H_{t+1},coordination)\]

즉 robot action뿐 아니라 인간의 다음 행동과 team coordination까지 예측해야 한다.

2025년 VLA 기반 collaborative robotic assistant 연구도 인간 손 pose와 collaborator intent를 별도로 예측하기 시작했다. 다만 특정 시연자에 대한 과적합을 주요 한계로 보고하므로 사람 간 일반화까지 입증한 것은 아니다. (arXiv [14])

3.4 World Model → Shared World Model

일반적인 Physical AI world model은

\[p(s_{t+1}\mid s_t,a_t)\]

을 학습한다.

Cosmos는 이를 Physical AI의 “digital twin of the world”로 표현하고, video/world foundation model을 이용해 physical AI training 및 simulation을 지원한다. (arXiv [15])

2026년 Physical AI world-model survey에서는 특히

  • uncertainty
  • long-horizon consistency
  • physical constraints
  • decision coupling
  • sim-to-real

이 핵심 연구문제로 정리된다. (Springer [1])

그러나 협업지능에서는 world model이 robot 자신의 환경모델만이어서는 안 된다.

Collaborative World Model

\[W_t= \{Environment,\ Robot,\ Human,\ Goal,\ Roles,\ Intent,\ Capability,\ Uncertainty\}\]

가 되어야 한다.

AAAI-26 Bridge Program 채택으로 표기된 2026년 Explicit World Models for Reliable Human-Robot Collaboration은 바로 이 관점에서 인간과 AI가 공유하는 accessible common-ground world model을 reliability의 핵심으로 제안한다. (arXiv [16])

이 방향은 매우 중요한다.

미래의 협업 로봇은 “세계를 이해하는 로봇”에서 “사람과 세계에 대한 이해를 공유할 수 있는 로봇”으로 진화해야 한다.

3.5 Shared Mental Model / Common Ground

협업에서는 인간도 robot을 이해해야 한다.

사람이 robot vision capability를 실제보다 과대평가하면 impossible task를 지시하거나 위험 상황에서 잘못 의존할 수 있다.

2026년 AR을 이용해 robot field-of-view를 사람에게 직접 보여주는 연구에서는 robot의 perceptual capability를 명시적으로 제시함으로써 인간의 mental model을 더 정확하게 맞추려 했다. (Springer [17])

따라서 협업지능의 목표는

\[M_H(R) \approx R_{\text{actual capability}}\]

뿐 아니라

\[M_R(H) \approx H_{\text{actual state/intention}}\]

도 성립시키는 것이다.

mutual mental-model alignment이다.

Embodied XAI는 그 정렬을 수행하는 중요한 mechanism이다.

3.6 Proactive Collaboration

Reactive robot:

사람이 행동한다 → 로봇이 반응한다.

Proactive collaborative robot:

사람의 다음 행동을 예측한다 → 필요한 행동을 먼저 준비한다.

2026년 LLM-powered CCM-FCC 연구는 cognition-centered AI agent가 operator state와 task semantics를 이용해 필요한 기능을 활성화하고 reinforcement learning으로 proactive collaboration을 수행하도록 설계했다. (ScienceDirect [18])

이는 다음의 변화를 나타냅니다.

Reactive HRC → Predictive HRC → Proactive HRC

3.7 Mutual Adaptation

진정한 협업에서는 robot만 human에게 적응해서도 충분하지 않는다.

2025년 co-transportation 연구는 인간의 preference uncertainty를 확률적으로 모델링하고 상황에 따라 robot이 leader 또는 follower로 전환하는 mutual adaptation 구조를 제안했다. (arXiv [19])

따라서 앞으로의 Physical AI는 고정적인

Human = leader / Robot = follower

구조가 아니라,

\[Role_t \in \{HumanLeader,\ RobotLeader,\ Shared\}\]

를 상황에 따라 바꾸는 방향으로 발전할 가능성이 큽니다.

3.8 Adaptive Shared Autonomy

Mutual adaptation이 더 발전하면 동적 제어권 배분(dynamic authority allocation) 문제가 된다.

사람의 visual/haptic compliance를 활용하는 공유 조향 제어 연구처럼 인간과 AI의 control authority를 조절하는 접근이 나타나고 있다. 위험도와 robot confidence를 함께 반영하는 아래 표현은 이를 협업지능 관점으로 확장한 개념적 설계이다. (ScienceDirect [20])

즉,

\[u_t= \alpha_tu^H_t+(1-\alpha_t)u^R_t\]

이며

\[\alpha_t=f( Risk, HumanState, HumanIntent, RobotUncertainty, Trust )\]

가 된다.

3.9 Embodied XAI and Decision Provenance

바로 이 지점에서 원자료의 연구안과 결합된다.

협업 로봇은 단순히

“제가 안전하다고 생각해서 멈췄습니다.”

가 아니라,

“오른쪽 작업자의 손이 예상경로로 진입할 가능성이 높아졌고, 제 예측 불확실성도 임계값을 넘었기 때문에 계획 B로 전환했습니다.”

라고 말할 수 있어야 한다.

이를 위해 원자료에서 설명한 연구안에서는

  • Multimodal Evidence
  • Perception–Decision–Plan–Action Trace
  • alternatives/counterfactuals
  • uncertainty propagation
  • Faithful Explanation Grounding

을 제안한다.

협업지능 관점에서는 이를 Joint Decision Provenance로 확장할 수 있다.

\[Human\ Action \rightarrow Robot\ Belief \rightarrow Joint\ Decision \rightarrow Robot\ Action \rightarrow Human\ Response\]

를 모두 저장하는 것이다.

3.10 Trust → Trust Calibration

2026년 HRC trust review의 중요한 결론 중 하나는 trust가 많을수록 좋은 것이 아니라 적절하게 calibrated되어야 한다는 것이다.

최신 연구는 gaze, intervention frequency, reaction time, physiological signal 등을 사용해 trust를 실시간으로 추정하고 이를 adaptive control 및 explanation에 연결하는 closed-loop architecture를 논의한다. (Springer [21])

따라서 목표함수는

\[\max Trust\]

가 아니라

\[\min|Trust_{Human}-Capability_{Robot}|\]

에 가깝다. 이 식은 신뢰와 능력의 정렬을 강조하는 도식이다. 실제 평가는 동일 과제와 비교 가능한 척도에서 과신·과소신뢰 및 개입 판단을 측정해야 한다.

3.11 Digital Human / Human Response Model

원자료에서 설명한 연구안의 Human Response Model은 Physical AI 협업지능 관점에서 상당히 중요한 요소이다.

문서는 사람의 상태와 맥락으로부터 설명의 이해·신뢰·개입 반응을 예측하고, Adaptive Explanation Policy에 연결하도록 제안한다.

2026년 Digital Human + RL 연구에서는 human motion data를 simulation 안에서 생성하여 collaborative robot control을 학습하는 접근이 제안되고 있다. (The Advanced Portfolio [22])

다음 단계는 motion digital human을 넘어,

Cognitive–Behavioral Digital Human

이다.

\[H_t= \{pose, intent, attention, fatigue, knowledge, trust, preference, response\}\]

를 모델링한다.

PART 03 / HUMAN–PHYSICAL AI

왜 지금 인간–로봇 공동지능인가

§04Introduction — 왜 지금 협업지능인가?

2025년 이후 세 기술흐름이 동시에 만났기 때문이다.

첫째, 로봇의 인지능력이 급격히 커졌다.

LLM/VLM/VLA가 natural-language instruction, open-world perception, long-horizon planning을 가능하게 만들고 있다. Smart-manufacturing HRC survey도 LLM/VLM이 task planning, navigation, manipulation, human–robot skill transfer를 통합하기 시작했다고 분석한다. (Springer [23])

둘째, Physical AI가 laboratory demonstration에서 실제 환경으로 이동하고 있다.

WEF는 human-centric Physical AI에서 seamless, safe, empathetic HRC를 향후 핵심 benchmark로 제시한다. 다만 이는 학술논문이라기보다 policy/industry perspective이다. (World Economic Forum [24])

Capgemini 조사에서도 기업의 관심이 robot model 자체에서 deployment와 human–robot collaboration으로 이동하고 있는 흐름이 확인된다. (Capgemini [25])

셋째, autonomy 증가가 새로운 human-factor 문제를 만들고 있다.

2026년 Cognitive HRC review는 현재 연구가 perception, learning, reasoning 같은 core cognitive capabilities에는 많이 투자하지만 social cognition은 부족하며, technical performance에 비해 trust, workload, ergonomics 평가가 부족하다고 지적한다. (Springer [3])

즉 Physical AI가 더 똑똑해질수록 HRI는 덜 중요해지는 것이 아니라 오히려 더 중요해진다.

§05Motivation and Background

연구흐름을 단계적으로 보면 변화가 더 분명한다.

Robot Automation 1.0

Human commands → Robot executes

HRC 2.0

Human works + Robot assists

Cognitive HRC 3.0

Robot recognizes intent → adapts action

Proactive Physical AI 4.0

Robot predicts human → anticipates need → reallocates task

Collaborative Physical Intelligence 5.0

Human ↔ Robot jointly perceive, predict, decide, adapt and explain

특히 2026년 cognitive robotics review는 perception, attention, memory, learning, reasoning, metacognition, prospection과 더불어 joint attention, communication, intention reading 같은 social cognitive abilities가 효과적 협업의 핵심이라고 본다. (Springer [3])

이것이 “협업지능”이라는 개념을 연구할 학술적 기반을 제공한다.

PART 04 / HUMAN–PHYSICAL AI

기술 난제와 14개 연구질문

§06Challenges

6.1 Human intent는 ground truth가 아니다

사람은 계획을 변경하고, 모호하게 행동하며, 동일 움직임도 여러 목적을 가질 수 있다.

따라서 deterministic intent classification 대신

\[P(I_{t:t+k}\mid H_{0:t},C_t)\]

형태의 probabilistic prediction이 필요한다.

6.2 World Model이 human dynamics를 충분히 모델링하지 못한다

현재 world model 대부분은 objects와 robot dynamics에 집중한다.

협업지능에는

Physical dynamics + Human behavioral dynamics + Social dynamics

가 동시에 필요한다.

6.3 VLA가 협업을 배웠다고 보기 어렵다

현재 VLA는 대부분

“instruction following”

을 잘한다.

그러나

“사람과 역할을 협상하고 서로의 계획을 수정한다”

는 더 어려운 문제이다.

6.4 인간의 행동은 multimodal이다

pose만으로는 부족한다.

같은 손동작도

  • gaze
  • speech
  • tool state
  • force
  • history

에 따라 의미가 달라진다.

6.5 Physical Safety와 Cognitive Safety를 동시에 해결해야 한다

전통적 safety는

collision을 방지했는가?

였지만 협업지능에서는

사람이 robot을 과신하지 않는가?

도 safety이다.

2026년 cognitive HRC review도 physical safety뿐 아니라 automation bias, reduced agency, psychological factors를 중요한 위험으로 지적한다. (Springer [3])

6.6 설명이 실시간 제어를 방해할 수 있다

모든 결정을 자세히 설명하면 cognitive load와 latency가 증가한다.

그래서

What to explain / When / To whom / How much / Through which modality

를 policy로 학습해야 한다.

6.7 Human adaptation과 Robot adaptation이 동시에 일어난다

이것은 non-stationary two-agent learning problem이다.

사람이 robot에게 적응하는 동안 robot도 다시 사람에게 적응하므로 기존 supervised learning의 fixed-distribution 가정이 깨진다.

6.8 Sim2Real보다 더 어려운 Sim2Human

물리적 trajectory는 simulation으로 만들 수 있지만,

실제 사람이 robot의 행동을 이해하고 신뢰하고 개입하는 방식

은 정확히 simulation하기 어렵다.

§07Research Questions

향후 연구과제로는 다음 RQ가 특히 중요한다.

RQ연구질문
RQ1multimodal signal에서 human intent와 latent state를 실시간으로 얼마나 정확히 추론할 수 있는가?
RQ2인간 행동의 multimodal future uncertainty를 world model에 어떻게 포함할 것인가?
RQ3robot-centric world model을 Human–Robot Shared World Model로 확장할 수 있는가?
RQ4human과 robot의 서로 다른 world model이 불일치할 때 이를 어떻게 탐지·교정할 것인가?
RQ5인간의 다음 행동을 예측해 robot이 언제 proactive action을 취해야 하는가?
RQ6언제 Human-led, Robot-led, Shared-control로 역할을 전환해야 하는가?
RQ7VLA가 instruction-following을 넘어 joint-task policy를 학습할 수 있는가?
RQ8tactile·vision·speech·gaze를 어떤 latent representation으로 융합하는 것이 좋은가?
RQ9joint decision의 causal provenance를 실시간으로 어떻게 생성할 것인가?
RQ10robot uncertainty를 사람이 올바르게 이해하도록 표현하는 최적 방식은 무엇인가?
RQ11Human Response Model이 trust, workload, intervention을 장기적으로 예측할 수 있는가?
RQ12synthetic digital human에서 학습한 collaboration policy가 실제 사람에게 transfer되는가?
RQ13서로 다른 robot embodiment에 하나의 collaboration policy를 적용할 수 있는가?
RQ14individual task performance보다 “team intelligence”를 어떻게 정량화할 것인가?
PART 05 / HUMAN–PHYSICAL AI

폐루프 구현 구조와 적용 분야

§08Approaches / Methods

이 연구들을 하나의 시스템으로 통합한다면 다음과 같은 Physical AI Collaborative Intelligence Loop가 가장 자연스럽다.

Layer 1. Multimodal Human–World Sensing

입력:

  • RGB/RGB-D
  • audio
  • gaze
  • skeleton
  • tactile/force
  • wearable
  • robot proprioception
  • environment sensors

출력:

\[O_t^{H,R,E}\]

Layer 2. Human State & Intent Model

Human Response Model을 확대하여

\[H_t= f(O_{0:t}, InteractionHistory)\]

를 만들고,

  • current intention
  • future action
  • attention
  • fatigue
  • competence
  • preference
  • trust
  • intervention likelihood

을 추정한다.

원자료에서 설명한 연구안의 Human Response Model을 협업지능용 Human Collaborative State Model로 확장하는 것이다.

Layer 3. Collaborative World Model

기존 robot world model에 human latent state를 추가한다.

\[W_t= \{E_t,R_t,H_t,G_t,U_t\}\]

그리고

\[\hat W_{t+k} = F(W_t, a^R, a^H)\]

를 예측한다.

즉 robot action뿐 아니라 human action에 따른 joint future rollout을 수행한다.

Layer 4. Joint Task and Role Planner

Planner는 단순히 robot action만 결정하지 않는다.

\[\pi_{team}: W_t \rightarrow (Role_H,Role_R, Task_H,Task_R)\]

Who should do what, when, and with what authority?

를 결정한다.

Layer 5. Adaptive Shared Control

실제 control에서는 human intent와 robot safety constraint를 공동 반영한다.

Human-led ↔ Shared ↔ Robot-led 사이에서 authority가 연속적으로 이동한다.

Layer 6. Safety & Uncertainty Supervisor

모든 단계에서

  • sensing uncertainty
  • intent uncertainty
  • world-model uncertainty
  • action uncertainty
  • safety risk

를 추적한다.

2026 world-model review 역시 uncertainty calibration과 planner exploitation을 실세계 Physical AI의 핵심 난제로 본다. (Springer [1])

Layer 7. Joint Decision Provenance

원자료의 Robot Decision Trace를 확장한다.

예:

Human gaze → inferred intention → predicted hand trajectory → collision risk → task reallocation → robot slowdown → explanation

을 하나의 graph로 보존한다.

이는 원자료에서 설명한 연구안의 Robot Decision Provenance Graph Engine과 매우 직접적으로 연결된다.

Layer 8. Embodied XAI Interaction

이하 1.2초 대기 발화는 설명 형식의 가상 예시이며 측정된 응답시간이 아니다.

결정 trace에서 evidence를 가져와

  • speech
  • gesture
  • display
  • AR
  • tactile cue
  • robot motion

중 적합한 channel로 설명한다.

“다음 부품을 먼저 가져오려 했지만, 오른손이 작업영역으로 진입할 가능성이 높아서 1.2초 대기합니다.”

와 같은 형태이다.

Layer 9. Human Response Closed Loop

설명 이후의

  • acceptance
  • correction
  • intervention
  • hesitation
  • gaze
  • task performance

를 다시 모델에 입력한다.

\[H_{t+1}= f(H_t,e_t,r_t)\]

이 단계가 명령–실행 시스템을 진정한 협업지능으로 바꾸는 핵심 루프이다.

§09Key Applications

① 인간–휴머노이드 공동 제조

사람은 dexterous manipulation과 문제 해결을 담당하고 robot은 lifting, repetitive manipulation, precision 작업을 담당하며 역할을 실시간으로 변경한다.

2026 pHHI review가 가장 직접적으로 연결된다. (Springer [6])

② Collaborative Assembly

손 움직임·의도를 예측하여 부품이나 도구를 미리 준비하는 proactive assembly가 대표적이다. (ScienceDirect [7])

③ 공동 운반과 Co-Manipulation

2026 human–human/human–robot carrying 연구는 physical forces 자체가 implicit communication channel 역할을 하며 mutual prediction과 shared representation이 중요함을 보여준다. (ScienceDirect [26])

④ 돌봄·재활·생활지원

고령자나 장애인과 장기적으로 상호작용하는 robot은 task accuracy 이상으로

  • intent
  • physical state
  • trust
  • explanation
  • preference

를 모델링해야 한다.

⑤ 건설·위험환경

multi-worker behavior anticipation을 이용해 robot이 사람에게 미리 assistance를 제공할 수 있다. (ScienceDirect [8])

⑥ Search & Rescue

사람–robot–robot 사이의 knowledge transfer와 decentralized collaborative intelligence가 특히 중요한다. (Frontiers [5])

⑦ 물류·다종 로봇

향후에는 Human–Robot 협업을 넘어

\[Human + Robot + Robot + EdgeAI\]

의 heterogeneous collaborative intelligence로 확장될 가능성이 높다.

PART 06 / HUMAN–PHYSICAL AI

미해결 문제와 아홉 가지 연구방향

§10Open Problems

10.1 Shared World Model의 정답은 무엇인가?

Human world model과 robot world model이 서로 다릅니다.

사람이 아는 정보를 robot이 모를 수도 있고 반대도 가능한다.

따라서 belief alignment 자체가 연구문제이다.

10.2 Human model을 어디까지 학습해야 하는가?

의도·피로·인지부하·신뢰를 모델링할수록 협업은 좋아질 수 있지만 privacy와 manipulation 문제가 커진다.

10.3 Collaborative VLA 데이터 부족

현재 대규모 robot dataset의 대부분은

\[Robot + Task\]

구조이다.

미래에는

\[Human + Robot + Joint Task + Interaction\]

dataset이 필요한다.

10.4 Joint Action Ground Truth 부재

“로봇의 행동이 정답이었는가?”보다

“인간과 robot의 공동 행동 전체가 최적이었는가?”

를 평가해야 한다.

10.5 Long-horizon human–robot adaptation

10분짜리 demonstration보다 수일·수주·수개월 collaborative learning이 훨씬 어렵다.

Trust review도 long-term trust trajectory와 repeated interaction 연구의 부족을 명시적으로 지적한다. (Springer [27])

10.6 Social cognition 부족

2026 systematic review의 중요한 결과는 40개 분석 논문 중 core cognition은 대부분 포함하지만 joint attention, intention reading 등 복수 social cognitive ability를 다룬 연구는 매우 적었다는 점이다. (Springer [3])

이는 향후 매우 큰 연구공백이다.

§11Future Directions

방향 1. Human-Aware World Model → Collaborative World Foundation Model

현재 Cosmos 같은 World Foundation Model을 다음과 같이 확장하는 방향이다.

\[WorldModel \rightarrow Human\text{-}AwareWorldModel \rightarrow CollaborativeWorldModel\]

물체 dynamics뿐 아니라

  • human intention
  • human motion
  • social affordance
  • interaction history
  • team goals

까지 모델링한다.

방향 2. VLA → Human–Robot Joint VLA

현재:

\[Vision + Language \rightarrow RobotAction\]

미래:

\[Vision+ Language+ HumanState+ SharedGoal \rightarrow JointAction\]

이다.

이것이 Physical AI 협업지능의 가장 중요한 foundation-model 연구방향 중 하나가 될 수 있다.

방향 3. Robot Decision Provenance → Joint Decision Provenance

원자료에서 설명한 연구안의 provenance 기술을 다음 단계로 확장한다.

\[Robot\ DecisionTrace \rightarrow Human\text{-}Robot\ JointDecisionTrace\]

즉,

누가 무엇을 보고, 무엇을 예상해서, 어떤 선택지를 버리고, 왜 역할을 바꾸었는가

를 추적한다.

방향 4. Robot World Model + Human Response Model의 결합

원자료에서 설명한 연구안에서는 Human Response Model과 Decision Trace가 별도 세부기술처럼 보이지만, Physical AI 협업지능에서는 이 둘을 강하게 결합해야 한다.

\[Collaborative\ State = WorldModel + HumanResponseModel + DecisionProvenance\]

가 되는 구조이다.

방향 5. Digital Human → Collaborative Digital Human

단순 motion simulator가 아니라

  • intention
  • attention
  • fatigue
  • trust
  • preference
  • intervention
  • learning/adaptation

까지 포함한다.

그러면 원자료에서 설명한 연구안의 HIL Simulation/Data Factory 는 단순 데이터 생성기를 넘어

Human–Physical-AI Collaboration Simulator

가 된다.

방향 6. Explainable AI → Negotiable Physical AI

다음 단계의 설명가능성은 설명만 하는 것이 아니다.

Human:

“왜 내가 아니라 네가 이 작업을 하려고 하지?”

Robot:

“현재 작업의 하중이 높아서 제가 수행하는 것이 안전합니다.”

Human:

“그렇지만 내가 먼저 잡을 테니 너는 지지해.”

Robot:

“제가 안전하다고 생각해서 멈췄습니다.”0

즉,

Explainability → Dialogue → Negotiation → Replanning

이다.

이것이 future Embodied XAI의 중요한 방향이 될 수 있다.

방향 7. Proactive Robot → Reciprocal Prediction

현재 proactive HRC:

\[Robot \ predicts \ Human\]

미래 협업지능:

\[Robot \ predicts \ Human \quad\land\quad Human \ predicts \ Robot\]

이다.

그리고 Embodied XAI는 인간의 robot prediction을 돕는 기술이다.

방향 8. Trust-aware → Epistemically Calibrated Collaboration

목표는 trust를 높이는 것이 아니라

사람이 robot의 능력·한계·불확실성을 정확히 이해하고 적절한 순간에 의존하거나 개입하도록 만드는 것

이다.

이는 협업지능의 안전성 KPI가 되어야 한다.

방향 9. Individual Intelligence Benchmark → Team Intelligence Benchmark

원자료에서 설명한 연구안은 AI–Human–Robot 3계층 평가체계를 제안한다.

이를 협업지능 관점에서 발전시키면 다음 지표가 중요한다.

영역미래 핵심 KPI
Robottask success, safety, uncertainty calibration
Humanworkload, intervention accuracy, comprehension
Interactioncommunication efficiency, conflict rate
Teamjoint task success, mutual predictability
Adaptationrecovery time, role-switch quality
XAIcausal faithfulness, mental-model alignment
Safetyphysical + cognitive safety
Generalizationunseen human/task/robot adaptability

여기에 새로운 지표로

\[\textbf{Collaborative Intelligence Gain} = Perf(H+R)-\max(Perf(H),Perf(R))\]

을 생각할 수 있다.

즉,

사람과 robot을 결합했을 때 실제로 개별 주체보다 얼마나 더 잘하는가?

를 측정하는 것이다. 비교에는 동일한 과제·시간·자원·안전 제약과 같은 성능 척도가 필요하다. 시간이나 오류처럼 작을수록 좋은 지표는 방향을 변환하거나 별도로 해석해야 한다. 협업에 투입한 추가 자원을 통제하지 않으면 결합 효과를 과대평가할 수 있다.

PART 07 / HUMAN–PHYSICAL AI

국가 공통기술체계로의 통합

§12전체 연구방향을 하나의 구조로 통합하면

앞서 원자료에서 설명한 연구안의 예상 산출물은

  1. Human Response & Interaction Foundation Model
  2. Robot Decision Provenance Graph Engine
  3. Simulation/Data Factory + Universal Runtime
  4. National Dataset + Benchmark + Open Evaluation Platform

이다.

Physical AI 협업지능 관점에서는 이를 다음 6계층 국가 공통기술체계로 발전시키는 것이 매우 자연스럽다.

① Human Collaborative Intelligence Layer

Human State / Intent / Response / Trust Model

② Collaborative World Model Layer

Human + Robot + Environment + Shared Goal + Uncertainty

③ Joint Reasoning & Planning Layer

Intent prediction / Task allocation / Role negotiation

④ Physical Collaboration Layer

VLA / tactile intelligence / shared control / proactive action

⑤ Embodied XAI & Provenance Layer

Why / Why-not / What-if / How-sure / Decision Trace

⑥ Simulation–Runtime–Benchmark Layer

Digital Human / HIL / Sim2Real / SDK / Evaluation / Standard

그리고 이 전체를 하나의 closed loop로 만든다.

\[\boxed{\begin{gathered} Sense \rightarrow Understand\ Human \rightarrow Build\ SharedWorld \rightarrow Predict \\ \rightarrow JointPlan \rightarrow Act \rightarrow Explain \\ \rightarrow ObserveHumanResponse \rightarrow Adapt \end{gathered}}\]

§13가장 중요한 연구적 결론

2025–2026 문헌을 Physical AI 관점에서 보면 중요한 변화는 “더 강한 robot model”만 만드는 것이 아니다.

현재 가장 큰 공백은 오히려 다음에 있다.

VLA, World Model, tactile intelligence, human intent modeling, shared control, Embodied XAI가 각각 발전하고 있지만, 이들을 인간–Physical AI의 공동 perception–prediction–planning–action–explanation loop로 통합하는 일반적 협업지능 architecture는 아직 확립되지 않았다.

2026 pHHI review도 control, intent estimation, computational human modeling의 통합 부족을 핵심 gap으로 본다. (Springer [6])

2026 cognitive HRC systematic review도 core cognitive ability에 비해 social cognition과 human-centered evaluation이 여전히 부족하다고 지적한다. (Springer [3])

2026 trust review 역시 trust modeling, explainability, adaptive control이 개별 연구로 분리되어 있고 이를 하나의 closed-loop HRC framework로 통합해야 한다고 결론 냅니다. (Springer [27])

따라서 앞서의 Embodied XAI 연구를 더 큰 국가 전략적 연구개념으로 발전시킨다면 가장 강력한 framing은 다음과 같다.

“Human–Physical AI Collaborative Intelligence”
인간과 Physical AI가 서로의 상태·의도·능력·불확실성을 이해하고, 공유 세계모델을 기반으로 미래를 함께 예측하며, 역할·계획·제어권을 동적으로 조정하고, 그 판단근거를 상호 설명하면서 함께 학습·적응하는 인간중심 폐루프 협업지능

그리고 Embodied XAI는 이 구조의 부가적 설명기능이 아니라 Human–Robot Shared Intelligence를 성립시키는 핵심 interface로 보는 것이 적절한다.

READING GUIDE

핵심 문헌과 종합 관점

학술적 출발점으로 특히 중요한 자료는 다음과 같다.

이 문헌들을 종합하면, 향후 가장 연구가치가 높은 축은 “범용 Physical AI 자체”보다 한 단계 위의 Human–Physical AI Collaborative World Model + Joint VLA + Embodied XAI + Mutual Adaptation 통합체계라고 볼 수 있다. 이는 기존 원자료에서 설명한 연구안의 네 세부과제를 유지하면서도, 연구의 중심을 “로봇이 사람에게 설명한다”에서 “사람과 로봇이 서로를 이해·예측·설명하면서 하나의 협업지능을 형성한다”로 확장하는 방향이다.

FULL BIBLIOGRAPHY · SOURCE STATUS

전체 참고자료

  1. A survey of world models for physical AI with uncertainty representation and control | Discover Artificial Intelligence | Springer Nature Link학술 논문·연구보고서 · 이번 편집에서 원문 페이지 확인
  2. Physical AI and humanoid robots | Deloitte Insights산업·정책 자료 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  3. Emerging Trends in Cognitive Abilities and Their Impact on Human-Robot Collaboration: A Systematic Literature Review | International Journal of Social Robotics | Springer Nature Link학술 논문·연구보고서 · 이번 편집에서 원문 페이지 확인
  4. Embodied artificial intelligence as a paradigm shift for human–robot collaboration학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  5. Frontiers | Learning-Enabled Collaborative Intelligence in Search and Rescue Robotics: A Survey of Paradigms, Models, and Challenges학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  6. Toward Seamless Physical Human-Humanoid Interaction: Insights from Control, Intent, and Modeling with a Vision for What Comes Next | Journal of Intelligent & Robotic Systems | Springer Nature Link학술 논문·연구보고서 · 이번 편집에서 원문 페이지 확인
  7. Proactive robot task sequencing through real-time hand motion prediction in human–robot collaboration - ScienceDirect학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  8. Uncertainty-Aware Long-Term Multi-Worker Behavior Anticipation for Proactive Human-Robot Collaboration in Construction - ScienceDirect학술 논문·연구보고서 · HKUST 공식 연구정보로 핵심 조건 확인
  9. A tactile reflex arc for physical human–robot interaction - ScienceDirect학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  10. Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization학술 논문·연구보고서 · 이번 편집에서 원문 페이지 확인
  11. Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications | IEEE Journals & Magazine | IEEE Xplore학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  12. Gemini Robotics: Bringing AI into the Physical World학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  13. π0.5: a Vision-Language-Action Model with Open-World Generalization학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  14. Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models학술 논문·연구보고서 · 이번 편집에서 원문 페이지 확인
  15. Cosmos World Foundation Model Platform for Physical AI학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  16. Explicit World Models for Reliable Human-Robot Collaboration학술 논문·연구보고서 · 이번 편집에서 원문 페이지 확인
  17. Indicating Robot Vision Capabilities with Augmented Reality | International Journal of Social Robotics | Springer Nature Link학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  18. CCM-FCC: LLM-powered cognition-centered AI agent framework for proactive human-robot collaboration - ScienceDirect학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  19. Mutual Adaptation in Human-Robot Co-Transportation with Human Preference Uncertainty학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  20. Shared steering control framework based on visual-haptic compliance information for mitigating human–machine conflict - ScienceDirect학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  21. Trust as a design principle in human–robot collaboration: a review of explainable and adaptive control | Artificial Intelligence Review | Springer Nature Link학술 논문·연구보고서 · 이번 편집에서 원문 페이지 확인
  22. Robotic Control for Human–Robot Collaborative Assembly Based on Digital Human Model and Reinforcement Learning - Yao - Advanced Robotics Research - Wiley Online Library학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  23. Vision-language model-based human-robot collaboration for smart manufacturing: A state-of-the-art survey | ENGINEERING Management | Springer Nature Link학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  24. Why the next decade of physical AI must be human-centric | World Economic Forum산업·정책 자료 · 이번 편집에서 원문 페이지 확인
  25. Two-thirds of organizations rate physical AI as a high priority for the next three to five years - Capgemini Australia산업·정책 자료 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  26. Human–human and human–robot co–manipulation: A biomechanical analysis of a joint carrying task - ScienceDirect학술 논문·연구보고서 · 첨부 보고서 인용 보존 · 이번 편집에서 개별 상세 결과 재검증하지 않음
  27. Trust as a design principle in human–robot collaboration: a review of explainable and adaptive control | Artificial Intelligence Review | Springer Nature Link학술 논문·연구보고서 · 이번 편집에서 원문 페이지 확인
  28. HKUST · Multi-Worker Behavior Anticipation3초 관찰 예시 및 2026년 출판 정보의 보완 출처

원자료: Physical-AI-HAI.md. 13개 주제, 핵심 개념 11개, 도전과제 8개, RQ 14개, 구현 9단계, 적용 7개, 미해결 문제 6개, 미래 방향 9개, 공통기술 6계층과 원문 참고번호 27개를 보존했다. [21]과 [27]은 같은 논문의 중복 인용이다.