01 · The supervision bottleneck
일반화 가능한 보상 모델은 무엇을 정답으로 배워야 하는가
generalist robot policy가 reinforcement learning으로 확장될수록, 모든 새로운 과제마다 사람이 reward를 설계하거나 preference pair를 만드는 방식은 병목이 된다. RynnValue 논문은 이 문제를 모델 구조보다 supervision의 좌표계 문제로 본다.
기존의 범용 robotic reward model은 trajectory preference, reference demonstration, local state comparison처럼 특정 trajectory 내부의 anchor를 사용하는 경우가 많다. 또 하나의 흔한 대안은 trajectory 시작과 끝을 0과 1로 정규화한 progress이다. 문제는 이런 값이 서로 다른 embodiment, 수행 속도, task duration을 가로질러 같은 의미를 유지하기 어렵다는 점이다.
논문이 제기하는 핵심 비판은 progress가 task 내부의 상대 좌표라는 데 있다. 같은 0.5 progress라도 4초짜리 동작과 40초짜리 동작에서 제어 이론적 의미는 다르다. 더구나 실패와 후퇴가 포함된 trajectory에서는 단순히 시간 순서가 뒤로 갈수록 progress가 커지는 가정 자체가 잘못될 수 있다.
02 · A reusable target
Temporal distance를 goal-conditioned cost-to-go로 읽는다
RynnValue는 observation에서 language-specified goal까지 남은 시간을 temporal distance로 정의한다. relabeled completion cutoff를 tG, 현재 observation의 timestamp를 ti라고 하면 absolute target은 다음과 같다.
즉 목표에 가까워질수록 값은 0에 가까워진다. minimum-time objective라는 가정 아래 이 값은 hitting-time cost-to-go로 해석할 수 있다. normalized progress와 달리 원래의 시간 scale을 보존하기 때문에 서로 다른 trajectory를 같은 단위로 표현할 수 있다는 것이 논문의 주장이다.
모델은 absolute distance만 배우지 않는다. presented sequence에서 이웃한 두 observation의 timestamp 차이도 relative target으로 학습한다.
학습 중 observation 순서를 섞기 때문에 이 값은 양수뿐 아니라 음수도 될 수 있다. 따라서 모델은 “영상에서 뒤에 나온 frame이 더 좋다”는 규칙을 외우기보다 실제 상태 변화가 goal에 접근했는지 후퇴했는지를 시각적으로 판단해야 한다.
03 · Architecture
하나의 multimodal backbone, 두 개의 temporal head, 별도의 language output
RynnValue는 embodied foundation model인 RynnBrain을 backbone으로 사용한다. main setting에서 한 trajectory clip당 K=8개의 observation을 샘플링하고, 각 temporal prediction마다 N=8개의 반복 query token을 둔다. 한 개의 query token에 object configuration, robot-object interaction, 중간 단계, completion evidence를 모두 압축시키지 않고 여러 query representation을 연결해 더 풍부한 readout을 만들려는 설계이다.
Absolute-value query
각 observation에서 completion cutoff까지 남은 temporal distance를 예측한다.
Relative-value query
presented sequence에서 인접 observation 사이의 signed temporal displacement를 예측한다.
Language branch
전체 video를 본 뒤 description, instruction match, success 판단을 autoregressive하게 생성한다.
absolute와 relative 값은 단순 scalar regression으로 학습하지 않는다. absolute 범위 [0, 512]초와 relative 범위 [-256, 256]초를 각각 256개의 symlog-spaced bin으로 나누고, continuous target을 인접 두 bin의 two-hot distribution으로 표현한다. inference에서는 bin probability의 기댓값을 symlog space에서 계산한 뒤 inverse symlog로 되돌린다.
논문은 이 방식을 큰 시간값의 dynamic range를 압축하면서 near-zero 영역의 정밀도를 유지하고, target magnitude가 gradient scale을 직접 지배하는 것을 줄이는 distributional regression으로 설명한다. 두 head는 BroNet residual MLP로 구현되며 hidden width 4096, depth 8을 사용한다.
04 · Shortcut suppression
Temporal value를 배우려다 시간 순서만 외우는 문제
multi-frame model에 timestamp-derived label을 주면 예상보다 쉬운 shortcut이 생긴다. frame을 일정 간격으로 뽑으면 모델은 시각 장면을 읽지 않고 query 위치만 보고 거의 등차수열인 value curve를 추정할 수 있다. chronological order를 항상 유지하면 뒤쪽 frame이 목표에 더 가깝다는 bias도 이용할 수 있다. 더 나쁘게는 앞 observation의 value-query representation을 참고해 다음 value를 extrapolate할 수도 있다.
RynnValue는 이 세 shortcut을 서로 다른 층에서 차단한다.
Random temporal sampling
8개 observation을 불규칙한 timestamp에서 뽑아 고정 간격을 없앤다.
Temporal-order shuffling
절반의 training sequence는 시간순 정렬 없이 샘플링하고, 나머지는 backward transition을 포함할 수 있는 forward-biased walk를 사용한다.
Value-isolation attention
서로 다른 observation의 temporal-query group이 서로의 value query를 보지 못하게 해 value extrapolation을 막는다.
value-isolation mask는 context token이 temporal-query token을 다시 읽는 간접 경로도 차단한다. 같은 observation 안의 반복 query끼리는 상호 attention을 허용하지만, 다른 observation의 query group으로 정보가 넘어가지는 않는다. inference에서도 이 mask를 유지한다.
추가로 10%의 training sample에서는 다른 trajectory의 instruction을 붙이는 instruction-mismatch augmentation을 사용한다. 이 경우 원래 completion cutoff가 잘못된 target이 되므로 absolute temporal loss는 mask하고, instruction과 무관한 relative loss는 유지한다. language branch는 Match: No, Success: No를 배우게 된다.
05 · Scaling data
7,000시간을 모으는 것보다 서로 다른 데이터를 같은 target으로 바꾸는 일이 핵심이다
학습 mixture는 real robot, simulation, egocentric trajectory를 포함하며 7,000시간을 넘는다. 원본 기준 1,671,313 episode를 subtask segmentation과 cutoff relabeling을 거쳐 3,091,969개의 instruction-conditioned segment로 확장하고, 표에 집계된 instruction 수는 223,395개이다.
데이터는 AgiBot, EgoDex, Galaxea Open-World, InternData-A1, Open X-Embodiment, RDT, RoboCOIN, RoboMIND, RoboTwin, Soft-FOLD 등 여러 source를 포함한다. 중요한 것은 각 dataset을 별도의 progress scale로 맞추지 않고, timestamp와 completion cutoff에서 남은 시간을 직접 계산한다는 점이다.
다만 “label-cheap”은 “annotation-free”와 같은 말이 아니다. raw episode boundary가 semantic completion과 맞지 않는 경우 subtask segmentation과 cutoff relabeling이 필요하고, 일부 source는 dataset-specific trimming을 사용한다. appendix에서는 malformed instruction, placeholder, data-quality metadata, pure-motion instruction을 제거하는 source-aware curation도 설명한다. 4개 주요 source의 별도 curation 통계에서는 trajectory unit의 83.35%와 unique instruction의 98.99%를 남긴다.
natural-language video description supervision은 Qwen3-VL-27B로 생성한 segment-level caption을 사용한다. 이 caption은 temporal target을 바꾸지는 않지만 language branch와 shared representation을 학습하는 데 참여한다.
06 · Reward interface
남은 시간을 음수 potential로 뒤집고 dense shaping reward를 만든다
RynnValue의 raw prediction은 reward가 아니라 non-negative remaining time이다. 일반적인 value convention처럼 큰 값이 좋도록 부호를 뒤집어 observation potential을 정의한다.
goal에 가까워질수록 potential은 0에 접근한다. downstream RL에서는 이 potential의 시간차를 potential-based shaping에 사용한다. 논문의 real-world experiment는 shaping reward에 더해 task completion 전에는 -1, 성공 transition에는 0인 sparse term도 유지한다. authors는 reward-model prediction noise가 있는 상황에서 clean success signal을 보존하기 위해 이 sparse term을 남긴다고 설명한다.
07 · RBM-EVAL-OOD
Preference label 없이 trajectory ranking에서 preference-supervised baseline을 넘는다
주요 intrinsic benchmark는 Robometer가 제안한 RBM-EVAL-OOD trajectory-ranking track이다. 서로 다른 기관, robot embodiment, viewpoint, task family를 포함한 6개 OOD dataset의 976개 trajectory를 대상으로, ground-truth quality order와 모델 score order 사이의 Kendall's τa를 측정한다.
RynnValue는 마지막 queried observation의 absolute temporal-distance potential -v_end를 trajectory score로 사용한다. relative head와 language output은 이 benchmark score 계산에 사용하지 않는다.
| Method | Supervision 특징 | Average Kendall's τa | 해석 |
|---|---|---|---|
| Robometer (Progress only) | progress-only | 0.292 | normalized progress만 사용할 때의 비교점 |
| RoboReward-4B | preference-free 계열 비교 | 0.502 | 기존 preference-free 결과 중 강한 비교점 |
| Robometer (RBM-1M) | progress + trajectory preference | 0.655 | fully preference-supervised baseline |
| RynnValue-4B | temporal distance, no trajectory preference | 0.670 | 4B에서도 8B와 근접 |
| RynnValue-8B | temporal distance, no trajectory preference | 0.675 | 논문에서 보고한 최고 평균 |
8B의 평균 0.675는 fully preference-supervised Robometer의 0.655보다 높고, progress-only ablation 0.292의 두 배를 넘는다. 4B가 0.670에 도달하고 8B와 차이가 작다는 점도 저자들이 강조하는 결과이다.
instruction-trajectory alignment 실험에서는 모든 instruction과 trajectory를 교차 pairing한 confusion matrix를 만든다. matched pair가 diagonal에 집중될수록 language goal에 잘 grounding된 모델이다. 동일 protocol로 공개 weight를 재평가했을 때 RynnValue의 normalized diagonal margin은 0.79이고, 가장 강한 baseline은 0.67이다.
08 · Ablation and scaling
성능의 중심은 모델 크기보다 shortcut을 제거한 데이터 구성에 가깝다
ablation은 논문의 설계 논리를 직접 시험한다. full model의 평균 Kendall's τa가 0.675일 때 temporal-order shuffling을 제거하면 0.189로 가장 크게 하락한다. random sampling을 uniform sampling으로 바꾸면 0.379, value-isolation attention을 제거하면 0.482이다.
| Variant | Average Kendall's τa | 관찰 |
|---|---|---|
| w/o Shuffle | 0.189 | sequence position shortcut이 가장 치명적임을 시사 |
| Uniform Sampling | 0.379 | 고정 temporal interval이 stereotyped value curve를 허용 |
| w/o Isolation | 0.482 | query 간 value extrapolation 차단이 중요 |
| w/o Language | 0.537 | description/match/success supervision이 semantic grounding에 기여 |
| w/o Relative | 0.627 | local forward/backward signal이 absolute target을 보완 |
| Full Model | 0.675 | 모든 구성요소 사용 |
scaling analysis도 흥미롭다. 저자들은 episode volume과 task diversity를 분리해, comparable episode count를 유지한 채 하나는 같은 task 내 episode를 늘리고 다른 하나는 task 수를 늘린다. unseen-task validation의 mean absolute temporal-distance error를 보면 within-task episode volume은 이른 구간에서 포화되는 반면, task diversity를 늘릴수록 error가 전체 범위에서 계속 낮아진다.
09 · Real-world policy learning
Value model이 실제 policy를 더 잘 학습시키는가
논문은 benchmark ranking을 넘어 dual-arm Franka system에서 네 가지 manipulation task를 평가한다. Bread Basket Placement, Steak Serving with a Spatula, Box-in-Drawer Placement, Bimanual Box Transfer이며, reward-model training corpus에 이 task와 object, workspace scene이 포함되지 않았고 target-domain fine-tuning도 하지 않았다고 보고한다.
offline RL은 IQL을, online RL은 DSRL을 사용한다. 평가에서는 task별로 policy를 20회 실행한다. online 조건의 세 task는 initial object configuration을 randomize하고, 정밀한 box-drawer alignment가 필요한 Box-in-Drawer는 fixed reset을 사용한다.
online에서 RynnValue는 72.5%, Robometer는 52.5%, sparse reward는 48.8%의 평균 성공률을 보인다. offline에서는 각각 82.5%, 63.8%이며 sparse reward는 22.5%, original SFT policy는 23.8%이다. 특히 Steak Serving과 Bimanual Box Transfer의 online 성공률은 Robometer 대비 각각 30, 35 percentage point 높다.
하지만 모든 task에서 큰 차이가 난 것은 아니다. Box-in-Drawer의 online 성공률은 shared initial checkpoint가 50%인 조건에서 Robometer 65%, RynnValue 70%에 그친다. 저자들은 reward model이 third-person RGB만 보고 있어 visually similar state 사이의 grasp stability와 precise alignment 차이를 충분히 판별하기 어렵다고 분석한다.
real-world reward inference에서는 training의 8 frame 대신 Robometer protocol에 맞춰 history를 4 frame으로 uniform subsampling한다. RynnValue는 이 경로에서 absolute temporal-distance head만 사용하고 language generation branch는 호출하지 않는다.
10 · Limits and cautions
Temporal distance는 좋은 공통 단위이지만 모든 가치의 정의는 아니다
첫째, objective가 시간 중심이다. temporal distance는 approximately minimum-time objective를 가정한다. 빠른 completion이 반드시 좋은 behavior인 것은 아니다. energy, safety, precision 같은 task-specific cost를 직접 표현하지 못하며, 논문도 이를 future work로 명시한다.
둘째, horizon이 아직 짧다. RynnValue는 sampled observation의 short window에서 temporal distance를 추정한다. 더 긴 horizon과 streaming inference는 향후 확장 과제로 남아 있다.
셋째, preference-free는 human-free가 아니다. preference pair는 만들지 않지만 completion cutoff와 subtask segmentation, source-aware data curation이 필요하다. downstream 실로봇 학습에서는 성공 label을 human operator가 기록한다. language supervision도 외부 VLM이 생성한 caption을 사용한다.
넷째, real-world evidence의 폭은 제한적이다. 네 task, task당 20회 evaluation이라는 실험은 reward interface의 실용성을 보여주지만, 모든 manipulation family와 embodiment에서의 일반성을 입증하기에는 범위가 좁다. 특히 precision-sensitive task에서는 RGB-only reward model의 state observability 한계가 직접 드러난다.
다섯째, downstream reward comparison은 완전히 동일한 scale이 아니다. RynnValue와 Robometer의 potential 자체가 다른 단위를 가지며 shaping coefficient도 0.1 대 1.0으로 다르다. 따라서 downstream success rate는 raw score calibration의 직접 비교가 아니라 각 reward interface의 tuned configuration이 policy learning에 미치는 효과로 해석해야 한다.
11 · Broader meaning
RynnValue가 던지는 더 큰 질문: foundation model의 label은 얼마나 재사용 가능한가
foundation model의 scaling은 흔히 backbone size나 raw data volume으로 설명된다. RynnValue는 다른 축을 강조한다. 서로 다른 dataset이 이미 가진 timestamp를 goal-conditioned cost-to-go로 재해석하면, preference pair를 새로 만들지 않고도 대규모 heterogeneous corpus를 value learning에 투입할 수 있다는 주장이다.
또 하나의 중요한 포인트는 value model의 실패가 단순한 capacity 부족이 아니라 shortcut learning에서 올 수 있다는 점이다. sequence position, uniform sampling interval, 다른 query가 노출한 value는 모두 model이 task semantics를 읽지 않고도 loss를 줄일 수 있게 한다. RynnValue의 strongest ablation이 temporal-order shuffling이라는 결과는 supervision target과 input presentation이 함께 설계되어야 한다는 사실을 보여준다.
12 · Key takeaways
핵심 정리
13 · References and resources
References
- Dongchi Huang, Hongyin Zhang, Bohan Hou, Siteng Huang, et al. RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance. arXiv:2608.09853v1, 2026.
- Anthony Liang, Yigit Korkmaz, Jiahui Zhang, et al. Robometer: Scaling general-purpose robotic reward models via trajectory comparisons. arXiv preprint arXiv:2603.02115, 2026.
- Tony Lee, Andrew Wagenmaker, Karl Pertsch, Percy Liang, Sergey Levine, Chelsea Finn. RoboReward: General-purpose vision-language reward models for robotics. arXiv preprint arXiv:2601.00675, 2026.
- Andrew Y. Ng, Daishi Harada, Stuart Russell. Policy invariance under reward transformations: Theory and application to reward shaping. ICML, 1999.
- Ilya Kostrikov, Ashvin Nair, Sergey Levine. Offline reinforcement learning with implicit Q-learning. ICLR, 2022.
논문이 제공한 공개 리소스
- RynnValue project page
- RynnValue GitHub repository
- RynnValue Hugging Face collection
- RynnValue ModelScope collection
이 글의 수치, 모델 구성, 실험 조건과 한계는 위 primary source의 본문과 appendix를 기준으로 정리했다. 해석과 추론은 별도의 callout으로 구분했다.