跳转到内容

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

Section titled “Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models”

⚠️ AI 生成 · 建议对照原文 本页为自动整理的学习笔记;关键数据与引用如需引用,请回查 PDF / 官方版本。

学习档位 中文笔记

类型 文献 · 更新 2026-07-19

所属 具身智能能力栈

中文学习笔记(自动生成,需核验)

Section titled “中文学习笔记(自动生成,需核验)”

Topic: embodied-capabilities · Tier: recent · Year: 2026 · Venue:
Evidence level: partial · 本地全文: 是 · 建议阅读: ~45 分钟
Paper: https://arxiv.org/abs/2606.11324
Code:
Generator: heuristic

与主题相关的代表性工作(启发式摘要,待来源核验):Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

— page 1 — Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Yifu Yuanv,# 1 , Yaoting Huang1 , Xianze Yao1 , Yutong Li1 , Shuoheng Zhang1 , Linqi Han1 , Pengyi Li1 , Jiangeng Sun1 , Wenting

待来源核验

待来源核验

  • Method Pick Coke Move Near Open/Close Drawer Open Top Drawer and Place Apple Overall RT-1-X 56.7 31.7 59.7 21.3 42.4 RT-2-X 78.7 77.9 25.0 3.7 46.3 OpenVLA 16.3 46.2 35.6 0.0 24.5 SpatialVLA 86.0 77.9

Method Pick Coke Move Near Open/Close Drawer Open Top Drawer and Place Apple Overall RT-1-X 56.7 31.7 59.7 21.3 42.4 RT-2-X 78.7 77.9 25.0 3.7 46.3 OpenVLA 16.3 46.2 35.6 0.0 24.5 SpatialVLA 86.0 77.9 57.4 0.0 55.3 OpenVLA-OFT 72.3 69.6 47.2 – 63.0 𝜋0 97.9 78.7 62.2 46.6 71.4 𝜋0 -FAST 75.3 67.5 42.9 0.0 46.4 𝜋0.5 – – – – 72.7 GR00T-N1.5 51.7 54.0 27.8 7.4 35.2 GR00T-N1.6 – – – – 67.7 Embodied-R1.5-VLA 92.3 93.8 86.1 97.2 92.4 pretraining, demonstrating that internalized embodied reasoning can ef

待来源核验

待来源核验

待来源核验(启发式提取未给出可靠数值)

待来源核验

与前序 / 同期 / 后续方法的关系

Section titled “与前序 / 同期 / 后续方法的关系”

待来源核验

待来源核验

先读 Abstract 与 Introduction,再读 Method,最后 Experiments。

  1. Q: 本文要解决的核心问题是什么? A: 待来源核验
  2. Q: 方法的关键表示是什么? A: 待来源核验
  3. Q: 主要数据集与指标? A: 待来源核验
  4. Q: 相对前序方法的差异? A: 待来源核验
  5. Q: 主要局限? A: 待来源核验
  • extract: — page 1 — Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Yifu Yuanv,# 1 , Yaoting Huang1 , Xianze Yao1 , Yutong Li1 , Shuoheng Zhang1 , Linqi Han1 , Pengyi Li1 , Jiangeng Sun1 , Wenting Jia1 , Zhao Zhang1 , Yuhao Liu1 , Ruihao Liao1 , Yucheng Hu
  • extract: Method Pick Coke Move Near Open/Close Drawer Open Top Drawer and Place Apple Overall RT-1-X 56.7 31.7 59.7 21.3 42.4 RT-2-X 78.7 77.9 25.0 3.7 46.3 OpenVLA 16.3 46.2 35.6 0.0 24.5 SpatialVLA 86.0 77.9 57.4 0.0 55.3 OpenVLA-OFT 72.3 69.6 47.2 – 63.0 𝜋0 97.9 78.7 62.2 46.6 71.4 𝜋0
  • topic: embodied-agents
  • sources: arxiv
  • retrieved_at: 2026-07-20
  • query: embodied AI agents foundation models
  • arxiv: 2606.11324
  • score_total: 39
  • suggested_tier: watch

(no prose relevance explanation — numeric score only or HTTP source)

(no snippet evidence in candidate pool)

生成:2026-07-21 · 来源条数 2 · 模型 heuristic · 需人工核验数字

围绕「Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models」的核心问题与动机(待结合全文核验)。

  • 见原文方法章节;以下为基于摘要/摘录的要点提示。
  • Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models …

与相近工作的关系待核验;请对照 related work。

  • 勿仅凭摘要推断未给出的数值指标。
  1. 这篇工作的输入/输出表示是什么?(Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models)
  2. 训练目标与评测协议各是什么?
  3. 主要失败模式或局限是什么?
flowchart LR
A["输入 / 观测"] --> B["表示 / 编码"]
B --> C["推理 / 解码"]
C --> D["输出 / 动作或检测"]
%% method sketch for: Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

方法结构示意(重绘;细节以原论文为准,待 PDF 核验)。

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models arch p.1

来源:原论文约 p.1(arch);学习用途摘录。

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models table p.15

来源:原论文约 p.15(table);学习用途摘录。

展开英文 Paper Card / AI deep analysis
Field Content
Year 2026
Authors Yifu Yuan, Yaoting Huang, Xianze Yao, Yutong Li, Shuoheng Zhang, Linqi Han, Pengyi Li, Jiangeng Sun, Wenting Jia, Zhao Zhang, Yuhao Liu, Ruihao Liao
arXiv 2606.11324
DOI
Topics embodied-capabilities
Paper https://arxiv.org/abs/2606.11324
展开 Extract / Selections / Local assets
  • embodied-capabilities: tier=recent rank=4 score=45 — auto refresh 2026-07-19 sources=arxiv
Embodied-R1.5: Evolving Physical Intelligence via
Embodied Foundation Models
Yifu Yuanv,# 1 , Yaoting Huang1 , Xianze Yao1 , Yutong Li1 , Shuoheng Zhang1 , Linqi Han1 ,
Pengyi Li1 , Jiangeng Sun1 , Wenting Jia1 , Zhao Zhang1 , Yuhao Liu1 , Ruihao Liao1 , Yucheng
Hu1 , Qiyu Wu1 , Yuxiao Li1 , Zibin Dong1 , Fei Ni1 , Yan Zheng1 , Shuyang Gu# 2 , Yi Mav,# 1 ,
Hongyao Tangv,# 1 , Han Hu2 , Jianye Hao# 1
1 Tianjin University, 2 Tencent Hunyuan
v Project Leader, # Corresponding Author (Contact: yuanyf@tju.edu.cn)
We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates
comprehensive embodied reasoning capabilities, spanning embodied cognition, task planning,
correction, and pointing, within a single architecture toward general physical intelligence. Lever-
aging three automated data construction pipelines to significantly expand the data coverage of
critical capabilities, we build a large-scale data system of over 15B tokens, and design a multi-task
balanced RL recipe to alleviate heterogeneous task conflicts. We further introduce a Planner-
Grounder-Corrector (PGC) closed-loop framework that enables a single model to autonomously
execute and self-correct over long-horizon tasks. With only 8B parameters, Embodied-R1.5
achieves SOTA on 16 out of 24 embodied VLM benchmarks, surpassing leading models like
arXiv:2606.11324v2 [cs.RO] 11 Jul 2026
Gemini-Robotics-ER-1.5 and GPT-5.4. Benefiting from the internalized embodied capabilities,
Embodied-R1.5 can be fine-tuned into a VLA with only a small amount of data, outperforming
leading VLA models like 𝜋0.5 across 4 popula