PAPER / ARXIV:2609.14973
DeepCybo Team , Yu Bin , Haipeng Cao , Zheng Chang , Kai Chen , Youning Chen , Kailin Deng , Yichao Du , Xiaotong Fu , Haoyang Ge , Yunlong Guo , Chenliu Hao , Jiyan He , Xuguo He , Yakun Hou , Kai Hu , Cong Huang , Tuopusen Huang , Yu Huang , Hong Li , Peize Li , Shijie Lian , Xiaopeng Lin , Yun Lin , Haibao Liu , Haochen Liu , Qiuzhi Liu , Shengcai Liu , Zhiqiang Liu , Tao Luo , Peng Ren , Shuo Ren , Chaoyi Ruan , Zhaolong Shen , Yukun Shi , Qiyuan Su , Yuxuan Tian , Yining Wang , Changti Wu , Hao Wu , Xueyin Xu , Ruoqi Yang , Zhaoyang Yang , Hang Yuan , Zhaoyang Zeng , Hanwen Zhang , Ruimeng Zhang , Yao Zhang , Yibo Zhang , Yuxiang Zhang , Zhirui Zhang , Ziyi Zhang , Zubin Zheng , Zishen Zhuang
RESUMO
We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.
NO MESMO MAPA