PAPER / ARXIV:2609.06078
Chang Liu , Henghui Ding , Lingyi Hong , Ning Xu , Linjie Yang , Yuchen Fan , Canyang Wu , Jinrong Zhang , Xusheng He , Ce Bian , Xianjing Han , Jianlong Wu , Mingqi Gao , Sijie Li , Jungong Han , JeongRae Kim , Chaehyun Kim , Changwon Lim , Jungyoon Lee , Gyuil Lim , Doeon Kim , Seong-heum Kim , Pranjal Aggarwal , Sean Welleck , Yiwen Ren , Jianing Liu , Yingxin Wang , Kexin Zhang , Licheng Jiao , Lingling Li , Xu Liu , Jinxing Zhou , Suiyi Zhao , Yanghao Zhou , Ruohao Guo , Liangtao Shi , Jinxia Xie , Xiantao Hu , Ting Liu
RESUMO
This report summarizes the 8th Large-scale Video Object Segmentation (LSVOS) Challenge, held in conjunction with ECCV 2026. The challenge evaluates video segmentation in three complementary settings: complex semi-supervised video object segmentation on MOSEv2, text-guided referring video object segmentation on MeViSv2-Text, and audio-guided referring video object segmentation on MeViSv2-Audio. We describe the tasks and evaluation protocols and review the methods of the top three teams in each track. Across the nine leading solutions, foundation segmentation models are combined with target-aware memory, multimodal reasoning, explicit target-existence verification, agentic interaction, and corrective tracking. These systems illustrate a broader transition from single-model mask propagation toward modular pipelines that reason about object identity, query validity, and temporal reliability.
NO MESMO MAPA