PAPER / ARXIV:2609.20715
Zhang, J.; Makhija, D.; Arivazhagan, M.G.
RESUMO
Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning applies loss only to agent-authored action tokens, using environment observations as context but not prediction targets. We ask whether this convention provides the best initialization for subsequent reinforcement learning. We introduce ActObs, which also supervises observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy model to consider action consequences without adding data, parameters, sequence length, or forward passes. The methods perform similarly after SFT but diverge after GRPO. On Qwen3-4B, GRPO from ActObs achieves higher pass@k at every evaluated sampling budget than its action-only counterpart on Terminal-Bench 2.0. On Qwen3-8B, it trades some pass@1 reliability for solving more distinct tasks. The advantage extends to code editing on aider-polyglot. ActObs retains more entropy during RL while requiring less policy movement, leaving the final policy closer to its initialization. Our analysis traces this difference to SFT: action observation gradients rapidly become orthogonal, while training leaves a large residual gradient that degrades prediction below the base model. Joint supervision prevents one-sided specialization, preserving consequence prediction and preparing the policy for downstream exploration.
NO MESMO MAPA