PAPER / ARXIV:2609.20625
Tisha Chawla, Susheem Koul
RESUMO
Large language model responses are non-deterministic, so failures in LLM agents hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and a multi-step trajectory re-run rarely repeats. Record-and-replay makes a run reproducible but existing agent tooling records runs only to trace or score them, not test code change against them. We present Chronicle, which an agent at its non-deterministic boundaries as immutable envelopes and replays it from the record. Its central operation, cut-point replay, serves a chosen subset of boundaries record and executes the complementary subset live with new code, turning a recorded incident into a regression test continuous integration. On a benchmark of 6 simulated model boundaries, recording adds 23 μs per crossing (0.008% of assumed 300 ms model call), full replay issues zero model calls and bit-stable across 20 repetitions, and tests fail on faulty code pass guarded and benign changes for all incidents. In a mutation study of tools, tests catch every mutant Chronicle lets the unsafe action through, while a baseline stubs every boundary, using the same assertion, catches none. Chronicle and publicly available https://github.com/theagentplane/chronicle.
NO MESMO MAPA