PAPER / ARXIV:2609.20548
Sivachandran, K.; Paleja, R.
RESUMO
Reinforcement learning agents trained to maximize their own reward in repeated interactions can converge to supra-competitive outcomes resembling explicit collusion, without communication or shared design. Existing mitigation approaches are largely tied to specific economic settings, like two-sided platforms and auctions, leaving open how to design interventions for general games. We address this gap by formalizing the connection between empirical observations from prior work on Q-learning collusion and classical theory of Simple Penal Codes (SPCs). We show any non-trivial SPC induces a quantifiable conditional dependence between agents' policies, detectable via total variation distance between an agent's action distributions across cooperation and defection histories. Building on this connection, we propose CURB (Collusion Unwinding Reward shaping Belief injection), a reward-shaping framework that penalizes the Total Variation (TV) signal during Q-learning and is guaranteed to convert any fixed point of the dynamics into a trivial one, thus precluding collusive equilibria sustained by punishment threats. Empirically, CURB substantially reduces collusion in both Bertrand and Cournot Competition Repeated Games. We further demonstrate that CURB extends to deep Q-network agents in continuous competition, suggesting the mechanism generalizes beyond tabular Q-learning.
NO MESMO MAPA