PAPER / ARXIV:2609.20779
Sarah Wyer, Sue Black, Noura Al Moubayed
RESUMO
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this harm laundering. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed do not. The pattern is most visible at GPT-5: Topic 5 (1,997 documents) frames breast cancer as a men's rights debate, with zero equivalent clusters appearing in women-directed output. Three independent classifiers score non-toxic. Sentiment scores invert at GPT-4: early models demean women; later models over-correct. Topic diversity falls to 36% relative to men at the GPT-4 alignment boundary (W/M = 0.58, from 0.91 at GPT-2). REGARD representational disparity correlates with release date (ρ = +0.55, p = .034) while Detoxify does not (ρ = -0.23, p = .42): toxicity fall as representational harm grows. We formalise harm laundering as a three-criteria test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, toxicity reduction is not sufficient proxy for harm reduction.
NO MESMO MAPA