PAPER / ARXIV:2609.19616
Hernandez, M.; Zhao, T.
RESUMO
Complexity measured from generated code is failure-dependent: a difficult prompt can yield a short failing program and be assigned low output complexity. We introduce a six-dimension prompt-side structural-complexity index scored before generation and kept separate from correctness. We select 5,000 Python prompts across six bands of a preliminary single-rater rubric. Four out-of-panel LLM raters rescore the locked prompts, giving 19,997 score rows; composite inter-rater reliability is ICC = 0.872 on the 4,998 prompts with all four ratings. We evaluate 21 models per prompt, yielding 105,000 generations. In the unadjusted mean-pooled analysis, pass rate has a nonmonotone breakpoint at 13.75, with 79.9% or below and 87.6% above. This is not universal failure cutoff. Task-type fixed effects shift breakpoint to 10.75 and cut regime gap from 7.6 to 2.1 points. A construction-frame control shifts it to 8.50 with raw gap -3.5 points, and neither frame alone reproduces pooled +7.6-point change. Model-specific fits include 16 upward and five downward changes. A 365-prompt audit-clean extension matches original five-model estimates in 15 bins but adds only 14 above bin 16. Among zero-pass generations with computable Lizard complexity, 28.5% pair a prompt with 8 or above complexity, most at 10. Human agreement is moderate and rater-dependent with disagreement-enriched calibration set; paraphrase and cross-language rescoring preserve score ordering. Overidentification tests reject joint restrictions on dimensions, so we treat composite as an index and make no causal interpretation of 2SLS estimates. The contribution is pre-generation measurement framework and bounded observational analysis of regimes.
NO MESMO MAPA