PAPER / ARXIV:2609.20068
Combe
RESUMO
This paper builds a theoretical bridge between the economic notion of marginal utility and two machine-learning constructs, matrix factorization and Key-Value cache transformer language models. The singular value spectrum of rating matrix is shown to be a diminishing marginal utility schedule for latent factors, the eigenvalue spectrum projected covariance operator is shown to be a diminishing marginal utility schedule for model's learned representation, cache eviction low-rank compression instances are shown to be constrained utility maximization under a memory budget. The three collapse into a single allocation rule: retain the top dimensions whose eigenvalue exceeds the shadow price of binding constraint. This framework is applied to automated extraction of structured information from geo-mining documents, where it motivates a multi-pass inference protocol, layer-wise TIES model merging procedure, a selection policy combining extraction quality, localization drift energy, scalarized with a Conditional Value-at-Risk term on drift. Two empirical contributions are reported. An 11.2-million-parameter hierarchical classifier, trained in about five minutes on GPU, reaches 90.0 per cent level-1 accuracy on held-out test set of 973-document uranium-exploration corpus, against 92.0 per cent of proprietary model on fifty-document human audit of same corpus, at latency of 2.62 ms per card against approximately 2,000 ms API and negligible cost. A diagnostic uniform-density TIES model merging exposes a reproducible degenerate mode in which merged model returns token-identical outputs across five geographically distinct districts while declaring high confidence; re-executing the merge under layer-wise calibrated densities removes that signature from sample. The full-scale extraction benchmark, including LoRA fine-tuning, is reported as projected rather than measured and remains an empirical extension of this work.
NO MESMO MAPA