PAPER / ARXIV:2609.19928
Mitsuhashi, Morimura, Ito
RESUMO
Per-user LLM inference on transaction histories binds the budget linearly to user count, which becomes prohibitive at applied scale. We re-cast attribute inference from per-user per-transaction-pattern. The pipeline runs in three phases: Resolve abstracts item names with optional web grounding, Profile infers attributes for each frequent pattern, and Tag clusters free-text attributes into a queryable database. In Profile, a single LLM call per pattern emits predefined categorical labels, free-text attributes, and per-attribute prevalence estimates. Because inference runs over patterns rather than users, budget grows with pattern count rather than user count. On the public Open e-commerce corpus, database is statistically indistinguishable from an LLM that reads each user's raw history directly in AUC across evaluated attributes. Estimates carry discriminative signal between positive and negative users. The pipeline is deployed in major Japanese bank profiling order of tens of millions of users, close a three-order-of-magnitude reduction in targets versus a pipeline. The code is publicly available https://github.com/CyberAgentAILab/profiling-agent-open-ecommerce.
NO MESMO MAPA