PAPER / ARXIV:2609.15067
Jiajun Peng , Fengrui Liu , Xinyu Liu , Feng Liu
RESUMO
Procedural audio has emerged as a viable source for transferable audio representation learning, but its design principles remain this http URL revisit two questions: how a procedural source should be scaled, and whether training choices developed on natural audio should transfer unchanged to procedural this http URL a controlled source, we separate scale into formula-class coverage C and within-class rendering diversity this http URL with FDSL and AudioMAE show that these two forms of scale provide different benefits and depend on the learning formulation and downstream task. A matched AudioMAE study further shows that procedural audio favors low mask ratios (10%--25%), whereas AudioSet-28K favors 50%--75%. Shared-codebook analysis reveals lower patch diversity and stronger temporal predictability in procedural audio. These results motivate source-aware procedural pre-training, where source scaling and learning configuration are considered this http URL is available at this https URL .
NO MESMO MAPA