Research
VoicEmu: Simulating the Tails of Human Speech
Aug 7, 2026
Share
Frontier models struggle with accents and emotions. VoicEmu pairs speech generation with PCA-whitened embeddings for controllable voice agent stress tests.
TL;DR
Multi-modal foundation models fail at classifying paralinguistic attributes — their predictions are biased and drift unpredictably between model checkpoints. We show that simple embedding geometry (PCA-whitened centroids) can be used to classify accents at ~93% accuracy where frontier models perform at ~39%. We then use the same embeddings backbone to drive a generation pipeline that stress-tests voice agents against long-tail users before they reach production.


