Research

VoicEmu: Simulating the Tails of Human Speech

Aug 7, 2026
Share
VoicEmu: Simulating the Tails of Human Speech

Frontier models struggle with accents and emotions. VoicEmu pairs speech generation with PCA-whitened embeddings for controllable voice agent stress tests.

TL;DR

Multi-modal foundation models fail at classifying paralinguistic attributes — their predictions are biased and drift unpredictably between model checkpoints. We show that simple embedding geometry (PCA-whitened centroids) can be used to classify accents at ~93% accuracy where frontier models perform at ~39%. We then use the same embeddings backbone to drive a generation pipeline that stress-tests voice agents against long-tail users before they reach production.

Related articles

A Systems View of the Space
ResearchNov 21, 2025

A Systems View of the Space

Distyl Takes #1 Spot on BIRD Benchmark (Leading Text-to-SQL Benchmark)
ResearchJul 25, 2024

Distyl Takes #1 Spot on BIRD Benchmark (Leading Text-to-SQL Benchmark)

Lattice: Building Self-Correcting Guardrails for Conversational Agents
ResearchFeb 20, 2026

Lattice: Building Self-Correcting Guardrails for Conversational Agents