ServiceNow researchers introduced EVA, an approach to evaluating voice agents that considers task success alongside conversational quality. Recognition errors, an awkwardly long answer, or a delay can damage an interaction even when the underlying language model is capable.
Context
A good written answer does not necessarily work aloud. People cannot skim a spoken list in the same way they scan text. Product evaluations should therefore include pacing, confirmation, correction, and whether the user gets an understandable answer without repeatedly restating the question. The interface is part of the result.
Sources & authors
- A New Framework for Evaluating Voice Agents (EVA)Hugging Face · March 24, 2026



