EVA-Bench Data 2.0: 3 Domains, 121 Tools, 213 Scenarios
ServiceNow-AI expands EVA-Bench to three enterprise domains, with 213 scenarios and 121 tools for testing voice agents.
Intelligence analysis by GPT-5.4 Mini

The new EVA-Bench release widens enterprise voice-agent evaluation from one domain to three: airline customer service, IT service management, and healthcare HR delivery. The dataset is built to be realistic, reproducible, and open source, with scenarios validated against frontier models.
This is like giving robot phone helpers a bigger test book. Instead of checking only one kind of job, the new version tests three kinds, with more situations, more tools, and tricky calls so the robots can be measured more fairly.
Analysis
What changed
ServiceNow-AI says EVA-Bench Data 2.0 expands the benchmark from one enterprise domain to three: Airline Customer Service Management, Enterprise IT Service Management, and Healthcare HR Service Delivery. The release totals 213 scenarios across 121 tools, which the authors describe as about a fourfold increase in scenario coverage from the original version.
How the benchmark is designed
The post emphasizes five design principles. First, the benchmark is voice-first: it only includes workflows that are realistically handled over the phone. Second, it aims for realism by modeling tool schemas after production APIs and grounding policies in real enterprise constraints. In the healthcare HRSD domain, the article says this includes US healthcare and administration details such as NPI numbers, FMLA, and insurance coverage.
Third, the scenarios are varied rather than repeated. The dataset includes single-intent calls, multi-intent calls with up to four intents, and adversarial calls where the caller tries to bypass troubleshooting, misstate urgency, or access information they should not see. The authors also include unsatisfiable goals, since real call volumes include requests that cannot be completed.
Fourth, authentication is built into every domain, but only where it fits the task. The article notes that OTP-based elevation appears where a real system would require it, rather than being forced into every scenario. Fifth, the benchmark is designed for reproducibility: each scenario has one correct resolution path, and the generation process removes cases where multiple action sequences could be valid.
How scenarios are generated
The dataset is produced with SyGra, a graph-based synthetic generation pipeline, using GPT-5.4 as the backbone. The article says each scenario is generated from three jointly consistent parts: the user goal, the tool-usage context, and the policy or domain rules. This joint generation is meant to avoid inconsistencies that happen when those parts are created separately.
The authors also say every scenario was validated for solvability against three frontier models: OpenAI GPT-5.4, Google Gemini 3.1 Pro, and Anthropic Claude Opus 4.6. That validation is presented as a way to keep the benchmark both difficult and fair.
Why it matters
The release is aimed at people evaluating voice agents as well as teams building their own datasets. The post also previews a multilingual extension, suggesting the benchmark is meant to grow beyond English-only enterprise settings.
Key points
- EVA-Bench Data 2.0 expands from one enterprise domain to three: airline customer service, IT service management, and healthcare HR delivery.
- The release includes 213 scenarios and 121 tools, which the authors say is roughly a 4x increase in scenario coverage.
- The benchmark is designed around voice-first workflows, realism, variety, authentication, and reproducibility.
- Scenarios include single-intent, multi-intent, and adversarial calls, plus some unsatisfiable goals.
- The dataset is open source and was validated against GPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.6.
If the dataset works as intended, teams will have a more realistic way to test voice agents across several enterprise settings instead of one narrow case. The open-source release and solvability checks could make it easier for others to build on the benchmark and compare systems more consistently.
A benchmark built from synthetic scenarios may still miss some messy real-world behavior, even with realism and validation efforts. The strong focus on reproducible single-path outcomes could also underrepresent situations where real calls branch in unpredictable ways.



