Inaudible Audio Attacks Can Hijack AI Voice Models, Study Finds
Researchers say hidden commands in audio can steer AI voice models with high success, including some commercial systems.
Intelligence analysis by GPT-5.4 Mini

A Zhejiang University study describes AudioHijack, a way to bury commands inside audio that humans cannot hear but voice models can follow. The researchers say it worked across open models and some commercial systems, while common defenses blocked only a small share of attempts.
Some scientists found a trick for hiding secret instructions inside sound that people cannot hear. A voice AI can still pick up those instructions, like a walkie-talkie hearing a code hidden in static.
The team says this trick worked very well on several systems, even some made by big companies. They also said many normal safety checks did not stop it.
That means voice helpers may need stronger locks, because the sound itself can be used like a sneaky key. If the wrong audio gets in, the machine may follow the hidden order instead of the obvious one.
Analysis
What the study claims
Researchers at Zhejiang University say they built an attack called AudioHijack that hides instructions inside audio clips at levels people cannot hear. Those hidden commands can change how large audio-language models behave, with the paper reporting success rates ranging from 79% to 96%.
How far it reached
The team says the method transferred from open models to commercial voice AI systems from Microsoft and Mistral. That matters because it suggests the problem is not limited to one open-source stack or one training recipe. The article also says standard defenses stopped only a small fraction of the attempts, which suggests current filters and safety checks may not be enough on their own.
What comes next
The researchers are now looking at whether the same technique can reach closed models from OpenAI and Anthropic through shared open-source audio components. That is an important detail: even when a model itself is proprietary, the surrounding audio pipeline may still expose a path for abuse. The broader takeaway is that voice AI can be manipulated through the signal itself, not just through text prompts, which raises the bar for anyone deploying these systems in sensitive settings.
Key points
- Zhejiang University researchers described AudioHijack, an attack that hides commands in audio that humans cannot hear.
- The paper reports success rates of 79% to 96% against large audio-language models.
- The attack reportedly transferred from open models to commercial voice AI from Microsoft and Mistral.
- The article says common defenses blocked only a small portion of the attempts.
- The researchers are testing whether the approach can also affect closed models from OpenAI and Anthropic through shared audio components.



