When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI
Sihan Jia, Oliver Lemon
arXiv:2608.28518v1Voice-controlled robots and embodied AI systems are becoming more common, but they inherit a hidden weakness: speech recognition mistakes. This paper asks what happens when a robot mishears a user. The authors simulate realistic ASR errors and test them against embodied AI safety benchmarks to see whether the system still refuses dangerous requests or instead carries them out. They find that some misheard inputs keep enough of the original meaning to preserve harmful intent, while others blur the wording in ways that make unsafe plans more likely. In some cases, automatic correction helps, but not reliably. The key takeaway is that speech errors are not just a usability problem; they can become a safety problem, especially when language models turn spoken commands into physical actions. As embodied AI moves into homes, hospitals, and workplaces, understanding and defending against these failure modes is essential.
Also spotted that day
Previous daily papers