Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary, Maike Züfle, Sara Rajaee, Pinzhen Chen, et al.
arXiv:2609.04173v1Today’s paper is about a problem that every machine translation benchmark eventually runs into: once models get strong enough, the benchmark stops being challenging, and the metrics stop telling us anything useful. The authors introduce the Last Translation Benchmark, a live collection of human-authored examples designed to break state-of-the-art translation systems across text, images, audio, and video. Instead of relying only on automatic scores or one-off human judgments, each example comes with handcrafted verification rules that spell out the specific failure modes to check. That makes evaluation more reproducible, more objective, and much more actionable. The broader idea matters because translation is no longer just text-to-text; modern systems have to handle multimodal input and subtle edge cases that standard benchmarks miss. By focusing on hard failures rather than average performance, this benchmark aims to keep pace with rapidly improving models and give researchers a clearer target for real progress.
Also spotted that day
Previous daily papers