Translation metrics cannot judge what their tokeniser deletes
Our paper has been accepted to the WMT 2026 Proceedings, another one after our EMNLP 2026 Main Conference paper. It is about blind spots in machine translation evaluation: Translation Metrics Cannot Judge What Their Tokeniser Deletes.
The finding
We found that certain invisible Unicode corruptions can get normalised away during tokenisation, making a corrupted translation and the clean version identical to the metric at the input level.
LIGATUR and AEGIS
We explore this through LIGATUR and AEGIS, targeting the kinds of errors standard neural MT metrics can completely miss. LIGATUR is a contrastive challenge set; AEGIS is a metric designed to account for these blind spots.
The code is in the Palimpsest repository on GitHub.
Thanks
Thanks to Krishang Sharma for being in the trenches with me on this one, and to Sonia Khetarpaul for the constant guidance and push. We will be presenting this work at WMT 2026. See you in Budapest.
