Translation metrics cannot judge what their tokeniser deletes

Listen
Share

Our paper has been accepted to the WMT 2026 Proceedings, another one after our EMNLP 2026 Main Conference paper. It is about blind spots in machine translation evaluation: Translation Metrics Cannot Judge What Their Tokeniser Deletes.

The finding

We found that certain invisible Unicode corruptions can get normalised away during tokenisation, making a corrupted translation and the clean version identical to the metric at the input level.

LIGATUR and AEGIS

We explore this through LIGATUR and AEGIS, targeting the kinds of errors standard neural MT metrics can completely miss. LIGATUR is a contrastive challenge set; AEGIS is a metric designed to account for these blind spots.

The code is in the Palimpsest repository on GitHub.

Thanks

Thanks to Krishang Sharma for being in the trenches with me on this one, and to Sonia Khetarpaul for the constant guidance and push. We will be presenting this work at WMT 2026. See you in Budapest.

All articles

Research and building cool stuff

© 2026, Animesh Mishra

GitHub|LinkedIn

New Delhi, India