ML/NLP Researcher
Open to research rolesNew Delhi
Animesh Mishra

NLP Researcher, New Delhi
Paper at EMNLP 2026 Main (A* venue) on retrieval-augmented generation for science, plus a WMT 2026 paper on blind spots in machine translation evaluation. Presenting both in Budapest.
Evaluation, Retrieval and Scientific AI
How models behave on real scientific tasks, and what they are actually responding to: 88% of 1,072 molecules got inconsistent predictions depending on notation, retrieval that keeps conflicting evidence apart, and translation metrics that cannot see what a tokeniser deletes.
Open Source and Fast Teams
Contributed to open-source work at Nous Research, researched at DRDO and Complexity Science Hub Vienna, and co-founded ClerkTree. I like fast, high-output environments over slow academic ones. Looking for my next research role before a PhD.
Research statement
Venues & Places
Published and presented
Worked and studied with
Research you can run
Selected Work
Open benchmarks, metrics and research code. Disagreement-aware retrieval for science, translation metrics that
notice what tokenisers delete, and chemistry language models probed from the inside, all public on GitHub.
EVIRAG-Bench
Notation Matters

Notation matters is out in Digital Discovery (Royal Society of Chemistry), Gold Open Access. The same molecule written as SMILES, IUPAC, InChI or SELFIES gave inconsistent predictions for 88% of 1,072 molecules.

Wired Google's fruit-fly connectome into a crypto trading system on Binance Spot Testnet, with dopamine-gated plasticity driven by realised P&L. Built with Krishang Sharma and open-sourced.

Global Fintech Fest 2026: left with better questions than I came in with, on AI exposed to real conversations across languages, voice agents, and where verification should happen when AI creates and acts in the same loop.

WMT 2026: Translation Metrics Cannot Judge What Their Tokeniser Deletes. LIGATUR and AEGIS target invisible Unicode corruptions that get normalised away before a metric sees them. See you in Budapest.

EMNLP 2026 Main: Beyond Epistemic Collapse, Disagreement-Aware Scientific Retrieval-Augmented Generation. EVIRAG and EVIRAG-BENCH, a 1,250-query benchmark across five scientific domains.
Machina, open machine intelligence from ClerkTree: bearing-fault classification, remaining useful life, visual inspection and evidence-grounded industrial reasoning. Signals stay legible, models stay portable, actions stay governed.

Our agent Juris ranked #2 globally on the IBM VAKRA benchmark for tool selection, using a 36B model and a capability-specific routing stack over roughly 8,000 APIs in 62 domains.





