Student researcher, May to September 2025, Vienna (remote)

Student Researcher, Complexity Science Hub

Part-time research on auditing gender-bias debiasing frameworks for LLMs at the Complexity Science Hub.

What I did

  • Found that the standard metric these frameworks are evaluated on can be gamed by responding less: refusing to answer looked like debiasing even when the model's actual bias had not changed.
  • Replaced it with partial-identification bounds instead of point estimates, and showed that under evaluations that allow non-response the uncertainty interval for a model's bias rate is exactly as wide as its non-response rate, however much data there is.
  • Reanalysed 96 published results across two standard bias benchmarks and found most reported improvements came from models answering less, not from being less biased.

Get in touch

Send a message

Email animeshmishra0567@gmail.com if you work in any of these areas and want to talk shop:

  • NLP and ML evaluation
  • Scientific AI and retrieval-augmented generation
  • Research roles ahead of a PhD

Research and building cool stuff

© 2026, Animesh Mishra

GitHub|LinkedIn

New Delhi, India