Research project, 2026

Measurement Validity of Hidden-Bias Audits

Partial identification and calibrated instrument verification for LLM fairness evaluation. Manuscript in preparation. Advisor: Dr. Sonia Khetarpaul, Dept. of Computer Science and Engineering, Shiv Nadar Institution of Eminence.

What I did

  • Audited a cross-model gender-bias debiasing framework and found its headline metric maximised by degenerate, uninformative response policies; replaced it with a signed bias-gap statistic and Manski-style partial-identification bounds.
  • Showed that when an evaluation permits non-response, the identified interval for a model's stereotype rate has width exactly equal to the non-response rate, independent of sample size.
  • Reanalysed 96 published method-model results on BBQ and UNQOVER and found that a majority of reported bias-score reductions came from increased abstention rather than a change in committed-answer behaviour.
  • Designed and synthetically verified a calibration framework, a cross-instrument compatibility certificate with an identified blind spot for shared measurement error, for auditing whether bias-elicitation methods reveal or rewrite model decisions.

Get in touch

Send a message

Email animeshmishra0567@gmail.com if you work in any of these areas and want to talk shop:

  • NLP and ML evaluation
  • Scientific AI and retrieval-augmented generation
  • Research roles ahead of a PhD

Research and building cool stuff

© 2026, Animesh Mishra

GitHub|LinkedIn

New Delhi, India