Research project, 2026
Measurement Validity of Hidden-Bias Audits
Partial identification and calibrated instrument verification for LLM fairness evaluation. Manuscript in preparation. Advisor: Dr. Sonia Khetarpaul, Dept. of Computer Science and Engineering, Shiv Nadar Institution of Eminence.
What I did
- Audited a cross-model gender-bias debiasing framework and found its headline metric maximised by degenerate, uninformative response policies; replaced it with a signed bias-gap statistic and Manski-style partial-identification bounds.
- Showed that when an evaluation permits non-response, the identified interval for a model's stereotype rate has width exactly equal to the non-response rate, independent of sample size.
- Reanalysed 96 published method-model results on BBQ and UNQOVER and found that a majority of reported bias-score reductions came from increased abstention rather than a change in committed-answer behaviour.
- Designed and synthetically verified a calibration framework, a cross-instrument compatibility certificate with an identified blind spot for shared measurement error, for auditing whether bias-elicitation methods reveal or rewrite model decisions.
Email animeshmishra0567@gmail.com if you work in any of these areas and want to talk shop:
- NLP and ML evaluation
- Scientific AI and retrieval-augmented generation
- Research roles ahead of a PhD