Where does politeness live in Hindi?
Where does a model keep respect?
TL;DR. Hindi encodes respect in its grammar: choose aap, tum or tu and the verb must follow. This paper asks whether Gemma’s residual stream already holds that distinction before any pronoun appears, whether a sparse autoencoder feature carries it, and whether that feature can steer generation. The honest answer is mixed: the information is there and linearly readable (AUROC 0.722), but none of 50 candidate SAE features survives a pronoun-specificity audit, and a small steering pilot shows no causal control.
Models deployed in Indian-language settings routinely misjudge register, and the standard remedies each fall short: zero-shot prompting is inconsistent over long generations, fine-tuning is expensive per language, and machine-translation style honorific control does not carry over to interactive generation. Without knowing whether the model already encodes honorific level internally, every intervention stays at the surface. Mechanistic interpretability had, until now, almost no intersection with Indic NLP; this study is, to the authors’ knowledge, the first to combine public SAEs with causal interventions for an Indic sociolinguistic phenomenon.
Design: make the easy answer impossible
The obvious way to get a high score is to detect the pronoun. A feature that fires on aap is not evidence of an abstract honorific representation, so the whole design is built to take that route away. Gemma-2-2B (about 5 GB in bf16; model plus SAE peaks at 14.4 GB) is read through transformer-lens, and Gemma Scope JumpReLU SAEs are loaded through sae-lens.
| Dataset part | What it contains |
|---|---|
| Hindi contrastive pairs | 600 minimal pairs differing only in honorific level (aap vs tum, aap vs tu, tum vs tu) across six domains: social, professional, family, public service, medical, customer service (90–102 pairs each) |
| Pronoun-stripped condition | every sentence has a variant with all second-person pronouns removed, keeping verb agreement and context |
| Context-only, pre-marker condition | a social scenario that fixes who is speaking to whom, ending before any pronoun or agreement-bearing form; the residual state at the last context token must predict the register of the next utterance |
| Japanese comparison | 200 desu/masu vs plain-form pairs, same minimal-pair design |
| Construction | templates varying only the marker; labels follow deterministically from the template; built by the first author, a native Hindi speaker; naturalness checked on a held-out sample |
| Confound | Control |
|---|---|
| Pronoun token | pronoun-stripped probe, plus the pre-marker context-only probe; a pre-specified Ahon / Apron audit for SAE features |
| Length and content | minimal pairs: members differ only in the honorific marker |
| Pair leakage | cross-validation grouped by pair; a held-out-addressee-role split tests a stronger shift |
| Topic and register | six domains with 90–102 pairs each |
| Verb-agreement suffix | kept in the stripped diagnostic set but absent from the primary pre-marker input; stated as a residual limitation of the feature audit |
RQ1: the register is readable before it is spoken
The strictest condition ends the input before any second-person pronoun or agreement-bearing form. The activation at the final context token must then predict which register comes next. A character n-gram classifier, a strong surface baseline, is near chance on it, as is the layer-0 representation.
RQ2: the best-looking feature is a pronoun detector
Having found the information in the residual stream, the next question is whether a Gemma Scope SAE feature isolates it. Ranking features by raw effect size is the natural move, and it is exactly what fails here: in a task where one token is so predictive, the top of the ranking is full of features that detect the token. The paper therefore fixes a two-way, orientation-invariant audit before looking: honorific discrimination on pronoun-stripped pairs (Ahon) must be at least 0.80 while pronoun discrimination (Apron) stays at or below 0.55.
This does not show that no honorific feature exists: the search covers 50 raw-ranked candidates in one 16k-feature SAE at layer 2, an operational choice and not a claim that layer 2 is special. It shows that the usual ranking recipe surfaces pronoun-related features rather than a validated honorific one, and that the audit itself is reusable for other closed-class address-term systems such as Japanese, Korean or Javanese.
RQ3: steering does not establish control
For the top candidate (f646, layer 2), the paper adds the feature’s unit-normalised decoder direction during generation (sufficiency) and subtracts it or zeroes the encoder activation (necessity), scoring outputs with a deterministic morphology proxy.
| Condition (n = 5) | Mean proxy | Δ vs baseline |
|---|---|---|
| Baseline (informal prompts) | −0.100 | – |
| Steering α = 2 | −0.091 | +0.009 |
| Steering α = 4 | −0.091 | +0.009 |
| Steering α = 8 | −0.072 | +0.028 |
| Steering α = 16 | −0.216 | −0.116 |
| Negative ablation (formal prompts) | +0.216 | no reduction |
| Zero ablation (formal prompts) | +0.245 | no reduction |
A quick look at Japanese
| Hindi | Japanese | |
|---|---|---|
| Completed-sentence probe AUROC (best layer) | 0.817 | 0.999 |
| AUROC spread over 26 layers | 0.021 | 0.001 |
| Clear AUROC peak? | No | No |
Limits and what comes next
- Context correlations. The context-only probe can still exploit social-role information correlated with the label; held-out-role and naturalistic-prefix checks reduce but do not eliminate this.
- Scope. One model (Gemma-2-2B), one SAE and layer for the feature analysis, a templated dataset; a probe-only replication covers nine models in the supplement.
- Underpowered pilot. One layer, five prompts per condition, one generation per prompt: descriptive, not inferential.
- No human evaluation yet. A pre-registered study is planned: three Hindi speakers blind-rating baseline, steering, ablation and random-feature generations on six 0–3 axes (addressee honorific, verb agreement, register consistency, context appropriateness, plus fluency and content preservation as controls).
Planned follow-ups: other Indic languages (Bengali, Tamil, Marathi), graded register axes beyond a binary honorific, instruction-tuned Gemma-2 variants, the full causal triangle on the top-5 and top-20 directions, and a test of whether Japanese keigo’s one-dimensional grammatical program steers more easily than Hindi’s coordinated one.


