Where does politeness live in Hindi?

Listen
Share

Where does a model keep respect?

TL;DR. Hindi encodes respect in its grammar: choose aap, tum or tu and the verb must follow. This paper asks whether Gemma’s residual stream already holds that distinction before any pronoun appears, whether a sparse autoencoder feature carries it, and whether that feature can steer generation. The honest answer is mixed: the information is there and linearly readable (AUROC 0.722), but none of 50 candidate SAE features survives a pronoun-specificity audit, and a small steering pilot shows no causal control.

600contrastive Hindi pairs (plus 200 Japanese)
0.722pre-marker probe AUROC (layer-0: 0.535)
0 / 50SAE candidates passing the audit
24 GBa single GPU runs the whole pipeline

Models deployed in Indian-language settings routinely misjudge register, and the standard remedies each fall short: zero-shot prompting is inconsistent over long generations, fine-tuning is expensive per language, and machine-translation style honorific control does not carry over to interactive generation. Without knowing whether the model already encodes honorific level internally, every intervention stays at the surface. Mechanistic interpretability had, until now, almost no intersection with Indic NLP; this study is, to the authors’ knowledge, the first to combine public SAEs with causal interventions for an Indic sociolinguistic phenomenon.

Hindi honorific register is carried by pronoun and verb agreementThree second-person pronouns, each with its own verb agreement: tu with karta hai, tum usually the same singular form, aap with the plural-respectful karte hain.तू tuintimate or dismissiveतू करता हैtu karta haisingularतुम tumfamiliar, peersतुम करते होtum karte hoagreeing formआप aapformal, respectfulआप करते हैंaap karte hainplural-respectfulRegister is distributed: the pronoun and the verb must agree, across the whole utterance.Japanese, by contrast, marks it mainly in the verb suffix (-masu vs plain form).
Figure 1 Why Hindi is an interesting test. A wrong register is a grammatical and social error no competent speaker makes (addressing an elder with tu, or a singular verb where the plural-respectful is obligatory), and it matters in customer-service, healthcare and voice-assistant deployments. The paper separates “honorific” from general formality: it means the morphosyntactically encoded speaker–addressee relation. Source: paper · §1–2.2

Design: make the easy answer impossible

The obvious way to get a high score is to detect the pronoun. A feature that fires on aap is not evidence of an abstract honorific representation, so the whole design is built to take that route away. Gemma-2-2B (about 5 GB in bf16; model plus SAE peaks at 14.4 GB) is read through transformer-lens, and Gemma Scope JumpReLU SAEs are loaded through sae-lens.

Analysis pipeline and the causal triangleContrastive Hindi pairs feed Gemma-2-2B; residual activations go to a linear probe for correlation, a Gemma Scope SAE audit for features, and steering plus ablation for sufficiency and necessity.600 Hindi pairscontrastive,six domainsGemma-2-2Bresidual stream,26 layersLinear probeRQ1correlationGemma Scope SAERQ2feature auditSteer + ablateRQ3suff. / necessityThe “causal triangle” (Geiger et al.): a feature earns a causal claim only if it iscorrelatedprobe / effect sizesufficientadding it raises honorific markingnecessaryremoving it lowers marking
Figure 2 The study design (after the paper’s Figure 1). Each stage answers one research question: RQ1 does a linear probe decode honorific level after controlling for pronoun and agreement markers; RQ2 which Gemma Scope SAE features keep honorific specificity once both surface channels are audited; RQ3 do the tested directions steer or ablate the behaviour. Source: paper · Figure 1, §1
Dataset partWhat it contains
Hindi contrastive pairs600 minimal pairs differing only in honorific level (aap vs tum, aap vs tu, tum vs tu) across six domains: social, professional, family, public service, medical, customer service (90–102 pairs each)
Pronoun-stripped conditionevery sentence has a variant with all second-person pronouns removed, keeping verb agreement and context
Context-only, pre-marker conditiona social scenario that fixes who is speaking to whom, ending before any pronoun or agreement-bearing form; the residual state at the last context token must predict the register of the next utterance
Japanese comparison200 desu/masu vs plain-form pairs, same minimal-pair design
Constructiontemplates varying only the marker; labels follow deterministically from the template; built by the first author, a native Hindi speaker; naturalness checked on a held-out sample
The dataset, released under CC-BY-SA with the paper, together with a six-axis honorific rubric for a future native-speaker evaluation. Source: paper §4.2, §4.6
ConfoundControl
Pronoun tokenpronoun-stripped probe, plus the pre-marker context-only probe; a pre-specified Ahon / Apron audit for SAE features
Length and contentminimal pairs: members differ only in the honorific marker
Pair leakagecross-validation grouped by pair; a held-out-addressee-role split tests a stronger shift
Topic and registersix domains with 90–102 pairs each
Verb-agreement suffixkept in the stripped diagnostic set but absent from the primary pre-marker input; stated as a residual limitation of the feature audit
How surface cues are controlled. Source: paper Table 2

RQ1: the register is readable before it is spoken

The strictest condition ends the input before any second-person pronoun or agreement-bearing form. The activation at the final context token must then predict which register comes next. A character n-gram classifier, a strong surface baseline, is near chance on it, as is the layer-0 representation.

Probe AUROC on the context-only, pre-marker conditionProbe AUROC on the context-only, pre-marker conditionLayer 0 representation0.535Character n-grams0.550Word TF-IDF0.610Naturalistic pre-marker prefixes0.655Held-out addressee role0.692Best middle layer (L2 probe)0.722
Figure 3 The central result. Before a pronoun or agreement marker has appeared, a linear probe on a middle layer predicts the register at AUROC 0.722: +0.187 over layer 0, +0.172 over character n-grams and +0.112 over the stronger word-TF-IDF baseline (chance is 0.5). It holds when addressee roles are held out (0.692) and on naturalistic prefixes (0.655). The paper reads this as contextual decodability beyond the two most direct surface channels, not as proof of a wholly abstract representation. Source: paper · Table 3, §5.1

RQ2: the best-looking feature is a pronoun detector

Having found the information in the residual stream, the next question is whether a Gemma Scope SAE feature isolates it. Ranking features by raw effect size is the natural move, and it is exactly what fails here: in a task where one token is so predictive, the top of the ranking is full of features that detect the token. The paper therefore fixes a two-way, orientation-invariant audit before looking: honorific discrimination on pronoun-stripped pairs (Ahon) must be at least 0.80 while pronoun discrimination (Apron) stays at or below 0.55.

The SAE specificity auditA plot of honorific discrimination against pronoun discrimination; the passing region is high honorific, low pronoun; the top feature sits far from it.passing regionf646000.250.250.50.50.750.7511A_pron: raw vs stripped sentence discrimination (want ≤ 0.55)A_honRule fixed in advanceA_hon ≥ 0.80A_pron ≤ 0.5550 candidates, ranked byraw Cohen’s d, all auditedTop feature f646d = 2.209A_hon = 0.514A_pron = 0.808All 50: A_hon ≤ 0.518,A_pron 0.656–0.808
Figure 4 A negative result, reported as one. A feature can look strongly different between honorific levels (Cohen’s d = 2.2) simply because it detects the pronoun. Audited with orientation-invariant scores, none of the 50 candidates (layer-2 SAE, 16k features, JumpReLU) comes near the pre-specified rule: the best has chance-level honorific discrimination on pronoun-stripped pairs (0.514) and high pronoun discrimination (0.808), the signature of a pronoun detector. The small box shows where all 50 fall: the paper reports ranges (A_hon at most 0.518, and it cannot be below 0.5 by construction), not individual points. Source: paper · Table 4, §5.2

This does not show that no honorific feature exists: the search covers 50 raw-ranked candidates in one 16k-feature SAE at layer 2, an operational choice and not a claim that layer 2 is special. It shows that the usual ranking recipe surfaces pronoun-related features rather than a validated honorific one, and that the audit itself is reusable for other closed-class address-term systems such as Japanese, Korean or Javanese.

RQ3: steering does not establish control

For the top candidate (f646, layer 2), the paper adds the feature’s unit-normalised decoder direction during generation (sufficiency) and subtracts it or zeroes the encoder activation (necessity), scoring outputs with a deterministic morphology proxy.

Honorific proxy under steering (informal prompts, n = 5)Honorific proxy under steering (informal prompts, n = 5)-0.25-0.19-0.12-0.060.00baselineα = 2α = 4α = 8α = 16proxy meansteering strength α (unit-normalised decoder direction added at generated positions)
Figure 5 The exploratory intervention does not show control. Adding feature f646 moves the proxy by at most +0.028 (α = 8) and then reverses at α = 16 (−0.116 versus baseline): a non-monotonic dose response. On formal prompts, negative-direction ablation and zero-ablation fail to reduce honorific marking (the proxy goes up instead). Five prompts and one generation each is too little for statistics, so the paper treats this as a failure to establish causal control, not as evidence the feature is neither sufficient nor necessary. Source: paper · Table 5, §5.3
Condition (n = 5)Mean proxyΔ vs baseline
Baseline (informal prompts)−0.100–
Steering α = 2−0.091+0.009
Steering α = 4−0.091+0.009
Steering α = 8−0.072+0.028
Steering α = 16−0.216−0.116
Negative ablation (formal prompts)+0.216no reduction
Zero ablation (formal prompts)+0.245no reduction
The pilot in full. The proxy is a deterministic morphology score, hon(x) = (n_aap − n_tu) / n_tokens, counting aap-paradigm pronouns with their plural-respectful verb forms against tu-paradigm ones; it ranges roughly over [−1, 1]. Source: paper Table 5, §4.5

A quick look at Japanese

HindiJapanese
Completed-sentence probe AUROC (best layer)0.8170.999
AUROC spread over 26 layers0.0210.001
Clear AUROC peak?NoNo
A lightweight comparison with Japanese keigo. The Hindi figure is the diagnostic pronoun-stripped condition, not the 0.722 pre-marker result. Both languages are linearly decodable at essentially every layer, and Japanese almost perfectly, consistent with its desu/masu suffix being an even more regular surface cue than Hindi’s pronoun-plus-agreement. The comparison shows no depth trend in either language; it says nothing about whether the representation is abstract. Source: paper Table 6, §7

Limits and what comes next

  • Context correlations. The context-only probe can still exploit social-role information correlated with the label; held-out-role and naturalistic-prefix checks reduce but do not eliminate this.
  • Scope. One model (Gemma-2-2B), one SAE and layer for the feature analysis, a templated dataset; a probe-only replication covers nine models in the supplement.
  • Underpowered pilot. One layer, five prompts per condition, one generation per prompt: descriptive, not inferential.
  • No human evaluation yet. A pre-registered study is planned: three Hindi speakers blind-rating baseline, steering, ablation and random-feature generations on six 0–3 axes (addressee honorific, verb agreement, register consistency, context appropriateness, plus fluency and content preservation as controls).

Planned follow-ups: other Indic languages (Bengali, Tamil, Marathi), graded register axes beyond a binary honorific, instruction-tuned Gemma-2 variants, the full causal triangle on the top-5 and top-20 directions, and a test of whether Japanese keigo’s one-dimensional grammatical program steers more easily than Hindi’s coordinated one.

All articles

Research and building cool stuff

© 2026, Animesh Mishra

GitHub|LinkedIn

New Delhi, India