LegalNexus: finding precedent in a hierarchy

Listen
Share

Why law needs geometry

TL;DR. Legal authority is a tree (apex courts at the top, a flood of lower-court judgments below) and legal truth is a dispute (precedents get distinguished, even overruled). LegalNexus embeds the citation network in hyperbolic space so that radius means authority, uses three debating agents to clean contradictions out of the graph, and ends with a simulated courtroom that explains why a case was retrieved. It is a preprint under review at Engineering Applications of Artificial Intelligence.

49,633Indian court cases, 127,891 citations
0.88precision@5 (best baseline 0.85)
94%citation conflicts resolved by the swarm
0.42Gromov δ: the graph is tree-like

A flat embedding treats a Supreme Court judgment and a tribunal order as two points that happen to be near each other in meaning. A lawyer knows one binds the other. Existing retrieval systems, from BM25 to graph networks such as CaseGNN, largely ignore that structure, or model the citation graph as static even though precedents are constantly distinguished and overruled. LegalNexus is an attempt to put both properties into one system.

LegalNexus architectureThree phases: a hyperbolic embedding stage, a game-theoretic agent swarm that cleans the citation graph, and an adversarial hybrid retrieval stage that ends in a simulated courtroom.PHASE 1 · HYPERBOLIC FRAMEWORKCase textslegal casesdatabaseEmbeddingsGemini768-dHGCNℝ⁷⁶⁸ → 𝔻¹²⁸→ 𝔻⁶⁴Fermi-Diracdecoder:citation prob.Hierarchyradius =authorityPHASE 2 · GAME-THEORETIC MULTI-AGENT SWARMLinkerproposescitationsInterpreterlabels each edge:follow · overruleConflictcycles andcontradictionsNashpayoffsstabiliseKG storageconsistentgraphrefinement roundsPHASE 3 · ADVERSARIAL STRATEGY ENGINEQueryGemini 2.5Flash intentWeightsadapt tointent5 retrieversGNN · semantic · text· graph · citationCourtroomprosecutor ·defense · judgeAnswertop-k + ruling
Figure 1 The system, redrawn from the manuscript’s Figure 1. Geometry organises the space, agents keep the citation graph consistent, and an adversarial layer explains the answer. The authors describe the first two as the “map” and the “compass”. Source: LegalNexus manuscript · Figures 1–3

Hyperbolic space: a place for trees

A tree with a branching factor above one grows exponentially with depth, but a Euclidean plane grows only polynomially, so embedding a deep hierarchy into flat space forces distortion. Hyperbolic space expands exponentially, which is why it has become the geometry of choice for taxonomies. The manuscript checks that the citation graph really is tree-like before relying on that, using Gromov’s δ.

Gromov δ-hyperbolicity (lower = more tree-like)Gromov δ-hyperbolicity (lower = more tree-like)Perfect binary tree0.00Legal citation network0.42Barabási–Albert scale-free1.23Erdős–Rényi random1.87
Figure 2 Is a citation network really a tree? Gromov’s δ measures how far a graph is from one (0 is a perfect tree). The legal citation network scores 0.42, well below a random graph (1.87) and a scale-free one (1.23): moderately hyperbolic, which is the paper’s justification for a Poincaré-ball embedding. Source: LegalNexus manuscript · Table 1

The model is a hyperbolic graph convolutional network: 768-dimensional Gemini embeddings of each case are mapped into the Poincaré ball, passed through two hyperbolic layers, and trained to predict citations with a Fermi-Dirac decoder and a contrastive loss that pulls cited cases together and pushes negatives apart. Nothing in training mentions court level.

the hyperbolic layer, as written in the manuscript
# log map to the tangent space, then Möbius aggregation (Poincaré ball, curvature c)
log0_c(x) = (1/√c) · artanh(√c‖x‖) · x/‖x‖
h_v⁽ˡ⁺¹⁾ = σ⊗c ( Aggregate{ W⁽ˡ⁾ ⊗c h_u⁽ˡ⁾ : u ∈ N(v) } )

# Fermi-Dirac decoder: probability that a citation edge exists
P(e_uv = 1 | h_u, h_v) = 1 / ( exp((d_c(h_u, h_v) − r) / t) + 1 )

# contrastive loss: pull cited pairs together, push sampled negatives past margin m
L = Σ_{(u,v)∈E⁺} d_c² + Σ_{(u,v)∈E⁻} max(0, m − d_c²)
SettingValue
Input768-d Gemini text embeddings
NetworkHGCN: layer 1 ℝ⁷⁶⁸ → 𝔻¹²⁸ (log map + linear), layer 2 𝔻¹²⁸ → 𝔻⁶⁴ (hyperbolic aggregation)
Library and optimisergeoopt · Riemannian Adam · learning rate 0.001 · weight decay 0.0005 · dropout 0.5
Losscontrastive, margin m = 2.0, 5 sampled negative edges per positive
Training5 epochs, early stopping (patience 20), batch of 256 edges
Costabout 12 minutes on one NVIDIA A100
HGCN training setup. Source: manuscript §4.2, §7.1
Learned radius by court levelA Poincare disk with three rings at the mean radii of Supreme Court, High Court and lower-court cases.schematic: ring radii are the reported meansmean radius rSupreme Court0.08 ± 0.03rootHigh Courts0.15 ± 0.05intermediateLower courts0.28 ± 0.08leaf94% of Supreme Court cases lie within r < 0.1087% of lower-court cases lie beyond r > 0.20no court labels were used in training
Figure 3 The headline geometric result: hierarchy emerges without supervision. Trained only to predict citations, the model places Supreme Court cases near the origin of the Poincaré ball, High Courts at intermediate radii and lower courts toward the boundary. The dots are an illustration, not the 49,633 real embeddings. Source: LegalNexus manuscript · Table 7, §4.3

A swarm that argues the graph into consistency

Geometry alone cannot tell a valid precedent from an overruled one. The manuscript therefore adds three agents that debate over the citation graph in rounds, using hyperbolic distance as a trust prior.

The debate-refine loopLinker proposes edges, Interpreter labels them, Conflict agent finds cycles and contradictions, and the graph is refined until no conflicts remain or the round limit is hit.Linkermaximises recall:regex + LLM proposecitationsInterpretermaximises precision:FOLLOW · DISTINGUISH· OVERRULEConflict agentenforces consistency:cycles andcontradictionsrefine: remove edges with confidence < 0.5, downgrade ambiguous ×0.8, reclassify contradictionsStop when no cycles or contradictions remain, or at the round limit.Φ(S) = α·|E_valid| − λ·|C_conflicts| with α = 1.0, λ = 1.5
Figure 4 Three agents with different jobs, so no one model has to find, verify and consistency-check citations at once (the authors’ answer to “hallucination loops”). The hyperbolic radius acts as a trust prior: citations flowing from low-radius (high-authority) to high-radius nodes start with more confidence; anti-hierarchical ones need strong textual evidence. Source: LegalNexus manuscript · §5, Algorithm 1
AgentRoleImplementation (as reported)Payoff
Linkerproposerregex patterns for citation formats (e.g. “AIR 2020 SC 1234”) plus LLM validation (Gemma-2-2B, temperature 0.3)u_L = Φ
InterpreteranalystBERT-based legal classifier trained on 10,000 annotated pairs; keyword heuristics and LLM for ambiguous casesu_I = Φ + C_I
Conflictcriticdepth-first search for citation cycles and contradictory edge types; emits REMOVE, DOWNGRADE or RECLASSIFYu_C = Φ
The three agents. Because each payoff differs from the shared potential Φ only by a term that does not depend on the agent’s own move, the game is an exact potential game, which guarantees best-response dynamics converge to a Nash equilibrium. Source: manuscript §5
Citation edge types in the corpus (127,891 citations)Citation edge types in the corpus (127,891 citations)FOLLOW89,523DISTINGUISH31,972OVERRULE6,396
Figure 5 Overrulings are rare (5%) but decisive: citing an overruled case as good law is exactly the kind of error that destroys trust. The Interpreter reaches 87% accuracy on 1,000 verified citation pairs, with precision 0.89 for OVERRULE. Source: LegalNexus manuscript · Table 3, §7.5

Across rounds the graph stabilises: the paper reports resolving 94% of detected citation cycles and contradictions, with a mean of roughly three to four rounds (the text gives 3.2 in one place and 4.2 in another, with a round limit of three to five). The authors are careful to say this is an empirical stable state and not a formal proof of equilibrium.

Retrieval that explains itself

The retrieval layer combines five scorers whose weights adapt to the type of query, adds a temporal score so landmark cases are not penalised for age, extracts Toulmin-style arguments, and finally stages a short courtroom debate over the top results.

RetrieverScore
Semanticcosine between query and case embeddings in hyperbolic space
Knowledge-graph traversalPageRank(c) × path similarity; follows FOLLOW edges and avoids OVERRULE paths (Neo4j / Cypher)
Text patternBM25-style score on the case text
Citation networkauthority-weighted citation importance
GNN link predictionlink probability from the trained HGCN
The five scorers combined as Score(q, c) = Σ wᵢ(q)·sᵢ(q, c). Source: manuscript §6.1
Retrieval weights: default and for a statute-interpretation queryRetrieval weights: default and for a statute-interpretation query0.000.120.250.380.500.350.40semantic0.250.28graph0.200.17text0.150.12citation0.050.03GNNdefaultadapted (Section 65B query)
Figure 6 Five retrievers, weighted by intent. A query-analysis step (Gemini 2.5 Flash) classifies the query as precedent search, fact-finding or statute interpretation and shifts the weights: precedent queries raise the citation weight by 0.15, fact-finding raises text by 0.15, constitutional matters raise semantic and graph by 0.10 and 0.05. Weights are renormalised to sum to one. The adapted example is the worked case study. Source: LegalNexus manuscript · §6.1, Table 9

Time matters in law: an old case can be semantically perfect and legally obsolete, yet a 1950s landmark that is still cited today should not be punished for age. The temporal score has a base decay and a “resurrection” term for recent citations; on 500 test queries it reduced obsolete recommendations by 34% while keeping actively cited historical precedents.

temporal scoring (§6.2)
T(c) = 1 / log(age(c) + 2) · (1 + R(c))        # base decay × resurrection
R(c) = Σ_{t ∈ Cites(c)}  1 / (age(t) + 1)       # recent citations revive an old case
Toulmin partExtracted from Anvar P.V. v. P.K. Basheer (2014)
ClaimElectronic records without a Section 65B certificate are inadmissible.
DataThe appellant produced CDs without the certificate required by Section 65B(4).
WarrantSection 65B is a special provision that overrides the general rules of secondary evidence.
BackingThe Evidence Act’s purpose is to ensure authenticity; electronic records are susceptible to tampering.
Argument extraction, one worked example. Arguments are stored in a graph and retrieval can follow chains of support instead of surface similarity; the paper reports 85% extraction accuracy on 200 validation cases and about +8% on complex multi-step queries. Source: manuscript Table 2, §6.3

The courtroom

Three LLM agents turn a ranked list into an explanation. A Prosecutor argues for the position using the top cases, a Defense agent finds distinguishing precedents and mitigating factors, and a Judge weighs both using case authority (a hyperbolic radius below 0.10 carries more weight) and citation support, then writes a balanced ruling. Each side is limited to three sentences over a single round. A counterfactual engine then perturbs the query’s facts to find the pivots on which the result turns.

Counterfactual pivotsChanging one fact in the query and measuring how much the result set changes; larger changes mark outcome-determinative facts.Query: “drunk driving at night resulting in pedestrian death”was drunk → was sober0.82 PIVOTpedestrian death → no fatality0.75 PIVOTat night → during daylight0.31 stablepivot threshold 0.501
Figure 7 “What fact would I need to change to win this case?” A shadow agent re-runs retrieval on the counterfactual query; impact = 1 − Jaccard(original results, counterfactual results); facts above 0.5 are flagged as pivots. The three example values are from the manuscript. Source: LegalNexus manuscript · §6.5, Eq. 14

Results

ItemValue
Cases49,633 Indian court cases, 1950–2024 (Supreme Court 19.8%, High Courts 50.0%, lower courts 30.2%)
Citations127,891 (2.6 per case on average); 8,432 cases contain explicit overruling or distinguishing statements
Split70 / 15 / 15 (34,743 / 7,445 / 7,445) with no temporal leakage: training cases predate test cases
Queries2,165 gold-standard queries, 3–8 relevant cases each (mean 4.7), relevance verified by ensemble LLM voting
Query mixprecedent search 35% · fact pattern 30% · statute interpretation 20% · procedural 15%
SourceIndian Kanoon, via the NyayAnumana dataset (available from its authors on request)
The evaluation corpus. Source: manuscript §7.1, Tables 3–4
Retrieval quality: precision@5 and recall@10Retrieval quality: precision@5 and recall@100.000.250.500.751.000.620.58TF-IDF0.680.64BM250.820.79Legal-BERT0.840.81Longformer-Legal0.790.74Hier-SPCNet0.850.83CaseGNN0.880.86LegalNexusP@5R@10
Figure 8 Main comparison. LegalNexus reaches P@5 0.88 and R@10 0.86 against the strongest baseline, CaseGNN, at 0.85 and 0.83 (MAP 0.87 vs 0.84, NDCG 0.89 vs 0.87). All baselines were evaluated under identical conditions on the same corpus, with Legal-BERT fine-tuned on it. Source: LegalNexus manuscript · Table 5
MethodP@5R@10MAPNDCG
TF-IDF0.620.580.600.64
BM250.680.640.660.70
Legal-BERT0.820.790.800.83
Longformer-Legal0.840.810.820.85
Hier-SPCNet0.790.740.760.80
CaseGNN0.850.830.840.87
LegalNexus0.880.860.870.89
Full retrieval table. Source: manuscript Table 5
Precision@5 by query typePrecision@5 by query type0.000.250.500.751.000.91Precedentsearch0.91Factpattern0.89Statuteinterpretation0.88Procedural
Figure 9 Where the gain is. The advantage over Legal-BERT is largest on precedent search (+17.3%, 758 queries), where court hierarchy matters most, and smallest on procedural questions (+8.6%, 325 queries). Overall P@5 is 0.88, +13.6% over Legal-BERT on all 2,165 queries. Source: LegalNexus manuscript · Table 6

What each component contributes

Drop in precision@5 when a component is removed (%)Drop in precision@5 when a component is removed (%)w/o Interpreter−10.2%w/o citation network−10.2%Euclidean GCN (no hyperbolic)−7.9%w/o Linker−6.8%w/o Conflict agent−4.5%Hyperbolic only (no agents)−3.4%Static weighting−3.4%w/o resurrection−3.4%
Figure 10 Ablations. Every component earns its place. Removing the Interpreter (edge labelling: follow vs overrule) hurts most, which the authors read as “correct edge labeling matters more than raw connectivity”; swapping the hyperbolic network for a Euclidean GCN costs 7.9%. Full system: P@5 0.88. Source: LegalNexus manuscript · Table 8

The authors summarise the ablation as geometry giving the “map” and agents the “compass”: the hyperbolic network separates court levels without the crowding a Euclidean network shows, and the agents clean the noisy graph the geometry depends on. Removing the Conflict agent also increases contradictory outputs by 23%.

Failure and cost

Why queries fail (100 queries with P@5 < 0.5)Why queries fail (100 queries with P@5 < 0.5)Ambiguous query34%No precedent exists22%Missing citation18%Lexical gap13%Multi-domain8%Other5%
Figure 11 Error analysis. A third of failures are queries with no legal terminology (“Can I film police?”) and another fifth are emerging doctrines with no clear precedent. Source: LegalNexus manuscript · Fig. 7, §7.8
Latency breakdownA single stacked bar: adversarial simulation takes 74.2 percent of the 11.4 second average response time, the rest is retrieval.adversarial simulation · 8.46 s · 74.2%retrieval 2.46 s (all five retrievers + query analysis)average 11.4 s per query over 2,165 queries; the LLM courtroom debate dominates
Figure 12 Where the time goes. Retrieval is fast (2.46 s, comparable to a search engine); the explainable courtroom simulation costs 8.46 s. The authors list speculative decoding and cached debate patterns as ways to cut this below 2 s, but those are projections, not measurements. Source: LegalNexus manuscript · Table 10
StageOutput for “Electronic evidence admissibility WhatsApp messages Section 65B”
1 Query expansionconcepts: Section 65B (Evidence Act), electronic records, certificate requirement; intent: statute interpretation; domain: evidence law
2 Dynamic weightssemantic +0.05, graph +0.03 → semantic 0.40, graph 0.28, text 0.17, citation 0.12, GNN 0.03
3 Top casesAnvar P.V. v. P.K. Basheer (2014, SC, r = 0.08) · Arjun Panditrao v. Kailash (2020, SC, r = 0.09) · Shafhi Mohammad v. State of HP (2018) · State v. Mohd. Afzal (2003) · Tomaso Bruno v. State of UP (2015)
4 CourtroomProsecutor: certificate mandatory (Anvar); Defense: exception for contemporaneous records (Arjun Panditrao); Judge: certificate generally required, narrow exception subject to verification
5 CheckAnvar P.V. ranked first, consistent with established doctrine
6 Counterfactualremoving “certificate requirement” shifts focus to Section 3; Jaccard shift 0.72
A full pass through the system on one statutory-interpretation query. Source: manuscript Table 9, §7.7

Limits, and why it is a decision-support tool

  • One dominant hierarchy. India has a unitary court hierarchy; federal systems like the US need multi-hierarchy (product-manifold) embeddings.
  • Manual payoff tuning. Agent parameters (α = 1.0, λ = 1.5) are set by hand and do not adapt to the query.
  • Latency. The explainable debate takes 8.46 s of an 11.4 s query, limiting real-time use.
  • Scope. Evaluated only on Indian case law, in English, and dependent on the underlying LLMs; some components (the embedding model, the GNN link predictor) remain partly opaque.
  • Evidence. Practitioner feedback on the debate-style explanations is described as preliminary; formal user studies are future work.

The authors argue the framework extends beyond law to any hierarchical citation network (academic papers, patents) and that the agent-debate construction is a general way to build consistent knowledge graphs from conflicting sources. That is the thread to the rest of my work: structured knowledge, with disagreement represented explicitly instead of averaged away, as a basis for explainable decision support.

Read the preprint: SSRN 6678264.

All articles

Research and building cool stuff

© 2026, Animesh Mishra

GitHub|LinkedIn

New Delhi, India