LegalNexus: finding precedent in a hierarchy
Why law needs geometry
TL;DR. Legal authority is a tree (apex courts at the top, a flood of lower-court judgments below) and legal truth is a dispute (precedents get distinguished, even overruled). LegalNexus embeds the citation network in hyperbolic space so that radius means authority, uses three debating agents to clean contradictions out of the graph, and ends with a simulated courtroom that explains why a case was retrieved. It is a preprint under review at Engineering Applications of Artificial Intelligence.
A flat embedding treats a Supreme Court judgment and a tribunal order as two points that happen to be near each other in meaning. A lawyer knows one binds the other. Existing retrieval systems, from BM25 to graph networks such as CaseGNN, largely ignore that structure, or model the citation graph as static even though precedents are constantly distinguished and overruled. LegalNexus is an attempt to put both properties into one system.
Hyperbolic space: a place for trees
A tree with a branching factor above one grows exponentially with depth, but a Euclidean plane grows only polynomially, so embedding a deep hierarchy into flat space forces distortion. Hyperbolic space expands exponentially, which is why it has become the geometry of choice for taxonomies. The manuscript checks that the citation graph really is tree-like before relying on that, using Gromov’s δ.
The model is a hyperbolic graph convolutional network: 768-dimensional Gemini embeddings of each case are mapped into the Poincaré ball, passed through two hyperbolic layers, and trained to predict citations with a Fermi-Dirac decoder and a contrastive loss that pulls cited cases together and pushes negatives apart. Nothing in training mentions court level.
# log map to the tangent space, then Möbius aggregation (Poincaré ball, curvature c)
log0_c(x) = (1/√c) · artanh(√c‖x‖) · x/‖x‖
h_v⁽ˡ⁺¹⁾ = σ⊗c ( Aggregate{ W⁽ˡ⁾ ⊗c h_u⁽ˡ⁾ : u ∈ N(v) } )
# Fermi-Dirac decoder: probability that a citation edge exists
P(e_uv = 1 | h_u, h_v) = 1 / ( exp((d_c(h_u, h_v) − r) / t) + 1 )
# contrastive loss: pull cited pairs together, push sampled negatives past margin m
L = Σ_{(u,v)∈E⁺} d_c² + Σ_{(u,v)∈E⁻} max(0, m − d_c²)| Setting | Value |
|---|---|
| Input | 768-d Gemini text embeddings |
| Network | HGCN: layer 1 ℝ⁷⁶⁸ → 𝔻¹²⁸ (log map + linear), layer 2 𝔻¹²⁸ → 𝔻⁶⁴ (hyperbolic aggregation) |
| Library and optimiser | geoopt · Riemannian Adam · learning rate 0.001 · weight decay 0.0005 · dropout 0.5 |
| Loss | contrastive, margin m = 2.0, 5 sampled negative edges per positive |
| Training | 5 epochs, early stopping (patience 20), batch of 256 edges |
| Cost | about 12 minutes on one NVIDIA A100 |
A swarm that argues the graph into consistency
Geometry alone cannot tell a valid precedent from an overruled one. The manuscript therefore adds three agents that debate over the citation graph in rounds, using hyperbolic distance as a trust prior.
| Agent | Role | Implementation (as reported) | Payoff |
|---|---|---|---|
| Linker | proposer | regex patterns for citation formats (e.g. “AIR 2020 SC 1234”) plus LLM validation (Gemma-2-2B, temperature 0.3) | u_L = Φ |
| Interpreter | analyst | BERT-based legal classifier trained on 10,000 annotated pairs; keyword heuristics and LLM for ambiguous cases | u_I = Φ + C_I |
| Conflict | critic | depth-first search for citation cycles and contradictory edge types; emits REMOVE, DOWNGRADE or RECLASSIFY | u_C = Φ |
Across rounds the graph stabilises: the paper reports resolving 94% of detected citation cycles and contradictions, with a mean of roughly three to four rounds (the text gives 3.2 in one place and 4.2 in another, with a round limit of three to five). The authors are careful to say this is an empirical stable state and not a formal proof of equilibrium.
Retrieval that explains itself
The retrieval layer combines five scorers whose weights adapt to the type of query, adds a temporal score so landmark cases are not penalised for age, extracts Toulmin-style arguments, and finally stages a short courtroom debate over the top results.
| Retriever | Score |
|---|---|
| Semantic | cosine between query and case embeddings in hyperbolic space |
| Knowledge-graph traversal | PageRank(c) × path similarity; follows FOLLOW edges and avoids OVERRULE paths (Neo4j / Cypher) |
| Text pattern | BM25-style score on the case text |
| Citation network | authority-weighted citation importance |
| GNN link prediction | link probability from the trained HGCN |
Time matters in law: an old case can be semantically perfect and legally obsolete, yet a 1950s landmark that is still cited today should not be punished for age. The temporal score has a base decay and a “resurrection” term for recent citations; on 500 test queries it reduced obsolete recommendations by 34% while keeping actively cited historical precedents.
T(c) = 1 / log(age(c) + 2) · (1 + R(c)) # base decay × resurrection
R(c) = Σ_{t ∈ Cites(c)} 1 / (age(t) + 1) # recent citations revive an old case| Toulmin part | Extracted from Anvar P.V. v. P.K. Basheer (2014) |
|---|---|
| Claim | Electronic records without a Section 65B certificate are inadmissible. |
| Data | The appellant produced CDs without the certificate required by Section 65B(4). |
| Warrant | Section 65B is a special provision that overrides the general rules of secondary evidence. |
| Backing | The Evidence Act’s purpose is to ensure authenticity; electronic records are susceptible to tampering. |
The courtroom
Three LLM agents turn a ranked list into an explanation. A Prosecutor argues for the position using the top cases, a Defense agent finds distinguishing precedents and mitigating factors, and a Judge weighs both using case authority (a hyperbolic radius below 0.10 carries more weight) and citation support, then writes a balanced ruling. Each side is limited to three sentences over a single round. A counterfactual engine then perturbs the query’s facts to find the pivots on which the result turns.
Results
| Item | Value |
|---|---|
| Cases | 49,633 Indian court cases, 1950–2024 (Supreme Court 19.8%, High Courts 50.0%, lower courts 30.2%) |
| Citations | 127,891 (2.6 per case on average); 8,432 cases contain explicit overruling or distinguishing statements |
| Split | 70 / 15 / 15 (34,743 / 7,445 / 7,445) with no temporal leakage: training cases predate test cases |
| Queries | 2,165 gold-standard queries, 3–8 relevant cases each (mean 4.7), relevance verified by ensemble LLM voting |
| Query mix | precedent search 35% · fact pattern 30% · statute interpretation 20% · procedural 15% |
| Source | Indian Kanoon, via the NyayAnumana dataset (available from its authors on request) |
| Method | P@5 | R@10 | MAP | NDCG |
|---|---|---|---|---|
| TF-IDF | 0.62 | 0.58 | 0.60 | 0.64 |
| BM25 | 0.68 | 0.64 | 0.66 | 0.70 |
| Legal-BERT | 0.82 | 0.79 | 0.80 | 0.83 |
| Longformer-Legal | 0.84 | 0.81 | 0.82 | 0.85 |
| Hier-SPCNet | 0.79 | 0.74 | 0.76 | 0.80 |
| CaseGNN | 0.85 | 0.83 | 0.84 | 0.87 |
| LegalNexus | 0.88 | 0.86 | 0.87 | 0.89 |
What each component contributes
The authors summarise the ablation as geometry giving the “map” and agents the “compass”: the hyperbolic network separates court levels without the crowding a Euclidean network shows, and the agents clean the noisy graph the geometry depends on. Removing the Conflict agent also increases contradictory outputs by 23%.
Failure and cost
| Stage | Output for “Electronic evidence admissibility WhatsApp messages Section 65B” |
|---|---|
| 1 Query expansion | concepts: Section 65B (Evidence Act), electronic records, certificate requirement; intent: statute interpretation; domain: evidence law |
| 2 Dynamic weights | semantic +0.05, graph +0.03 → semantic 0.40, graph 0.28, text 0.17, citation 0.12, GNN 0.03 |
| 3 Top cases | Anvar P.V. v. P.K. Basheer (2014, SC, r = 0.08) · Arjun Panditrao v. Kailash (2020, SC, r = 0.09) · Shafhi Mohammad v. State of HP (2018) · State v. Mohd. Afzal (2003) · Tomaso Bruno v. State of UP (2015) |
| 4 Courtroom | Prosecutor: certificate mandatory (Anvar); Defense: exception for contemporaneous records (Arjun Panditrao); Judge: certificate generally required, narrow exception subject to verification |
| 5 Check | Anvar P.V. ranked first, consistent with established doctrine |
| 6 Counterfactual | removing “certificate requirement” shifts focus to Section 3; Jaccard shift 0.72 |
Limits, and why it is a decision-support tool
- One dominant hierarchy. India has a unitary court hierarchy; federal systems like the US need multi-hierarchy (product-manifold) embeddings.
- Manual payoff tuning. Agent parameters (α = 1.0, λ = 1.5) are set by hand and do not adapt to the query.
- Latency. The explainable debate takes 8.46 s of an 11.4 s query, limiting real-time use.
- Scope. Evaluated only on Indian case law, in English, and dependent on the underlying LLMs; some components (the embedding model, the GNN link predictor) remain partly opaque.
- Evidence. Practitioner feedback on the debate-style explanations is described as preliminary; formal user studies are future work.
The authors argue the framework extends beyond law to any hierarchical citation network (academic papers, patents) and that the agent-debate construction is a general way to build consistent knowledge graphs from conflicting sources. That is the thread to the rest of my work: structured knowledge, with disagreement represented explicitly instead of averaged away, as a basis for explainable decision support.
Read the preprint: SSRN 6678264.


