Interpretability
Discovering Cross-Language Reasoning Invariance in LLMs
Shared geometry across languages may reveal common reasoning structure, but representational similarity alone does not establish causal interchangeability.
- Venue
- Workshop paperICML 2026 Workshop on Mechanistic Interpretability
- Role
- First author · method · analysis pipeline
- Themes
- Interpretability
Pipeline: parallel reasoning prompts in several languages produce residual-stream activations at a fixed layer. A top-K sparse autoencoder with an InfoNCE contrastive term aligns feature activations for the same problem across languages. From the shared dictionary, geometry metrics measure alignment and causal feature patching swaps shared feature values, measuring output disruption via KL divergence. The joint reading classifies model-layer observations as convergent, saturated, or low-sharing.
- GI-SAE
- geometry-invariant sparse autoencoder
- top-K SAE with an InfoNCE contrastive term aligning features across languages
- Causal
- feature patching
- shared-feature value swaps with output disruption measured by KL divergence
- 3
- sharing profiles
- convergent, saturated, and low-sharing model-layer observations
01
The problem
If a language model reasons the same way in two languages, its internal representations of that reasoning should be shared. Most evidence for sharing is geometric: activations for parallel inputs line up after a rotation or projection. But geometric alignment is a weak test. Two sets of features can be similar in shape and still do different work when the model actually reasons.
This project asks a sharper question: are the features that carry reasoning in one language functionally interchangeable with those in another? Answering it requires a feature basis that is not an artefact of one language's geometry, and an intervention that tests function rather than similarity.
Held constant across languages: the reasoning task, the prompt structure, and the model layer under analysis. Varied: the language of the prompt and, in interventions, the source language of the features injected.
02
System and method
The GI-SAE analysis pipeline. Parallel reasoning prompts in several languages produce residual-stream activations. A top-K sparse autoencoder with an InfoNCE contrastive term pulls feature activations for the same problem together across languages. Geometry metrics quantify alignment. Causal feature patching swaps shared feature values and measures output disruption via KL divergence. The joint reading classifies model-layer observations into convergent, saturated, and low-sharing profiles.
03
Experimental design
- Research question
- Do LLMs reuse common reasoning structure across languages, and is it causally interchangeable?
- Method
- Top-K sparse autoencoders with an InfoNCE contrastive alignment term, trained over parallel multilingual reasoning activations
- Language and model setup
- Parallel reasoning prompts across multiple languages on open-weight models · the exact language and model list is specified in the paper
- Geometry metrics
- Alignment and similarity measures between per-language feature geometries
- Causal validation
- Shared-feature value swaps with output disruption measured by KL divergence
- Profiles
- Model-layer observations classified as convergent, saturated, or low-sharing by baseline shared fraction
04
Main findings
- 01
Geometric alignment and causal interchangeability are different questions
Representations can align closely in geometry while shared-feature swaps still disrupt outputs, and vice versa. The two measurements disagree often enough that neither can stand in for the other.
Why it mattersClaims of 'shared multilingual reasoning' based on similarity alone should be read as hypotheses, not conclusions.
- 02
Sharing is organised into profiles rather than a single degree
Model-layer observations exhibit convergent, saturated, and low-sharing profiles that differ in both alignment and causal transferability.
Why it mattersA model can be multilingual in some reasoning components and language-specific in others. Interpretability tooling should report which is which.
- 03
Contrastive training yields a usable cross-language dictionary
The contrastive alignment term produces a shared dictionary whose features can be compared and patched across languages without hand-aligning bases.
Why it mattersGI-SAE makes causal cross-language probing tractable as a routine analysis step.
05
What I built
The engineering behind the result. Every number above was produced by systems I designed and implemented.
Activation capture
Hooks for collecting residual-stream activations on parallel multilingual reasoning prompts.
GI-SAE training
Top-K SAE training with the InfoNCE contrastive alignment term on GPU.
Geometry metrics
Alignment and similarity computations between per-language feature geometries.
Causal patching harness
Shared-feature value swaps with KL-divergence evaluation of output disruption.
Analysis pipeline
Regime assignment and figure generation from intervention results.
06
Limitations
08
Related research
- EvaluationEvaluation of Multi-Turn Consistency in LLM AgentsReliability is temporal: models fail at different stages and produce systematically different narratives before doing so.Open project
- ArchitectureContext, Reasoning, and Hierarchy in Compound LLM AgentsWhat an agent sees can matter more than how long it deliberates.Open project
Contact
Building reliable AI systems requires both research and engineering.
I’m interested in research engineering, applied research, agent infrastructure, evaluation, reliability, interpretability, and research-to-production work.