← All projects

Mechanistic Hallucination Tracing.

A mechanistic interpretability workflow for tracing where factual hallucinations change inside Llama-2-7B.

PythonPyTorchLlama-2-7BInterpretability
View source ↗

What it does

Mechanistic Hallucination Tracing is a pilot interpretability study that asks where a factual hallucination begins to change inside Llama-2-7B. It starts with paired prompts: one reproduces a known hallucination, while a minimally edited version moves the model toward a verified fact. The pair must pass a validation gate before it can enter the tracing pipeline.

After validation, the workflow aligns the first token where the hallucinated and corrected continuations diverge. It captures hidden states and logits, patches activations at the layer level, narrows the search to individual attention heads, and then ablates candidate heads to measure whether they weaken the hallucinated continuation.

The completed pilot contains two fully traced examples. Both showed their strongest repeated restoration signal in late layers around L29–L30, while their strongest individual heads differed. The project therefore reports evidence for an example-level pattern without claiming a universal hallucination circuit.

Core capabilities

What I built

  • Built a reproducible Python and PyTorch pipeline from pair validation through head ablation.
  • Fully traced two validated examples, with repeated restoration signals around layers L29–L30.
  • Documented example-specific head results and kept conclusions within the limits of the pilot.

Frameworks and tools

Language
Python
Modeling
PyTorch, Hugging Face Transformers, Llama-2-7B-Chat
Data and configuration
pandas, CSV, JSON, JSONL, PyYAML
Research workflow
CLI scripts, Jupyter notebooks, run-scoped artifacts
Evaluation methods
Forward-pass alignment, activation patching, head patching, head ablation

How it works

  1. 01

    Register candidate hallucination examples and check whether the labeled span maps cleanly to model tokens.

  2. 02

    Author a hallucinating prompt and a minimally corrected prompt, then validate the expected output flip.

  3. 03

    Find the first divergent output token and capture the corresponding logits and hidden states.

  4. 04

    Patch layers first, then rank attention heads within the strongest restoring layers.

  5. 05

    Ablate candidate heads and save manifests, summaries, JSONL results, and written findings for review.

Explore all projects →