Project Sherlock

Artificial Intelligence · Deep Learning

Interpretability & Mechanistic Analysis

A topic within Deep Learning, itself one of 14 topics in that field and part of Artificial Intelligence.

Reading on Interpretability & Mechanistic Analysis

2

2 works

Essay2017

Feature Visualization

Olah, Mordvintsev & Schubert

Argues the clearest way to see what a neuron in a trained network has learned is to synthesise, by gradient ascent, the input image that most excites it — and works through the practical tricks needed to make that optimisation produce interpretable images rather than noise.

≈6,000 wordslink checked 17 Sept 2026
Paper2022

Toy Models of Superposition

Elhage et al.

Uses small, fully-understood toy networks to show why individual neurons often represent no single human-interpretable concept: networks pack more features than they have dimensions by representing them as overlapping directions, a phenomenon (superposition) that limits interpretability at the neuron level.

≈15,000 wordslink checked 17 Sept 2026

Other topics in Deep Learning