Back to Lab
VisualizerStable

Attention-Head Visualizer

Interactive real-time visualization of multi-head self-attention weights and query-key matrix dot products.

Limitation: Simulates 8-head transformer layers with 128-dim embeddings; full 4096-dim model heads run on GPU.

Model Controls

Attention Head:Head 1 of 8
Softmax Temp ($\tau$):0.7
CURRENT LAYERLayer 12 (Self-Attn)
HEAD INDEXHead 1
AVG ENTROPY1.42 nats
MATRIX SIZE7 × 7