A curated list of resources for activation engineering
-
Updated
Oct 2, 2025
A curated list of resources for activation engineering
Official code for Steering Large Language Models using Conceptors, presented at the NeurIPS 2024 MINT Workshop.
Runtime control of LLM agent behaviors through activation steering vectors. More calibrated than prompting.
🔓 Ablate — directional ablation (abliteration) toolkit for open-source LLMs. Automatic censorship/refusal removal via residual-stream direction ablation, with KL-guided search, an LLM-judge harness, and one-call push to the Hub. pip install ablate-llm
Iterative Sparse Matrix Steering: Closed-Form Subspace Alignment for Multi-Layer LLM Control (No SGD required).
How meaning moves through a transformer - found, traced, and tested across four model scales.
A closed-loop control system for Large Language Models that steers internal activation states in real-time to prevent mode collapse and toxicity
Add a description, image, and links to the activation-engineering topic page so that developers can more easily learn about it.
To associate your repository with the activation-engineering topic, visit your repo's landing page and select "manage topics."