Official code for Steering Large Language Models using Conceptors, presented at the NeurIPS 2024 MINT Workshop.
-
Updated
Mar 13, 2025 - Jupyter Notebook
Official code for Steering Large Language Models using Conceptors, presented at the NeurIPS 2024 MINT Workshop.
Official implementation of "CSKS: Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models" (EMNLP 2025)
Policy-enforcing inference for open-weight LLMs. Self-hosted and OpenAI-compatible: enforces your policy at the residual-stream level, so it survives obfuscation, role-play and jailbreak wrappers. Byte-for-byte no-op on in-policy requests.
Lobopy is a lightweight PyTorch/HuggingFace library for analysing, steering/abliteration of causal language models.
Guided undergraduate notebooks on diffusion, flow matching, deterministic sampling, and post-hoc generative-model steering, with a validated CIFAR-10 EDM bridge.
Goodfire is an AI interpretability research lab building tools to understand, debug, and intentionally design neural networks by surfacing the internal features (via sparse autoencoders) that drive model behavior.
Envariant is building the control layer for foundation models — an AI interpretability SDK that lets teams inspect, steer, and control model behavior. The SDK exposes a compact set of primitives: detect and causally trace behaviors like hallucinations or invariant violations, reason inductively and steer outputs programmatically, extract…
Add a description, image, and links to the model-steering topic page so that developers can more easily learn about it.
To associate your repository with the model-steering topic, visit your repo's landing page and select "manage topics."