Pinned Loading
-
sycophancy-construct-validity
sycophancy-construct-validity PublicConstruct validity of sycophancy interventions in Llama-3.1-8B — reading a concept is not controlling it
Python
-
safety-concept-vectors
safety-concept-vectors PublicSafety-concept directions in Qwen2.5-7B: contrastive extraction and linear probes
Python 1
-
eval-awareness-detection
eval-awareness-detection PublicPrototype: mechanistic eval-awareness detection in open-weight LLMs (March 2026)
Python 1
-
does-quantization-kill-interpretability
does-quantization-kill-interpretability PublicDoes Quantization Kill Interpretability? Scaling study across 5 models (124M-2.8B): RTN destroys induction heads in small models, GPTQ preserves them at all scales.
Python 1
-
If the problem persists, check the GitHub status page or contact support.