identifying and ablating the activation-space directions that enable jailbreaks in large language models