The Hidden Geometry of Safety: Mapping Latent Safety Representations and Real-Time Activation Steering in Large Language Models
DOI:
https://doi.org/10.66021/Keywords:
Activation Steering, AI Safety, Large Language Models, Latent Representations, Mechanistic Interpretability, Real-Time Inference, Representation Engineering, Safety Alignment, Transformer Models.Abstract
While Large Language Models (LLMs) have already found their way into high-stakes applications, concerns have been raised about the potential for the generation of unsafe or harmful content, even with emerging alignment improvements including Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI. This study examines if safety-relevant information is stored in the internal models and can be manipulated when making inferences without retraining. Hidden-state activations were extracted from 20 benign prompts using TinyLlama-1.1B-Chat, and then combined with a latent safety vector that was placed through forward-hook interventions in the transformer layers, with different depths and steering strengths. The methodology was tested against 20 cybersecurity related prompts ranging from malware, phishing, credential theft, ransomware to offensive cyber operations. Experimental results show that the effectiveness of activation steering is highly dependent on both the magnitude of the intervention and its position in the layers; the effectiveness is higher for early and middle layers than for the layers at the end. The generation of harmful responses was strongly reduced with moderate interventions but not strong interventions, suggesting a trade-off between the safety of the intervention and the coherence of the model. The results indicate that latent activation is a mechanism to encode behaviorally relevant safety information and call for understanding the feasibility of manipulating representations at inference time as a lightweight alignment mechanism. But the observed behavior also suggests that the extracted vector is a direction related to safety, but not necessarily a safety circuit that has been causally verified, suggesting significant challenges and opportunities in future research about discovering safety circuits and real-time alignment.