The Hidden Geometry of Safety: Mapping Latent Safety Representations and Real-Time Activation Steering in Large Language Models

Authors

  • Shoaib Ali Faculty of Science and Technology, Iqra University Karachi, Pakistan, Karachi, Sindh, Pakistan Author
  • Muhammad Irfan Anis Faculty of Science and Technology, Iqra University Karachi, Pakistan, Karachi, Sindh, Pakistan Author
  • Sajid Ali Faculty of Science and Technology, Iqra University Karachi, Pakistan, Karachi, Sindh, Pakistan Author
  • Muhammad Shafiq Department of Computer Science, Nazeer Hussain University Karachi, Sindh, Pakistan Author

DOI:

https://doi.org/10.66021/

Keywords:

Activation Steering, AI Safety, Large Language Models, Latent Representations, Mechanistic Interpretability, Real-Time Inference, Representation Engineering, Safety Alignment, Transformer Models.

Abstract

While Large Language Models (LLMs) have already found their way into high-stakes applications, concerns have been raised about the potential for the generation of unsafe or harmful content, even with emerging alignment improvements including Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI. This study examines if safety-relevant information is stored in the internal models and can be manipulated when making inferences without retraining. Hidden-state activations were extracted from 20 benign prompts using TinyLlama-1.1B-Chat, and then combined with a latent safety vector that was placed through forward-hook interventions in the transformer layers, with different depths and steering strengths. The methodology was tested against 20 cybersecurity related prompts ranging from malware, phishing, credential theft, ransomware to offensive cyber operations. Experimental results show that the effectiveness of activation steering is highly dependent on both the magnitude of the intervention and its position in the layers; the effectiveness is higher for early and middle layers than for the layers at the end. The generation of harmful responses was strongly reduced with moderate interventions but not strong interventions, suggesting a trade-off between the safety of the intervention and the coherence of the model. The results indicate that latent activation is a mechanism to encode behaviorally relevant safety information and call for understanding the feasibility of manipulating representations at inference time as a lightweight alignment mechanism. But the observed behavior also suggests that the extracted vector is a direction related to safety, but not necessarily a safety circuit that has been causally verified, suggesting significant challenges and opportunities in future research about discovering safety circuits and real-time alignment.

Downloads

Download data is not yet available.

Author Biographies

  • Muhammad Irfan Anis, Faculty of Science and Technology, Iqra University Karachi, Pakistan, Karachi, Sindh, Pakistan

     

     

     

     

  • Sajid Ali, Faculty of Science and Technology, Iqra University Karachi, Pakistan, Karachi, Sindh, Pakistan

     

     

     

     

  • Muhammad Shafiq, Department of Computer Science, Nazeer Hussain University Karachi, Sindh, Pakistan

     

     

Downloads

Published

2026-06-22

How to Cite

The Hidden Geometry of Safety: Mapping Latent Safety Representations and Real-Time Activation Steering in Large Language Models. (2026). Annual Methodological Archive Research Review, 4(6), 202-225. https://doi.org/10.66021/

Similar Articles

21-30 of 1720

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)