Sabrina Sadiekh

Researcher at AikyamLab · R&D Research Lead at HiveTrace · Lecturer at HSE University

prof_pic.png

sadsobr7@gmail.com

Whoami?

Hi! I am an AI researcher with a background in mathematics, working on the internal structure of language models — how representations form, why they break, and what that reveals about their geometry and robustness.

My research lies at the intersection of Explainable AI and Mechanistic Interpretability and Mathematics; selected papers and my open-source projects are below.

Open-source

Homines dum docent discunt, Seneca the Younger

I have been trying to understand interpretability since 2023. I think the best way to understand something is to try to explain it. I like to show the beauty of the field to others, so a significant part of my work is publicly available:

AI Interpretability School

I co-founded the AI Interpretability School together with Elena Ericheva — a course, that teaches classic XAI and mechanistic interpretability as a single subject.

Writing

I write about XAI and interpretability on Telegram, ru in Russian, and recently started writing in English on Substack and Medium.

Collaborations and work

I collaborate with an international group at AikyamLab on mechanistic interpretability and AI safety. I teach M.Sc. courses at HSE University. I also lead R&D Lab at HiveTrace, where I run a proposal-driven lab spanning three tracks — XAI in Security, AI Security, and AI Safety.

AI models are fascinating entities in our world, and if they get out of control, *I want to explain how and why*. This is my main motivation. Feel free to contact me =)

selected publications

  1. AAAI
    Polarity-Aware Probing for Quantifying Latent Alignment in Language Models
    Sabrina Sadiekh, Elena Ericheva, and Chirag Agarwal
    In Proceedings of the AAAI Conference on Artificial Intelligence, 2026
  2. ACL
    Towards Understanding the Robustness of Sparse Autoencoders
    Ahson Saiyed, Sabrina Sadiekh, and Chirag Agarwal
    In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, 2026