Sabrina Sadiekh
Researcher at AikyamLab · R&D Research Lead at HiveTrace · Lecturer at HSE University
sadsobr7@gmail.com
Whoami?
Hi! I am an AI researcher with a background in mathematics, working on the internal structure of language models — how representations form, why they break, and what that reveals about their geometry and robustness.
My research lies at the intersection of Explainable AI and Mechanistic Interpretability and Mathematics; selected papers and my open-source projects are below.
Open-source
Homines dum docent discunt, Seneca the Younger
I have been trying to understand interpretability since 2023. I think the best way to understand something is to try to explain it. I like to show the beauty of the field to others, so a significant part of my work is publicly available:
- a bank of open XAI tutorials — hands-on and code-first, maintained since 2023;
- the first Russian-language Explainable AI course — 600+ students;
- an interactive table of XAI frameworks — a navigation map of the tooling landscape, kept up to date since 2023.
AI Interpretability School
I co-founded the AI Interpretability School together with Elena Ericheva — a course, that teaches classic XAI and mechanistic interpretability as a single subject.
Writing
I write about XAI and interpretability on Telegram, ru in Russian, and recently started writing in English on Substack and Medium.
Collaborations and work
I collaborate with an international group at AikyamLab on mechanistic interpretability and AI safety. I teach M.Sc. courses at HSE University. I also lead R&D Lab at HiveTrace, where I run a proposal-driven lab spanning three tracks — XAI in Security, AI Security, and AI Safety.
AI models are fascinating entities in our world, and if they get out of control, *I want to explain how and why*. This is my main motivation. Feel free to contact me =)
selected publications
- AAAIPolarity-Aware Probing for Quantifying Latent Alignment in Language ModelsIn Proceedings of the AAAI Conference on Artificial Intelligence, 2026
- ACLTowards Understanding the Robustness of Sparse AutoencodersIn Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, 2026