publications

publications by categories in reversed chronological order. generated by jekyll-scholar.

2026

  1. AAAI
    Polarity-Aware Probing for Quantifying Latent Alignment in Language Models
    Sabrina Sadiekh, Elena Ericheva, and Chirag Agarwal
    In Proceedings of the AAAI Conference on Artificial Intelligence, 2026
  2. ACL
    Towards Understanding the Robustness of Sparse Autoencoders
    Ahson Saiyed, Sabrina Sadiekh, and Chirag Agarwal
    In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, 2026
  3. arXiv
    GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy
    B. Minko, Sabrina Sadiekh, and Evgeniy Kokuykin
    arXiv preprint, 2026
  4. arXiv
    Cross-Lingual Jailbreak Detection via Semantic Codebooks
    Shirin Alanova, B. Minko, Sabrina Sadiekh, and 1 more author
    arXiv preprint, 2026