Chijioke Ugwuanyi

AI Safety Researcher | Carnegie Mellon University

recent.jpg

Jinesis AI Lab

University of Toronto (Remote)

I am Chijioke Ugwuanyi, an AI safety researcher with an M.Sc. in Information Technology from Carnegie Mellon University. I currently work part-time at the Jinesis AI Lab at the University of Toronto, remotely.

My current research focuses on alignment faking in large language models, which is the phenomenon where models behave safely under observation but pursue different objectives when unmonitored, as well as on agentic misalignment. I designed AF-Arena, a multi-axis evaluation arena for alignment faking (accepted as an oral at the ICML 2026 Workshop on Agents in the Wild), and co-authored AM-Bench, a taxonomy and evaluation suite for agentic misalignment.

More broadly, I am interested in the empirical study of AI deception, evaluation, and the interpretability of AI systems.

If you’d like to connect or collaborate, reach me via email, LinkedIn, or Twitter.

selected publications

  1. ICML 2026
    AF-Arena: An Evaluation Arena for Multi-Axis Alignment Faking in Frontier Language Models
    Chijioke Ugwuanyi, Terry Jingchen Zhang, Bernhard Schölkopf, and 1 more author
    Workshop on Agents in the Wild (AIWILD), ICML 2026 (Oral). Under review at NeurIPS 2026 (Datasets & Benchmarks Track)., 2026
  2. Preprint
    AM-Bench: A Unified Taxonomy and Evaluation Suite for Agentic Misalignment
    Eric Zhang, Terry Jingchen Zhang, Chijioke Ugwuanyi, and 3 more authors
    Submitted to NeurIPS 2026, 2026
  3. LessWrong
    From 8B to Frontier: How System Prompts Control Whether AI Agents Blackmail, Leak, and Kill
    Chijioke Ugwuanyi
    2026