Weiliang Zhao

I'm a first-year PhD student at Columbia University, advised by Professor Junfeng Yang and Professor Zhou Yu. My research interests mainly include Alignment of LLMs, Continual Learning and Mechanistic Interpretability.

I completed my Master’s in Computer Science in the Department of Computer Science at Columbia University, advised by Prof. Junfeng Yang and Prof. Chengzhi Mao.

I hold a BSc in Mathematics from the University of Edinburgh, where I was advised by Prof. Burak Buke.

📮Email  /  🔗LinkedIn  /  🎓Google Scholar  /  📃CV

profile photo
Preprints
Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives
Zheyuan Liu, Weiliang Zhao, Xiangchi Yuan, Ningshan Ma, Yue Huang, Meng Jiang
Under Review
arXiv

Introduces KnownLieBench, a benchmark that tests whether LLM agents in customer-service settings make false claims when incentivized to deny user entitlements they demonstrably know are valid. Across eighteen models deception varies widely, and honesty-focused fine-tuning reduces it.

Publications
Jailbreaking Jailbreaks: A Proactive Defense for LLMs
Weiliang Zhao, JinJun Peng, Daniel Ben-Levi, Junfeng Yang, Chengzhi Mao,
EMNLP, 2026
arXiv

A proactive defense framework that injects strategically crafted spurious outputs to mislead attackers’ optimization loops, prematurely collapsing multi-turn jailbreak searches and dramatically reducing LLM vulnerability.

Mirage Probes: How Vision Models Fake Visual Understanding
Daniel Ben-Levi, Judah Goldfeder, Weiliang Zhao, Raz Lapid, Amit LeVi, Allen G. Roush, Ravid Shwartz-Ziv, Hod Lipson
Mechanistic Interpretability Workshop at ICML, 2026
arXiv

Vision-language models often answer image questions correctly without ever looking at the image. A contrastive probing framework shows this "mirage" behavior is linearly decodable from internal activations, separating genuine visual grounding from language-prior shortcuts.

Diversity Helps Jailbreak Large Language Models
Weiliang Zhao, Daniel Ben-Levi, Junfeng Yang, Chengzhi Mao,
NAACL, 2025, Oral
arXiv

A Generalised jailbreaking technique by encouraging higher levels of diversification and adjacent obfuscated prompting to evaluate the vulnerabilities of LLMs.

Learning to Rewrite: Generalized LLM-Generated Text Detection
Wei Hao, Ran Li , Weiliang Zhao, Junfeng Yang, Chengzhi Mao,
ACL, 2025
arXiv

We propose a method designed to enhance the detection of LLM-generated text by learning to rewrite more on LLM-generated inputs and less on human generated inputs.

Other Papers
CheatBench: Measuring Reward Gaming in AI Agents
Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika, Wenyu Zhang, Zheyuan Liu, Richard Ren, Jingxiang Meng, Yaoteng Tan, Weiliang Zhao, Addison Wu, Matei Anghel, Dan Hendrycks
arXiv

Introduces CheatBench, a benchmark that measures reward gaming — how AI agents pursue goals through dishonest shortcuts on difficult tasks — across multiple domains, to help measure and mitigate deceptive behavior as agents take on higher-stakes work.

Experience

Center for AI Safety, San Francisco, CA — Research Scientist Intern
May – August 2026

Acknowledgement

I would like to acknowledge the Thinker Research Grants support from Thinking Machine.

Travel

Beyond research, I love traveling and scuba diving 🤿 — here are my footprints around the world 🌍.


Feel free to steal this website's source code. Do not scrape the HTML from this page itself, as it includes analytics tags that you do not want on your own website — use the github code instead. Also, consider using Leonid Keselman's Jekyll fork of this page.