Jingyu (Jack) Zhang

张景昱

Logo

jzhan237@jhu.edu

Google Scholar

LinkedIn

Twitter

CV

I am a PhD candidate in Computer Science at Johns Hopkins University, proudly advised by Daniel Khashabi and Benjamin Van Durme. My research is supported by the Amazon AI PhD Fellowship.

My research focuses on post-training and alignment for LLM agents, with an emphasis on ensuring safe and reliable model behavior. I study training and evaluation methods that enable robust generalization through reasoning over behavioral specifications and learning from interaction. Along these lines, my recent work includes reinforcement learning for multi-agent collaboration, enhancing safety controllability, and stress-testing instruction hierarchy as complexity scales. My overarching research goal is to build AI systems that can effectively collaborate with humans, reliably complete economically valuable tasks, and accelerate scientific discovery.

I was an intern at Apple AIML, where I worked on Apple Foundation Model alignment working with Joseph Yitan Cheng and Shruti Palaskar, a student researcher at Meta Superintelligence Labs collaborating with Hongyuan Zhan and Jason Weston, and a research intern and student researcher at Microsoft from 2024-2025 working with Ahmed Elgohary Ghoneim.

I completed my B.S. also from JHU with majors in Computer Science, Mathematics, Applied Mathematics, and minor in Economics. GO HOP! 💙🤍💙 During my undergrad, I collaborated with Mark Dredze at JHU CLSP, Yulia Tsvetkov and Tianxing He at the University of Washington, and Jim Glass at MIT CSAIL.

I’m always excited about collaborations. If you are interested in working together, please feel free to drop me an email: jzhan237[at]jhu.edu!

Selected Works

Image description Many-Tier Instruction Hierarchy in LLM Agents
Jingyu Zhang, Tianjian Li, William Jurayj, Hongyuan Zhan, Benjamin Van Durme, Daniel Khashabi.
Findings of EMNLP 2026

LLM agents receive instructions from many sources with varying levels of trust, but existing instruction hierarchy approaches assume only a few rigid role labels. We propose Many-Tier Instruction Hierarchy (ManyIH) for resolving conflicts among arbitrarily many privilege levels, and introduce ManyIH-Bench, the first benchmark for ManyIH spanning up to 12 levels of conflicting instructions.

Image description The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
Jingyu Zhang, Haozhu Wang, Eric Michael Smith, Sid Wang, Amr Sharaf, Mahesh Pasupuleti, Benjamin Van Durme, Daniel Khashabi, Jason Weston, Hongyuan Zhan.
ICLR 2026

We introduce WaltzRL, a multi-agent RL framework that frames LLM safety as a positive-sum game between a conversation agent and a feedback agent. We introduce a novel Dynamic Improvement Reward to jointly train two agents to collaborate, and give feedback adaptively at inference. WaltzRL improves safety & reduces overrefusals without degrading general capabilities.

Image description Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
Jingyu Zhang, Ahmed Elgohary, Ahmed Magooda, Daniel Khashabi, Benjamin Van Durme.
ICLR 2025

The current paradigm for safety alignment of large language models (LLMs) follows a one-size-fits-all approach and lacks flexibility in the face of varying social norms across cultures, and diverse user needs. We propose Controllable Safety Alignment, a framework that adapt models to diverse safety requirements without re-training.

Image description Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
Jingyu Zhang, Marc Marone, Tianjian Li, Benjamin Van Durme, Daniel Khashabi.
NAACL 2025 (oral)

To trust the fluent generations of large language models, humans must be able to verify their correctness against trusted external sources. We trivialize the verification process by developing models that quote verbatim statements from trusted sources in their pre-training data.

Image description SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation
Abe Bohan Hou*, Jingyu Zhang*, Tianxing He*, Yichen Wang, Yung-Sung Chuang, Hongwei Wang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, Yulia Tsvetkov.
NAACL 2024

Existing watermarking algorithms are vulnerable to paraphrase attacks because of their token-level design. To address this issue, we propose SemStamp, a robust sentence-level semantic watermarking algorithm based on locality-sensitive hashing (LSH), which partitions the semantic space of sentences.

All Publications

Taha Entesari, Jingyu Zhang, Daniel Khashabi, Mahyar Fazlyab. Minimally Invasive Steering of Language Models. NeurIPS 2026.

Kaiser Sun, Bernal Jiménez Gutiérrez, Hongjun Liu, Jingyu Zhang, Jie Gao, Mark Dredze, Daniel Khashabi. Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict. Findings of EMNLP 2026.

Tianjian Li, Jingyu Zhang, William Jurayj, Xi Wang, Chuanyang Jin, Mehrdad Farajtabar, Eric Nalisnick, Daniel Khashabi. Self-Compacting Language Model Agents. NeurIPS 2026.

Alexander K. Saeri, Jess Graham, Michael Noetel, Peter Slattery, …, Jingyu Zhang, … (188 authors). Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts. arXiv preprint.

Jingyu Zhang, Tianjian Li, William Jurayj, Hongyuan Zhan, Benjamin Van Durme, Daniel Khashabi. Many-Tier Instruction Hierarchy in LLM Agents. Findings of EMNLP 2026.

Guangyao Dou, Luis Brena, Akhil Deo, William Jurayj, Jingyu Zhang, Nils Holzenberger, Benjamin Van Durme. DeonticBench: A Benchmark for Reasoning over Rules. AAAI 2027 submission.

Hexuan Wang, Jingyu Zhang, Benjamin Van Durme, Daniel Khashabi. Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation. ACL ARR submission.

Pranjal Aggarwal, Marjan Ghazvininejad, Seungone Kim, Ilia Kulikov, Jack Lanchantin, Xian Li, Tianjian Li, Bo Liu, Graham Neubig, Anaelia Ovalle, Swarnadeep Saha, Sainbayar Sukhbaatar, Sean Welleck, Jason Weston, Chenxi Whitehouse, Adina Williams, Jing Xu, Ping Yu, Weizhe Yuan, Jingyu Zhang, Wenting Zhao. Reasoning over Mathematical Objects: On-Policy Reward Modeling and Test Time Aggregation. arXiv preprint.

Hoang Phan, Xianjun Yang, Kevin Yao, Jingyu Zhang, Shengjie Bi, Xiaocheng Tang, Madian Khabsa, Lijuan Liu, Deren Lei. Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models. Findings of ACL 2026.

Jingyu Zhang, Haozhu Wang, Eric Michael Smith, Sid Wang, Amr Sharaf, Mahesh Pasupuleti, Benjamin Van Durme, Daniel Khashabi, Jason Weston, Hongyuan Zhan. The Alignment Waltz: Jointly Training Agents to Collaborate for Safety. ICLR 2026.

Jingyu Zhang, Ahmed Elgohary, Xiawei Wang, A S M Iftekhar, Ahmed Magooda, Benjamin Van Durme, Daniel Khashabi, Kyle Jackson. Jailbreak Distillation: Renewable Safety Benchmarking. Findings of EMNLP 2025.

Jingyu Zhang, Jiacan Yu, Marc Marone, Benjamin Van Durme, Daniel Khashabi. Certified Mitigation of Worst-Case LLM Copyright Infringement. EMNLP 2025.

Abe Bohan Hou, Hongru Du, Yichen Wang, Jingyu Zhang, Zixiao Wang, Paul Pu Liang, Daniel Khashabi, Lauren Gardner, Tianxing He. Can A Society of Generative Agents Simulate Human Behavior and Inform Public Health Policy? A Case Study on Vaccine Hesitancy. COLM 2025

Jingyu Zhang, Ahmed Elgohary, Ahmed Magooda, Daniel Khashabi, Benjamin Van Durme. Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements. ICLR 2025.

Dongwei Jiang, Guoxuan Wang, Yining Lu, Andrew Wang, Jingyu Zhang, Chuyu Liu, Benjamin Van Durme, Daniel Khashabi. Rationalyst: Pre-training Process-Supervision for Improving Reasoning. ACL 2025.

Zhengping Jiang, Jingyu Zhang, Nathaniel Weir, Seth Ebner, Miriam Wanner, Kate Sanders, Daniel Khashabi, Anqi Liu, Benjamin Van Durme. Core: Robust Factual Precision Scoring with Informative Sub-Claim Identification. Findings of ACL 2025.

Dongwei Jiang, Jingyu Zhang, Orion Weller, Nathaniel Weir, Benjamin Van Durme, Daniel Khashabi. Self-(In)Correct: LLMs Struggle with Discriminating Self-Generated Responses. AAAI 2025.

Jingyu Zhang, Marc Marone, Tianjian Li, Benjamin Van Durme, Daniel Khashabi. Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data. NAACL 2025 (oral).

Kevin Xu, Yeganeh Kordi, Kate Sanders, Yizhong Wang, Adam Byerly, Jingyu Zhang, Benjamin Van Durme, Daniel Khashabi. TurkingBench: A Challenge Benchmark for Web Agents. NAACL 2025.

Weiting Tan, Jingyu Zhang, Lingfeng Shen, Daniel Khashabi, Philipp Koehn. DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation. NeurIPS 2024.

Abe Bohan Hou, Jingyu Zhang, Yichen Wang, Daniel Khashabi, Tianxing He. k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text. Findings of ACL 2024.

Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, Daniel Khashabi. The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts. Findings of ACL 2024.

Abe Bohan Hou*, Jingyu Zhang*, Tianxing He*, Yichen Wang, Yung-Sung Chuang, Hongwei Wang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, Yulia Tsvetkov. SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation. NAACL 2024.

Xiao Pu, Jingyu Zhang, Xiaochuang Han, Yulia Tsvetkov, Tianxing He. On the Zero-Shot Generalization of Machine-Generated Text Detectors. Findings of EMNLP 2023.

Tianxing He*, Jingyu Zhang*, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James Glass, Yulia Tsvetkov. On the Blind Spots of Model-Based Evaluation Metrics for Text Generation. ACL 2023 (oral).

Jingyu Zhang, Alexandra DeLucia, Chenyu Zhang, Mark Dredze. Geo-Seq2seq: Twitter User Geolocation on Noisy Data through Sequence to Sequence Learning. Findings of ACL 2023.

Jingyu Zhang, James Glass, Tianxing He. PCFG-based Natural Language Interface Improves Generalization for Controlled Text Generation. *SEM 2023. Preliminary version accepted at 2nd Workshop on Efficient Natural Language and Speech Processing (ENLSP), NeurIPS 2022. Best Paper Award.

Jingyu Zhang, Alexandra DeLucia, Mark Dredze. Changes in Tweet Geolocation over Time: A Study with Carmen 2.0. Proceedings of the 8th Workshop on Noisy User-generated Text (W-NUT), COLING 2022.

Abhinav Chinta*, Jingyu Zhang*, Alexandra DeLucia, Anna L. Buczak, Mark Dredze. Study of Manifestation of Civil Unrest on Twitter. Proceedings of the 7th Workshop on Noisy User-generated Text (W-NUT), EMNLP 2021.

*Equal Contribution

Teaching & Mentorship

Service

Misc

🏎️🏎️🏎️ In my free time, I enjoy go-karting and sim racing. I’m a big car enthusiast and love watching motorsports such as formula 1. My favorite driver is Zhou Guanyu, the first ever Chinese driver to compete in F1.