Hi, I am Sitao Cheng, a Ph.D. student at the University of Waterloo, fortunate to be advised by Prof. Victor Zhong. I am currently a Student Researcher at Samaya AI, working with Yuhao Zhang. Before that, I was a research scholar in the UCSB NLP Group with Prof. William Wang, and a research intern at Microsoft Research Asia. I received my Master’s degree from Nanjing University.

My research asks how reasoning in large language models can generalize efficiently: how a model can autonomously compose the skills it already has (parametric) with newly acquired ones (contextual), and recursively improve itself. I work across language agents, reinforcement learning, retrieval-augmented generation (RAG), and neural-symbolic reasoning. Three questions drive my current work:

  1. Language agents. What makes reasoning transfer to open, real-world environments rather than to curated benchmarks?
  2. Automatic reward modeling. Can reward functions be discovered rather than hand-designed, through differentiable evolutionary meta-rewards?
  3. RL and compositional generalization. Which training strategies produce models that compose skills they were never explicitly taught, and what does that reveal about how RL works? I study this concretely by building robust GUI agents.

I am always glad to talk about research, so please feel free to reach out. You can also read my CV.

Recent News

  • 2026-08 Joined Samaya AI as a Student Researcher.
  • 2026-04 Attended ICLR 2026 in Rio de Janeiro.
  • 2026-04 One paper accepted to ACL 2026.
  • 2025-12 Gave talks on compositional generalization at Peking University and Nanjing University.
  • 2025-11 Received the TD Layer 6 Graduate Scholarship in Data & AI for Fall 2025.
  • 2025-09 Began my Ph.D. at the University of Waterloo.
  • 2025-06 Two papers accepted to EMNLP 2025.
  • 2025-03 Three papers accepted to ACL 2025.
  • 2024-11 Attended SoCal NLP 2024 in San Diego.
  • 2024-11 Attended EMNLP 2024 in Miami.
  • 2024-09 One paper accepted to EMNLP 2024.
  • 2024-08 Attended and volunteered at ACL 2024 in Bangkok.
  • 2024-07 Joined the UC Santa Barbara NLP Group.
  • 2024-05 Two papers accepted to ACL 2024.
  • 2023-12 Attended EMNLP 2023 in Singapore.
  • 2023-10 One paper accepted to EMNLP 2023.
  • 2022-11 One paper accepted to AAAI 2023.

Preprints

  • From Atomic to Composite: Reinforcement Learning Enables Generalization in Complementary Reasoning
    Sitao Cheng, Xunjian Yin, Ruiwen Zhou, Yuxuan Li, Xinyi Wang, Liangming Pan, William Yang Wang, Victor Zhong
    paper code data

  • Differentiable Evolutionary Reinforcement Learning
    Sitao Cheng*, Tianle Li*, Xuhan Huang*, Xunjian Yin, Difan Zou
    paper code model

  • Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems
    Ruiwen Zhou*, Maojia Song*, Xiaobao Wu, Sitao Cheng, Xunjian Yin, Yuxi Xie, Zoey Hao, Wenyue Hua, Liangming Pan, Soujanya Poria, Min-Yen Kan
    paper code

Selected Publications

  • [KnowFM @ ACL’25 Oral] Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models
    Sitao Cheng, Liangming Pan, Xunjian Yin, Xinyi Wang, William Yang Wang
    paper code

  • [ACL’24 Findings] Call Me When Necessary: LLMs can Efficiently and Faithfully Reason over Structured Environments
    Sitao Cheng, Ziyuan Zhuang, Yong Xu, Fangkai Yang, Chaoyun Zhang, Xiaoting Qin, Xiang Huang, Ling Chen, Qingwei Lin, Dongmei Zhang, Saravan Rajmohan, Qi Zhang
    paper code

  • [ACL’26 Oral] LEDOM: Reverse Language Model
    Xunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu, Li Lin, Xinyi Wang, Liangming Pan, William Yang Wang, Xiaojun Wan
    paper model

  • [ACL’24 Oral] QueryAgent: A Reliable and Efficient Reasoning Framework with Environmental Feedback based Self-Correction
    Xiang Huang*, Sitao Cheng*, Shanshan Huang, Jiayu Shen, Yong Xu, Chaoyun Zhang, Yuzhong Qu
    paper code

  • [ACL’25] Disentangling Memory and Reasoning Ability in Large Language Models
    Mingyu Jin, Weidi Luo, Sitao Cheng, Xinyi Wang, Wenyue Hua, Ruixiang Tang, William Yang Wang, Yongfeng Zhang
    paper code

  • [EMNLP’24] EfficientRAG: Efficient Retriever for Multi-Hop Question Answering
    Ziyuan Zhuang, Zhiyang Zhang, Sitao Cheng, Fangkai Yang, Jia Liu, Shujian Huang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang
    paper code

  • [ACL’25] RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
    Ruiwen Zhou, Wenyue Hua, Liangming Pan, Sitao Cheng, Xiaobao Wu, En Yu, William Yang Wang
    paper code

Teaching

  • CS240 Data Structures and Data Management, Teaching Assistant, Winter & Spring 2026, UWaterloo
  • CS115 Introduction to Computer Science 1, Teaching Assistant, Spring 2025, UWaterloo

Services

  • Reviewer: ARR, ICLR, ICML, NeurIPS
  • Volunteer: ACL 2024