Hi, I am Sitao Cheng, a Ph.D. student at the University of Waterloo, fortunate to be advised by Prof. Victor Zhong. I am currently a Student Researcher at Samaya AI, working with Yuhao Zhang. Before that, I was a research scholar in the UCSB NLP Group with Prof. William Wang, and a research intern at Microsoft Research Asia. I received my Master’s degree from Nanjing University.
My research asks how reasoning in large language models can generalize efficiently: how a model can autonomously compose the skills it already has (parametric) with newly acquired ones (contextual), and recursively improve itself. I work across language agents, reinforcement learning, retrieval-augmented generation (RAG), and neural-symbolic reasoning. Three questions drive my current work:
- Language agents. What makes reasoning transfer to open, real-world environments rather than to curated benchmarks?
- Automatic reward modeling. Can reward functions be discovered rather than hand-designed, through differentiable evolutionary meta-rewards?
- RL and compositional generalization. Which training strategies produce models that compose skills they were never explicitly taught, and what does that reveal about how RL works? I study this concretely by building robust GUI agents.
I am always glad to talk about research, so please feel free to reach out. You can also read my CV.
Recent News
- 2026-08 Joined Samaya AI as a Student Researcher.
- 2026-04 Attended ICLR 2026 in Rio de Janeiro.
- 2026-04 One paper accepted to ACL 2026.
- 2025-12 Gave talks on compositional generalization at Peking University and Nanjing University.
- 2025-11 Received the TD Layer 6 Graduate Scholarship in Data & AI for Fall 2025.
- 2025-09 Began my Ph.D. at the University of Waterloo.
- 2025-06 Two papers accepted to EMNLP 2025.
- 2025-03 Three papers accepted to ACL 2025.
- 2024-11 Attended SoCal NLP 2024 in San Diego.
- 2024-11 Attended EMNLP 2024 in Miami.
- 2024-09 One paper accepted to EMNLP 2024.
- 2024-08 Attended and volunteered at ACL 2024 in Bangkok.
- 2024-07 Joined the UC Santa Barbara NLP Group.
- 2024-05 Two papers accepted to ACL 2024.
- 2023-12 Attended EMNLP 2023 in Singapore.
- 2023-10 One paper accepted to EMNLP 2023.
- 2022-11 One paper accepted to AAAI 2023.
Preprints
From Atomic to Composite: Reinforcement Learning Enables Generalization in Complementary Reasoning
paper code dataDifferentiable Evolutionary Reinforcement Learning
paper code modelEpistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems
paper code
Selected Publications
[KnowFM @ ACL’25 Oral] Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models
paper code[ACL’24 Findings] Call Me When Necessary: LLMs can Efficiently and Faithfully Reason over Structured Environments
paper code[ACL’24 Oral] QueryAgent: A Reliable and Efficient Reasoning Framework with Environmental Feedback based Self-Correction
paper code[ACL’25] Disentangling Memory and Reasoning Ability in Large Language Models
paper code[EMNLP’24] EfficientRAG: Efficient Retriever for Multi-Hop Question Answering
paper code[ACL’25] RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
paper code
Teaching
- CS240 Data Structures and Data Management, Teaching Assistant, Winter & Spring 2026, UWaterloo
- CS115 Introduction to Computer Science 1, Teaching Assistant, Spring 2025, UWaterloo
Services
- Reviewer: ARR, ICLR, ICML, NeurIPS
- Volunteer: ACL 2024
