Hi! I am Sizhe Tang, a second-year Ph.D. candidate at The George Washington University, advised by Prof. Tian Lan.
My research focuses on reinforcement learning and self-evolving agents, with an emphasis on tree-search and planning for multi-agent decision-making (the Zero series), computer-use agents, and multi-turn agent policy optimization.
You can find my publications on Google Scholar.
🔥 News
- 2026.06: 🎉 Agent Alpha accepted to COLM 2026 (score: 778).
- 2026.05: 🎉 NonZero accepted to ICML 2026 (Spotlight).
- 2026.05: 🎉 HFPS accepted to ICML 2026.
- 2026.04: 🎉 I will join Amazon AWS AI as an Applied Scientist Intern (Summer 2026, Santa Clara, CA).
- 2026.03: 🎉 T-STAR accepted to ACL 2026 Findings.
- 2025.09: 🎉 MALinZero accepted to NeurIPS 2025.
📝 Selected Publications
RL Tree search
Sizhe Tang, Zuyuan Zhang, Mahdi Imani, Tian Lan. NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search. ICML 2026 (Spotlight).
An interaction-guided exploration rule that scales Monte Carlo Tree Search to multi-agent planning.
RL Tree search
Sizhe Tang, Jiayu Chen, Tian Lan. MALinZero: Efficient Low-Dimensional Search for Mastering Complex Multi-Agent Planning. NeurIPS 2025.
A low-dimensional linear-bandit (LinUCT) search that makes joint-action MCTS efficient for complex multi-agent planning.
Computer-Use Agent Tree search
Sizhe Tang, Rongqian Chen, Tian Lan. Agent Alpha: Tree Search Unifying Generation, Exploration and Evaluation for Computer-Use Agents. COLM 2026.
A tree-search framework unifying generation, exploration, and evaluation for computer-use agents.
RL
Zuyuan Zhang, Sizhe Tang, Tian Lan. Cochain Perspectives on Temporal-Difference Signals for Learning Beyond Markov Dynamics. ICML 2026.
A cochain view of temporal-difference signals that extends RL to non-Markovian dynamics.
Post-training
Yu Li, Sizhe Tang, Tian Lan. Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization. ACL 2026 Findings.
Multi-turn agent policy optimization that self-rectifies reasoning chains and grafts good branches via tree search.
💼 Experience
- 2026.05 - 2026.08, Applied Scientist Intern, Amazon AWS AI, Santa Clara, CA.
- Working on web and coding agents (LLM-based autonomous agents).
📖 Educations
- 2024 - 2028 (expected), Ph.D. in Computer Engineering, The George Washington University, Washington, D.C.