Hi! I am Sizhe Tang, a second-year Ph.D. candidate at The George Washington University, advised by Prof. Tian Lan.

My research focuses on reinforcement learning and self-evolving agents, with an emphasis on tree-search and planning for multi-agent decision-making (the Zero series), computer-use agents, and multi-turn agent policy optimization.

You can find my publications on Google Scholar.

🔥 News

  • 2026.06:  🎉 Agent Alpha first author accepted to COLM 2026 (score: 778).
  • 2026.05:  🎉 NonZero first author accepted to ICML 2026 (Spotlight).
  • 2026.05:  🎉 HFPS accepted to ICML 2026.
  • 2026.04:  🎉 I will join Amazon AWS AI as an Applied Scientist Intern (Summer 2026, Santa Clara, CA).
  • 2026.03:  🎉 T-STAR accepted to ACL 2026 Findings.
  • 2025.09:  🎉 MALinZero first author accepted to NeurIPS 2025.

📝 Selected Publications

RL Tree search

Sizhe Tang, Zuyuan Zhang, Mahdi Imani, Tian Lan. NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search. ICML 2026 (Spotlight).

An interaction-guided exploration rule that scales Monte Carlo Tree Search to multi-agent planning.

RL Tree search

Sizhe Tang, Jiayu Chen, Tian Lan. MALinZero: Efficient Low-Dimensional Search for Mastering Complex Multi-Agent Planning. NeurIPS 2025.

A low-dimensional linear-bandit (LinUCT) search that makes joint-action MCTS efficient for complex multi-agent planning.

Computer-Use Agent Tree search

Sizhe Tang, Rongqian Chen, Tian Lan. Agent Alpha: Tree Search Unifying Generation, Exploration and Evaluation for Computer-Use Agents. COLM 2026.

A tree-search framework unifying generation, exploration, and evaluation for computer-use agents.

RL

Zuyuan Zhang, Sizhe Tang, Tian Lan. Cochain Perspectives on Temporal-Difference Signals for Learning Beyond Markov Dynamics. ICML 2026.

A cochain view of temporal-difference signals that extends RL to non-Markovian dynamics.

Post-training

Yu Li, Sizhe Tang, Tian Lan. Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization. ACL 2026 Findings.

Multi-turn agent policy optimization that self-rectifies reasoning chains and grafts good branches via tree search.

💼 Experience

  • 2026.05 - 2026.08, Applied Scientist Intern, Amazon AWS AI, Santa Clara, CA.
    • Working on web and coding agents (LLM-based autonomous agents).

📖 Educations

  • 2024 - 2028 (expected), Ph.D. in Computer Engineering, The George Washington University, Washington, D.C.