I am a Ph.D. candidate at
The University of Hong Kong, advised by Prof. Ben Kao, and also work with Prof. Lingpeng Kong at HKUNLP. Iβm currently a research scientist intern at
Google Research, Mountain View. My research focuses on (1) Computer-using agents and (2) Code intelligence. I completed my masterβs at
National University of Singapore, where I interned at A*STAR with Dr. Xiaoli Li. After that, I spent time at
Shanghai AI Lab with Dr. Zhiyong Wu and Dr. Qipeng Guo. Earlier, I completed my B.Eng with distinction at
East China Normal University, where I was privileged to be instructed by Prof. Weining Qian, Prof. Xuesong Lu, and Prof. Xiang Li.
π Office hours: I am holding office hours (1~2 hours per week) dedicated to offering consultation for COMP7607 / COMP3270 students and mentorship programs. If you want to have a chat (whether or not itβs about research), please book me through this link (and send me an e-mail.)
π Collaboration: I am looking for motivated collaborators interested in the above topics. If you would like to explore these directions together, feel free to contact me. UG/MSc students are also welcomed! π±
π₯ News
- 2026.07βοΈπ We release OSReward and OS-Shepherd to advance computer-use reward models!
- 2026.06π³ Attending ACL 2026 in San Diego ποΈ 2 oral talks: OS-Sentinel and CodeEvo!
- 2026.04π Excited to be joining
Google Research as a Student Researcher Intern! - 2026.04π OS-Sentinel received the Best Paper Award at AIWILD @ ICLR 2026!
- 2026.04π Four papers are accepted by ACL 2026! See you in San Diego πΊπΈ!
- 2026.03π§βοΈ Excited to receive the Tencent Qingyun Scholarship (Student Travel Grant).
- 2026.02ππ¬ Introducing
Intern-S1-Pro, a trillion-scale scientific multimodal foundation model! - 2026.01π JanusCoder, ScienceBoard, and ScaleCUA (Oral) are accepted by ICLR'26! See you in Rio π§π·!
- 2025.11β±οΈ Attending EMNLP 2025 in Suzhou π¨π³
- 2025.10π‘οΈπ€ We launch OS-Sentinel and MobileRisk to advance the safety research of mobile agents.
- 2025.10βοΈπ€ Introducing JanusCoder-7B/14B and JanusCoderV-7B/8B, new foundation models for multimodal code intelligence.
- 2025.08πποΈ Start serving as an Area Chair for ACL Rolling Review.
- 2025.06ποΈπ ScienceBoard will be presented as an oral paper at ICML 2025 Workshop on Computer Use Agents π¨π¦!
- 2025.05π¬π§ͺ We release ScienceBoard to advance computer-using agents in scientific workflows!
- 2025.05π Six papers are accepted by ACL 2025! Bis bald in Wien π¦πΉ!
- 2025.03ποΈ Will attend ICLR 2025! See you at Singapore πΈπ¬!
- 2024.12π€ We release OS-Genesis (ACL'25), OS-Atlas (ICLR'25) and AgentStore (ACL'25) to advance GUI agents!
- 2024.08βοΈ (Physically) started my PhD at
The University of Hong Kong ππ°! - 2024.07π One paper get accepted by COLM 2024! See you at Upenn πΊπΈ!
- 2024.05π₯ Four papers are accepted by ACL 2024! See you in Bangkok πΉπ!
- 2024.03π Check out our Code Intelligence Survey Paper π₯
- 2024.02π Graduated from
National University of Singapore. - 2023.12β±οΈ Attending EMNLP 2023 in SG πΈπ¬
- 2023.07β¨ Started my research intern at NLP Group,
Shanghai AI Lab
- 2023.05π HugNLP Framework (CIKM'23 Best Demo Paper) is ready for use! Please check our Paper, Repo and Blogs
- 2023.05π We release SelfAware for benchmarking LLMs' self-knowledge
- 2023.01π Started my research intern at
I2R, A*STAR, Singapore
- 2022.12π Our team won second prize (100k RMB) in the International Algorithm Case Competition: PLM Tuning Track.
- 2022.08π Started my master's studies at National University of Singapore. πΈπ¬
- 2022.07π Graduated from
School of Data Science and Engineering of
ECNU! - 2022.07π₯ Awarded outstanding UG thesis and Shanghai Outstanding Graduate.
- 2021.09π Started serving as a TA for Deep Learning for Computer Vision course this semester.
- 2021.05π Led my team to win the Finalist Award in the Mathematical and Interdisciplinary Contest in Modeling!
- 2021.02βοΈ Attending Data Science Winter School at Imperial College London π¬π§.
π Selected Publications
[New!] OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models βοΈπ₯οΈ
Qiushi Sunβ , Kanzhi Chengβ , Yian Wang, Bowen Yang, Hang Yan, Liheng Chen, Fangzhi Xu, Zichen Ding, Nuo Chen, Jialin Cao, Xingdong Gong, Zehao Li, Kaiming Jin, Xinfeng Yuan, Zhoumianze Liu, Jingyang Gong, Zhangyue Yin, Jiahui Gao, Zhiyong Wu, Tianbao Xie, Jianbing Zhang, Ben Kao, Lingpeng Kong
Paper Project OSReward
OS-Shepherd-9B/35B Data Code BIB
- First standardized evaluation revealing how well VLMs judge computer-use agents across platforms βοΈ
- The most comprehensive study to date, with extensive insights into effective reward models for CUAs π
- OS-Shepherd-100K: a large-scale corpus of trajectory judgments built on our findings π
- OS-Shepherd-9B/35B: open reward models rivaling commercial judges at a fraction of the cost π
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows π§ͺπ¬
Qiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding, Fangzhi Xu, Zhangyue Yin, Haiteng Zhao, Zhenyu Wu, Kanzhi Cheng, Zhaoyang Liu, Jianing Wang, Qintong Li, Xiangru Tang, Tianbao Xie, Xiachong Feng, Xiang Li, Ben Kao, Wenhai Wang, Biqing Qi, Lingpeng Kong, Zhiyong Wu
Paper Slide Project HF Env Code BIB
- First to apply computer-using agents to assist scientific exploration π
- Dynamic environment & benchmark for realistic scientific workflows π
- Comprehensive evaluation of SOTA LLM/VLM agents π§
JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence 
Qiushi Sun*, Jingyang Gong*, Yang Liu*, Qiaosheng Chen*, Lei Li, Kai Chen, Qipeng Guo, Ben Kao, Fei Yuan
Paper Slide Project HF Data Code BIB
- JanusCoder series: foundational models establishing a unified visual-programmatic interface. βοΈ
- A versatile data synthesis toolkit for multimodal code intelligence. π οΈ
- Superior performance on diverse text- and vision-centric tasks. π§
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows π‘οΈπ§
Qiushi Sun*, Mukai Li*, Zhoumianze Liu*, Zhihui Xie*, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, Lingpeng Kong
Paper Slide Project HF Env Code BIB
π AIWILD @ ICLR 2026 Best Paper Award
- MobileRisk-Live & MobileRisk, a dynamic environment and benchmark for realistic mobile agent safety π±
- OS-Sentinel, a hybrid detection framework combining formal verification with contextual judgment π‘οΈ
- Advanced mobile agent safety at both the step-level and trajectory-level π§
CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback π§¬βοΈ
Qiushi Sun, Jingyang Gong, Lei Li, Qipeng Guo and Fei Yuan
- Interaction-driven synthesis via CoderβReviewer agents in a closed loop π€
- Hybrid feedback fusing NL critique with compiler signals π
- Iterative refinement yields high-quality trajectories π
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis π§π€
Qiushi Sun*, Kanzhi Cheng*, Zichen Ding*, Chuanyang Jin*, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, Zhiyong Wu
Paper Slide Project HF Data Code BIB
- Shift from task-driven to interaction-driven GUI data synthesis π€
- A manual-free pipeline for constructing diverse GUI agent trajectories π§¬
- Great performance on online mobile/web benchmarks π
Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration π€π
Qiushi Sun, Zhangyue Yin, Xiang Li, Zhiyong Wu, Xipeng Qiu and Lingpeng Kong
- Among the earliest works to have agents collaborate through reasoning and coding to solve complex tasks π§©
- Three collaboration modes: Discuss, Review, and Retrieve π€
- Multi-agent collaboration outperforms prompting single-model with brute forceπ
A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond π₯π₯
Qiushi Sun, Zhirui Chen, Fangzhi Xu, Chang Ma, Kanzhi Cheng, Zhangyue Yin, Jianing Wang, Chengcheng Han, Renyu Zhu, Shuai Yuan, Pengcheng Yin, Qipeng Guo, Xipeng Qiu, Xiaoli Li, Fei Yuan, Lingpeng Kong, Xiang Li, Zhiyong Wu
- Follow LMs for code as a thread to trace the fieldβs development π
- Explore cross-domain synergies and opportunities π±
- Present a broad array of promising research avenues π‘
Technical ReportIntern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale, Intern-S1-Pro Team.ICLR 2025 (Spotlight)OS-ATLAS: A Foundation Action Model For Generalist GUI Agents, Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang and Yu Qiao.ACL 2025Dynamic and Generalizable Process Reward Modeling, Zhangyue Yin, Qiushi Sun, Zhiyuan Zeng, Qinyuan Cheng, Xipeng Qiu and Xuanjing Huang.ACL 2024SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents, Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang and Zhiyong Wu. [LLMAgents @ ICLR 2024]COLING 2024TransCoder: Towards Unified Transferable Code Representation Learning Inspired by Human Skills, Qiushi Sun, Nuo Chen, Jianing Wang, Xiang Li and Ming Gao.COLING 2024Make Prompt-based Black-Box Tuning Colorful: Boosting Model Generalization from Three Orthogonal Perspectives, Qiushi Sun, Chengcheng Han, Nuo Chen, Renyu Zhu, Jingyang Gong, Xiang Li and Ming Gao. | π₯ 100K RMB Award-winning SolutionCIKM 2023 (Demo)HugNLP: A Unified and Comprehensive Library for Natural Language Processing, Jianing Wang, Nuo Chen, Qiushi Sun, Wenkang Huang, Chengyu Wang and Ming Gao. | π Best Paper Award
*Denotes equal contribution, β denotes corresponding author, more working drafts / preprints under review will be released later βοΈ
π Honors and Awards
- 2026.04 Best Paper Award, AIWILD @ ICLR 2026
- 2026.03 Tencent Qingyun Scholarship (Student Travel Grant)
- 2023.10 Best Paper Award (Demo Track), CIKM 2023, ACM & SIGIR
- 2022.12 International Algorithm Case Competition: PLM Tuning, Second Prize, Guangdong-Hong Kong-Macao Greater Bay Area
- 2022.07 Shanghai Outstanding Graduates, Shanghai Municipal Education Commission
More π
- 2022.05 Excellent Bachelor Thesis Award, East China Normal University
- 2021.11 Outstanding Student, East China Normal University
- 2021.10 First Class Scholarship, East China Normal University
- 2021.05 Finalist (Special Class Prize), Mathematical and Interdisciplinary Contest in Modeling
π¬ Invited Talks
- 2026.5 Towards Generalist Computer-using Agents: Models, Data, Safety & Beyond, HKUST-GZ Slide
- 2026.4 Towards Versatile Computer Agents: Cross-Domain Frontiers, Security Guardrails, and the Open-Source Landscape (with Kanzhi Cheng), WebAgentLab & QingKe AI Slide
- 2026.4 Towards Generalist Computer-using Agents: Models, Data & Beyond, HKU ECE Slide
- 2025.11 Computer-using Agent Panel with WebAgentLab, EMNLPβ25, Suzhou, China
- 2025.7 From Digital Agents to AI Co-Scientists,
Slide
- 2025.7 From Digital Agents to AI Co-Scientists, THU Slide
- 2025.7 Towards Generalist Computer-using Agents: Models, Data, and Beyond
Slide - 2025.6 Building GUI Agents with OS-Genesis, NLP Academic Exchange Platform
Slide - 2025.2 Constructing Trajectory Data for Generalist GUI Agents, ModelScope Seminar
Slide
π§βπ« Teaching
I serve(d) as a teaching assistant for:
- COMP3270: Introduction to Artificial Intelligence (for UG), HKU, 2025 Fall. Instructor: Lingpeng Kong
- COMP7607: Natural Language Processing (for masters), HKU, 2024 Fall. Instructor: Lingpeng Kong
- Deep Learning for Computer Vision (for UG), ECNU, 2021 Fall. Instructor: Xuesong Lu
π Educations
- 2024.08 - Present, Ph.D, The University of Hong Kong
, Hong Kong SAR. - 2022.08 - 2024.01, Master, National University of Singapore
, Singapore.
- 2018.09 - 2022.07, Undergraduate, School of Data Science and Engineering, East China Normal University
, Shanghai, China. - 2011.09 - 2018.07, Middle School, Pudong Foreign Languages School
, SISU, Shanghai, China.
π Services
- I serve(d) as an area chair for the following venues:
- ACL Rolling Review (ARR), ACL, AACL-IJCNLP
- AI4Science Workshop, ICMLβ26
- I serve(d) as a reviewer / program committee member for the following conferences, journals, and workshops:
- Conferences: EMNLPβ22, ACLβ23, CIKMβ23, EMNLPβ23 (Best Reviewer Award), ICLRβ24, NeurIPSβ24, NLPCCβ24, ICLRβ25, COLINGβ25, NAACLβ25, ICMLβ25, COLMβ25, NeurIPSβ25, ICLRβ26, CVPRβ26, CPALβ26, COLMβ26
- Journals: Transactions on Machine Learning Research, Frontiers of Computer Science, Knowledge-Based Systems
- Workshops: WiNLP @ EMNLPβ24, DL4C @ ICLRβ25 / NeurIPSβ25 / ICMLβ25, WCUA @ ICMLβ25, ACL Student Research Workshop.
π Recent readings
- Fading Victory: The Diary of Admiral Matome Ugaki, 1941-1945 by Matome Ugaki, 1991
- The Fourth Paradigm: Data-Intensive Scientific Discovery edited by Tony Hey, Stewart Tansley, and Kristin Tolle, 2009
- Sternstunden der Menschheit. Vierzehn historische Miniaturen by Stefan Zweig, 1927
- Science, the Endless Frontier by Vannevar Bush, 1945
More π
- Pearl Harbor: From Infamy to Greatness by Craig Nelson, 2016
- Ulysses S. Grant: The Unlikely Hero by Michael Korda, 2004
- General History of Global Technology (ε ¨ηη§ζιε²) by Jun Wu
βWhatβs past is prologue.β β William Shakespeare (The Tempest)