πŸ“ Selected Publications

Preprint
sym

[New!] OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models βš–οΈπŸ–₯️

Qiushi Sun†, Kanzhi Cheng†, Yian Wang, Bowen Yang, Hang Yan, Liheng Chen, Fangzhi Xu, Zichen Ding, Nuo Chen, Jialin Cao, Xingdong Gong, Zehao Li, Kaiming Jin, Xinfeng Yuan, Zhoumianze Liu, Jingyang Gong, Zhangyue Yin, Jiahui Gao, Zhiyong Wu, Tianbao Xie, Jianbing Zhang, Ben Kao, Lingpeng Kong

Paper Project OSReward OS-Shepherd-9B/35B Data Code BIB

  • First standardized evaluation revealing how well VLMs judge computer-use agents across platforms βš–οΈ
  • The most comprehensive study to date, with extensive insights into effective reward models for CUAs πŸ”
  • OS-Shepherd-100K: a large-scale corpus of trajectory judgments built on our findings πŸ“Š
  • OS-Shepherd-9B/35B: open reward models rivaling commercial judges at a fraction of the cost πŸš€
ICLR 2026
sym

ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows πŸ§ͺπŸ”¬

Qiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding, Fangzhi Xu, Zhangyue Yin, Haiteng Zhao, Zhenyu Wu, Kanzhi Cheng, Zhaoyang Liu, Jianing Wang, Qintong Li, Xiangru Tang, Tianbao Xie, Xiachong Feng, Xiang Li, Ben Kao, Wenhai Wang, Biqing Qi, Lingpeng Kong, Zhiyong Wu

Paper Slide Project HF Env Code BIB

  • First to apply computer-using agents to assist scientific exploration 🌌
  • Dynamic environment & benchmark for realistic scientific workflows 🌍
  • Comprehensive evaluation of SOTA LLM/VLM agents 🧭
ICLR 2026
sym

JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence

Qiushi Sun*, Jingyang Gong*, Yang Liu*, Qiaosheng Chen*, Lei Li, Kai Chen, Qipeng Guo, Ben Kao, Fei Yuan

Paper Slide Project HF Data Code BIB

  • JanusCoder series: foundational models establishing a unified visual-programmatic interface. βš™οΈ
  • A versatile data synthesis toolkit for multimodal code intelligence. πŸ› οΈ
  • Superior performance on diverse text- and vision-centric tasks. 🧭
ACL 2026 Oral
sym

OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows πŸ›‘οΈπŸ§

Qiushi Sun*, Mukai Li*, Zhoumianze Liu*, Zhihui Xie*, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, Lingpeng Kong

Paper Slide Project HF Env Code BIB

πŸ† AIWILD @ ICLR 2026 Best Paper Award

  • MobileRisk-Live & MobileRisk, a dynamic environment and benchmark for realistic mobile agent safety πŸ“±
  • OS-Sentinel, a hybrid detection framework combining formal verification with contextual judgment πŸ›‘οΈ
  • Advanced mobile agent safety at both the step-level and trajectory-level 🧭
ACL 2026 Oral
sym

CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback πŸ§¬βš™οΈ

Qiushi Sun, Jingyang Gong, Lei Li, Qipeng Guo and Fei Yuan

Paper Slide Code Data BIB

  • Interaction-driven synthesis via Coder–Reviewer agents in a closed loop πŸ€–
  • Hybrid feedback fusing NL critique with compiler signals πŸ”
  • Iterative refinement yields high-quality trajectories 🌠
ACL 2025
sym

OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis πŸ’§πŸ€–

Qiushi Sun*, Kanzhi Cheng*, Zichen Ding*, Chuanyang Jin*, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, Zhiyong Wu

Paper Slide Project HF Data Code BIB

  • Shift from task-driven to interaction-driven GUI data synthesis πŸ€–
  • A manual-free pipeline for constructing diverse GUI agent trajectories 🧬
  • Great performance on online mobile/web benchmarks 🌟
COLM 2024
sym

Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration πŸ€–πŸ’­

Qiushi Sun, Zhangyue Yin, Xiang Li, Zhiyong Wu, Xipeng Qiu and Lingpeng Kong

Paper Slide Code BIB

  • Among the earliest works to have agents collaborate through reasoning and coding to solve complex tasks 🧩
  • Three collaboration modes: Discuss, Review, and Retrieve 🀝
  • Multi-agent collaboration outperforms prompting single-model with brute forceπŸ“ˆ
Survey
sym

A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond πŸ”₯πŸ”₯

Qiushi Sun, Zhirui Chen, Fangzhi Xu, Chang Ma, Kanzhi Cheng, Zhangyue Yin, Jianing Wang, Chengcheng Han, Renyu Zhu, Shuai Yuan, Pengcheng Yin, Qipeng Guo, Xipeng Qiu, Xiaoli Li, Fei Yuan, Lingpeng Kong, Xiang Li, Zhiyong Wu

Paper Slide Project Code BIB

  • Follow LMs for code as a thread to trace the field’s development πŸš€
  • Explore cross-domain synergies and opportunities 🌱
  • Present a broad array of promising research avenues πŸ’‘

*Denotes equal contribution, βœ‰ denotes corresponding author, more working drafts / preprints under review will be released later βŒ›οΈ