🔭 Research Overview
My research aims at building agentic methods that operate real-world software, along two threads:
Computer-Using Agents — agents that operate computers, phones, and browsers the way people do, toward digital automation. My work spans GUI grounding (SeeClick, OS-Atlas), trajectory synthesis (OS-Genesis), evaluation in realistic workflows (ScienceBoard), safety (OS-Sentinel), and reward modeling (OSReward & OS-Shepherd).
Code Intelligence — language models that understand and generate code, and code as an interface for agents. My work charts the field (NCI Survey), synthesizes code-centric data via agent interaction (CodeEvo), and builds multimodal code models (JanusCoder) for generative UI and beyond.
🧵 Publication Threads
* equal contribution · † project co-lead
-
Preprint
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward ModelsA standardized benchmark for computer-use reward models, with OS-Shepherd judges trained on 100K trajectory judgments. -
ICLR'26

-
ICLR'26

-
ACL'26 Oral

-
ACL'26 Oral

-
ACL'25

-
COLM'26
OpenMobile: Building Open Mobile Agents with Task and Trajectory SynthesisAn open recipe for mobile agents that synthesizes tasks from environment memory and rolls out trajectories with policy switching. -
ACL'26
OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using AgentsOrchestrates tool, grounding, and reflection-memory agents for robust computer use across operating systems. -
ICLR'26 Oral
ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform DataScales open computer-use agents with a cross-platform corpus spanning six operating systems. -
ICLR'25 Spotlight
OS-ATLAS: A Foundation Action Model for Generalist GUI AgentsA foundation action model unifying GUI grounding, action, and agent modes across platforms. -
ACL'24
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsPioneers screenshot-only GUI agents via grounding pre-training and introduces the ScreenSpot benchmark. -
COLM'24
Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model CollaborationMulti-model collaboration that pushes complex reasoning beyond single-model prompting. -
Preprint
OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive InteractionsBenchmarks LLMs on long-horizon, active, and inductive interaction in live environments. -
ACL'25
Dynamic and Generalizable Process Reward ModelingAutomatically designs process rewards and allocates them dynamically via reward trees and Pareto-optimal selection. -
ACL'25
Interactive Evolution: A Neural-Symbolic Self-Training Framework for Large Language ModelsA neural-symbolic self-training loop that improves LLMs from environment feedback without human annotation. -
ACL'24
Boosting Language Models Reasoning with Chain-of-Knowledge PromptingPrompts models with structured knowledge triples to ground multi-step reasoning and curb hallucination. -
COLING'24
Make Prompt-based Black-Box Tuning Colorful: Boosting Model Generalization from Three Orthogonal PerspectivesBoosts gradient-free prompt tuning with two-stage optimizers, multi-verbalizers, and better initialization. -
Survey

OS-Shepherd-9B/35B