System-Level Evaluation for Code Agents
Investigating the gap between unit-level correctness and system-level behavioral correctness in end-to-end software generation.
I work on data-centric AI, code security for code agents, and agentic reinforcement learning. My current interests center on how software engineering agents can be evaluated, improved, and aligned with realistic end-to-end development workflows.
Investigating the gap between unit-level correctness and system-level behavioral correctness in end-to-end software generation.
Studying model-specific harness optimization for coding agents, including prompts, tool scheduling, retry strategies, and workflow control.
Exploring how code agents should interact with humans under ambiguous or underspecified requirements.
Collaborative research project
A multi-agent framework that transforms sparse CVE metadata into executable agentic code-security tasks. I contributed to data-quality auditing, the false-positive taxonomy, and LLM-as-Judge automatic auditing.
First-author research project
A closed-loop self-evolution framework for software engineering agents that extracts reusable skills from solving traces and uses them for dynamic curriculum generation.