Yu Kang 康昱

Senior Expert / Research Scientist Huawei · Software Engineering Application Technology Lab Previously Principal Research Manager, Microsoft DKI (2018 – 2026)

I work at the intersection of AI, software engineering, and systems. My research centers on agent evolution and recursive self-improvement (RSI): AI agents that keep getting better at real software work by learning from their own experience. A large part of this is LLM training, especially agentic reinforcement learning, on executable and verifiable environments at scale.

What sets my work apart is where it starts: the complexity of real products. Industrial codebases are huge, heterogeneous, and constantly evolving, with internal toolchains, build systems, and task types that public benchmarks never see. I model that complexity as scientific problems, then carry the results back into practice. I pioneered turning industrial product repositories and product tasks into executable data and environments for evaluating, tuning, and training coding agents, so that LLMs and coding agents attend to internal products while they are being trained and developed, and can be optimized for them.

Before Huawei, I spent eight years at Microsoft DKI (Data, Knowledge, Intelligence). There I led SWE-bench-Live and RepoLaunch (RepoLaunch builds agentic training environments for frontier open models such as GLM-5), contributed to UFO, TaskWeaver, and Agent Lightning, and worked with 10+ product teams to ship 20+ research results into products, including AIOps technologies in the fundamental cloud services behind Azure and Microsoft 365. I am also an adjunct master’s supervisor at the School of Computer Science, Fudan University. I received my PhD from The Chinese University of Hong Kong under Prof. Michael R. Lyu.

YK
Agent Evolution RSI Agentic RL Coding Agents AIOps GUI Agents
Core researchScenarios
0CitationsGoogle Scholar
0h-indexi10-index 59
0Papers2011 – 2026
0GitHub starsAgent Lightning · UFO · TaskWeaver
NeurIPSICLRICMLACLEMNLPNAACLTMLRICSEESEC/FSEEuroSysUSENIX ATCWWWICDEISSREDSNECAIIEEE CLOUDICWS NeurIPSICLRICMLACLEMNLPNAACLTMLRICSEESEC/FSEEuroSysUSENIX ATCWWWICDEISSREDSNECAIIEEE CLOUDICWS
01 — Research

Agents that improve themselves on real software

At the intersection of AI, software engineering, and systems: two core directions, built on one foundation, carried into three application scenarios.

North star

Self-evolving agents for real-world software

Agents that do not stop improving at release. They learn from their own trajectories and failures, and they are trained against the products they will actually work on, not only against public benchmarks.

Core direction 2

LLM Training & Agentic RL

Training coding models and agents with reinforcement learning inside real agent harnesses, on executable, verifiable environments at scale. This covers environment construction, rewards and verification, trajectory data, and training frameworks. Agent Lightning v1.0 lifts Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4% with 6K examples.

built on
Foundation

Product-grounded environments, data & evaluation

Both directions need a steady supply of executable tasks with reliable verification. I turn product repositories, product tasks, and developer history into exactly that, so the same infrastructure serves evaluation, tuning, and training.

From product complexity to scientific problems

My research style: start from what makes a real product hard, state it as a research problem, and take the answer back into production.

Product complexityScientific problemResult in practice
Every repository builds differentlyMany languages, platforms, and toolchains; no two setups alike.
→
Environment construction as an agent taskCan an agent set up, build, and test any repository on its own?
→
RepoLaunch → GLM-5Executable environments for frontier-model agentic training.
Product code evolves every dayInternal tasks live in commit history, not in public benchmarks.
→
Task reconstruction from repository historyTurn merged changes into verified tasks on today’s code.
→
Change2Task79.6% verified construction across five task families; data for training and evaluation.
Products must move across platformsMillion-line codebases tied to one OS and one language.
→
Repository-level translation with verificationTranslate the skeleton first, then fill in code under test-driven checks.
→
Skeleton-guided translationHelped 10+ Microsoft products migrate million-line repositories from Windows to Linux.
Cloud incidents at hyperscaleThousands of services, noisy signals, costly outages.
→
LLM reasoning over diagnosticsRoot-cause analysis, triage, and mitigation as learning problems.
→
RCACopilot & AIOpsIntegrated into the cloud services behind Azure and Microsoft 365.

Application scenarios

Where the methods are applied and tested.

Current focus

Coding Agents

Issue resolution, test generation, code translation, dependency upgrades, and repository understanding, on both open-source and industrial code.

SWE-bench-LiveTestExploraExeCoderUpdate from Hell
Computer use

GUI Agents

Desktop AgentOS and multi-device agent systems, plus a widely cited survey of LLM-brained GUI agents.

UFOUFO²UFO³GUI Agent Survey
Cloud operations

AIOps

Incident detection, triage, root-cause analysis, logging, and monitoring for hyperscale cloud, deployed at Microsoft.

RCACopilotXpertUniLogUniParser
02 — Impact

Research that frontier labs build on

Environments, benchmarks, frameworks, and systems used outside the lab: in frontier-model training and evaluation, by large open-source communities, and in hyperscale production.

Agentic RL training · GLM-5

RepoLaunch builds executable environments for GLM-5

“We employ an environment setup pipeline based on the RepoLaunch framework that scales the construction of executable environments from real-world SWE issues.”

— GLM-5 Technical Report, Zhipu AI / Z.ai (arXiv 2602.15763)

RepoLaunch is the first agent that resolves dependencies, compiles code, and extracts test results for any language on any OS: C/C++, C#, Python, Java, JS/TS, Go, and Rust, on Linux and Windows. It was used to build 10,000+ Docker environments for large-scale agentic training.

Evaluation · Qwen3-Coder

SWE-bench-Live in Qwen3-Coder’s headline evaluation

Qwen3-Coder release benchmark table including an SWE-bench Live row

Qwen3-Coder’s release reports SWE-bench-Live next to Kimi-K2, DeepSeek-V3, Claude Sonnet 4, and GPT-4.1.

Agentic RL framework

Agent Lightning

“Our original Agent Lightning introduced a disaggregated architecture that connects arbitrary agents to RL training through an LLM endpoint proxy, an approach later adopted by frameworks such as verl Uni-Agent, AReaL 2.0, slime, and Polar.”

— Agent Lightning v1.0 (arXiv 2608.17528)

Production · Microsoft

Research shipped into products

20+ results shipped with 10+ Microsoft product teams (Azure, Microsoft 365, Windows, DevDiv, Copilot): AIOps across the incident lifecycle, and resource-management algorithms for Azure and Microsoft 365 that save hundreds of millions of dollars a year. 10+ patents.

RCACopilotXpertUniLogUniParserMonitorAssistant
03 — Projects

Selected projects

A short selection, chosen for adoption, citations, or top-venue publication. Led marks projects I led.

SWE-bench-Live pipeline: issue crawling, RepoLaunch environment setup, task validation
LedNeurIPS 2025 · D&BUsed by GLM-5 · Qwen3-Coder

SWE-bench-Live & RepoLaunch

A continuously updated issue-resolution benchmark, built by an automated pipeline. RepoLaunch is an agent that sets up, builds, and tests any repository, in any language, on Linux or Windows, turning real issues into verified, reproducible Docker tasks with no human in the loop. The benchmark now has MultiLang (1,077 tasks · 8 languages) and Windows splits, and the pipeline reaches ≥98% build/test success at 82% lower LLM cost.

2026LLM4Code @ ICSE 2026

Change2Task & enterprise benchmarks

Product repositories become a renewable source of agent tasks. Change2Task converts merged changes into verified tasks on healthy modern revisions (bug fix, feature addition, test generation, API migration, security repair), reaching 79.6% verified construction and 29.2% more tasks than a PR-based baseline. A companion line of work generates continuous benchmarks for enterprise-scale agents.

2026★ 18.4k

Agent Lightning v1.0

A lightweight framework for harnessed agentic RL: the deploy-time agent harness owns the interaction loop and takes part in post-training directly. It addresses retokenization, sample merging, advantage calculation, and loss normalization, and ships a full coding-agent RL pipeline: 41.8% → 56.4% on SWE-bench Verified for Qwen3.5-9B. At Microsoft I drove its large-scale distributed deployment on internal clusters, with rollout and training decoupled.

DoVer intervention-driven debugging pipeline
ICLR 2026

DoVer

Automatic debugging for LLM multi-agent systems, a step toward agents that repair themselves. DoVer segments the trace, forms failure hypotheses, intervenes at the orchestrator, then replays to validate the fix.

Evolution of LLM-brained GUI agents
NAACL 2025ICML 2025TMLR 2025★ 9.8k550+ citations

UFO series: Desktop AgentOS

UFO was a pioneering UI-focused agent for Windows. UFO² turned it into a multi-agent Desktop AgentOS (HostAgent + AppAgents, hybrid GUI–API actions), and UFO³ connects agents across devices. Related work includes the survey of LLM-brained GUI agents (230+ citations) and a study of how API and GUI agents diverge and converge (ICML 2025).

TaskWeaver architecture
★ 6.2k120+ citations

TaskWeaver

A code-first agent framework. User requests become executable code: a Planner and a Code Interpreter work together over stateful data structures, with domain plugins for data analytics.

RCACopilot pipeline
EuroSys 2024370+ citations

RCACopilot

One of the first LLM systems for root-cause analysis of cloud incidents: incident-specific handlers collect diagnostics, and an LLM predicts and explains the root cause. Evaluated on a year of real Microsoft incidents.

UniLog in-context logging
ICSE 2024 ×2WWW 2022

LLMs for cloud observability

UniLog decides where and what to log through in-context learning, UniParser parses heterogeneous logs, and Xpert writes KQL diagnostic queries for new incidents.

04 — Publications

Selected publications

Papers chosen for top-venue publication, citations, or adoption. See the full list or Google Scholar.

05 — Experience

Experience & education

  1. 2026 —

    Senior Expert / Research Scientist

    Software Engineering Application Technology Lab, Huawei

    Agent evolution and RSI, LLM training and agentic RL, and product-grounded environments for coding agents.

  2. 2018 — 2026

    Principal Research Manager

    Microsoft DKI (Data, Knowledge, Intelligence; formerly part of Microsoft Research Asia)

    Led SWE-bench-Live and RepoLaunch. Contributed to Agent Lightning, UFO, TaskWeaver, and a line of AIOps and resource-management systems shipped with 10+ product teams (Azure, Microsoft 365, Windows, DevDiv, Copilot).

  3. 2020 —

    Adjunct Master’s Supervisor

    School of Computer Science, Fudan University

  4. 2016 — 2018

    Postdoctoral Researcher

    School of Computer Science, Fudan University

    Co-advisor: Prof. Xin Wang

  5. 2016 —

    Honorary Research Associate

    The Chinese University of Hong Kong

  6. 2012 — 2013

    Research Intern

    Software Analytics Group, Microsoft Research Asia

    Star of Tomorrow Award.

  7. 2012

    Visiting Student

    Computer Science, Duke University

    Host: Prof. Kishor S. Trivedi

  8. 2010 — 2016

    Ph.D., Computer Science & Engineering

    The Chinese University of Hong Kong

    Advisor: Prof. Michael R. Lyu

    Mobile app performance diagnosis (DiagDroid, FSE’16), cloud service deployment, and software security.

  9. 2006 — 2010

    B.S., Computer Science & Technology

    Fudan University

    Top 5%; first-class scholarship three years in a row.

06 — Service & honors

Service

  • ICSE 2027
    Proceedings Co-Chair (Organizing Committee); Program Committee, Research Track
  • FSE 2025
    Program Committee, Industry Papers
  • 2021 – 2025
    Cloud Intelligence / AIOps Workshop: Proceedings Chair (2021, 2023, 2025); Program Committee (2024)
  • ICSE 2022
    Mentor, SMeW Student Mentoring Workshop
  • FSE 2021
    Session Chair, Architectures & Design — Cloud Computing

Honors & grants

  • 2022
    ACM SIGSOFT Distinguished Paper Award, ESEC/FSE 2022 (SPINE)
  • 2021
    Best Research Paper Award, ISSRE 2021 (incident time-to-mitigation prediction)
  • 2016, 2017
    Third Prize, Software Research Prototype Competition, NASAC (CCF Technical Committee on Software Engineering); 2016 for DiagDroid
  • 2017
    Third Prize, Next-Generation Internet Technology Innovation Competition (CERNET)
  • 2013
    Star of Tomorrow Award, Microsoft Research Asia
  • Grants
    PI, NSFC Young Scientists Fund; China Postdoctoral Science Foundation
  • Patents
    10+ patents
07 — Join us

We are hiring

Huawei · SE Application Technology Lab

Build agents that improve themselves

Our research team is growing. If you want to work on agent evolution, agentic RL, and coding agents that meet the full complexity of real products, I would like to hear from you.

  • Full-time researchers & engineersAgentic RL and LLM training, agent evolution, environments and evaluation for coding agents.
  • Research internsPhD and master’s students; internships can lead to top-venue papers and to systems used in real products.
  • Research collaborationJoint projects with universities and industry labs on self-improving agents and agentic training.