Qijia He (何其佳)

profile photo

I am a final-year PhD student in the Department of Statistics at the University of Washington, advised by Prof. Bo Zhang. I collaborate with Prof. Ranjay Krishna at the UW CSE RAIVN Lab on multimodal AI research. Before graduate school, I received my B.S. in Statistics from Sun Yat-sen University. I have also interned at Google, TikTok and Apple.

I build LLM & VLM agents through post-training, spanning on-policy distillation, agentic reinforcement learning, and capability-aware data curation. My research focuses on enabling agents to learn stably and efficiently from their own experience, toward reliable long-horizon reasoning, tool use, and multimodal generation in open-ended, real-world environments.

I am currently seeking full-time industry opportunities starting in 2027. Please feel free to reach out for any relevant opportunities!

Contact me: heqj3@uw.edu

Selected Research

LLM / VLM

DTDD: Divergence-Triggered Dynamic Distillation for Reliable On-Policy Supervision
Qijia He, Zequn Li, Spencer Luo, Tianle Chen, Pradyumna Narayana, Zixian Ma, Ranjay Krishna, Lan Nie, Sugato Basu
Preprint, 2026

TL;DR: When a student trajectory drifts too far from the teacher, on-policy distillation becomes unstable. DTDD detects the drift via segment-level divergence and hands off to teacher recovery. Across long-reasoning math, base-model distillation, and ALFWorld agents, it cuts the peak gradient norm from 12.0 to 1.4, reaches OPD-level accuracy in about 2.5× fewer updates, and raises success from 76.5% to 81.3%.

Agent Capability Shapes the Value of Post-Training Data
Jiayi Cheng, Yunshu Wu, Chenqian Le, Qijia He, Runhao Li, Michael Yue, Xupeng Chen, Michal Mankowski
Preprint, 2026

TL;DR: Which experiences should an agent learn from? We select post-training data by the agent's own prediction gaps: for SFT, traces whose execution outcomes the agent mispredicts yield 12.8% higher mean reward than random selection; for RL, tasks whose rollout rewards it mispredicts improve final reward by 33.3%, and cheap predicted-reward proxies work without extra verification.

What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents
Chenqian Le, Jiayi Cheng, Qijia He, Runhao Li, Yinghao Li, Xupeng Chen
Preprint, 2026 [Paper]

TL;DR: Multi-harness RL for coding agents both trains across several harnesses and compares their rewards inside one GRPO group. Replaying identical Qwen3-8B trajectories from Aider, OpenHands, Qwen Code, and SWE-agent under per-harness vs. cross-harness grouping, we find the evaluation harness moves SWE-bench Verified solve rate by 4.3×, while the grouping rule has no detectable effect on an unseen held-out harness.

VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
Qijia He, Xunmei Liu, Hammaad Memon, Ziang Li, Zixian Ma, Jaemin Cho, Zhongzheng Ren, Dan Weld, Ranjay Krishna
NeurIPS 2026 (Evaluations & Datasets Track) [Paper] [Project Page] [Code]

TL;DR: VFIG is the first large-scale ecosystem for scientific figure-to-SVG generation, introducing a 66K figure–SVG dataset, a new benchmark, and a family of VLMs trained with SFT+RL that achieve state-of-the-art open-source performance and parity with GPT-5.2.

ARCHER: Adaptive Recovery Routing with Conformal-risk Calibration for Coding Agents
Qijia He, Jiayi Cheng, Chenqian Le, Rui Wang, Xunmei Liu, Yixian Chen, Jie Mei, Zhihao Wang, Xupeng Chen, Yuhuan Chen, Tao Wang
Efficient Reasoning @ COLM 2026 Spotlight [Paper] [Code]

TL;DR: When a coding agent fails, should it attempt cheap recovery or escalate to a stronger model? ARCHER casts this as a routing problem over heterogeneous recovery actions trained on execution rollouts, with a conformal risk control layer that adapts to any budget without retraining—exceeding always-escalate solve rates at 35% of its recovery cost.

Rethinking Human Preference Evaluation of LLM Rationales
Ziang Li, Manasi Ganti, Zixian Ma, Helena Vasconcelos, Qijia He, Ranjay Krishna
XLLM-Reason-Plan @ COLM 2025 Best Paper Award (Honorable Mention) [Paper]

TL;DR: Large language models (LLMs) generate rationales that improve reasoning and interpretability, but existing evaluations using binary human or LLM preferences are limited and opaque. We introduce an attribute-based evaluation framework that defines key rationale qualities, explains human preferences, and enables more nuanced model comparisons through attribute-specific analysis.

Causal Inference

Dealing with positivity violations in mediation analysis via weighted controlled effects
Qijia He, Bo Zhang
arxiv, 2026 [Paper] [Code]

TL;DR: This paper introduces a weighted controlled risk approach to causal mediation analysis, addressing scenarios where traditional fixed-level interventions violate the positivity assumption by targeting subpopulations with a high probability of attaining specific mediator levels.

Role of placebo samples in observational studies
Ting Ye, Qijia He, Shuxiao Chen, Bo Zhang
Journal of Causal Inference, 2025 [Paper] [Supplement]

TL;DR: We proposed a framework for using placebo samples in observational studies to detect and correct for unmeasured confounding bias. It develops identification assumptions and estimation methods—including regression, weighting, and doubly robust approaches—and validates them through simulations and an applied case study on tax credits and infant health.

Generalizing the Intention-to-Treat Effect of an Active Control from Historical Placebo-Controlled Trials
Qijia He, Fei Gao, Oliver Dukes, Sinead Delany-Moretlwe, Bo Zhang
Journal of the American Statistical Association, 2024 [Paper] [Supplement]

TL;DR: We developed a potential outcomes framework to estimate the ITT effect of an active control versus placebo in active-controlled trials, using historical placebo-controlled data. Our method enables ITT estimation when the placebo arm is unavailable and accounts for unmeasured confounders using instrumental variables.

Estimating individualized treatment rules by optimizing the adjusted probability of a longer survival
Qijia He, Shixiao Zhang, Michael L LeBlanc, Yingqi Zhao
Statistical Methods in Medical Research, 2024 [Paper] [Supplement]

TL;DR: We introduced a new criterion for individualized treatment rules based on the adjusted probability of longer survival, offering a clear and clinically relevant objective. Our method, optimal adjusted probability learning, constructs the best treatment rule by maximizing this nonparametric survival benefit.

Industry Experience

Software Engineer Intern (PhD), Summer 2026

Google, New York, NY
Research on on-policy distillation in LLMs

Machine Learning Intern, Spring 2026

TikTok, Bellevue, WA
Worked on agentic AI systems for e-commerce applications

Machine Learning Intern, Summer 2025

Apple, Cupertino, CA
Worked on causal machine learning methods for ads marketplace

Teaching Experience

Department of Statistics, University of Washington
  • STAT 311 Elements of Statistical Methods (Winter 2024, Autumn 2025)
  • STAT 513 Statistical Inference (Winter 2026)
  • STAT 435 Introduction to Statistical Machine Learning (Spring 2026)
Academic tutoring center, School of Mathematics, Sun Yat-sen University

Tutor in Mathematical analysis (Fall 2018)

TAL Education Group

Teaching Assistant in primary-school Olympiad Mathematics (2017-2018)

  • Rated S (top 5% performance)