Disentangling to Re-couple: Resolving the Similarity-Controllability Paradox in Subject-Driven Text-to-Image Generation
A disentangle-and-recouple framework combining a task-specific reward model with GRPO for subject-driven generation.
Researcher · Tencent Beijing, China
I work on AIGC and vision-language model (VLM) training at Tencent, which I joined through the Tencent Qingyun Talent Program. My research spans computer vision and multimodal generation, with a particular interest in visual reasoning and model alignment. I received my Ph.D. in Computer Science from Tsinghua University in 2025. I was a member of the TSAIL Group, led by Prof. Bo Zhang and Prof. Jun Zhu, and was supervised by Prof. Xiaolin Hu. Before that, I received my B.S. in Mathematics and Physics from Tsinghua in 2020.
DSH-Bench, our benchmark for subject-driven text-to-image generation, was accepted to ECCV 2026.
DisCo, our work on subject-driven image generation, was accepted to CVPR 2026.
FGNet, which transfers SAM2 representations to 3D EM neuron segmentation, was accepted to AAAI 2026.
Joined Tencent through the Tencent Qingyun Talent Program.
Received my Ph.D. in Computer Science from Tsinghua University.
Affinity-Guided Queries was accepted to ICLR 2025.
* Equal contribution. My name is in bold.
A disentangle-and-recouple framework combining a task-specific reward model with GRPO for subject-driven generation.
A difficulty- and scenario-aware benchmark covering 58 fine-grained subject categories, with a new identity-consistency metric and diagnostic evaluation of 19 leading models.
Transfers visual foundation-model representations to the EM domain through feature-guided attention and a dual-affinity decoder.
A lightweight query-based framework for large-scale EM neuron segmentation with 2–3× inference acceleration.
Shows that adversarially pretrained backbones are essential for robust detection and introduces an effective fast adversarial fine-tuning recipe.
Dynamically routes images through transformer heads of different complexity, combining budget-aware switching with online head distillation.
Brain-inspired modeling for robust audio-visual speech separation.
Encourages consistent pixel representations within each instance and improves multiple segmentation baselines without inference overhead.
Dynamic boundary patch selection and refinement for high-quality instance segmentation.
A data-centric robust logo detection system for long-tailed e-commerce imagery, ranked fifth in the ACM MM 2021 Grand Challenge.
AIGC and vision-language model training.
Exploratory research in multimodal interaction and GUI agents.
Visual perception for autonomous driving.
TSAIL. Advised by Prof. Xiaolin Hu. Research in computer vision and efficient visual modeling.
Minor in Computer Science and Technology. Academic Excellence Award.