VLDB 2026 Research / reviewers in the wild / expert
Xiao He 0014
dblp:02/2315-14
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-6597-5058ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mixture of Ranks with Degradation-Aware Routing for One-Step Real-World Image Super-ResolutionabstractThe demonstrated success of sparsely-gated Mixture-of-Experts (MoE) architectures, exemplified by models such as DeepSeek and Grok, has motivated researchers to investigate their adaptation to diverse domains. In real-world image super-resolution (Real-ISR), existing approaches mainly rely on fine-tuning pre-trained diffusion models through Low-Rank Adaptation (LoRA) module to reconstruct high-resolution (HR) images. However, these dense Real-ISR models are limited in their ability to adaptively capture the heterogeneous characteristics of complex real-world degraded samples or enable knowledge sharing between inputs under equivalent computational budgets. To address this, we investigate the integration of sparse MoE into Real-ISR and propose a Mixture-of-Ranks (MoR) architecture for single-step image super-resolution. We introduce a fine-grained expert partitioning strategy that treats each rank in LoRA as an independent expert. This design enables flexible knowledge recombination while isolating fixed-position ranks as shared experts to preserve common-sense features and minimize routing redundancy. Furthermore, we develop a degradation estimation module leveraging CLIP embeddings and predefined positive-negative text pairs to compute relative degradation scores, dynamically guiding expert activation. To better accommodate varying sample complexities, we incorporate zero-expert slots and propose a degradation-aware load-balancing loss, which dynamically adjusts the number of active experts based on degradation severity, ensuring optimal computational resource allocation. Comprehensive experiments validate our framework's effectiveness and state-of-the-art performance. Xiao He 0014, Zhijun Tu, Mingrui Zhu, Jie Hu 0021, Nannan Wang 0001, Xinbo Gao 0001 |
AAAI | 1 |
| 2026 | One Step Diffusion-Based Super-Resolution With Time-Aware Distillationabstractiffusion-based image super-resolution (SR) has shown strong potential in recovering high-fidelity details from low-resolution inputs. However, the need for tens or hundreds of sampling steps leads to substantial inference latency. Recent works attempt to accelerate this process via knowledge distillation, but often rely solely on pixel-level loss or overlook the fact that diffusion models capture different information across time steps. To address this, we propose TAD-SR, a time-aware diffusion distillation framework. Specifically, we introduce a novel score distillation strategy to align the score functions between the outputs of the student and teacher models after minor noise perturbation. This distillation strategy eliminates the inherent bias in score distillation sampling (SDS) and enables the student models to focus more on high-frequency image details by sampling at smaller time steps. We further introduce a time-aware discriminator that exploits the teacher’s knowledge to differentiate real and synthetic samples across different noise scales, using explicit temporal conditioning. Extensive experiments on SR tasks demonstrate that TAD-SR outperforms existing singl-estep diffusion methods and achieves performance on par with multi-step state-of-the-art models.iffusion-based image super-resolution (SR) has shown strong potential in recovering highfidelity details from low-resolution inputs. However, the need for tens or hundreds of sampling steps leads to substantial inference latency. Recent works attempt to accelerate this process via knowledge distillation, but often rely solely on pixel-level loss or overlook the fact that diffusion models capture different information across time steps. To address this, we propose TADSR, a time-aware diffusion distillation framework. Specifically, we introduce a novel score distillation strategy to align the score functions between the outputs of the student and teacher models after minor noise perturbation. This distillation strategy eliminates the inherent bias in score distillation sampling (SDS) and enables the student models to focus more on highf-requency image details by sampling at smaller time steps. We further introduce a time-aware discriminator that exploits the teacher’s knowledge to differentiate real and synthetic samples across different noise scales, using explicit temporal conditioning. Extensive experiments on SR tasks demonstrate that TAD-SR outperforms existing single-step diffusion methods and achieves performance on par with multi-step state-of-the-art models D. Xiao He 0014, Huaao Tang, Zhijun Tu, Hanting Chen, Mingrui Zhu, Jie Hu 0021, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Effective Diffusion Transformer Architecture for Image Super-ResolutionabstractRecent advances indicate that diffusion model holds great promise in image super-resolution. While latest methods are primarily based on latent diffusion models with convolutional neural networks, there are few attempts to explore transformers, which have demonstrated remarkable performance in image generation. In this work, we design an effective diffusion transformer for image super resolution (DiT-SR) that achieves the visual quality of prior-based methods, but through a training-from-scratch manner. In practice, DiT-SR leverages an overall U-shaped architecture, and adopts uniform isotropic design for all the transformer blocks across different stages. The former facilitates multi-scale hierarchical feature extraction, while the latter reallocate the computational resources to critical layers to further enhance performance. Moreover, we thoroughly analyze the limitation of the widely used AdaLN, and present a frequency-adaptive time-step conditioning module, enhancing the model's capacity to process distinct frequency information at different time steps. Extensive experiments demonstrate that DiT-SR outperforms the existing training-from-scratch diffusion-based SR methods significantly, and even beats some of the prior-based methods on pretrained Stable Diffusion, proving the superiority of diffusion transformer in image super resolution. Zhijun Tu, Xiao He 0014, Liyu Chen, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001, Jie Hu 0021 |
AAAI | 4 |
| 2025 | Diff-MoE: Diffusion Transformer with Time-Aware and Space-Adaptive ExpertsabstractDiffusion models have transformed generative modeling but suffer from scalability limitations due to computational overhead and inflexible architectures that process all generative stages and tokens uniformly. In this work, we introduce Diff-MoE, a novel framework that combines Diffusion Transformers with Mixture-of-Experts to exploit both temporarily adaptability and spatial flexibility. Our design incorporates expert-specific timestep conditioning, allowing each expert to process different spatial tokens while adapting to the generative stage, to dynamically allocate resources based on both the temporal and spatial characteristics of the generative task. Additionally, we propose a globally-aware feature recalibration mechanism that amplifies the representational capacity of expert modules by dynamically adjusting feature contributions based on input relevance. Extensive experiments on image generation benchmarks demonstrate that Diff-MoE significantly outperforms state-of-the-art methods. Our work demonstrates the potential of integrating diffusion models with expert-based designs, offering a scalable and effective framework for advanced generative modeling. Xiao He 0014, Zhijun Tu, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001, Jie Hu 0021 |
ICML | 2 |
| 2024 | Diff-Privacy: Diffusion-Based Face Privacy ProtectionabstractPrivacy protection has become a top priority due to the widespread collection and misuse of personal data. Anonymization and visual identity information hiding are two crucial tasks in face privacy protection, both striving to alter identifying characteristics from face images to prevent privacy information leakage. However, the goals of the two are not entirely the same. Consequently, training a model to simultaneously perform both tasks proves challenging. In this paper, we propose Diff-Privacy, a novel face privacy protection method based on diffusion models that unifies the task of anonymization and visual identity information hiding. Specifically, we present a Multi-Scale image Inversion module (MSI) that, through training, generates a set of Stable Diffusion (SD) format conditional embeddings for the original image. With these conditional embeddings, we design corresponding embedding scheduling strategies and formulate distinct energy functions during the inference process to achieve anonymization and visual identity information hiding, respectively. Extensive experiments demonstrate the effectiveness of the proposed method in protecting face privacy. Xiao He 0014, Mingrui Zhu, Dongxin Chen, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Few-Shot Font Generation by Learning Style Difference and SimilarityabstractFew-shot font generation (FFG) aims to preserve the underlying global structure of the original character while generating target fonts by referring to a few samples. It has been applied to font library creation, a personalized signature, and other scenarios. Existing FFG methods explicitly disentangle content and style of reference glyphs universally or component-wisely. However, they ignore the difference between glyphs in different styles and the similarity of glyphs in the same style, which results in artifacts such as local distortions and style inconsistency. To address this issue, we propose a novel font generation approach by learning the Difference between different styles and the Similarity of the same style (DS-Font). We introduce contrastive learning to consider the positive and negative relationship between styles. Specifically, we propose a multi-layer style projector (MSP) for style encoding and realize a distinctive style representation via our proposed Cluster-level Contrastive Style (CCS) loss. The MSP module is employed to assist the generator during training to enhance the style consistency between the generated glyph and the reference glyphs. In addition, we design a glyph-independent patch discriminator, which comprehensively considers different areas of the image and ensures that each style can be distinguished independently. We conduct qualitative and quantitative evaluations comprehensively to demonstrate that our approach achieves significantly better results than state-of-the-art methods. Xiao He 0014, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | All-to-key Attention for Arbitrary Style TransferabstractAttention-based arbitrary style transfer studies have shown promising performance in synthesizing vivid local style details. They typically use the all-to-all attention mechanism—each position of content features is fully matched to all positions of style features. However, all-to-all attention tends to generate distorted style patterns and has quadratic complexity, limiting the effectiveness and efficiency of arbitrary style transfer. In this paper, we propose a novel all-to-key attention mechanism—each position of content features is matched to stable key positions of style features—that is more in line with the characteristics of style transfer. Specifically, it integrates two newly proposed attention forms: distributed and progressive attention. Distributed attention assigns attention to key style representations that depict the style distribution of local regions; Progressive attention pays attention from coarse-grained regions to fine-grained key positions. The resultant module, dubbed StyA2K, shows extraordinary performance in preserving the semantic structure and rendering consistent style patterns. Qualitative and quantitative comparisons with state-of-the-art methods demonstrate the superior performance of our approach. Codes and models are available on https://github.com/LearningHx/StyA2K. Mingrui Zhu, Xiao He 0014, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001 |
ICCV | 2 |
| 2023 | BiTGAN: bilateral generative adversarial networks for Chinese ink wash painting style transfer
Xiao He 0014, Mingrui Zhu, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001 |
Sci. China Inf. Sci. | 1 |