VLDB 2026 Research / reviewers in the wild / expert
Qichao Sun
dblp:83/524
· DBLP profile ↗
7ranked-venue papers
2as first author
2since 2021 · last 2025
0009-0008-0598-5693ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 100% | |
| Artificial intelligence
2 papers |
Generative modeling · 67% Face, body and person analysis · 33% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Visual content generation and editing
image generation |
1.7 | 2 | 2025 | DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning · SIGGRAPH Asia 2025 AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models · CVPR 2025 |
Machine learning › Generative modeling › diffusion model › controllable generation
controllable image generation |
0.9 | 1 | 2025 | AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models · CVPR 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models · CVPR 2025 |
Computer vision › Face, body and person analysis › face manipulation
face swapping |
0.9 | 1 | 2025 | DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning · SIGGRAPH Asia 2025 |
Visual content generation and editing › controllable generation
identity-preserving generation |
0.9 | 1 | 2025 | DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning · SIGGRAPH Asia 2025 |
Visual content generation and editing
virtual try-on |
0.9 | 1 | 2025 | AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
triplet ID group learning · 1.7latent diffusion · 1.7feature extraction · 1.7diffusion model · 1.7attention mechanism · 1.7ID adapter · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion ModelsabstractRecent advances in garment-centric image generation from text and image prompts based on diffusion models are impressive. However, existing methods lack support for various combinations of attire, and struggle to preserve the garment details while maintaining faithfulness to the text prompts, limiting their performance across diverse scenarios. In this paper, we focus on a new task, i.e., Multi-Garment Virtual Dressing, and we propose a novel Any-Dressing method for customizing characters conditioned on any combination of garments and any personalized text prompts. AnyDressing comprises two primary networks named GarmentsNet and DressingNet, which are respectively dedicated to extracting detailed clothing features and generating customized images. Specifically, we propose an efficient and scalable module called Garment-Specific Feature Extractor in GarmentsNet to individually encode garment textures in parallel. This design prevents garment confusion while ensuring network efficiency. Meanwhile, we design an adaptive Dressing-Attention mechanism and a novel Instance-Level Garment Localization Learning strategy in DressingNet to accurately inject multi-garment features into their corresponding regions. This approach efficiently integrates multi-garment texture cues into generated images and further enhances text-image consistency. Additionally, we introduce a Garment-Enhanced Texture Learning strategy to improve the fine-grained texture details of garments. Thanks to our well-craft design, Any-Dressing can serve as a plug-in module to easily integrate with any community control extensions for diffusion models, improving the diversity and controllability of synthesized images. Extensive experiments show that AnyDressing achieves state-of-the-art results. Xinghui Li, Qichao Sun, Pengze Zhang, Fulong Ye, Zhichao Liao, Wanquan Feng, Songtao Zhao |
CVPR | 2 |
| 2025 | DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group LearningabstractIn this paper, we introduce DreamID, a diffusion-based face swapping model that achieves high levels of ID similarity, attribute preservation, image fidelity, and fast inference speed. Unlike the typical face swapping training process, which often relies on implicit supervision and struggles to achieve satisfactory results. DreamID establishes explicit supervision for face swapping by constructing Triplet ID Group data, significantly enhancing identity similarity and attribute preservation. The iterative nature of diffusion models poses challenges for utilizing efficient image-space loss functions, as performing time-consuming multi-step sampling to obtain the generated image during training is impractical. To address this issue, we leverage the accelerated diffusion model SD Turbo, reducing the inference steps to a single iteration, enabling efficient pixel-level end-to-end training with explicit Triplet ID Group supervision. Additionally, we propose an improved diffusion-based model architecture comprising SwapNet, FaceNet, and ID Adapter. This robust architecture fully unlocks the power of the Triplet ID Group explicit supervision. Finally, to further extend our method, we explicitly modify the Triplet ID Group data during training to fine-tune and preserve specific attributes, such as glasses and face shape. Extensive experiments demonstrate that DreamID outperforms state-of-the-art methods in terms of identity similarity, pose and expression preservation, and image fidelity. Overall, DreamID achieves high-quality face swapping results at 512×512 resolution in just 0.6 seconds and performs exceptionally well in challenging scenarios such as complex lighting, large angles, and occlusions. Our project: https://superhero-7.github.io/DreamID/. Fulong Ye, Miao Hua, Pengze Zhang, Xinghui Li, Qichao Sun, Songtao Zhao |
SIGGRAPH Asia | 5 |
| 2008 | L-shaped segmentations in motion-compensated prediction of H.264abstractPlaying the role of removing temporal correlation in video signals, motion-compensated prediction aims at the ideal prediction and eliminates the compensation. However, the block-based video coding infrastructure at present was not competent in exactly present the various shapes of moving objects in actual video. Contrast to that, an exact segmentation of coding area according to the actual moving objects in video is also not easy to achieve and will introduce significant computational complexity into the video coding scheme. Concerned with this conflict, this paper proposes a new set of segmentations of macroblocks, in which one macroblock splits into one L-shaped and one square segment. Experiments of this named "L-shaped segmentations" technique show that it can outperform the H.264 P-picture coding by 0.11 dB. Qichao Sun, Xiaoyang Wu 0001, Lu Yu 0003 |
ISCAS | 2 |
| 2007 | Modeling Natural Image for Estimating DCT Coefficient Properties of Intra PredictionabstractAn elliptical symmetric model in spatial domain of natural image is proposed in this paper. Based on the model, the statistical properties of intra prediction error are studied. Furthermore, using the linear characteristic of Discrete Cosine Transform (DCT) and the proposed model, we offer a mathematical analysis of the properties of DCT coefficient for intra prediction error. The horizontal intra prediction mode is selected as the example. DCT coefficient properties of the rest prediction modes can be derived by the same method. Xiaoyang Wu 0001, Qichao Sun, Lu Yu 0003 |
ICME | 2 |
| 2007 | A Content-adaptive Fast Multiple Reference Frames Motion Estimation in H.264abstractThe H.264/AVC video coding standard adopts various advanced coding techniques such as variable block sizes and multiple reference frames motion compensation (MRF-MC). The computational burden of motion estimation increases as the number of reference frames searched. In this paper, we present a content-adaptive fast motion estimation algorithm. This algorithm fully uses the temporal and spatial content information at frame, MB and block levels to speed up the search process for multiple reference frames. Simulation results show that the proposed algorithm can be 3.6 times faster than the multiple reference frames UMHexagonS scheme, with negligible degradation of video quality. Qichao Sun, Xin-Hao Chen, Xiaoyang Wu 0001, Lu Yu 0003 |
ISCAS | 1 |
| 2007 | An efficient multi-frame dynamic search range motion estimation for H.264abstractH.264/AVC achieves higher compression efficiency than previous video coding standards. However, this comes at the cost of increased complexity due to the use of variable block size motion estimation and long-term memory motion compensated prediction (LTMCP). In this paper, an efficient multi-frame dynamic search range motion estimation algorithm is proposed. This algorithm can adjust the spatial search range and temporal search range according to the video content dynamically. This algorithm can be on the top of many other fast motion estimation (Fast ME) algorithms. Compared with the constant search range scheme used by multi-frame UMHexagonS algorithm, the proposed algorithm can be 4.86 time faster, with negligible degradation of video quality. Qichao Sun, Xin-Hao Chen |
VCIP | 1 |
| 2006 | Low-Complexity Tools in AVS Part 7
Feng Yi, Qichao Sun, Jie Dong 0001, Lu Yu 0003 |
J. Comput. Sci. Technol. | 2 |