Yunqi Miao

dblp:207/3366 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 4 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Face, body and person analysis · 43% Generative modeling · 29% Image recognition and object detection · 12%
Computer graphics and multimedia
5 papers
Image and video processing · 81% Visual content generation and editing · 13% Rendering · 5%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.622025
CaricatureBooth: Data-Free Interactive Caricature Generation in a Photo Booth · CVPR 2025
WaveFace: Authentic Face Restoration with Efficient Frequency Recovery · CVPR 2024
Image and video processing › image restoration › face restoration
blind face restoration
1.622025
Unlocking the Potential of Diffusion Priors in Blind Face Restoration · ICCV 2025
WaveFace: Authentic Face Restoration with Efficient Frequency Recovery · CVPR 2024
Image and video processing
image restoration
1.622025
Unlocking the Potential of Diffusion Priors in Blind Face Restoration · ICCV 2025
WaveFace: Authentic Face Restoration with Efficient Frequency Recovery · CVPR 2024
Visual content generation and editing › image generation › stylized image generation
caricature generation
0.912025
CaricatureBooth: Data-Free Interactive Caricature Generation in a Photo Booth · CVPR 2025
Image and video processing › image restoration
degradation modeling
0.912025
Unlocking the Potential of Diffusion Priors in Blind Face Restoration · ICCV 2025
Computer vision › Face, body and person analysis
person re-identification
0.812024
Confidence-Guided Centroids for Unsupervised Person Re-Identification · IEEE Trans. Inf. Forensics Secur. 2024
Computer vision › Segmentation and scene understanding › pseudo-label learning
pseudo-label refinement
0.812024
Confidence-Guided Centroids for Unsupervised Person Re-Identification · IEEE Trans. Inf. Forensics Secur. 2024
Computer vision › Face, body and person analysis › person re-identification
unsupervised person re-identification
0.812024
Confidence-Guided Centroids for Unsupervised Person Re-Identification · IEEE Trans. Inf. Forensics Secur. 2024
Computer vision › Face, body and person analysis › face recognition
cross-spectral face recognition
0.612022
Physically-Based Face Rendering for NIR-VIS Face Recognition · NeurIPS 2022
Computer vision › Face, body and person analysis
face recognition
0.612022
Physically-Based Face Rendering for NIR-VIS Face Recognition · NeurIPS 2022
Machine learning › Generative modeling
face synthesis
0.612022
Physically-Based Face Rendering for NIR-VIS Face Recognition · NeurIPS 2022
Computer vision › Face, body and person analysis › face recognition › cross-spectral face recognition
NIR-VIS face recognition
0.612022
Physically-Based Face Rendering for NIR-VIS Face Recognition · NeurIPS 2022
Image and video processing › texture analysis
local binary pattern
0.512021
Learning Transformation-Invariant Local Descriptors With Low-Coupling Binary Codes · IEEE Trans. Image Process. 2021
Image and video processing › feature extraction › feature descriptor
local feature descriptor
0.512021
Learning Transformation-Invariant Local Descriptors With Low-Coupling Binary Codes · IEEE Trans. Image Process. 2021
Computer vision › Image recognition and object detection › object counting
crowd counting
0.412020
Shallow Feature Based Dense Attention Network for Crowd Counting · AAAI 2020
Computer vision › Image recognition and object detection › object counting › crowd counting
density map estimation
0.412020
Shallow Feature Based Dense Attention Network for Crowd Counting · AAAI 2020
Machine learning › Deep learning architectures and training
multi-scale feature fusion
0.412020
Shallow Feature Based Dense Attention Network for Crowd Counting · AAAI 2020
Rendering
face rendering
0.212022
Physically-Based Face Rendering for NIR-VIS Face Recognition · NeurIPS 2022
Rendering
physically based rendering
0.212022
Physically-Based Face Rendering for NIR-VIS Face Recognition · NeurIPS 2022
Image and video processing
image matching
0.112021
Learning Transformation-Invariant Local Descriptors With Low-Coupling Binary Codes · IEEE Trans. Image Process. 2021

Methods — techniques the papers use, named apart from their topics

diffusion model · 4.1thin plate spline deformation · 1.7wavelet transformation · 1.5physically-based rendering · 1.1identity-based maximum mean discrepancy loss · 1.13d face reconstruction · 1.1face embedding · 0.9confidence-guided centroids · 0.8clustering · 0.8wasserstein loss · 0.5unsupervised learning · 0.5adversarial constraint module · 0.5convolutional neural network · 0.4attention mechanism · 0.4
YearPublicationVenuePosition
2025 CaricatureBooth: Data-Free Interactive Caricature Generation in a Photo Booth
abstract
We present CaricatureBooth, a system that transforms caricature creation into a simple interactive experience – as easy as using a photo booth! A key challenge in caricature generation is two-fold: the scarcity of high-quality caricature data and the difficulty in enabling precise creative control over the exaggeration process while maintaining identity. Prior approaches either require large-scale caricature and photo data or lack intuitive mechanisms for users to guide the deformation without losing identity. We address the data scarcity by synthesising training data through Thin Plate Spline (TPS) deformation of standard face images. For creative control, we design a Bézier curve interface where users can easily manipulate facial features, with these edits then driving TPS transformations at inference time. When combined with a pre-trained ID-preserving diffusion model, our system maintains both identity preservation and creative flexibility. Through extensive experiments, we demonstrate that CaricatureBooth achieves state-of-the-art quality while making the joy of caricature creation as accessible as taking a photo – just walk in and walk out with your personalised caricature! Code is available at https://github.com/WinKawaks/CaricatureBooth.
Zhiyu Qu, Yunqi Miao, Zhensong Zhang, Jifei Song, Jiankang Deng, Yi-Zhe Song
CVPR2
2025 Unlocking the Potential of Diffusion Priors in Blind Face Restoration
abstract
Although diffusion prior is rising as a powerful solution for blind face restoration (BFR), the inherent gap between the vanilla diffusion model and BFR settings hinders its seamless adaptation. The gap mainly stems from the discrepancy between 1) high-quality (HQ) and low-quality (LQ) images and 2) synthesized and real-world images. The vanilla diffusion model is trained on images with no or less degradations, whereas BFR handles moderately to severely degraded images. Additionally, LQ images used for training are synthesized by a naive degradation model with limited degradation patterns, which fails to simulate complex and unknown degradations in real-world scenarios. In this work, we use a unified network FLIPNET that switches between two modes to resolve specific gaps. In Restoration mode, the model gradually integrates BFR-oriented features and face embeddings from LQ images to achieve authentic and faithful face restoration. In Degradation mode, the model synthesizes real-world like degraded images based on the knowledge learned from real-world degradation datasets. Extensive evaluations on benchmark datasets show that our model 1) outperforms previous diffusion prior based BFR methods in terms of authenticity and fidelity, and 2) outperforms the naive degradation model in modeling the real-world degradations.
Yunqi Miao, Zhiyu Qu, Mingqi Gao 0003, Changrui Chen, Jifei Song, Jungong Han, Jiankang Deng
ICCV1
2024 WaveFace: Authentic Face Restoration with Efficient Frequency Recovery
abstract
Although diffusion models are rising as a powerful solution for blind face restoration, they are criticized for two problems: 1) slow training and inference speed, and 2)failure in preserving identity and recovering fine-grained facial details. In this work, we propose WaveFace to solve the problems in the frequency domain, where low- and high-frequency components decomposed by wavelet transformation are considered individually to maximize authenticity as well as efficiency. The diffusion model is applied to recover the low-frequency component only, which presents general information of the original image but 1/16 in size. To preserve the original identity, the generation is conditioned on the low-frequency component of low-quality images at each denoising step. Meanwhile, high-frequency components at multiple decomposition levels are handled by a unified network, which recovers complex facial details in a single step. Evaluations on four benchmark datasets show that: 1) WaveFace outperforms state-of-the-art methods in authenticity, especially in terms of identity preservation, and 2) authentic images are restored with the efficiency 10x faster than existing diffusion model-based BFR methods.
Yunqi Miao, Jiankang Deng, Jungong Han
CVPR1
2024 Confidence-Guided Centroids for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (ReID) aims to train a feature extractor for identity retrieval without exploiting identity labels. Due to the no-reference trust in imperfect clustering results, the learning is inevitably misled by unreliable pseudo labels. Albeit the pseudo label refinement has been investigated by previous works, they generally leverage auxiliary information such as camera IDs and body part predictions. This work explores the internal characteristics of clusters to refine pseudo labels. To this end, Confidence-Guided Centroids (CGC) are proposed to provide reliable cluster-wise prototypes for feature learning. Since samples with high confidence are exclusively involved in the formation of centroids, the identity information of low-confidence samples, i.e., boundary samples, are NOT likely to contribute to the corresponding centroid. Given the new centroids, the current learning scheme, where samples are forced to learn from their assigned centroids solely, is unwise. To remedy the situation, we propose to use Confidence-Guided pseudo Label (CGL), which enables samples to approach not only the originally assigned centroid but also other centroids that are potentially embedded with their identity information. Empowered by confidence-guided centroids and labels, our method yields comparable performance with, or even outperforms, state-of-the-art pseudo label refinement works that largely leverage auxiliary information.
Yunqi Miao, Jiankang Deng, Guiguang Ding, Jungong Han
IEEE Trans. Inf. Forensics Secur.1
2023 On exploring pose estimation as an auxiliary learning task for Visible-Infrared Person Re-identification
Yunqi Miao, Nianchang Huang, Xiao Ma 0013, Qiang Zhang 0020, Jungong Han
Neurocomputing1
2022 Physically-Based Face Rendering for NIR-VIS Face Recognition
abstract
Near infrared (NIR) to Visible (VIS) face matching is challenging due to the significant domain gaps as well as a lack of sufficient data for cross-modality model training. To overcome this problem, we propose a novel method for paired NIR-VIS facial image generation. Specifically, we reconstruct 3D face shape and reflectance from a large 2D facial dataset and introduce a novel method of transforming the VIS reflectance to NIR reflectance. We then use a physically-based renderer to generate a vast, high-resolution and photorealistic dataset consisting of various poses and identities in the NIR and VIS spectra. Moreover, to facilitate the identity feature learning, we propose an IDentity-based Maximum Mean Discrepancy (ID-MMD) loss, which not only reduces the modality gap between NIR and VIS images at the domain level but encourages the network to focus on the identity features instead of facial details, such as poses and accessories. Extensive experiments conducted on four challenging NIR-VIS face recognition benchmarks demonstrate that the proposed method can achieve comparable performance with the state-of-the-art (SOTA) methods without requiring any existing NIR-VIS face recognition datasets. With slightly fine-tuning on the target NIR-VIS face recognition datasets, our method can significantly surpass the SOTA performance. Code and pretrained models are released under the insightface GitHub.
Yunqi Miao, Alexander Lattas, Jiankang Deng, Jungong Han, Stefanos Zafeiriou
NeurIPS1
2021 Learning Transformation-Invariant Local Descriptors With Low-Coupling Binary Codes
abstract
Despite the great success achieved by prevailing binary local descriptors, they are still suffering from two problems: 1) vulnerable to the geometric transformations; 2) lack of an effective treatment to the highly-correlated bits that are generated by directly applying the scheme of image hashing. To tackle both limitations, we propose an unsupervised Transformation-invariant Binary Local Descriptor learning method (TBLD). Specifically, the transformation invariance of binary local descriptors is ensured by projecting the original patches and their transformed counterparts into an identical high-dimensional feature space and an identical low-dimensional descriptor space simultaneously. Meanwhile, it enforces the dissimilar image patches to have distinctive binary local descriptors. Moreover, to reduce high correlations between bits, we propose a bottom-up learning strategy, termed Adversarial Constraint Module, where low-coupling binary codes are introduced externally to guide the learning of binary local descriptors. With the aid of the Wasserstein loss, the framework is optimized to encourage the distribution of the generated binary local descriptors to mimic that of the introduced low-coupling binary codes, eventually making the former more low-coupling. Experimental results on three benchmark datasets well demonstrate the superiority of the proposed method over the state-of-the-art methods. The project page is available at https://github.com/yoqim/TBLD.
Yunqi Miao, Zijia Lin, Xiao Ma 0013, Guiguang Ding, Jungong Han
IEEE Trans. Image Process.1
2020 Shallow Feature Based Dense Attention Network for Crowd Counting
abstract
While the performance of crowd counting via deep learning has been improved dramatically in the recent years, it remains an ingrained problem due to cluttered backgrounds and varying scales of people within an image. In this paper, we propose a Shallow feature based Dense Attention Network (SDANet) for crowd counting from still images, which diminishes the impact of backgrounds via involving a shallow feature based attention model, and meanwhile, captures multi-scale information via densely connecting hierarchical image features. Specifically, inspired by the observation that backgrounds and human crowds generally have noticeably different responses in shallow features, we decide to build our attention model upon shallow-feature maps, which results in accurate background-pixel detection. Moreover, considering that the most representative features of people across different scales can appear in different layers of a feature extraction network, to better keep them all, we propose to densely connect hierarchical image features of different layers and subsequently encode them for estimating crowd density. Experimental results on three benchmark datasets clearly demonstrate the superiority of SDANet when dealing with different scenarios. Particularly, on the challenging UCF_CC_50 dataset, our method outperforms other existing methods by a large margin, as is evident from a remarkable 11.9% Mean Absolute Error (MAE) drop of our SDANet.
Yunqi Miao, Zijia Lin, Guiguang Ding, Jungong Han
AAAI1
2019 Convolutional Attention in Ensemble With Knowledge Transferred for Remote Sensing Image Classification
abstract
Ensemble learning is one of the hottest topics in machine learning. In this letter, we develop a convolutional attention in ensemble (CAE) method, which, for the first time, introduces attention-based weighting scheme into ensemble learning. The knowledge contained in base classifiers is transferred into the final classifier, by which the base classifier with a higher performance could be given much more attention. In particular, we employ convolutional attention models to develop an efficient ensemble classifier for image classification. Our CAE can leverage the representation capacity of convolutional neural networks to enhance the performance of ensemble classifiers. We apply our method to remote sensing image classification tasks, which achieves much better performance than the state of the arts.
Hainan Wang, Yunqi Miao, Hongren Wang 0001, Baochang Zhang 0001
IEEE Geosci. Remote. Sens. Lett.2
2019 ST-CNN: Spatial-Temporal Convolutional Neural Network for crowd counting in videos
Yunqi Miao, Jungong Han, Yongsheng Gao 0001, Baochang Zhang 0001
Pattern Recognit. Lett.1
2018 The random boosting ensemble classifier for land-use image classification
Hainan Wang, Yunqi Miao
Multim. Tools Appl.2