Miao Hua

dblp:140/0202 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0000-8249-0899ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning
abstract
In this paper, we introduce DreamID, a diffusion-based face swapping model that achieves high levels of ID similarity, attribute preservation, image fidelity, and fast inference speed. Unlike the typical face swapping training process, which often relies on implicit supervision and struggles to achieve satisfactory results. DreamID establishes explicit supervision for face swapping by constructing Triplet ID Group data, significantly enhancing identity similarity and attribute preservation. The iterative nature of diffusion models poses challenges for utilizing efficient image-space loss functions, as performing time-consuming multi-step sampling to obtain the generated image during training is impractical. To address this issue, we leverage the accelerated diffusion model SD Turbo, reducing the inference steps to a single iteration, enabling efficient pixel-level end-to-end training with explicit Triplet ID Group supervision. Additionally, we propose an improved diffusion-based model architecture comprising SwapNet, FaceNet, and ID Adapter. This robust architecture fully unlocks the power of the Triplet ID Group explicit supervision. Finally, to further extend our method, we explicitly modify the Triplet ID Group data during training to fine-tune and preserve specific attributes, such as glasses and face shape. Extensive experiments demonstrate that DreamID outperforms state-of-the-art methods in terms of identity similarity, pose and expression preservation, and image fidelity. Overall, DreamID achieves high-quality face swapping results at 512×512 resolution in just 0.6 seconds and performs exceptionally well in challenging scenarios such as complex lighting, large angles, and occlusions. Our project: https://superhero-7.github.io/DreamID/.
Fulong Ye, Miao Hua, Pengze Zhang, Xinghui Li, Qichao Sun, Songtao Zhao
SIGGRAPH Asia2
2024 Key-point-guided adaptive convolution and instance normalization for continuous transitive face reenactment of any person
abstract
Abstract Face reenactment technology is widely applied in various applications. However, the reconstruction effects of existing methods are often not quite realistic enough. Thus, this paper proposes a progressive face reenactment method. First, to make full use of the key information, we propose adaptive convolution and instance normalization to encode the key information into all learnable parameters in the network, including the weights of the convolution kernels and the means and variances in the normalization layer. Second, we present continuous transitive facial expression generation according to all the weights of the network generated by the key points, resulting in the continuous change of the image generated by the network. Third, in contrast to classical convolution, we apply the combination of depth‐ and point‐wise convolutions, which can greatly reduce the number of weights and improve the efficiency of training. Finally, we extend the proposed face reenactment method to the face editing application. Comprehensive experiments demonstrate the effectiveness of the proposed method, which can generate a clearer and more realistic face from any person and is more generic and applicable than other methods.
Shibiao Xu, Miao Hua, Jiguang Zhang, Zhaohui Zhang 0002, Xiaopeng Zhang 0001
Comput. Animat. Virtual Worlds2
2023 ReGANIE: Rectifying GAN Inversion Errors for Accurate Real Image Editing
abstract
The StyleGAN family succeed in high-fidelity image generation and allow for flexible and plausible editing of generated images by manipulating the semantic-rich latent style space. However, projecting a real image into its latent space encounters an inherent trade-off between inversion quality and editability. Existing encoder-based or optimization-based StyleGAN inversion methods attempt to mitigate the trade-off but see limited performance. To fundamentally resolve this problem, we propose a novel two-phase framework by designating two separate networks to tackle editing and reconstruction respectively, instead of balancing the two. Specifically, in Phase I, a W-space-oriented StyleGAN inversion network is trained and used to perform image inversion and edit- ing, which assures the editability but sacrifices reconstruction quality. In Phase II, a carefully designed rectifying network is utilized to rectify the inversion errors and perform ideal reconstruction. Experimental results show that our approach yields near-perfect reconstructions without sacrificing the editability, thus allowing accurate manipulation of real images. Further, we evaluate the performance of our rectifying net- work, and see great generalizability towards unseen manipulation types and out-of-domain images.
Bingchuan Li, Tianxiang Ma, Miao Hua, Zili Yi
AAAI4
2023 CFFT-GAN: Cross-Domain Feature Fusion Transformer for Exemplar-Based Image Translation
abstract
Exemplar-based image translation refers to the task of generating images with the desired style, while conditioning on certain input image. Most of the current methods learn the correspondence between two input domains and lack the mining of information within the domain. In this paper, we propose a more general learning approach by considering two domain features as a whole and learning both inter-domain correspondence and intra-domain potential information interactions. Specifically, we propose a Cross-domain Feature Fusion Transformer (CFFT) to learn inter- and intra-domain feature fusion. Based on CFFT, the proposed CFFT-GAN works well on exemplar-based image translation. Moreover, CFFT-GAN is able to decouple and fuse features from multiple domains by cascading CFFT modules. We conduct rich quantitative and qualitative experiments on several image translation tasks, and the results demonstrate the superiority of our approach compared to state-of-the-art methods. Ablation studies show the importance of our proposed CFFT. Application experimental results reflect the potential of our method.
Tianxiang Ma, Bingchuan Li, Wei Liu 0035, Miao Hua, Jing Dong 0003, Tieniu Tan
AAAI4
2023 DyStyle: Dynamic Neural Network for Multi-Attribute-Conditioned Style Editings
abstract
The semantic controllability of StyleGAN is enhanced by unremitting research. Although the existing weak supervision methods work well in manipulating the style codes along one attribute, the accuracy of manipulating multiple attributes is neglected. Multi-attribute representations are prone to entanglement in the StyleGAN latent space, while sequential editing leads to error accumulation. To address these limitations, we design a Dynamic Style Manipulation Network (DyStyle) whose structure and parameters vary by input samples, to perform nonlinear and adaptive manipulation of latent codes for flexible and precise attribute control. In order to efficient and stable optimization of the DyStyle network, we propose a Dynamic Multi-Attribute Contrastive Learning (DmaCL) method: including dynamic multi-attribute contrastor and dynamic multi-attribute contrastive loss, which simultaneously disentangle a variety of attributes from the generative image and latent space of model. As a result, our approach demonstrates fine-grained disentangled edits along multiple numeric and binary attributes. Qualitative and quantitative comparisons with existing style manipulation methods verify the superiority of our method in terms of the multi-attribute control accuracy and identity preservation without compromising photorealism.
Bingchuan Li, Shaofei Cai, Miao Hua, Zili Yi
WACV6
2022 Region-Aware Face Swapping
abstract
This paper presents a novel Region-Aware Face Swapping (RAFSwap) network to achieve identity-consistent harmonious high-resolution face generation in a local-global manner: 1) Local Facial Region-Aware (FRA) branch augments local identity-relevant features by introducing the Transformer to effectively model misaligned crossscale semantic interaction. 2) Global Source Feature-Adaptive (SFA) branch further complements global identity-relevant cues for generating identity-consistent swapped faces. Besides, we propose a Face Mask Predictor (FMP) module incorporated with StyleGAN2 to predict identity-relevant soft facial masks in an unsupervised manner that is more practical for generating harmonious high-resolution faces. Abundant experiments qualitatively and quantitatively demonstrate the superiority of our method for generating more identity-consistent high-resolution swapped faces over SOTA methods, e.g., obtaining 96.70 ID retrieval that outperforms SOTA MegaFS by$5.87\uparrow$.
Chao Xu 0023, Jiangning Zhang, Miao Hua, Zili Yi, Yong Liu 0007
CVPR3
2016 Enhanced Use of Mattes for Easy Image Composition
abstract
Existing matting methods focus on improving matte quality to produce high-quality composites. This generally requires significant manual interaction, a tedious task for the user. Despite these efforts, the composites may still exhibit evident artifacts, especially in the case of transparent and complicated objects as their related pixels always contain percentage of the background. In this paper, we focus on the enhanced use of mattes to produce satisfactory composites by suppressing the discrepancies around objects of interest. This approach is motivated by cloning methods but overcomes their shortcoming of ineffective treatment of the over-included regions around objects of interest. For this, we present an enhanced matting function by including a term to smooth the local contrasts for seamless composition, and meanwhile, we develop a novel algorithm to generate mattes with reduced user interaction and improved usability. As a result, we reduce the composite's dependence on the user's input and only require the user to drag a box to enclose the objects of interest. As shown in the user studies and the experimental results, our method requires many times less user interaction than the existing matting methods and cloning methods. Our method is more effective in producing good composites in a simple interactive manner, especially when treating transparent and complicated objects, thereby providing a superior approach for image composition.
Wencheng Wang 0001, Xiaohui Bie, Miao Hua
IEEE Trans. Image Process.4
2015 Distinguishing Local and Global Edits for Their Simultaneous Propagation in a Uniform Framework
abstract
In propagating edits for image editing, some edits are intended to affect limited local regions, while others act globally over the entire image. However, the ambiguity problem in propagating edits is not adequately addressed in existing methods. Thus, tedious user input requirements remain since the user must densely or repeatedly input control samples to suppress ambiguity. In this paper, we address this challenge to propagate edits suitably by marking edits for local or global propagation and determining their reasonable propagation scopes automatically. Thus, our approach avoids propagation conflicts, effectively resolving the ambiguity problem. With the reduction of ambiguity, our method allows fewer and less-precise control samples than existing methods. Furthermore, we provide a uniform framework to propagate local and global edits simultaneously, helping the user to quickly obtain the intended results with reduced labor. With our unified framework, the potentially ambiguous interaction between local and global edits (evident in existing methods that propagate these two edit types in series) is resolved. We experimentally demonstrate the effectiveness of our method compared with existing methods.
Wencheng Wang 0001, Miao Hua, Minying Zhang, Xiaohui Bie
IEEE Trans. Image Process.4
2015 Effective structure restoration for image completion using internet resources
Miao Hua, Wencheng Wang 0001
Vis. Comput.1
2014 Edge-Aware Gradient Domain Optimization Framework for Image Filtering by Local Propagation
abstract
Gradient domain methods are popular for image processing. However, these methods even the edge-preserving ones cannot preserve edges well in some cases. In this paper, we present new constraints explicitly to better preserve edges for general gradient domain image filtering and theoretically analyse why these constraints are edge-aware. Our edge-aware constraints are easy to implement, fast to compute and can be seamlessly integrated into the general gradient domain optimization framework. The improved framework can better preserve edges while maintaining similar image filtering effects as the original image filters. We also demonstrate the strength of our edge-aware constraints on various applications such as image smoothing, image colorization and Poisson image cloning.
Miao Hua, Xiaohui Bie, Minying Zhang, Wencheng Wang 0001
CVPR1
2013 Extracting Dominant Textures in Real Time With Multi-Scale Hue-Saturation-Intensity Histograms
abstract
It is very important to extract high quality texture features from images. This is, however, often laborious, because the randomness in the color distribution patterns for texture elements makes texture measurement very difficult, despite these elements having a very similar visual appearance. In this paper, we propose the use of multi-scale color histograms to measure the effect of color distribution patterns efficiently and without having to compute the actual patterns, which saves considerable effort. Meanwhile, the hue-saturation-intensity color model is mainly adopted to take the advantage of human visual experiences in texture recognition. We discuss and validate the effectiveness and efficiency of our method by applying to various benchmarks. The results show that we can extract quality dominant textures automatically in real time, and faster by several orders of magnitude than existing methods.
Wencheng Wang 0001, Miao Hua
IEEE Trans. Image Process.2