Mingrui Zhu

dblp:94/2339 · DBLP profile ↗
← Back
42ranked-venue papers
16as first author
34since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 10 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 18 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Computer networks · 1Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mixture of Ranks with Degradation-Aware Routing for One-Step Real-World Image Super-Resolution
abstract
The demonstrated success of sparsely-gated Mixture-of-Experts (MoE) architectures, exemplified by models such as DeepSeek and Grok, has motivated researchers to investigate their adaptation to diverse domains. In real-world image super-resolution (Real-ISR), existing approaches mainly rely on fine-tuning pre-trained diffusion models through Low-Rank Adaptation (LoRA) module to reconstruct high-resolution (HR) images. However, these dense Real-ISR models are limited in their ability to adaptively capture the heterogeneous characteristics of complex real-world degraded samples or enable knowledge sharing between inputs under equivalent computational budgets. To address this, we investigate the integration of sparse MoE into Real-ISR and propose a Mixture-of-Ranks (MoR) architecture for single-step image super-resolution. We introduce a fine-grained expert partitioning strategy that treats each rank in LoRA as an independent expert. This design enables flexible knowledge recombination while isolating fixed-position ranks as shared experts to preserve common-sense features and minimize routing redundancy. Furthermore, we develop a degradation estimation module leveraging CLIP embeddings and predefined positive-negative text pairs to compute relative degradation scores, dynamically guiding expert activation. To better accommodate varying sample complexities, we incorporate zero-expert slots and propose a degradation-aware load-balancing loss, which dynamically adjusts the number of active experts based on degradation severity, ensuring optimal computational resource allocation. Comprehensive experiments validate our framework's effectiveness and state-of-the-art performance.
Xiao He 0014, Zhijun Tu, Mingrui Zhu, Jie Hu 0021, Nannan Wang 0001, Xinbo Gao 0001
AAAI4
2026 DIVER: Unlocking Diversity in Ad Headline Generation with Large Language Models
abstract
While Large Language Models (LLMs) possess remarkable generative capabilities, generating diversified and engaging ad headlines in industrial applications remains challenging. Conventional training paradigms often suffer from mode collapse, converging on dominant data patterns and yielding homogeneous outputs. Meanwhile, existing diversity-enhancing techniques like stochastic decoding frequently compromise semantic coherence and controllability. To break this trade-off, we propose DIVER, an automated training framework that internalizes diversity as an intrinsic model capability. DIVER employs an automatic data pipeline to synthesize high-quality, multi-faceted training pairs and utilizes multi-objective reinforcement learning to effectively co-optimize diversity with advertising metrics such as faithfulness and click-through rate (CTR). Unlike personalized approaches, our framework generates diverse content for general users without relying on heavy and costly user-behavior modeling, ensuring efficient inference for large-scale real-time systems. Real-world deployment on Xiaohongshu's Explore Feed demonstrates significant commercial impact, increasing advertiser value (ADVV) by 4.0% and CTR by 1.4%.
Depeng Yuan, Yuqi Chen 0018, Yanhua Huang, Yuanhang Zheng, Yinqi Zhang, Kedi Chen, Mingrui Zhu, Ruiwen Xu
SIGIR10
2026 Augmentation-free dynamic graph contrastive learning based on transformers
Mingrui Zhu
Expert Syst. Appl.1
2026 AdaAlign: A unified solution for traditional and modern zero-shot sketch-based image retrieval
Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
Neural Networks1
2026 Fine-Detailed Facial Sketch-to-Photo Synthesis With Detail-Enhanced Codebook Priors
abstract
Generating high-quality facial photos from fine-detailed sketches is a long-standing research topic that remains unsolved. The scarcity of large-scale paired data due to the cost of acquiring hand-drawn sketches poses a major challenge. Existing methods either lose identity information with oversimplified representations, or rely on costly inversion and strict alignment when using StyleGAN-based priors, limiting their practical applicability. Our primary finding in this work is that the discrete codebook and decoder trained through self-reconstruction in the photo domain can learn rich priors, helping to reduce ambiguity in cross-domain mapping even with current small-scale paired datasets. Based on this, a cross-domain mapping network can be directly constructed. However, empirical findings indicate that using the discrete codebook for cross-domain mapping often results in unrealistic textures and distorted spatial layouts. Therefore, we propose a Hierarchical Adaptive Texture-Spatial Correction (HATSC) module to correct the flaws in texture and spatial layouts. Besides, we introduce a Saliency-based Key Details Enhancement (SKDE) module to further enhance the synthesis quality. Overall, we present a “reconstruct-cross-enhance” pipeline for synthesizing facial photos from fine-detailed sketches. Experiments demonstrate that our method generates high-quality facial photos and significantly outperforms previous approaches across a wide range of challenging benchmarks. The code is publicly available at: https://github.com/Gardenia-chen/DECP.
Mingrui Zhu, Jianhang Chen, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 One Step Diffusion-Based Super-Resolution With Time-Aware Distillation
abstract
iffusion-based image super-resolution (SR) has shown strong potential in recovering high-fidelity details from low-resolution inputs. However, the need for tens or hundreds of sampling steps leads to substantial inference latency. Recent works attempt to accelerate this process via knowledge distillation, but often rely solely on pixel-level loss or overlook the fact that diffusion models capture different information across time steps. To address this, we propose TAD-SR, a time-aware diffusion distillation framework. Specifically, we introduce a novel score distillation strategy to align the score functions between the outputs of the student and teacher models after minor noise perturbation. This distillation strategy eliminates the inherent bias in score distillation sampling (SDS) and enables the student models to focus more on high-frequency image details by sampling at smaller time steps. We further introduce a time-aware discriminator that exploits the teacher’s knowledge to differentiate real and synthetic samples across different noise scales, using explicit temporal conditioning. Extensive experiments on SR tasks demonstrate that TAD-SR outperforms existing singl-estep diffusion methods and achieves performance on par with multi-step state-of-the-art models.iffusion-based image super-resolution (SR) has shown strong potential in recovering highfidelity details from low-resolution inputs. However, the need for tens or hundreds of sampling steps leads to substantial inference latency. Recent works attempt to accelerate this process via knowledge distillation, but often rely solely on pixel-level loss or overlook the fact that diffusion models capture different information across time steps. To address this, we propose TADSR, a time-aware diffusion distillation framework. Specifically, we introduce a novel score distillation strategy to align the score functions between the outputs of the student and teacher models after minor noise perturbation. This distillation strategy eliminates the inherent bias in score distillation sampling (SDS) and enables the student models to focus more on highf-requency image details by sampling at smaller time steps. We further introduce a time-aware discriminator that exploits the teacher’s knowledge to differentiate real and synthetic samples across different noise scales, using explicit temporal conditioning. Extensive experiments on SR tasks demonstrate that TAD-SR outperforms existing single-step diffusion methods and achieves performance on par with multi-step state-of-the-art models D.
Xiao He 0014, Huaao Tang, Zhijun Tu, Hanting Chen, Mingrui Zhu, Jie Hu 0021, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.8
2025 Effective Diffusion Transformer Architecture for Image Super-Resolution
abstract
Recent advances indicate that diffusion model holds great promise in image super-resolution. While latest methods are primarily based on latent diffusion models with convolutional neural networks, there are few attempts to explore transformers, which have demonstrated remarkable performance in image generation. In this work, we design an effective diffusion transformer for image super resolution (DiT-SR) that achieves the visual quality of prior-based methods, but through a training-from-scratch manner. In practice, DiT-SR leverages an overall U-shaped architecture, and adopts uniform isotropic design for all the transformer blocks across different stages. The former facilitates multi-scale hierarchical feature extraction, while the latter reallocate the computational resources to critical layers to further enhance performance. Moreover, we thoroughly analyze the limitation of the widely used AdaLN, and present a frequency-adaptive time-step conditioning module, enhancing the model's capacity to process distinct frequency information at different time steps. Extensive experiments demonstrate that DiT-SR outperforms the existing training-from-scratch diffusion-based SR methods significantly, and even beats some of the prior-based methods on pretrained Stable Diffusion, proving the superiority of diffusion transformer in image super resolution.
Zhijun Tu, Xiao He 0014, Liyu Chen, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001, Jie Hu 0021
AAAI7
2025 3D Test-Time Adaptation via Graph Spectral Driven Point Shift
abstract
While test-time adaptation (TTA) methods effectively address domain shifts by dynamically adapting pre-trained models to target domain data during online inference, their application to 3D point clouds is hindered by their irregular and unordered structure. Current 3D TTA methods often rely on computationally expensive spatial-domain optimizations and may require additional training data. In contrast, we propose Graph Spectral Domain Test-Time Adaptation (GSDTTA), a novel approach for 3D point cloud classification that shifts adaptation to the graph spectral domain, enabling more efficient adaptation by capturing global structural properties with fewer parameters. Point clouds in target domain are represented as outlier-aware graphs and transformed into graph spectral domain by Graph Fourier Transform (GFT). For efficiency, adaptation is performed by optimizing only the lowest 10% of frequency components, which capture the majority of the point cloud's energy. An inverse GFT (IGFT) is then applied to reconstruct the adapted point cloud with the graph spectral-driven point shift. This process is enhanced by an eigenmap-guided self-training strategy that iteratively refines both the spectral adjustments and the model parameters. Experimental results and ablation studies on benchmark datasets demonstrate the effectiveness of GSDTTA, outperforming existing TTA methods for 3D point cloud classification.
Yijie Fang, Mingrui Zhu
ICCV4
2025 Diff-MoE: Diffusion Transformer with Time-Aware and Space-Adaptive Experts
abstract
Diffusion models have transformed generative modeling but suffer from scalability limitations due to computational overhead and inflexible architectures that process all generative stages and tokens uniformly. In this work, we introduce Diff-MoE, a novel framework that combines Diffusion Transformers with Mixture-of-Experts to exploit both temporarily adaptability and spatial flexibility. Our design incorporates expert-specific timestep conditioning, allowing each expert to process different spatial tokens while adapting to the generative stage, to dynamically allocate resources based on both the temporal and spatial characteristics of the generative task. Additionally, we propose a globally-aware feature recalibration mechanism that amplifies the representational capacity of expert modules by dynamically adjusting feature contributions based on input relevance. Extensive experiments on image generation benchmarks demonstrate that Diff-MoE significantly outperforms state-of-the-art methods. Our work demonstrates the potential of integrating diffusion models with expert-based designs, offering a scalable and effective framework for advanced generative modeling.
Xiao He 0014, Zhijun Tu, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001, Jie Hu 0021
ICML5
2025 Abductive learning-guided uncertainty modeling for time series anomaly detection
Qi Zhang 0099, Mingrui Zhu, Jie Li 0068, Jinsong Bao, Dan Zhang 0006
Knowl. Based Syst.2
2025 Dynamic graph contrastive learning based on learnable view generators
Mingrui Zhu
Knowl. Based Syst.1
2025 Boosting Semi-Supervised Facial Attribute Recognition With Dynamic Threshold Pairs
abstract
Semi-supervised learning (SSL) has proven effective in assigning a pseudo-label to a confident sample whose largest class probability is above a fixed threshold. However, in the context of semi-supervised facial attribute recognition (SSFAR), where a sample is associated with multiple presence and absence pseudo-labels, directly applying existing SSL methods is challenging due to two issues: 1) the lack of a clear boundary between presence and absence predictions for an attribute makes it difficult to distinguish them using a single threshold; 2) the learning difficulty varies across attributes, so the fixed strategy fails to adaptively learn different attributes. To address these challenges, we propose Dynamic thrEShold Pairs (DESP), a simple yet effective method to handle the SSFAR problem. Specifically, during each training stage, we derive two sets for each attribute from labeled samples, which contain the predicted probabilities of presence and absence, respectively. We then compute the mid-ranges of the two sets as paired presence and absence thresholds. Finally, we assign a presence or absence pseudo-label for the attribute to an unlabeled sample when its prediction exceeds the presence threshold or falls below the absence threshold. Extensive experiments on the CelebA and LFWA datasets demonstrate that DESP achieves superior performance compared to state-of-the-art methods, especially in the case of scarce labeled samples. Also, DESP performs well on multi-label datasets such as Pascal VOC and MS-COCO. The code will be publicly available athttps://github.com/yihanxxu/DESP.
Hangyu Li 0001, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 CatVersion: Concatenating Embeddings for Diffusion-Based Text-to-Image Personalization
abstract
We propose CatVersion, an inversion-based method that learns the personalized concept through a handful of examples. Subsequently, users can utilize text prompts to generate images that embody the personalized concept, thereby achieving text-to-image personalization. In contrast to existing approaches that emphasize word embedding learning or parameter fine-tuning for the diffusion model, which potentially causes concept dilution or overfitting, our method concatenates embeddings on the feature-dense space of the text encoder in the diffusion model to learn the gap between the personalized concept and its base class, aiming to maximize the preservation of prior knowledge in diffusion models while restoring the personalized concepts. To this end, we first dissect the text encoder’s integration in the image generation process to identify the feature-dense space of the encoder. Afterward, we concatenate embeddings on the Keys and Values in this space to learn the gap between the personalized concept and its base class. In this way, the concatenated embeddings ultimately manifest as a residual on the original attention output. To more accurately and unbiasedly quantify the results of personalized image generation, we improve the CLIP image alignment score based on masks. Qualitatively and quantitatively, CatVersion helps to restore personalization concepts more faithfully and enables more robust editing.
Mingrui Zhu, Shiyin Dong, De Cheng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Disentangle Before Anonymize: A Two-Stage Framework for Attribute-Preserved and Occlusion-Robust De-Identification
abstract
In an era where personal photos are easily leaked and collected, face de-identification is a crucial method for protecting identity privacy. However, current face de-identification techniques face challenges in preserving attribute details and often produce anonymized results with reduced realistic. These shortcomings are particularly evident when handling occlusions, frequently resulting in noticeable editing artifacts. Our primary finding in this work is that simultaneous training of identity disentanglement and anonymization hinders their respective effectiveness. Therefore, we propose “Disentangle Before Anonymize”, a novel two-stage Framework (DBAF) designed for attribute-preserved and occlusion-robust de-identification. This framework includes a Contrastive Identity Disentanglement (CID) module and a Key-authorized Reversible Identity Anonymization (KRIA) module, achieving faithful attribute preservation and high-quality identity anonymization edits. Additionally, we introduce a Multi-scale Attentional Attribute Retention (MAAR) module to address the issue of reduced anonymization quality under occlusions. Extensive experiments demonstrate that our method outperforms state-of-the-art de-identification approaches, delivering superior quality, enhanced detail fidelity, improved attribute preservation performance, and greater robustness to occlusions. The code ispublicly available at: https://github.com/mrzhu-cool/DBAF.
Mingrui Zhu, Dongxin Chen, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.1
2025 Foodfusion: A Novel Approach for Food Image Composition via Diffusion Models
abstract
Food image composition requires the use of existing dish images and background images to synthesize a natural new image, while diffusion models have made significant advancements in image generation, enabling the construction of end-to-end architectures that yield promising results. However, existing diffusion models face challenges in processing and fusing information from multiple images and lack access to high-quality publicly available datasets, which prevents the application of diffusion models in food image composition. In this paper, we introduce a large-scale, high-quality food image composite dataset,FC22 k, which comprises 22,000 foreground, background, and ground truth ternary image pairs. Additionally, we propose a novel food image composition method,Foodfusion, which leverages the capabilities of the pre-trained diffusion models and incorporates a Fusion Module for processing and integrating foreground and background information. This fused information aligns the foreground features with the background structure by merging the global structural information at the cross-attention layer of the denoising UNet. To further enhance the content and structure of the background, we also integrate a Content-Structure Control Module. Extensive experiments demonstrate the effectiveness and scalability of our proposed method.
Chaohua Shi, Xuan Wang 0009, Xule Wang, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.5
2025 PStyle-3D: Example-Based 3-D-Aware Portrait Style Domain Adaptation
abstract
The creation of high-quality artistic portraits is a critical and desirable task in the field of computer vision. While recent3-D generative models have achieved impressive results in generating images with view consistency and intricate 3-D shapes, their application for generating artistic portraits is often more challenging than 2-D generative models due to the potentially destructive impact of 3-D structures on human faces. This article introduces a novel approach that leverages a meticulously designed domain feature extraction module to extract the specific feature information from both the source natural face domain and the target artistic portrait domain. These extracted features are seamlessly integrated into a 3-D representation, generating multiview consistent 3-D artistic portraits. To fuse the features of the source and target domains better, we propose a new module for domain adaptation. This module adds a path to the style path established by StyleGAN to introduce the artistic portrait domain information and regulate the target domain's feature information in $\mathcal {S}$ space. Our domain adaptation module is implemented in each StyleBlock of the 3-D representation generator to integrate the target domain information with the original facial information. Experimental results demonstrate that our approach generates high-quality 3-D artistic portraits that outperform existing approaches in preserving 3-D geometric information and multiview consistency.
Chaohua Shi, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Few-Shot Face Stylization via GAN Prior Distillation
abstract
Face stylization has made notable progress in recent years. However, when training on limited data, the performance of existing approaches significantly declines. Although some studies have attempted to tackle this problem, they either failed to achieve the few-shot setting (less than 10) or can only get suboptimal results. In this article, we propose GAN Prior Distillation (GPD) to enable effective few-shot face stylization. GPD contains two models: a teacher network with GAN Prior and a student network that fulfills end-to-end translation. Specifically, we adapt the teacher network trained on large-scale data in the source domain to the target domain using a handful of samples, where it can learn the target domain's knowledge. Then, we can achieve few-shot augmentation by generating source domain and target domain images simultaneously with the same latent codes. We propose an anchor-based knowledge distillation module that can fully use the difference between the training and the augmented data to distill the knowledge of the teacher network into the student network. The trained student network achieves excellent generalization performance with the absorption of additional knowledge. Qualitative and quantitative experiments demonstrate that our method achieves superior results than state-of-the-art approaches in a few-shot setting.
Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 TraSculptor: Visual Analytics for Enhanced Decision-Making in Road Traffic Planning
abstract
The design of urban road networks significantly influences traffic conditions, underscoring the importance of informed traffic planning. Traffic planning experts rely on specialized platforms to simulate traffic systems, assessing the efficacy of the road network across various states of modifications. Nevertheless, a prevailing issue persists: many existing traffic planning platforms exhibit inefficiencies in flexibly interacting with the road network's structure and attributes and intuitively comparing multiple states during the iterative planning process. This paper introduces TraSculptor, an interactive planning decision-making system. To develop TraSculptor, we identify and address two challenges: interactive modification of road networks and intuitive comparison of multiple network states. For the first challenge, we establish flexible interactions to enable experts to easily and directly modify the road network on the map. For the second challenge, we design a comparison view with a history tree of multiple states and a road-state matrix to facilitate intuitive comparison of road network states. To evaluate TraSculptor, we provided a usage scenario where the Braess's paradox was showcased, invited experts to perform a case study on the Sioux Falls network, and collected expert feedback through interviews.
Zikun Deng, Yuanbang Liu, Mingrui Zhu, Da Xiang, Zicheng Su, Qing-Long Lu, Tobias Schreck, Yi Cai 0001
IEEE Trans. Vis. Comput. Graph.3
2024 On the Analysis of GAN-based Image-to-Image Translation with Gaussian Noise Injection
abstract
Image-to-image (I2I) translation is vital in computer vision tasks like style transfer and domain adaptation. While recent advances in GAN have enabled high-quality sample generation, real-world challenges such as noise and distortion remain significant obstacles. Although Gaussian noise injection during training has been utilized, its theoretical underpinnings have been unclear. This work provides a robust theoretical framework elucidating the role of Gaussian noise injection in I2I translation models. We address critical questions on the influence of noise variance on distribution divergence, resilience to unseen noise types, and optimal noise intensity selection. Our contributions include connecting $f$-divergence and score matching, unveiling insights into the impact of Gaussian noise on aligning probability distributions, and demonstrating generalized robustness implications. We also explore choosing an optimal training noise level for consistent performance in noisy environments. Extensive experiments validate our theoretical findings, showing substantial improvements over various I2I baseline models in noisy settings. Our research rigorously grounds Gaussian noise injection for I2I translation, offering a sophisticated theoretical understanding beyond heuristic applications.
Chaohua Shi, Lu Gan 0002, Hongqing Liu 0001, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
ICLR5
2024 Bridging Generative and Discriminative Models for Unified Visual Perception with Diffusion Priors
Shiyin Dong, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IJCAI2
2024 Diff-Privacy: Diffusion-Based Face Privacy Protection
abstract
Privacy protection has become a top priority due to the widespread collection and misuse of personal data. Anonymization and visual identity information hiding are two crucial tasks in face privacy protection, both striving to alter identifying characteristics from face images to prevent privacy information leakage. However, the goals of the two are not entirely the same. Consequently, training a model to simultaneously perform both tasks proves challenging. In this paper, we propose Diff-Privacy, a novel face privacy protection method based on diffusion models that unifies the task of anonymization and visual identity information hiding. Specifically, we present a Multi-Scale image Inversion module (MSI) that, through training, generates a set of Stable Diffusion (SD) format conditional embeddings for the original image. With these conditional embeddings, we design corresponding embedding scheduling strategies and formulate distinct energy functions during the inference process to achieve anonymization and visual identity information hiding, respectively. Extensive experiments demonstrate the effectiveness of the proposed method in protecting face privacy.
Xiao He 0014, Mingrui Zhu, Dongxin Chen, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Few-Shot Font Generation by Learning Style Difference and Similarity
abstract
Few-shot font generation (FFG) aims to preserve the underlying global structure of the original character while generating target fonts by referring to a few samples. It has been applied to font library creation, a personalized signature, and other scenarios. Existing FFG methods explicitly disentangle content and style of reference glyphs universally or component-wisely. However, they ignore the difference between glyphs in different styles and the similarity of glyphs in the same style, which results in artifacts such as local distortions and style inconsistency. To address this issue, we propose a novel font generation approach by learning the Difference between different styles and the Similarity of the same style (DS-Font). We introduce contrastive learning to consider the positive and negative relationship between styles. Specifically, we propose a multi-layer style projector (MSP) for style encoding and realize a distinctive style representation via our proposed Cluster-level Contrastive Style (CCS) loss. The MSP module is employed to assist the generator during training to enhance the style consistency between the generated glyph and the reference glyphs. In addition, we design a glyph-independent patch discriminator, which comprehensively considers different areas of the image and ensures that each style can be distinguished independently. We conduct qualitative and quantitative evaluations comprehensively to demonstrate that our approach achieves significantly better results than state-of-the-art methods.
Xiao He 0014, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 PMSGAN: Parallel Multistage GANs for Face Image Translation
abstract
In this article, we address the face image translation task, which aims to translate a face image of a source domain to a target domain. Although significant progress has been made by recent studies, face image translation is still a challenging task because it has more strict requirements for texture details: even a few artifacts will greatly affect the impression of generated face images. Targeting to synthesize high-quality face images with admirable visual appearance, we revisit the coarse-to-fine strategy and propose a novel p arallel m ultistage architecture on the basis of g enerative a dversarial n etworks (PMSGAN). More specifically, PMSGAN progressively learns the translation function by disintegrating the general synthesis process into multiple parallel stages that take images with gradually decreasing spatial resolution as inputs. To prompt the information exchange between various stages, a cross-stage atrous spatial pyramid (CSASP) structure is specially designed to receive and fuse the contextual information from other stages. At the end of the parallel model, we introduce a novel attention-based module that leverages multistage decoded outputs as in situ supervised attention to refine the final activations and yield the target image. Extensive experiments on several face image translation benchmarks show that PMSGAN performs considerably better than state-of-the-art approaches.
Changcheng Liang, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 All-to-key Attention for Arbitrary Style Transfer
abstract
Attention-based arbitrary style transfer studies have shown promising performance in synthesizing vivid local style details. They typically use the all-to-all attention mechanism—each position of content features is fully matched to all positions of style features. However, all-to-all attention tends to generate distorted style patterns and has quadratic complexity, limiting the effectiveness and efficiency of arbitrary style transfer. In this paper, we propose a novel all-to-key attention mechanism—each position of content features is matched to stable key positions of style features—that is more in line with the characteristics of style transfer. Specifically, it integrates two newly proposed attention forms: distributed and progressive attention. Distributed attention assigns attention to key style representations that depict the style distribution of local regions; Progressive attention pays attention from coarse-grained regions to fine-grained key positions. The resultant module, dubbed StyA2K, shows extraordinary performance in preserving the semantic structure and rendering consistent style patterns. Qualitative and quantitative comparisons with state-of-the-art methods demonstrate the superior performance of our approach. Codes and models are available on https://github.com/LearningHx/StyA2K.
Mingrui Zhu, Xiao He 0014, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
ICCV1
2023 Controllable Face Sketch-Photo Synthesis with Flexible Generative Priors
abstract
Current face sketch-photo synthesis researches generally embrace an image-to-image (I2I) translation pipeline. However, these methods ignore the one-to-many mapping problem (i.e., multiple plausible photo results can correspond to a single input sketch) in sketch-to-photo synthesis task, resulting in significant performance degradation on diverse datasets. Besides, generating high-quality images on limited data is also a challenge for this task. To address these challenges, we propose a dual-path framework that introduces generative priors to better perform cross-domain reconstruction on limited data. The coarse path uses a layer-swapped pre-trained generator to achieve coarse cross-domain reconstruction, and the refinement path further improves the structure and texture details. To align the feature maps between the two paths, we introduce a spatial feature calibration module. Despite this, our framework still struggles to handle diverse datasets. Thanks to the flexibility of generative priors, we can extend the framework to achieve exemplar-guided I2I translation by incorporating an exemplar with style mixing and a proposed semantic-aware style refinement strategy, which addresses the one-to-many mapping problem in sketch-to-photo synthesis task. Furthermore, our framework can perform cross-domain editing by employing off-the-shelf editing methods based on the latent space, achieving fine-grained control. Extensive experiments on diverse datasets demonstrate the superiority of our framework over other state-of-the-art methods.
Mingrui Zhu, Nannan Wang 0001, Guozhang Li, Xiaoyu Wang 0002, Xinbo Gao 0001
ACM Multimedia2
2023 Traceability of abnormal energy consumption modes in grinding systems based on evolution analysis of causal network structure
Mingrui Zhu, Yangjian Ji
Adv. Eng. Informatics1
2023 Energy consumption mode identification and monitoring method of process industry system under unstable working conditions
Mingrui Zhu, Yangjian Ji, Xiaoyang Zhu, Kai Ren 0004
Adv. Eng. Informatics1
2023 BiTGAN: bilateral generative adversarial networks for Chinese ink wash painting style transfer
Xiao He 0014, Mingrui Zhu, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
Sci. China Inf. Sci.2
2023 Dual Conditional Normalization Pyramid Network for Face Photo-Sketch Synthesis
abstract
Face photo-sketch synthesis has undergone remarkable progress with the rapid development of deep learning techniques. Cutting-edge methods directly learn the cross-domain mapping between photos and sketches, which ignores the available reference samples. We argue that the reference samples can provide adequate prior information on texture and content in this task and improve the visual performance of synthetic images. This paper proposes a Dual Conditional Normalization Pyramid (DCNP) network with a multi-scale pyramid structure. The core of the DCNP network is a Dual Conditional Normalization (DCN) based architecture, which can obtain prior information on different semantics from reference samples. Specifically, DCN contains two conditional normalization branches. The first branch allows for spatially-adaptive normalization of the reference image conditioned on the semantic mask of the input image. The second branch enables adaptive instance normalization of the input image conditioned on the reference image. DCN can emphasize the isolated importance of textural and spatial factors by disintegrating the entire cross-domain mapping into two branches. To avoid information redundancy and improve the final performance, we propose a Gated Channel Attention Fusion (GCAF) module to distill and fuse the helpful information of the two branches. Qualitative and quantitative experimental results demonstrate the superior performance of the proposed method over the state-of-the-art approaches in structural information preservation and realistic texture generation. The code is public inhttps://github.com/Tony0720/DCNP.
Mingrui Zhu, Zicheng Wu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 An Efficient Transformer Based on Global and Local Self-Attention for Face Photo-Sketch Synthesis
abstract
Face photo-sketch synthesis tasks have been dominated by convolutional neural networks (CNNs), especially CNN-based generative adversarial networks (GANs), because of their strong texture modeling capabilities and thus their ability to generate more realistic face photos/sketches beyond traditional methods. However, due to CNNs' locality and spatial invariance properties, there have weaknesses in capturing the global and structural information which are extremely important for face images. Inspired by the recent phenomenal success of the Transformer in vision tasks, we propose replacing CNNs with Transformers that are able to model long-range dependencies to synthesize more structured and realistic face images. However, the existing vision Transformers are mainly designed for high-level vision tasks and lack the dense prediction ability to generate high resolution images due to the quadratic computational complexity of their self-attention mechanism. In addition, the original Transformer is not capable of modeling local correlations which is an important skill for image generation. To address these challenges, we propose two types of memory-friendly Transformer encoders, one for processing local correlations via local self-attention and another for modeling global information via global self-attention. By integrating the two proposed Transformer encoders, we present an efficient GL-Transformer for face photo-sketch synthesis, which can synthesize realistic face photo/sketch images from coarse to fine. Extensive experiments demonstrate that our model achieves a comparable or better performance beyond the state-of-the-art CNN-based methods both qualitatively and quantitatively.
Wangbo Yu, Mingrui Zhu, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Image Process.2
2022 VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the Wild
abstract
We present VideoReTalking, a new system to edit the faces of a real-world talking head video according to input audio, producing a high-quality and lip-syncing output video even with a different emotion. Our system disentangles this objective into three sequential tasks: (1) face video generation with a canonical expression; (2) audio-driven lip-sync; and (3) face enhancement for improving photo-realism. Given a talking-head video, we first modify the expression of each frame according to the same expression template using the expression editing network, resulting in a video with the canonical expression. This video, together with the given audio, is then fed into the lip-sync network to generate a lip-syncing video. Finally, we improve the photo-realism of the synthesized faces through an identity-aware face enhancement network and post-processing. We use learning-based approaches for all three steps and all our modules can be tackled in a sequential pipeline without any user intervention. Furthermore, our system is a generic approach that does not need to be retrained to a specific person. Evaluations on two widely-used datasets and in-the-wild examples demonstrate the superiority of our framework over other state-of-the-art methods in terms of lip-sync accuracy and visual quality.
Xiaodong Cun, Yong Zhang 0034, Menghan Xia, Mingrui Zhu, Xuan Wang 0009, Jue Wang 0001, Nannan Wang 0001
SIGGRAPH Asia6
2022 Knowledge Distillation for Face Photo-Sketch Synthesis
abstract
Significant progress has been made with face photo-sketch synthesis in recent years due to the development of deep convolutional neural networks, particularly generative adversarial networks (GANs). However, the performance of existing methods is still limited because of the lack of training data (photo-sketch pairs). To address this challenge, we investigate the effect of knowledge distillation (KD) on training neural networks for the face photo-sketch synthesis task and propose an effective KD model to improve the performance of synthetic images. In particular, we utilize a teacher network trained on a large amount of data in a related task to separately learn knowledge of the face photo and knowledge of the face sketch and simultaneously transfer this knowledge to two student networks designed for the face photo-sketch synthesis task. In addition to assimilating the knowledge from the teacher network, the two student networks can mutually transfer their own knowledge to further enhance their learning. To further enhance the perception quality of the synthetic image, we propose a KD+ model that combines GANs with KD. The generator can produce images with more realistic textures and less noise under the guide of knowledge. Extensive experiments and a user study demonstrate the superiority of our models over the state-of-the-art methods.
Mingrui Zhu, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 A Sketch-Transformer Network for Face Photo-Sketch Synthesis
abstract
We present a face photo-sketch synthesis model, which converts a face photo into an artistic face sketch or recover a photo-realistic facial image from a sketch portrait. Recent progress has been made by convolutional neural networks (CNNs) and generative adversarial networks (GANs), so that promising results can be obtained through real-time end-to-end architectures. However, convolutional architectures tend to focus on local information and neglect long-range spatial dependency, which limits the ability of existing approaches in keeping global structural information. In this paper, we propose a Sketch-Transformer network for face photo-sketch synthesis, which consists of three closely-related modules, including a multi-scale feature and position encoder for patch-level feature and position embedding, a self-attention module for capturing long-range spatial dependency, and a multi-scale spatially-adaptive de-normalization decoder for image reconstruction. Such a design enables the model to generate reasonable detail texture while maintaining global structural information. Extensive experiments show that the proposed method achieves significant improvements over state-of-the-art approaches on both quantitative and qualitative evaluations.
Mingrui Zhu, Changcheng Liang, Nannan Wang 0001, Xiaoyu Wang 0002, Zhifeng Li 0001, Xinbo Gao 0001
IJCAI1
2021 Learning Deep Patch representation for Probabilistic Graphical Model-Based Face Sketch Synthesis
Mingrui Zhu, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
Int. J. Comput. Vis.1
2020 Community detection in complex network based on APT method
Qingfeng Chen, Yulu Qiao, Fang Hu 0001, Mingrui Zhu, Chengqi Zhang
Pattern Recognit. Lett.6
2019 Face Photo-Sketch Synthesis via Knowledge Transfer
abstract
Despite deep neural networks have demonstrated strong power in face photo-sketch synthesis task, their performance, however, are still limited by the lack of training data (photo-sketch pairs). Knowledge Transfer (KT), which aims at training a smaller and fast student network with the information learned from a larger and accurate teacher network, has attracted much attention recently due to its superior performance in the acceleration and compression of deep neural networks. This work has brought us great inspiration that we can train a relatively small student network on very few training data by transferring knowledge from a larger teacher model trained on enough training data for other tasks. Therefore, we propose a novel knowledge transfer framework to synthesize face photos from face sketches or synthesize face sketches from face photos. Particularly, we utilize two teacher networks trained on large amount of data in related task to learn the knowledge of face photos and face sketches separately and transfer them to two student networks simultaneously. In addition, the two student networks, one for photo ? sketch task and the other for sketch ? photo task, can transfer their knowledge mutually. With the proposed method, we can train our model which has superior performance using a small set of photo-sketch pairs. We validate the effectiveness of our method across several datasets. Quantitative and qualitative evaluations illustrate that our model outperforms other state-of-the-art methods in generating face sketches (or photos) with high visual quality and recognition ability.
Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Zhifeng Li 0001
IJCAI1
2019 A Deep Collaborative Framework for Face Photo-Sketch Synthesis
abstract
Great breakthroughs have been made in the accuracy and speed of face photo-sketch synthesis in recent years. Regression-based methods have gained increasing attention, which benefit from deeper and faster end-to-end convolutional neural networks. However, most of these models typically formulate the mapping from photo domain X to sketch domain Y as a unidirectional feedforward mapping, G: X → Y , and vice versa, F: Y → X ; thus, the utilization of mutual interaction between two opposite mappings is lacking. Therefore, we proposed a collaborative framework for face photo-sketch synthesis. The concept behind our model was that a middle latent domain ~Z between the photo domain X and the sketch domain Y can be learned during the learning procedure of G: X → Y and F: Y → X by introducing a collaborative loss that makes full use of two opposite mappings. This strategy can constrain the two opposite mappings and make them more symmetrical, thus making the network more suitable for the photo-sketch synthesis task and obtaining higher quality generated images. Qualitative and quantitative experiments demonstrated the superior performance of our model in comparison with the existing state-of-the-art solutions.
Mingrui Zhu, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2017 Deep Graphical Feature Learning for Face Sketch Synthesis
abstract
The exemplar-based face sketch synthesis method generally contains two steps: neighbor selection and reconstruction weight representation. Pixel intensities are widely used as features by most of the existing exemplar-based methods, which lacks of representation ability and robustness to light variations and clutter backgrounds. We present a novel face sketch synthesis method combining generative exemplar-based method and discriminatively trained deep convolutional neural networks (dCNNs) via a deep graphical feature learning framework. Our method works in both two steps by using deep discriminative representations derived from dCNNs. Instead of using it directly, we boost its representation capability by a deep graphical feature learning framework. Finally, the optimal weights of deep representations and optimal reconstruction weights for face sketch synthesis can be obtained simultaneously. With the optimal reconstruction weights, we can synthesize high quality sketches which is robust against light variations and clutter backgrounds. Extensive experiments on public face sketch databases show that our method outperforms state-of-the-art methods, in terms of both synthesis quality and recognition ability.
Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001
IJCAI1
2017 Data-driven vs. model-driven: Fast face sketch synthesis
Nannan Wang 0001, Mingrui Zhu, Jie Li 0001, Bin Song 0001, Zan Li 0001
Neurocomputing2
2014 Multi-way Theta-Join Based on CMD Storage Method
Lei Li 0003, Hong Gao 0001, Mingrui Zhu, Zhaonian Zou
DASFAA (1)3
2007 Bounds on the Expansion Properties of Tanner Graphs
abstract
This work focuses on the expansion properties of a Tanner Graph because they are known to be related to the performance of associated iterative message-passing algorithms over various channels. By analyzing the eigenvalues and corresponding eigenvectors of the normalized incidence matrix representing a Tanner Graph, lower bounds on these expansion properties are derived. Specifically, for the binary erasure channel, these results lead to two lower bounds on stopping distance for any given binary linear code and an upper bound on stopping redundancy for the family of difference-set codes (type-I 2-D projective geometry low-density parity-check (LDPC) codes).
Mingrui Zhu, Keith M. Chugg
IEEE Trans. Inf. Theory1
2005 A new approach to rapid PN code acquisition using iterative message passing techniques
abstract
Iterative message passing algorithms on graphs, which are generalized from the well-known turbo decoding algorithm, have been studied intensively in recent years because they can provide near-optimal performance and significant complexity reduction. In this paper, we demonstrate that this technique can be applied to pseudorandom code acquisition problems as well. To do this, we represent good pseudonoise (PN) patterns using sparse graphical models, then apply the standard iterative message passing algorithms over these graphs to approximate maximum-likelihood synchronization. Simulation results show that the proposed algorithm achieves better performance than both serial and hybrid search strategies in that it works at low signal-to-noise ratios and is much faster. Compared with full parallel search, this approach typically provides significant complexity reduction.
Keith M. Chugg, Mingrui Zhu
IEEE J. Sel. Areas Commun.2