EDBT 2026 Demo / reviewers in the wild / expert
Muyi Sun
dblp:217/1037
· DBLP profile ↗
33ranked-venue papers
3as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 17 since 2021Artificial intelligence and machine learning · 17 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Security and privacy · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive InteractionabstractGenerating responsive listener head dynamics with nuanced emotions and expressive reactions is crucial for dialogue modeling in various virtual avatar animations. Previous studies mainly focus on the direct short-term production of listener behavior. They overlook the fine-grained control over motion variations and emotional intensity, especially in long-sequence modeling. Moreover, the lack of long-term and large-scale paired speaker-listener corpora incorporating head dynamics and fine-grained multi-modality annotations limits the application of dialogue modeling. Therefore, we first newly collect a large-scale multi-turn dataset of 3D dyadic conversation containing more than 1.4M valid frames for multi-modal responsive interaction, dubbed ListenerX. Additionally, we propose VividListener, a novel framework enabling fine-grained, expressive, and controllable listener dynamics modeling. This framework leverages multi-modal conditions as guiding principles for fostering coherent interactions between speakers and listeners. Specifically, we design the Responsive Interaction Module (RIM) to adaptively represent the multi-modal interactive embeddings. RIM ensures the listener dynamics achieve fine-grained semantic coordination with textual descriptions and adjustments, while preserving expressive reaction with speaker behavior. Meanwhile, we propose the Emotional Intensity Tags (EIT) for emotion intensity editing with multi-modal information integration, applying to both text descriptions and listener motion amplitude. Extensive experiments conducted on our newly collected ListenerX dataset demonstrate that VividListener achieves state-of-the-art performance, realizing expressive and controllable listener dynamics. Xingqun Qi, Bingkun Yang, Weile Chen, Zezhao Tian, Muyi Sun, Man Zhang 0005, Zhenan Sun |
AAAI | 6 |
| 2026 | Rating-aware argument generation for movie reviews with multimodal large language models and a new dataset
Wenjie Hua, Quan Fang, Muyi Sun, Shibiao Xu, Man Zhang 0005 |
Multim. Syst. | 3 |
| 2026 | Pest manager: A systematic framework for precise pest monitoring in invisible grain pile storage environments
Chuanyang Ma, Xingqun Qi, Muyi Sun, Huiling Zhou |
Pattern Recognit. | 4 |
| 2025 | DanceEditor: Towards Iterative Editable Music-Driven Dance Generation with Open-Vocabulary DescriptionsabstractGenerating coherent and diverse human dances from music signals has gained tremendous progress in animating virtual avatars. While existing methods support direct dance synthesis, they fail to recognize that enabling users to edit dance movements is far more practical in real-world choreography scenarios. Moreover, the lack of high-quality dance datasets incorporating iterative editing also limits addressing this challenge. To achieve this goal, we first construct DanceRemix, a large-scale multiturn editable dance dataset comprising the prompt featuring over 25.3 M dance frames and 84.5 K pairs. In addition, we propose a novel framework for iterative and editable dance generation coherently aligned with given music signals, namely DanceEditor. Considering the dance motion should be both musical rhythmic and enable iterative editing by user descriptions, our framework is built upon a prediction-then-editing paradigm unifying multimodal conditions. At the initial prediction stage, our framework improves the authority of generated results by directly modeling dance movements from tailored, aligned music. Moreover, at the subsequent iterative editing stages, we incorporate text descriptions as conditioning information to draw the editable results through a specifically designed Cross-modality Editing Module (CEM). Specifically, CEM adaptively integrates the initial prediction with music and text prompts as temporal motion cues to guide the synthesized sequences. Thereby, the results display music harmonics while preserving fine-grained semantic alignment with text descriptions. Extensive experiments demonstrate that our method outperforms the state-of-the-art models on our newly collected DanceRemix dataset. Code is available at https://lzvsdy.github.io/DanceEditor/. Xingqun Qi, Muyi Sun, Siye Wang, Man Zhang 0005, Sirui Han |
ICCV | 5 |
| 2025 | ReMeREC: Relation-aware and Multi-entity Referring Expression ComprehensionabstractReferring Expression Comprehension (REC) aims to localize specified entities or regions from the source image according to the given natural language descriptions. While existing methods enable single-entity localization, they overlook modeling the complex inter-entity relationship in more practical multi-entity scenes, which limits their ability to produce accurate and reliable results. Moreover, the lack of high-quality multi-entity datasets incorporating fine-grained and paired image-text-relation annotations also limits addressing this challenge. To achieve this task, we first manually construct a relation-aware multi-entity REC dataset with fine-grained relation and text annotations, namely ReMeX. Additionally, we propose ReMeREC, a novel framework that effectively integrates textual and visual cues to localize multiple entities while capturing their inter-relationship. Specifically, to mitigate the semantic ambiguity arising from the absence of explicit entity boundaries in the source natural language description, we introduce a novel Text-adaptive Multi-entity Perceptron (TMP). TMP dynamically infers both the quantity and span of entities from corresponding fine-grained text cues, thus deriving representations that preserve the unique characteristics of each entity. Meanwhile, we design the Entity Inter-relationship Reasoner (EIR) to enhance semantic distinctiveness relationship modeling, leading to a more profound perception of the global scene. Furthermore, to better capture the fine-grained linguistic prompts for delineating multiple entity boundaries and inter-relationship, we leverage LLMs to generate a small-scale textual dataset, dubbed EntityText, which serves as an effective auxiliary resource and further improves the textual understanding. Extensive experiments conducted on four benchmark datasets demonstrate the superior performance of our framework. Remarkably, ReMeREC achieves outstanding results in multi-entity grounding and complex relationship prediction, outperforming other counterparts by a large margin. Yizhi Hu, Zezhao Tian, Xingqun Qi, Bingkun Yang, Junhui Yin, Muyi Sun, Man Zhang 0005, Zhenan Sun |
ACM Multimedia | 7 |
| 2025 | Correction: Open-Vocabulary Text-Driven Human Image Generation
Kaiduo Zhang, Muyi Sun, Jianxin Sun 0003, Kunbo Zhang, Zhenan Sun, Tieniu Tan |
Int. J. Comput. Vis. | 2 |
| 2025 | AnyFace++: A Unified Framework for Free-Style Text-to-Face Synthesis and ManipulationabstractHuman faces contain rich semantic information that could hardly be described without a large vocabulary and complex sentence patterns. However, most existing text-to-image synthesis methods could only generate meaningful results based on limited sentence templates with words contained in the training set, which heavily impairs the generalization ability of these models. In this paper, we define a novel 'free-style' text-to-face generation and manipulation problem, and propose an effective solution, named AnyFace++, which is applicable to a much wider range of open-world scenarios. The CLIP model is involved in AnyFace++ for learning an aligned language-vision feature space, which also expands the range of acceptable vocabulary as it is trained on a large-scale dataset. To further improve the granularity of semantic alignment between text and images, a memory module is incorporated to convert the description with arbitrary length, format, and modality into regularized latent embeddings representing discriminative attributes of the target face. Moreover, the diversity and semantic consistency of generation results are improved by a novel semi-supervised training scheme and a series of newly proposed objective functions. Compared to state-of-the-art methods, AnyFace++ is capable of synthesizing and manipulating face images based on more flexible descriptions and producing realistic images with higher diversity. Jianxin Sun 0003, Qiyao Deng, Qi Li 0005, Muyi Sun, Yunfan Liu 0001, Zhenan Sun |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Multi-Scale Semantic-Guidance Networks: Robust Blind Face Restoration Against Adversarial AttacksabstractImage processing networks are known to be vulnerable to adversarial examples, where adding carefully crafted adversarial perturbations to the inputs can mislead the model. This paper addresses the problem of robust blind face restoration (BFR) against adversarial attacks. BFR refers to recovering the HQ images from the LQ images, which suffer from diverse unknown degradation, such as noise, blur, artifact removal, low resolution, etc. Although existing BFR methods exhibit good performance, they experience significant degradation when subtle distortions and perturbations are introduced into the input images. This paper is the first to investigate, improve comprehensively, and evaluate BFR methods towards adversarial attacks. Project Gradient Descent (PGD) is employed to generate adversarial examples, and multiple types of attacks were used to thoroughly assess the robustness of various BFR methods across different objectives, regions, and levels. We evaluate the robustness of multiple BFR methods and analyze the advantages of their structures and modules towards adversarial attacks. Experimental results demonstrate that the method utilizing latent feature encoding and pre-trained discrete HQ codebook achieves better robustness than other methods, with the latter outperforming the former. Similarly, multi-scale semantic guidance information also exhibits superior performance in enhancing robustness. Therefore, we propose a powerful BFR method to mitigate this issue while maintaining better performance. Extensive experiments on three real-world datasets demonstrate our method’s state-of-the-art robustness in different scenarios. Zhenyuan Zhang 0001, Xingqun Qi, Zhenbo Song, Zhiqin Yang, Jianfeng Lu 0003, Muyi Sun, Man Zhang 0005, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Biphasic Face Photo-Sketch Synthesis via Semantic-Driven Generative Adversarial Network With Graph Representation LearningabstractBiphasic face photo-sketch synthesis has significant practical value in wide-ranging fields such as digital entertainment and law enforcement. Previous approaches directly generate the photo-sketch in a global view, they always suffer from the low quality of sketches and complex photograph variations, leading to unnatural and low-fidelity results. In this article, we propose a novel semantic-driven generative adversarial network to address the above issues, cooperating with graph representation learning. Considering that human faces have distinct spatial structures, we first inject class-wise semantic layouts into the generator to provide style-based spatial information for synthesized face photographs and sketches. In addition, to enhance the authenticity of details in generated faces, we construct two types of representational graphs via semantic parsing maps upon input faces, dubbed the intraclass semantic graph (IASG) and the interclass structure graph (IRSG). Specifically, the IASG effectively models the intraclass semantic correlations of each facial semantic component, thus producing realistic facial details. To preserve the generated faces being more structure-coordinated, the IRSG models interclass structural relations among every facial component by graph representation learning. To further enhance the perceptual quality of synthesized images, we present a biphasic interactive cycle training strategy by fully taking advantage of the multilevel feature consistency between the photograph and sketch. Extensive experiments demonstrate that our method outperforms the state-of-the-art competitors on the CUHK Face Sketch (CUFS) and CUHK Face Sketch FERET (CUFSF) datasets. Xingqun Qi, Muyi Sun, Zijian Wang 0009, Jiaming Liu 0003, Qi Li 0005, Fang Zhao 0006, Shanghang Zhang, Caifeng Shan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-Wise Pruning Error MetricabstractVision-language pretrained models have achieved impressive performance on various downstream tasks. However, their large model sizes hinder their utilization on platforms with limited computational resources. We find that directly using smaller pretrained models and applying magnitude-based pruning on CLIP models leads to in-flexibility and inferior performance. Recent efforts for VLP compression either adopt uni-modal compression metrics resulting in limited performance or involve costly mask-search processes with learnable masks. In this paper, we first propose the Module-wise Pruning Error (MoPE) met-ric, accurately assessing CLIP module importance by performance decline on cross-modal tasks. Using the MoPE metric, we introduce a unified pruning framework applica-ble to both pretraining and task-specific fine-tuning compression stages. For pretraining, MoPE-CLIP effectively leverages knowledge from the teacher model, significantly reducing pretraining costs while maintaining strong zero-shot capabilities. For fine-tuning, consecutive pruning from width to depth yields highly competitive task-specific models. Extensive experiments in two stages demonstrate the effectiveness of the MoPE metric, and MoPE-CLIP outperforms previous state-of-the-art VLP compression methods. Haokun Lin, Haoli Bai, Zhili Liu, Lu Hou 0002, Muyi Sun, Linqi Song, Ying Wei 0001, Zhenan Sun |
CVPR | 5 |
| 2024 | PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the WildabstractVideo quality assessment (VQA) is a challenging problem due to the numerous factors that can affect the perceptual quality of a video, e.g., content attractiveness, distortion type, motion pattern, and level. However, annotating the Mean opinion score (MOS) for videos is expensive and time-consuming, which limits the scale of VQA datasets, and poses a significant obstacle for deep learning-based methods. In this paper, we propose a VQA method named PTM-VQA, which leverages PreTrained Models to transfer knowledge from models pretrained on various pre-tasks, enabling benefits for VQA from different aspects. Specifically, we extract features of videos from different pretrained models with frozen weights and integrate them to generate representation. Since these models possess var-ious fields of knowledge and are often trained with labels irrelevant to quality, we propose an Intra-Consistency and Inter-Divisibility (ICID) loss to impose constraints on features extracted by multiple pretrained models. The intra-consistency constraint ensures that features extracted by different pretrained models are in the same unified quality-aware latent space, while the inter-divisibility introduces pseudo clusters based on the annotation of samples and tries to separate features of samples from different clusters. Furthermore, with a constantly growing number of pretrained models, it is crucial to determine which models to use and how to use them. To address this problem, we propose an efficient scheme to select suitable candidates. Models with better clustering performance on VQA datasets are chosen to be our candidates. Extensive experiments demonstrate the effectiveness of the proposed method. Kun Yuan 0003, Mading Li, Muyi Sun, Ming Sun 0008, Jiachao Gong, Jinhua Hao, Chao Zhou 0003, Yansong Tang |
CVPR | 4 |
| 2024 | Contrmix: Progressive Mixed Contrastive Learning for Semi-Supervised Medical Image SegmentationabstractWhile medical image segmentation has achieved impressive progress, it usually being constrained by labor-intensive and costly pixel-wise annotations. The existing semi-supervised learning methods ignore the inherent imbalance and high similarity of different categories in medical images. To address the above issues, we present a Progressive Mixed Contrastive Learning (ContrMix) framework, which contains a Cycle-mix module and a mix-based Contrastive Learning module. In Cycle-mix, a progressive mixing strategy with a cycle loss is designed to enforce the consistency between the mixed segmentation and corresponding generated mixing samples, effectively enhancing the ability to learn geometric features of the imbalanced medical data. We also introduce a mix-based Contrastive Learning module that learns the inter-instance similarities between the mixed patches and the original ones, which encourages the model to learn background-invariant representations from samples under different distortions and improves the semantic discrimination of high similarity categories. We conduct extensive experiments on the ACDC dataset and LA dataset and our method outperforms other state-of-the-art semi-supervised approaches. Meisheng Zhang, Chenye Wang, Wenxuan Zou, Xingqun Qi, Muyi Sun |
ICASSP | 5 |
| 2024 | EditHuman: Fine-Grained Text-Driven Human Video EditingabstractRecently, video editing has made significant advances. Human character, as one of the core elements in video editing, has attracted great research attention. However, when editing characters with strong structural information, previous methods generally encounter blurring and distortion in the limbs. In this paper, we present EditHuman, a model to realize fine-grained text-driven human video editing tasks, which achieves continuous pose movements and high-quality limb expression. Considering complex body structures and continuity of motion, more precise designs are needed to obtain practical performance. Specifically, we propose a Cascaded UNet (CAU) to realize a coarse-to-fine denoising process and refined noise estimation. Meanwhile, we introduce two Heatmap-Centric Attention Modules called Key-Element Attention (KEA) and Key-Temporal Attention (KTA) to enhance the quality of human limb expression and inter-frame continuity. Moreover, we utilize the estimated heatmap to guide the noise prediction, which further refines the video quality. Extensive experiments show that EditHuman has achieved the SOTA performance. Kaiduo Zhang, Muyi Sun, Junxing Hu, Kunbo Zhang, Zhenan Sun |
IJCB | 2 |
| 2024 | Toward Steady Video Content Delivery for High-Spend VehiclesabstractDespite the development of mobile networks, it is still challenging for current mobile networks to cope with enormous amounts of vehicle video traffic. In this paper, we present a video delivery solution, which is designed to provide steady video content delivery for high-speed vehicles. Our solution improves the steadiness of the video deliveries to vehicle users by a multi-path concurrent delivery mechanism and a path-based content pre-caching mechanism. The concurrent content delivery and content pre-caching can shorten the whole delivery time. However, it sometimes brings heavy stress to the network if the advanced delivery is not controlled. We propose a distributed delivery rate control method to solve the above problem. By the proposed content pre-caching mechanism, each basestation caches appropriate video chunks such that these cached chunks can be used with high probability when the vehicle is within the coverage of the basestation. Xinchang Zhang 0001, Muyi Sun, Bingyu He |
VTC Spring | 3 |
| 2024 | Open-Vocabulary Text-Driven Human Image Generation
Kaiduo Zhang, Muyi Sun, Jianxin Sun 0003, Kunbo Zhang, Zhenan Sun, Tieniu Tan |
Int. J. Comput. Vis. | 2 |
| 2024 | An Automated Framework for Histopathological Nucleus Segmentation With Deep Attention Integrated NetworksabstractClinical management and accurate disease diagnosis are evolving from qualitative stage to the quantitative stage, particularly at the cellular level. However, the manual process of histopathological analysis is lab-intensive and time-consuming. Meanwhile, the accuracy is limited by the experience of the pathologist. Therefore, deep learning-empowered computer-aided diagnosis (CAD) is emerging as an important topic in digital pathology to streamline the standard process of automatic tissue analysis. Automated accurate nucleus segmentation can not only help pathologists make more accurate diagnosis, save time and labor, but also achieve consistent and efficient diagnosis results. However, nucleus segmentation is susceptible to staining variation, uneven nucleus intensity, background noises, and nucleus tissue differences in biopsy specimens. To solve these problems, we propose Deep Attention Integrated Networks (DAINets), which mainly built on self-attention based spatial attention module and channel attention module. In addition, we also introduce a feature fusion branch to fuse high-level representations with low-level features for multi-scale perception, and employ the mark-based watershed algorithm to refine the predicted segmentation maps. Furthermore, in the testing phase, we design Individual Color Normalization (ICN) to settle the dyeing variation problem in specimens. Quantitative evaluations on the multi-organ nucleus dataset indicate the priority of our automated nucleus segmentation framework. Muyi Sun, Wenxuan Zou, Song Wang 0006, Zhenan Sun |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | Exploring Generalizable Distillation for Efficient Medical Image SegmentationabstractEfficient medical image segmentation aims to provide accurate pixel-wise predictions with a lightweight implementation framework. However, existing lightweight networks generally overlook the generalizability of the cross-domain medical segmentation tasks. In this paper, we propose Generalizable Knowledge Distillation (GKD), a novel framework for enhancing the performance of lightweight networks on cross-domain medical segmentation by generalizable knowledge distillation from powerful teacher networks. Considering the domain gaps between different medical datasets, we propose the Model-Specific Alignment Networks (MSAN) to obtain the domain-invariant representations. Meanwhile, a customized Alignment Consistency Training (ACT) strategy is designed to promote the MSAN training. Based on the domain-invariant vectors in MSAN, we propose two generalizable distillation schemes, Dual Contrastive Graph Distillation (DCGD) and Domain-Invariant Cross Distillation (DICD). In DCGD, two implicit contrastive graphs are designed to model the intra-coupling and inter-coupling semantic correlations. Then, in DICD, the domain-invariant semantic vectors are reconstructed from two networks (i.e., teacher and student) with a crossover manner to achieve simultaneous generalization of lightweight networks, hierarchically. Moreover, a metric named Fréchet Semantic Distance (FSD) is tailored to verify the effectiveness of the regularized domain-invariant features. Extensive experiments conducted on the Liver, Retinal Vessel and Colonoscopy segmentation datasets demonstrate the superiority of our method, in terms of performance and generalization ability on lightweight networks. Xingqun Qi, Zhuojie Wu, Wenxuan Zou, Yifan Gao 0003, Muyi Sun, Shanghang Zhang, Caifeng Shan, Zhenan Sun |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Diverse 3D Hand Gesture Prediction from Body Dynamics by Bilateral Hand DisentanglementabstractPredicting natural and diverse 3D hand gestures from the upper body dynamics is a practical yet challenging task in virtual avatar creation. Previous works usually overlook the asymmetric motions between two hands and generate two hands in a holistic manner, leading to unnatural results. In this work, we introduce a novel bilateral hand disentanglement based two-stage 3D hand generation method to achieve natural and diverse 3D hand prediction from body dynamics. In the first stage, we intend to generate natural hand gestures by two hand-disentanglement branches. Considering the asymmetric gestures and motions of two hands, we introduce a Spatial-Residual Memory (SRM) module to model spatial interaction between the body and each hand by residual learning. To enhance the coordination of two hand motions wrt. body dynamics holistically, we then present a Temporal-Motion Memory (TMM) module. TMM can effectively model the temporal association between body dynamics and two hand motions. The second stage is built upon the insight that 3D hand predictions should be non-deterministic given the sequential body postures. Thus, we further diversify our 3D hand predictions based on the initial output from the stage one. Concretely, we propose a Prototypical-Memory Sampling Strategy (PSS) to generate the non-deterministic hand gestures by gradient-based Markov Chain Monte Carlo (MCMC) sampling. Extensive experiments demonstrate that our method outperforms the state-of-the-art models on the B2H dataset and our newly collected TED Hands dataset. The dataset and code are available at: https://github.com/XingqunQilab/Diverse-3D-Hand-Gesture-Prediction. Xingqun Qi, Chen Liu 0028, Muyi Sun, Lincheng Li, Changjie Fan, Xin Yu 0002 |
CVPR | 3 |
| 2023 | Lightvessel: Exploring Lightweight Coronary Artery Vessel Segmentation Via Similarity Knowledge DistillationabstractIn recent years, deep convolution neural networks (DCNNs) have achieved great prospects in coronary artery vessel segmentation. However, it is difficult to deploy complicated models in clinical scenarios since high-performance approaches have excessive parameters and high computation costs. To tackle this problem, we propose LightVessel, a Similarity Knowledge Distillation Framework, for lightweight coronary artery vessel segmentation. Primarily, we propose a Feature-wise Similarity Distillation (FSD) module for semantic-shift modeling. Specifically, we calculate the feature similarity between the symmetric layers from the encoder and decoder. Then the similarity is transferred as knowledge from a cumbersome teacher network to a non-trained lightweight student network. Meanwhile, for encouraging the student model to learn more pixel-wise semantic information, we introduce the Adversarial Similarity Distillation (ASD) module. Concretely, the ASD module aims to construct the spatial adversarial correlation between the annotation and prediction from the teacher and student models, respectively. Through the ASD module, the student model obtains fined-grained subtle edge segmented results of the coronary artery vessel. Extensive experiments conducted on Clinical Coronary Artery Vessel Dataset demonstrate that LightVessel outperforms various knowledge distillation counterparts. Hao Dang, Yuekai Zhang, Xingqun Qi, Muyi Sun |
ICASSP | 5 |
| 2023 | Sensing Micro-Motion Human Patterns using Multimodal mmRadar and Video Signal for Affective and Psychological IntelligenceabstractAffective and psychological perception are pivotal in human-machine interaction and essential domains within artificial intelligence. Existing physiological signal-based affective and psychological datasets primarily rely on contact-based sensors, potentially introducing extraneous affectives during the measurement process. Consequently, creating accurate non-contact affective and psychological perception datasets is crucial for overcoming these limitations and advancing affective intelligence. In this paper, we introduce the Remote Multimodal Affective and Psychological (ReMAP) dataset, for the first time, apply head micro-tremor (HMT) signals for affective and psychological perception. ReMAP features 68 participants and comprises two sub-datasets. The stimuli videos utilized for affective perception undergo rigorous screening to ensure the efficacy and universality of affective elicitation. Additionally, we propose a novel remote affective and psychological perception framework, leveraging multimodal complementarity and interrelationships to enhance affective and psychological perception capabilities. Extensive experiments demonstrate HMT as a "small yet powerful" physiological signal in psychological perception. Our method outperforms existing state-of-the-art approaches in remote affective recognition and psychological perception. The ReMAP dataset is publicly accessible at https://remap-dataset.github.io/ReMAP. Yiwei Ru, Peipei Li 0002, Muyi Sun, Yunlong Wang 0003, Kunbo Zhang, Qi Li 0005, Zhaofeng He 0001, Zhenan Sun |
ACM Multimedia | 3 |
| 2023 | Trustworthy Localization With EM-Based Federated Control Scheme for IIoTsabstractIndustrial Internet of Things (IIoTs) are significantly changing informative and manufacturing pattern in smart factories while it also brings security and trustworthiness issue. Concerning about trustworthiness issues and private preservation of tracking systems, a hierarchical framework with federated control theory is designed, which consists of a federated control center, network layer, and a federated control node. The framework combines a collaborative Cloud-Edge-End structure and machine learning-oriented localization, which further forms the EM-based federated scheme. On this basis, a trustworthy localization model is built with the untrustworthiness probability as a latent variable. By exploring expectation maximization (EM) of trustworthy localization, the local messages and aggression equations are derived in an iterative way of federated learning. The EM-based federated control scheme with machine learning-oriented localization is finally given. Experiments have been conducted to prove the localization accuracy and convergence of the proposed method with trustworthiness issue. The results show that trustworthy localization outperforms traditional methods without considering the security threats. Song Wang 0006, Zhiyao Zhao, Muyi Sun |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Graph Flow: Cross-Layer Graph Flow Distillation for Dual Efficient Medical Image SegmentationabstractWith the development of deep convolutional neural networks, medical image segmentation has achieved a series of breakthroughs in recent years. However, high-performance convolutional neural networks always mean numerous parameters and high computation costs, which will hinder the applications in resource-limited medical scenarios. Meanwhile, the scarceness of large-scale annotated medical image datasets further impedes the application of high-performance networks. To tackle these problems, we propose Graph Flow, a comprehensive knowledge distillation framework, for both network-efficiency and annotation-efficiency medical image segmentation. Specifically, the Graph Flow Distillation transfers the essence of cross-layer variations from a well-trained cumbersome teacher network to a non-trained compact student network. In addition, an unsupervised Paraphraser Module is integrated to purify the knowledge of the teacher, which is also beneficial for the training stabilization. Furthermore, we build a unified distillation framework by integrating the adversarial distillation and the vanilla logits distillation, which can further refine the final predictions of the compact network. With different teacher networks (traditional convolutional architecture or prevalent transformer architecture) and student networks, we conduct extensive experiments on four medical image datasets with different modalities (Gastric Cancer, Synapse, BUSI, and CVC-ClinicDB). We demonstrate the prominent ability of our method on these datasets, which achieves competitive performances. Moreover, we demonstrate the effectiveness of our Graph Flow through a novel semi-supervised paradigm for dual efficient medical image segmentation. Our code will be available at Graph Flow. Wenxuan Zou, Xingqun Qi, Muyi Sun, Zhenan Sun, Caifeng Shan |
IEEE Trans. Medical Imaging | 4 |
| 2022 | ShowFace: Coordinated Face Inpainting with Memory-Disentangled Refinement Networks
Zhuojie Wu, Xingqun Qi, Zijian Wang 0009, Kun Yuan 0003, Muyi Sun, Zhenan Sun |
BMVC | 6 |
| 2022 | AnyFace: Free-style Text-to-Face Synthesis and ManipulationabstractExisting text-to-image synthesis methods generally are only applicable to words in the training dataset. However, human faces are so variable to be described with limited words. So this paper proposes the first free-style text-to-face method namely AnyFace enabling much wider open world applications such as metaverse, social media, cosmetics, forensics, etc. AnyFace has a novel two-stream framework for face image synthesis and manipulation given arbitrary descriptions of the human face. Specifically, one stream performs text-to-face generation and the other conducts face image reconstruction. Facial text and image features are extracted using the CLIP (Contrastive Language-Image Pre-training) encoders. And a collaborative Cross Modal Distillation (CMD) module is designed to align the linguistic and visual features across these two streams. Furthermore, a Diverse Triplet Loss (DT loss) is developed to model fine-grained features and improve facial diversity. Extensive experiments on Multi-modal CelebA-HQ and CelebAText-HQ demonstrate significant advantages of AnyFace over state-of-the-art methods. AnyFace can achieve high-quality, high-resolution, and high-diversity face synthesis and manipulation results without any constraints on the number and content of input captions. Jianxin Sun 0003, Qiyao Deng, Qi Li 0005, Muyi Sun, Zhenan Sun |
CVPR | 4 |
| 2022 | Self-supervised Correlation Mining Network for Person Image GenerationabstractPerson image generation aims to perform non-rigid deformation on source images, which generally requires unaligned data pairs for training. Recently, self-supervised methods express great prospects in this task by merging the disentangled representations for self-reconstruction. However, such methods fail to exploit the spatial correlation between the disentangled features. In this paper, we propose a Self-supervised Correlation Mining Network (SCM-Net) to rearrange the source images in the feature space, in which two collaborative modules are integrated, Decomposed Style Encoder (DSE) and Correlation Mining Module (CMM). Specifically, the DSE first creates unaligned pairs at the feature level. Then, the CMM establishes the spatial correlation field for feature rearrangement. Eventually, a translation module transforms the rearranged features to realistic results. Meanwhile, for improving the fidelity of cross-scale pose transformation, we propose a graph based Body Structure Retaining Loss (BSR Loss) to preserve reasonable body structures on half body to full body generation. Extensive experiments conducted on DeepFashion dataset demonstrate the superiority of our method compared with other supervised and unsupervised approaches. Furthermore, satisfactory results on face generation show the versatility of our method in other deformation tasks. Zijian Wang 0009, Xingqun Qi, Kun Yuan 0003, Muyi Sun |
CVPR | 4 |
| 2022 | MOST-Net: A Memory Oriented Style Transfer Network for Face Sketch SynthesisabstractFace sketch synthesis has been widely used in multimedia entertainment and law enforcement. Despite the recent developments in deep neural networks, accurate and realistic face sketch synthesis is still a challenging task due to the diversity and complexity of human faces. Current image-to-image translation-based face sketch synthesis frequently encounters over-fitting problems when it comes to small-scale datasets. To tackle this problem, we present an end-to-end Memory Oriented Style Transfer Network (MOST-Net) for face sketch synthesis which can produce high-fidelity sketches with limited data. Specifically, an external self-supervised dynamic memory module is introduced to capture the domain alignment knowledge in the long term. In this way, our proposed model could obtain the domain-transfer ability by establishing the durable relationship between faces and corresponding sketches on the feature level. Furthermore, we design a novel Memory Refinement Loss (MR Loss) for feature alignment in the memory module, which enhances the accuracy of memory slots in an unsupervised manner. Extensive experiments on the CUFS and the CUFSF datasets show that our MOST-Net achieves state-of-the-art performance, especially in terms of the Structural Similarity Index(SSIM). Fan Ji, Muyi Sun, Xingqun Qi, Qi Li 0005, Zhenan Sun |
ICPR | 2 |
| 2022 | A Unified Framework for Biphasic Facial Age Translation With Noisy-Semantic Guided Generative Adversarial NetworksabstractBiphasic facial age translation aims at predicting the appearance of the input face at any age. Facial age translation has received considerable research attention in the last decade due to its practical value in cross-age face recognition and various entertainment applications. However, most existing methods model age changes between holistic images, regardless of the human face structure and the age-changing patterns of individual facial components. Consequently, the lack of semantic supervision will cause infidelity of generated faces in detail. To this end, we propose a unified framework for biphasic facial age translation with noisy-semantic guided generative adversarial networks. Structurally, we project the class-aware noisy semantic layouts to “soft” latent maps for the following injection operation on the individual facial parts. In particular, we introduce two sub-networks, ProjectionNet and ConstraintNet. ProjectionNet introduces the low-level structural semantic information with noise map and produces “soft” latent maps. ConstraintNet disentangles the high-level spatial features to constrain the “soft” latent maps, which endows more age-related context into the “soft” latent maps. Specifically, attention mechanism is employed in ConstraintNet for feature disentanglement. Meanwhile, in order to mine the strongest mapping ability of the network, we embed two types of learning strategies in the training procedure, supervised self-driven generation and unsupervised condition-driven cycle-consistent generation. As a result, extensive experiments conducted on MORPH and CACD datasets demonstrate the prominent ability of our proposed method which achieves state-of-the-art performance. Muyi Sun, Jianshu Li, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | PAENet: A Progressive Attention-Enhanced Network for 3D to 2D Retinal Vessel Segmentationabstract3D to 2D retinal vessel segmentation is a challenging problem in Optical Coherence Tomography Angiography (OCTA) images. Accurate retinal vessel segmentation is important for the diagnosis and prevention of ophthalmic diseases. However, making full use of the 3D data of OCTA volumes is a vital factor for obtaining satisfactory segmentation results. In this paper, we propose a Progressive Attention-Enhanced Network (PAENet) based on attention mechanisms to extract rich feature representation. Specifically, the framework consists of two main parts, the three-dimensional feature learning path and the two-dimensional segmentation path. In the three-dimensional feature learning path, we design a novel Adaptive Pooling Module (APM) and propose a new Quadruple Attention Module (QAM). The APM captures dependencies along the projection direction of volumes and learns a series of pooling coefficients for feature fusion, which efficiently reduces feature dimension. In addition, the QAM reweights the features by capturing four-group cross-dimension dependencies, which makes maximum use of 4D feature tensors. In the two-dimensional segmentation path, to acquire more detailed information, we propose a Feature Fusion Module (FFM) to inject 3D information into the 2D path. Meanwhile, we adopt the Polarized Self-Attention (PSA) block to model the semantic interdependencies in spatial and channel dimensions respectively. Experimentally, our extensive experiments on the OCTA-500 dataset show that our proposed algorithm achieves state-of-the-art performance compared with previous methods. Zhuojie Wu, Zijian Wang 0009, Wenxuan Zou, Fan Ji, Hao Dang, Muyi Sun |
BIBM | 7 |
| 2021 | CoCo DistillNet: a Cross-layer Correlation Distillation Network for Pathological Gastric Cancer SegmentationabstractIn recent years, deep convolutional neural networks have made significant advances in pathology image segmentation. However, pathology image segmentation encounters with a dilemma in which the higher-performance networks generally require more computational resources and storage. This phenomenon limits the employment of high-accuracy networks in real scenes due to the inherent high-resolution of pathological images. To tackle this problem, we propose CoCo DistillNet, a novel Cross-layer Correlation (CoCo) knowledge distillation network for pathological gastric cancer segmentation. Knowledge distillation, a general technique which aims at improving the performance of a compact network through knowledge transfer from a cumbersome network. Concretely, our CoCo DistillNet models the correlations of channel-mixed spatial similarity between different layers and then transfers this knowledge from a pre-trained cumbersome teacher network to a non-trained compact student network. In addition, we also utilize the adversarial learning strategy to further prompt the distilling procedure which is called Adversarial Distillation (AD). Furthermore, to stabilize our training procedure, we make the use of the unsupervised Paraphraser Module (PM) to boost the knowledge paraphrase in the teacher network. As a result, extensive experiments conducted on the Gastric Cancer Segmentation Dataset demonstrate the prominent ability of CoCo DistillNet which achieves state-of-the-art performance. Wenxuan Zou, Xingqun Qi, Zhuojie Wu, Zijian Wang 0009, Muyi Sun, Caifeng Shan |
BIBM | 5 |
| 2021 | Face Sketch Synthesis via Semantic-Driven Generative Adversarial NetworkabstractFace sketch synthesis has made significant progress with the development of deep neural networks in these years. The delicate depiction of sketch portraits facilitates a wide range of applications like digital entertainment and law enforcement. However, accurate and realistic face sketch generation is still a challenging task due to the illumination variations and complex backgrounds in the real scenes. To tackle these challenges, we propose a novel Semantic-Driven Generative Adversarial Network (SDGAN) which embeds global structure-level style injection and local class-level knowledge re-weighting. Specifically, we conduct facial saliency detection on the input face photos to provide overall facial texture structure, which could be used as a global type of prior information. In addition, we exploit face parsing layouts as the semantic-level spatial prior to enforce globally structural style injection in the generator of SDGAN. Furthermore, to enhance the realistic effect of the details, we propose a novel Adaptive Re-weighting Loss (ARLoss) which dedicates to balance the contributions of different semantic classes. Experimentally, our extensive experiments on CUFS and CUFSF datasets show that our proposed algorithm achieves state-of-the-art performance. Xingqun Qi, Muyi Sun, Weining Wang 0001, Xiaoxiao Dong, Qi Li 0005, Caifeng Shan |
IJCB | 2 |
| 2021 | Contextual information enhanced convolutional neural networks for retinal vessel segmentation in color fundus images
Muyi Sun, Kaiqi Li, Xingqun Qi, Hao Dang, Guanhong Zhang |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | Accurate Retinal Vessel Segmentation in Color Fundus Images via Fully Attention-Based NetworksabstractAutomatic retinal vessel segmentation is important for the diagnosis and prevention of ophthalmic diseases. The existing deep learning retinal vessel segmentation models always treat each pixel equally. However, the multi-scale vessel structure is a vital factor affecting the segmentation results, especially in thin vessels. To address this crucial gap, we propose a novel Fully Attention-based Network (FANet) based on attention mechanisms to adaptively learn rich feature representation and aggregate the multi-scale information. Specifically, the framework consists of the image pre-processing procedure and the semantic segmentation networks. Green channel extraction (GE) and contrast limited adaptive histogram equalization (CLAHE) are employed as pre-processing to enhance the texture and contrast of retinal blood images. Besides, the network combines two types of attention modules with the U-Net. We propose a lightweight dual-direction attention block to model global dependencies and reduce intra-class inconsistencies, in which the weights of feature maps are updated based on the semantic correlation between pixels. The dual-direction attention block utilizes horizontal and vertical pooling operations to produce the attention map. In this way, the network aggregates global contextual information from semantic-closer regions or a series of pixels belonging to the same object category. Meanwhile, we adopt the selective kernel (SK) unit to replace the standard convolution for obtaining multi-scale features of different receptive field sizes generated by soft attention. Furthermore, we demonstrate that the proposed model can effectively identify irregular, noisy, and multi-scale retinal vessels. The abundant experiments on DRIVE, STARE, and CHASE_DB1 datasets show that our method achieves state-of-the-art performance. Kaiqi Li, Xingqun Qi, Yiwen Luo, Zeyi Yao, Xiaoguang Zhou, Muyi Sun |
IEEE J. Biomed. Health Informatics | 6 |
| 2020 | Accurate Cell Segmentation in Digital Pathology Images via Attention Enforced NetworksabstractAutomatic cell segmentation is an essential step in the pipeline of computer-aided diagnosis (CAD), such as the detection and grading of breast cancer. Accurate segmentation of cells can not only assist the pathologists to make a more precise diagnosis, but also save much time and labor. However, this task suffers from stain variation, cell inhomogeneous intensities, background clutters and cells from different tissues. To address these issues, we propose an Attention Enforced Network (AENet), which is built on spatial attention module and channel attention module, to integrate local features with global dependencies and weight effective channels adaptively. Besides, we introduce a feature fusion branch to bridge high-level and low-level features. Finally, the marker controlled watershed algorithm is applied to post-process the predicted segmentation maps for reducing the fragmented regions. In the test stage, we present an individual color normalization method to deal with the stain variation problem. We evaluate this model on the MoNuSeg dataset. The quantitative comparisons against several prior methods demonstrate the superiority of our approach. Zeyi Yao, Kaiqi Li, Yiwen Luo, Xiaoguang Zhou, Muyi Sun, Guanhong Zhang |
ICPR | 5 |