EDBT 2026 Demo / reviewers in the wild / expert
Xingqun Qi
dblp:226/4403
· DBLP profile ↗
29ranked-venue papers
7as first author
29since 2021 · last 2026
0000-0002-9772-5707ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 18 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive InteractionabstractGenerating responsive listener head dynamics with nuanced emotions and expressive reactions is crucial for dialogue modeling in various virtual avatar animations. Previous studies mainly focus on the direct short-term production of listener behavior. They overlook the fine-grained control over motion variations and emotional intensity, especially in long-sequence modeling. Moreover, the lack of long-term and large-scale paired speaker-listener corpora incorporating head dynamics and fine-grained multi-modality annotations limits the application of dialogue modeling. Therefore, we first newly collect a large-scale multi-turn dataset of 3D dyadic conversation containing more than 1.4M valid frames for multi-modal responsive interaction, dubbed ListenerX. Additionally, we propose VividListener, a novel framework enabling fine-grained, expressive, and controllable listener dynamics modeling. This framework leverages multi-modal conditions as guiding principles for fostering coherent interactions between speakers and listeners. Specifically, we design the Responsive Interaction Module (RIM) to adaptively represent the multi-modal interactive embeddings. RIM ensures the listener dynamics achieve fine-grained semantic coordination with textual descriptions and adjustments, while preserving expressive reaction with speaker behavior. Meanwhile, we propose the Emotional Intensity Tags (EIT) for emotion intensity editing with multi-modal information integration, applying to both text descriptions and listener motion amplitude. Extensive experiments conducted on our newly collected ListenerX dataset demonstrate that VividListener achieves state-of-the-art performance, realizing expressive and controllable listener dynamics. Xingqun Qi, Bingkun Yang, Weile Chen, Zezhao Tian, Muyi Sun, Man Zhang 0005, Zhenan Sun |
AAAI | 2 |
| 2026 | Pest manager: A systematic framework for precise pest monitoring in invisible grain pile storage environments
Chuanyang Ma, Xingqun Qi, Muyi Sun, Huiling Zhou |
Pattern Recognit. | 3 |
| 2025 | PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit RemeshingabstractPhotorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, existing methods for monocular full-body reconstruction, typically relying on front and/or predicted back view, still struggle with satisfactory performance due to the ill-posed nature of the problem and sophisticated self-occlusions. In this paper, we propose PSHuman, a novel framework that explicitly reconstructs human meshes utilizing priors from the multiview diffusion model. It is found that directly applying multiview diffusion on single-view human images leads to severe geometric distortions, especially on generated faces. To address it, we propose a cross-scale diffusion that models the joint probability distribution of global full-body shape and local facial characteristics, enabling identity-preserved novel-view generation without geometric distortion. Moreover, to enhance cross-view body shape consistency of varied human poses, we condition the generative model on parametric models (SMPL-X), which provide body priors and prevent unnatural views inconsistent with human anatomy. Leveraging the generated multiview normal and color images, we present SMPLX-initialized explicit human carving to recover realistic textured human meshes efficiently. Extensive experiments on CAPE and THuman2.1 demonstrate PSHuman’s superiority in geometry details, texture fidelity, and generalization capability. Wangguandong Zheng, Yuan Liu 0025, Tao Yu 0007, Yangguang Li 0001, Xingqun Qi, Xiaowei Chi, Si-Yu Xia, Yan-Pei Cao 0001, Wei Xue 0002, Wenhan Luo, Yike Guo |
CVPR | 6 |
| 2025 | DanceEditor: Towards Iterative Editable Music-Driven Dance Generation with Open-Vocabulary DescriptionsabstractGenerating coherent and diverse human dances from music signals has gained tremendous progress in animating virtual avatars. While existing methods support direct dance synthesis, they fail to recognize that enabling users to edit dance movements is far more practical in real-world choreography scenarios. Moreover, the lack of high-quality dance datasets incorporating iterative editing also limits addressing this challenge. To achieve this goal, we first construct DanceRemix, a large-scale multiturn editable dance dataset comprising the prompt featuring over 25.3 M dance frames and 84.5 K pairs. In addition, we propose a novel framework for iterative and editable dance generation coherently aligned with given music signals, namely DanceEditor. Considering the dance motion should be both musical rhythmic and enable iterative editing by user descriptions, our framework is built upon a prediction-then-editing paradigm unifying multimodal conditions. At the initial prediction stage, our framework improves the authority of generated results by directly modeling dance movements from tailored, aligned music. Moreover, at the subsequent iterative editing stages, we incorporate text descriptions as conditioning information to draw the editable results through a specifically designed Cross-modality Editing Module (CEM). Specifically, CEM adaptively integrates the initial prediction with music and text prompts as temporal motion cues to guide the synthesized sequences. Thereby, the results display music harmonics while preserving fine-grained semantic alignment with text descriptions. Extensive experiments demonstrate that our method outperforms the state-of-the-art models on our newly collected DanceRemix dataset. Code is available at https://lzvsdy.github.io/DanceEditor/. Xingqun Qi, Muyi Sun, Siye Wang, Man Zhang 0005, Sirui Han |
ICCV | 3 |
| 2025 | Freedance: Towards Harmonic Free-Number Group Dance Generation Via a Unified Framework
Liting Wen, Xingqun Qi |
ICCV | 5 |
| 2025 | Co3Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion
Xingqun Qi, Yatian Wang, Wei Xue 0002, Shanghang Zhang, Wenhan Luo, Yike Guo |
ICLR | 1 |
| 2025 | Empowering World Models with Reflection for Embodied Video PredictionabstractVideo generation models have made significant progress in simulating future states, showcasing their potential as world simulators in embodied scenarios. However, existing models often lack robust understanding, limiting their ability to perform multi-step predictions or handle Out-of-Distribution (OOD) scenarios. To address this challenge, we propose the Reflection of Generation (RoG), a set of intermediate reasoning strategies designed to enhance video prediction. It leverages the complementary strengths of pre-trained vision-language and video generation models, enabling them to function as a world model in embodied scenarios. To support RoG, we introduce Embodied Video Anticipation Benchmark(EVA-Bench), a comprehensive benchmark that evaluates embodied world models across diverse tasks and scenarios, utilizing both in-domain and OOD datasets. Building on this foundation, we devise a world model, Embodied Video Anticipator (EVA), that follows a multistage training paradigm to generate high-fidelity video frames and apply an autoregressive strategy to enable adaptive generalization for longer video sequences. Extensive experiments demonstrate the efficacy of EVA in various downstream tasks like video generation and robotics, thereby paving the way for large-scale pre-trained models in real-world video prediction applications. The video demos are available at https://sites.google.com/view/icml-eva. Xiaowei Chi, Chun-Kai Fan, Xingqun Qi, Rongyu Zhang, Anthony Chen, Chi-Min Chan, Wei Xue 0002, Shanghang Zhang, Yike Guo |
ICML | 4 |
| 2025 | ReMeREC: Relation-aware and Multi-entity Referring Expression ComprehensionabstractReferring Expression Comprehension (REC) aims to localize specified entities or regions from the source image according to the given natural language descriptions. While existing methods enable single-entity localization, they overlook modeling the complex inter-entity relationship in more practical multi-entity scenes, which limits their ability to produce accurate and reliable results. Moreover, the lack of high-quality multi-entity datasets incorporating fine-grained and paired image-text-relation annotations also limits addressing this challenge. To achieve this task, we first manually construct a relation-aware multi-entity REC dataset with fine-grained relation and text annotations, namely ReMeX. Additionally, we propose ReMeREC, a novel framework that effectively integrates textual and visual cues to localize multiple entities while capturing their inter-relationship. Specifically, to mitigate the semantic ambiguity arising from the absence of explicit entity boundaries in the source natural language description, we introduce a novel Text-adaptive Multi-entity Perceptron (TMP). TMP dynamically infers both the quantity and span of entities from corresponding fine-grained text cues, thus deriving representations that preserve the unique characteristics of each entity. Meanwhile, we design the Entity Inter-relationship Reasoner (EIR) to enhance semantic distinctiveness relationship modeling, leading to a more profound perception of the global scene. Furthermore, to better capture the fine-grained linguistic prompts for delineating multiple entity boundaries and inter-relationship, we leverage LLMs to generate a small-scale textual dataset, dubbed EntityText, which serves as an effective auxiliary resource and further improves the textual understanding. Extensive experiments conducted on four benchmark datasets demonstrate the superior performance of our framework. Remarkably, ReMeREC achieves outstanding results in multi-entity grounding and complex relationship prediction, outperforming other counterparts by a large margin. Yizhi Hu, Zezhao Tian, Xingqun Qi, Bingkun Yang, Junhui Yin, Muyi Sun, Man Zhang 0005, Zhenan Sun |
ACM Multimedia | 3 |
| 2025 | Multi-Scale Semantic-Guidance Networks: Robust Blind Face Restoration Against Adversarial AttacksabstractImage processing networks are known to be vulnerable to adversarial examples, where adding carefully crafted adversarial perturbations to the inputs can mislead the model. This paper addresses the problem of robust blind face restoration (BFR) against adversarial attacks. BFR refers to recovering the HQ images from the LQ images, which suffer from diverse unknown degradation, such as noise, blur, artifact removal, low resolution, etc. Although existing BFR methods exhibit good performance, they experience significant degradation when subtle distortions and perturbations are introduced into the input images. This paper is the first to investigate, improve comprehensively, and evaluate BFR methods towards adversarial attacks. Project Gradient Descent (PGD) is employed to generate adversarial examples, and multiple types of attacks were used to thoroughly assess the robustness of various BFR methods across different objectives, regions, and levels. We evaluate the robustness of multiple BFR methods and analyze the advantages of their structures and modules towards adversarial attacks. Experimental results demonstrate that the method utilizing latent feature encoding and pre-trained discrete HQ codebook achieves better robustness than other methods, with the latter outperforming the former. Similarly, multi-scale semantic guidance information also exhibits superior performance in enhancing robustness. Therefore, we propose a powerful BFR method to mitigate this issue while maintaining better performance. Extensive experiments on three real-world datasets demonstrate our method’s state-of-the-art robustness in different scenarios. Zhenyuan Zhang 0001, Xingqun Qi, Zhenbo Song, Zhiqin Yang, Jianfeng Lu 0003, Muyi Sun, Man Zhang 0005, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Biphasic Face Photo-Sketch Synthesis via Semantic-Driven Generative Adversarial Network With Graph Representation LearningabstractBiphasic face photo-sketch synthesis has significant practical value in wide-ranging fields such as digital entertainment and law enforcement. Previous approaches directly generate the photo-sketch in a global view, they always suffer from the low quality of sketches and complex photograph variations, leading to unnatural and low-fidelity results. In this article, we propose a novel semantic-driven generative adversarial network to address the above issues, cooperating with graph representation learning. Considering that human faces have distinct spatial structures, we first inject class-wise semantic layouts into the generator to provide style-based spatial information for synthesized face photographs and sketches. In addition, to enhance the authenticity of details in generated faces, we construct two types of representational graphs via semantic parsing maps upon input faces, dubbed the intraclass semantic graph (IASG) and the interclass structure graph (IRSG). Specifically, the IASG effectively models the intraclass semantic correlations of each facial semantic component, thus producing realistic facial details. To preserve the generated faces being more structure-coordinated, the IRSG models interclass structural relations among every facial component by graph representation learning. To further enhance the perceptual quality of synthesized images, we present a biphasic interactive cycle training strategy by fully taking advantage of the multilevel feature consistency between the photograph and sketch. Extensive experiments demonstrate that our method outperforms the state-of-the-art competitors on the CUHK Face Sketch (CUFS) and CUHK Face Sketch FERET (CUFSF) datasets. Xingqun Qi, Muyi Sun, Zijian Wang 0009, Jiaming Liu 0003, Qi Li 0005, Fang Zhao 0006, Shanghang Zhang, Caifeng Shan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Weakly-Supervised Emotion Transition Learning for Diverse 3D Co-Speech Gesture GenerationabstractGenerating vivid and emotional 3D co-speech gestures is crucial for virtual avatar animation in human-machine interaction applications. While the existing methods enable generating the gestures to follow a single emotion label, they overlook that long gesture sequence modeling with emotion transition is more practical in real scenes. In addition, the lack of large-scale available datasets with emotional transition speech and corresponding 3D human gestures also limits the addressing of this task. To fulfill this goal, we first incorporate the ChatGPT-4 and an audio inpainting approach to construct the high-fidelity emotion transition human speeches. Considering obtaining the realistic 3D pose annotations corresponding to the dynamically inpainted emotion transition audio is extremely difficult, we propose a novel weakly supervised training strategy to encourage authority gesture transitions. Specifically, to enhance the coordination of transition gestures w. r. t. different emotional ones, we model the temporal association representation between two different emotional gesture sequences as style guidance and infuse it into the transition generation. We further devise an emotion mixture mechanism that provides weak supervision based on a learnable mixed emotion label for transition gestures. Last, we present a keyframe sampler to supply effective initial posture cues in long sequences, enabling us to generate diverse gestures. Extensive experiments demonstrate that our method outperforms the state-of-the-art models constructed by adapting single emotion-conditioned counterparts on our newly defined emotion transition task and datasets. Our code and dataset will be released on the project page: https://xingqunqi-lab.github.io/Emo-Transition-Gesture/. Xingqun Qi, Ruibin Yuan, Xiaowei Chi, Wenhan Luo, Wei Xue 0002, Shanghang Zhang, Yike Guo |
CVPR | 1 |
| 2024 | Contrmix: Progressive Mixed Contrastive Learning for Semi-Supervised Medical Image SegmentationabstractWhile medical image segmentation has achieved impressive progress, it usually being constrained by labor-intensive and costly pixel-wise annotations. The existing semi-supervised learning methods ignore the inherent imbalance and high similarity of different categories in medical images. To address the above issues, we present a Progressive Mixed Contrastive Learning (ContrMix) framework, which contains a Cycle-mix module and a mix-based Contrastive Learning module. In Cycle-mix, a progressive mixing strategy with a cycle loss is designed to enforce the consistency between the mixed segmentation and corresponding generated mixing samples, effectively enhancing the ability to learn geometric features of the imbalanced medical data. We also introduce a mix-based Contrastive Learning module that learns the inter-instance similarities between the mixed patches and the original ones, which encourages the model to learn background-invariant representations from samples under different distortions and improves the semantic discrimination of high similarity categories. We conduct extensive experiments on the ACDC dataset and LA dataset and our method outperforms other state-of-the-art semi-supervised approaches. Meisheng Zhang, Chenye Wang, Wenxuan Zou, Xingqun Qi, Muyi Sun |
ICASSP | 4 |
| 2024 | AS-NeRF: Learning Auxiliary Sampling for Generalizable Novel View Synthesis from Sparse ViewsabstractWe tackle the problem of novel view synthesis (NVS) which aims to generate realistic images at novel views. Unlike existing works that require either costly per-scene optimization or relatively dense views, we propose AS-NeRF, a neural rendering based approach that can achieve generalizable NVS from only sparse views. Considering the inherent spatial continuity of images, we design a novel sparse-attention based auxiliary sampling module (ASM). Given a 3D point on the ray, the ASM adaptively attends to a sparse set of view-specific 2D auxiliary locations around the 3D point’s original projection pixels, and dynamically computes the attention weights in a cross-attention manner. This enables our model to effectively exploit the local correlation among neighboring pixels, obtaining the enhanced features with more powerful representation. Extensive experiments show that our method outperforms the state-of-the-art on both real and synthetic data. Jilin Tang, Lincheng Li, Xingqun Qi, Changjie Fan, Xin Yu 0002 |
ICME | 3 |
| 2024 | Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise AttentionabstractIn this paper, we introduce **Era3D**, a novel multiview diffusion method that generates high-resolution multiview images from a single-view image. Despite significant advancements in multiview generation, existing methods still suffer from camera prior mismatch, inefficacy, and low resolution, resulting in poor-quality multiview images. Specifically, these methods assume that the input images should comply with a predefined camera type, e.g. a perspective camera with a fixed focal length, leading to distorted shapes when the assumption fails. Moreover, the full-image or dense multiview attention they employ leads to a dramatic explosion of computational complexity as image resolution increases, resulting in prohibitively expensive training costs. To bridge the gap between assumption and reality, Era3D first proposes a diffusion-based camera prediction module to estimate the focal length and elevation of the input image, which allows our method to generate images without shape distortions. Furthermore, a simple but efficient attention layer, named row-wise attention, is used to enforce epipolar priors in the multiview diffusion, facilitating efficient cross-view information fusion. Consequently, compared with state-of-the-art methods, Era3D generates high-quality multiview images with up to a 512×512 resolution while reducing computation complexity of multiview attention by 12x times. Comprehensive experiments demonstrate the superior generation power of Era3D- it can reconstruct high-quality and detailed 3D meshes from diverse single-view input images, significantly outperforming baseline multiview diffusion methods. Yuan Liu 0025, Xiaoxiao Long, Feihu Zhang, Cheng Lin 0001, Xingqun Qi, Shanghang Zhang, Wei Xue 0002, Wenhan Luo, Ping Tan 0002, Wenping Wang 0001, Yike Guo |
NeurIPS | 7 |
| 2024 | Exploring Generalizable Distillation for Efficient Medical Image SegmentationabstractEfficient medical image segmentation aims to provide accurate pixel-wise predictions with a lightweight implementation framework. However, existing lightweight networks generally overlook the generalizability of the cross-domain medical segmentation tasks. In this paper, we propose Generalizable Knowledge Distillation (GKD), a novel framework for enhancing the performance of lightweight networks on cross-domain medical segmentation by generalizable knowledge distillation from powerful teacher networks. Considering the domain gaps between different medical datasets, we propose the Model-Specific Alignment Networks (MSAN) to obtain the domain-invariant representations. Meanwhile, a customized Alignment Consistency Training (ACT) strategy is designed to promote the MSAN training. Based on the domain-invariant vectors in MSAN, we propose two generalizable distillation schemes, Dual Contrastive Graph Distillation (DCGD) and Domain-Invariant Cross Distillation (DICD). In DCGD, two implicit contrastive graphs are designed to model the intra-coupling and inter-coupling semantic correlations. Then, in DICD, the domain-invariant semantic vectors are reconstructed from two networks (i.e., teacher and student) with a crossover manner to achieve simultaneous generalization of lightweight networks, hierarchically. Moreover, a metric named Fréchet Semantic Distance (FSD) is tailored to verify the effectiveness of the regularized domain-invariant features. Extensive experiments conducted on the Liver, Retinal Vessel and Colonoscopy segmentation datasets demonstrate the superiority of our method, in terms of performance and generalization ability on lightweight networks. Xingqun Qi, Zhuojie Wu, Wenxuan Zou, Yifan Gao 0003, Muyi Sun, Shanghang Zhang, Caifeng Shan, Zhenan Sun |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture GenerationabstractGenerating vivid and diverse 3D co-speech gestures is crucial for various applications in animating virtual avatars. While most existing methods can generate gestures from audio directly, they usually overlook thatemotionis one of the key factors of authentic co-speech gesture generation. In this work, we proposeEmotionGesture, a novel framework for synthesizing vivid and diverse emotional co-speech 3D gestures from audio. Considering emotion is often entangled with the rhythmic beat in speech audio, we first develop an Emotion-Beat Mining module (EBM) to extract the emotion and audio beat features as well as model their correlation via a transcript-based visual-rhythm alignment. Then, we propose an initial pose based Spatial-Temporal Prompter (STP) to generate future gestures from the given initial poses. STP effectively models the spatial-temporal correlations between the initial poses and the future gestures, thus producing the spatial-temporal coherent pose prompt. Once we obtain pose prompts, emotion, and audio beat features, we will generate 3D co-speech gestures through a transformer architecture. However, considering the poses of existing datasets often contain jittering effects, this would lead to generating unstable gestures. To address this issue, we propose an effective objective function, dubbed Motion-Smooth Loss. Specifically, we model motion offset to compensate for jittering ground-truth by forcing gestures to be smooth. Last, we present an emotion-conditioned VAE to sample emotion features, enabling us to generate diverse emotional results. Extensive experiments demonstrate that our framework outperforms the state-of-the-art, achieving vivid and diverse emotional co-speech 3D gestures. Our code and dataset will be released on the project page:https://xingqunqi-lab.github.io/Emotion-Gesture-Web/ Xingqun Qi, Chen Liu 0028, Lincheng Li, Xin Yu 0002 |
IEEE Trans. Multim. | 1 |
| 2024 | Calligraphy Font Generation via Explicitly Modeling Location-Aware Glyph Component DeformationsabstractAutomatic font generation is a challenging and time-consuming task, particularly in languages that consist of large amounts of characters with complicated structures. Typical component-wise font generation methods decompose the source character into components and search for them from the reference glyph set as candidate components. These candidate components are then utilized to learn the local styles of the target glyph. However, these methods overlook that the same component at different locations may have different profiles. When the candidate components locate differently from their corresponding components in the target glyph, the style of a generated glyph will look inconsistent. It is observed that for arbitrary components at two specific locations, the deformation patterns are similar. Driven by this, we present a location-aware component-deformable font generation method. Specifically, we search for candidate components and their corresponding deformative component pairs from the reference glyph set. Each deformative component pair can accurately depict how to deform the candidate component to the desired profile in the target glyph. Hence, we introduce a location-dependent deformation module to perform component warping. In this way, we significantly improve the component deformation ability. Lastly, we integrate deformed components into target glyphs while enforcing their styles to be consistent with the reference ones. Extensive experiments demonstrate that our method produces target-font consistent glyphs and outperforms the state-of-the-art on both seen and unseen fonts. Minda Zhao, Xingqun Qi, Zhipeng Hu, Lincheng Li, Yongqiang Zhang 0003, Zi Huang, Xin Yu 0002 |
IEEE Trans. Multim. | 2 |
| 2023 | Diverse 3D Hand Gesture Prediction from Body Dynamics by Bilateral Hand DisentanglementabstractPredicting natural and diverse 3D hand gestures from the upper body dynamics is a practical yet challenging task in virtual avatar creation. Previous works usually overlook the asymmetric motions between two hands and generate two hands in a holistic manner, leading to unnatural results. In this work, we introduce a novel bilateral hand disentanglement based two-stage 3D hand generation method to achieve natural and diverse 3D hand prediction from body dynamics. In the first stage, we intend to generate natural hand gestures by two hand-disentanglement branches. Considering the asymmetric gestures and motions of two hands, we introduce a Spatial-Residual Memory (SRM) module to model spatial interaction between the body and each hand by residual learning. To enhance the coordination of two hand motions wrt. body dynamics holistically, we then present a Temporal-Motion Memory (TMM) module. TMM can effectively model the temporal association between body dynamics and two hand motions. The second stage is built upon the insight that 3D hand predictions should be non-deterministic given the sequential body postures. Thus, we further diversify our 3D hand predictions based on the initial output from the stage one. Concretely, we propose a Prototypical-Memory Sampling Strategy (PSS) to generate the non-deterministic hand gestures by gradient-based Markov Chain Monte Carlo (MCMC) sampling. Extensive experiments demonstrate that our method outperforms the state-of-the-art models on the B2H dataset and our newly collected TED Hands dataset. The dataset and code are available at: https://github.com/XingqunQilab/Diverse-3D-Hand-Gesture-Prediction. Xingqun Qi, Chen Liu 0028, Muyi Sun, Lincheng Li, Changjie Fan, Xin Yu 0002 |
CVPR | 1 |
| 2023 | Lightvessel: Exploring Lightweight Coronary Artery Vessel Segmentation Via Similarity Knowledge DistillationabstractIn recent years, deep convolution neural networks (DCNNs) have achieved great prospects in coronary artery vessel segmentation. However, it is difficult to deploy complicated models in clinical scenarios since high-performance approaches have excessive parameters and high computation costs. To tackle this problem, we propose LightVessel, a Similarity Knowledge Distillation Framework, for lightweight coronary artery vessel segmentation. Primarily, we propose a Feature-wise Similarity Distillation (FSD) module for semantic-shift modeling. Specifically, we calculate the feature similarity between the symmetric layers from the encoder and decoder. Then the similarity is transferred as knowledge from a cumbersome teacher network to a non-trained lightweight student network. Meanwhile, for encouraging the student model to learn more pixel-wise semantic information, we introduce the Adversarial Similarity Distillation (ASD) module. Concretely, the ASD module aims to construct the spatial adversarial correlation between the annotation and prediction from the teacher and student models, respectively. Through the ASD module, the student model obtains fined-grained subtle edge segmented results of the coronary artery vessel. Extensive experiments conducted on Clinical Coronary Artery Vessel Dataset demonstrate that LightVessel outperforms various knowledge distillation counterparts. Hao Dang, Yuekai Zhang, Xingqun Qi, Muyi Sun |
ICASSP | 3 |
| 2023 | Audio-Visual Segmentation by Exploring Cross-Modal Mutual SemanticsabstractThe audio-visual segmentation (AVS) task aims to segment sounding objects from a given video. Existing works mainly focus on fusing audio and visual features of a given video to achieve sounding object masks. However, we observed that prior arts are prone to segment a certain salient object in a video regardless of the audio information. This is because sounding objects are often the most salient ones in the AVS dataset. Thus, current AVS methods might fail to localize genuine sounding objects due to the dataset bias. In this work, we present an audio-visual instance-aware segmentation approach to overcome the dataset bias. In a nutshell, our method first localizes potential sounding objects in a video by an object segmentation network, and then associates the sounding object candidates with the given audio. We notice that an object could be a sounding object in one video but a silent one in another video. This would bring ambiguity in training our object segmentation network as only sounding objects have corresponding segmentation masks. We thus propose a silent object-aware segmentation objective to alleviate the ambiguity. Moreover, since the category information of audio is unknown, especially for multiple sounding sources, we propose to explore the audio-visual semantic correlation and then associate audio with potential objects. Specifically, we attend predicted audio category scores to potential instance masks and these scores will highlight corresponding sounding instances while suppressing inaudible ones. When we enforce the attended instance masks to resemble the ground-truth mask, we are able to establish audio-visual semantics correlation. Experimental results on the AVS benchmarks demonstrate that our method can effectively segment sounding objects without being biased to salient objects and also achieves state-of-the-art performance in both the single-source and multi-source scenarios. Chen Liu 0028, Peike Li, Xingqun Qi, Hu Zhang 0005, Lincheng Li, Dadong Wang, Xin Yu 0002 |
ACM Multimedia | 3 |
| 2023 | LVSegNet: A novel deep learning-based framework for left ventricle automatic segmentation using magnetic resonance imaging
Hao Dang, Xingxiang Tao, Ge Zhang 0008, Xingqun Qi |
Comput. Commun. | 5 |
| 2023 | Graph Flow: Cross-Layer Graph Flow Distillation for Dual Efficient Medical Image SegmentationabstractWith the development of deep convolutional neural networks, medical image segmentation has achieved a series of breakthroughs in recent years. However, high-performance convolutional neural networks always mean numerous parameters and high computation costs, which will hinder the applications in resource-limited medical scenarios. Meanwhile, the scarceness of large-scale annotated medical image datasets further impedes the application of high-performance networks. To tackle these problems, we propose Graph Flow, a comprehensive knowledge distillation framework, for both network-efficiency and annotation-efficiency medical image segmentation. Specifically, the Graph Flow Distillation transfers the essence of cross-layer variations from a well-trained cumbersome teacher network to a non-trained compact student network. In addition, an unsupervised Paraphraser Module is integrated to purify the knowledge of the teacher, which is also beneficial for the training stabilization. Furthermore, we build a unified distillation framework by integrating the adversarial distillation and the vanilla logits distillation, which can further refine the final predictions of the compact network. With different teacher networks (traditional convolutional architecture or prevalent transformer architecture) and student networks, we conduct extensive experiments on four medical image datasets with different modalities (Gastric Cancer, Synapse, BUSI, and CVC-ClinicDB). We demonstrate the prominent ability of our method on these datasets, which achieves competitive performances. Moreover, we demonstrate the effectiveness of our Graph Flow through a novel semi-supervised paradigm for dual efficient medical image segmentation. Our code will be available at Graph Flow. Wenxuan Zou, Xingqun Qi, Muyi Sun, Zhenan Sun, Caifeng Shan |
IEEE Trans. Medical Imaging | 2 |
| 2022 | ShowFace: Coordinated Face Inpainting with Memory-Disentangled Refinement Networks
Zhuojie Wu, Xingqun Qi, Zijian Wang 0009, Kun Yuan 0003, Muyi Sun, Zhenan Sun |
BMVC | 2 |
| 2022 | Self-supervised Correlation Mining Network for Person Image GenerationabstractPerson image generation aims to perform non-rigid deformation on source images, which generally requires unaligned data pairs for training. Recently, self-supervised methods express great prospects in this task by merging the disentangled representations for self-reconstruction. However, such methods fail to exploit the spatial correlation between the disentangled features. In this paper, we propose a Self-supervised Correlation Mining Network (SCM-Net) to rearrange the source images in the feature space, in which two collaborative modules are integrated, Decomposed Style Encoder (DSE) and Correlation Mining Module (CMM). Specifically, the DSE first creates unaligned pairs at the feature level. Then, the CMM establishes the spatial correlation field for feature rearrangement. Eventually, a translation module transforms the rearranged features to realistic results. Meanwhile, for improving the fidelity of cross-scale pose transformation, we propose a graph based Body Structure Retaining Loss (BSR Loss) to preserve reasonable body structures on half body to full body generation. Extensive experiments conducted on DeepFashion dataset demonstrate the superiority of our method compared with other supervised and unsupervised approaches. Furthermore, satisfactory results on face generation show the versatility of our method in other deformation tasks. Zijian Wang 0009, Xingqun Qi, Kun Yuan 0003, Muyi Sun |
CVPR | 2 |
| 2022 | MOST-Net: A Memory Oriented Style Transfer Network for Face Sketch SynthesisabstractFace sketch synthesis has been widely used in multimedia entertainment and law enforcement. Despite the recent developments in deep neural networks, accurate and realistic face sketch synthesis is still a challenging task due to the diversity and complexity of human faces. Current image-to-image translation-based face sketch synthesis frequently encounters over-fitting problems when it comes to small-scale datasets. To tackle this problem, we present an end-to-end Memory Oriented Style Transfer Network (MOST-Net) for face sketch synthesis which can produce high-fidelity sketches with limited data. Specifically, an external self-supervised dynamic memory module is introduced to capture the domain alignment knowledge in the long term. In this way, our proposed model could obtain the domain-transfer ability by establishing the durable relationship between faces and corresponding sketches on the feature level. Furthermore, we design a novel Memory Refinement Loss (MR Loss) for feature alignment in the memory module, which enhances the accuracy of memory slots in an unsupervised manner. Extensive experiments on the CUFS and the CUFSF datasets show that our MOST-Net achieves state-of-the-art performance, especially in terms of the Structural Similarity Index(SSIM). Fan Ji, Muyi Sun, Xingqun Qi, Qi Li 0005, Zhenan Sun |
ICPR | 3 |
| 2021 | CoCo DistillNet: a Cross-layer Correlation Distillation Network for Pathological Gastric Cancer SegmentationabstractIn recent years, deep convolutional neural networks have made significant advances in pathology image segmentation. However, pathology image segmentation encounters with a dilemma in which the higher-performance networks generally require more computational resources and storage. This phenomenon limits the employment of high-accuracy networks in real scenes due to the inherent high-resolution of pathological images. To tackle this problem, we propose CoCo DistillNet, a novel Cross-layer Correlation (CoCo) knowledge distillation network for pathological gastric cancer segmentation. Knowledge distillation, a general technique which aims at improving the performance of a compact network through knowledge transfer from a cumbersome network. Concretely, our CoCo DistillNet models the correlations of channel-mixed spatial similarity between different layers and then transfers this knowledge from a pre-trained cumbersome teacher network to a non-trained compact student network. In addition, we also utilize the adversarial learning strategy to further prompt the distilling procedure which is called Adversarial Distillation (AD). Furthermore, to stabilize our training procedure, we make the use of the unsupervised Paraphraser Module (PM) to boost the knowledge paraphrase in the teacher network. As a result, extensive experiments conducted on the Gastric Cancer Segmentation Dataset demonstrate the prominent ability of CoCo DistillNet which achieves state-of-the-art performance. Wenxuan Zou, Xingqun Qi, Zhuojie Wu, Zijian Wang 0009, Muyi Sun, Caifeng Shan |
BIBM | 2 |
| 2021 | Face Sketch Synthesis via Semantic-Driven Generative Adversarial NetworkabstractFace sketch synthesis has made significant progress with the development of deep neural networks in these years. The delicate depiction of sketch portraits facilitates a wide range of applications like digital entertainment and law enforcement. However, accurate and realistic face sketch generation is still a challenging task due to the illumination variations and complex backgrounds in the real scenes. To tackle these challenges, we propose a novel Semantic-Driven Generative Adversarial Network (SDGAN) which embeds global structure-level style injection and local class-level knowledge re-weighting. Specifically, we conduct facial saliency detection on the input face photos to provide overall facial texture structure, which could be used as a global type of prior information. In addition, we exploit face parsing layouts as the semantic-level spatial prior to enforce globally structural style injection in the generator of SDGAN. Furthermore, to enhance the realistic effect of the details, we propose a novel Adaptive Re-weighting Loss (ARLoss) which dedicates to balance the contributions of different semantic classes. Experimentally, our extensive experiments on CUFS and CUFSF datasets show that our proposed algorithm achieves state-of-the-art performance. Xingqun Qi, Muyi Sun, Weining Wang 0001, Xiaoxiao Dong, Qi Li 0005, Caifeng Shan |
IJCB | 1 |
| 2021 | Contextual information enhanced convolutional neural networks for retinal vessel segmentation in color fundus images
Muyi Sun, Kaiqi Li, Xingqun Qi, Hao Dang, Guanhong Zhang |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Accurate Retinal Vessel Segmentation in Color Fundus Images via Fully Attention-Based NetworksabstractAutomatic retinal vessel segmentation is important for the diagnosis and prevention of ophthalmic diseases. The existing deep learning retinal vessel segmentation models always treat each pixel equally. However, the multi-scale vessel structure is a vital factor affecting the segmentation results, especially in thin vessels. To address this crucial gap, we propose a novel Fully Attention-based Network (FANet) based on attention mechanisms to adaptively learn rich feature representation and aggregate the multi-scale information. Specifically, the framework consists of the image pre-processing procedure and the semantic segmentation networks. Green channel extraction (GE) and contrast limited adaptive histogram equalization (CLAHE) are employed as pre-processing to enhance the texture and contrast of retinal blood images. Besides, the network combines two types of attention modules with the U-Net. We propose a lightweight dual-direction attention block to model global dependencies and reduce intra-class inconsistencies, in which the weights of feature maps are updated based on the semantic correlation between pixels. The dual-direction attention block utilizes horizontal and vertical pooling operations to produce the attention map. In this way, the network aggregates global contextual information from semantic-closer regions or a series of pixels belonging to the same object category. Meanwhile, we adopt the selective kernel (SK) unit to replace the standard convolution for obtaining multi-scale features of different receptive field sizes generated by soft attention. Furthermore, we demonstrate that the proposed model can effectively identify irregular, noisy, and multi-scale retinal vessels. The abundant experiments on DRIVE, STARE, and CHASE_DB1 datasets show that our method achieves state-of-the-art performance. Kaiqi Li, Xingqun Qi, Yiwen Luo, Zeyi Yao, Xiaoguang Zhou, Muyi Sun |
IEEE J. Biomed. Health Informatics | 2 |