VLDB 2026 Research / reviewers in the wild / expert
Xi Wu 0004
dblp:37/4465-4
· DBLP profile ↗
116ranked-venue papers
1as first author
90since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 42 since 2021Artificial intelligence and machine learning · 53 · 40 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 1 first-author · 21 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PromptEmo: Learning Emotion with Bilateral Textual Prompts in Multi-Domain Open-set ScenariosabstractFacial Expression Recognition (FER) is crucial to human-computer interaction. Existing cross-domain FER (CD-FER) methods mainly focus on single-source closed-set scenarios, transferring knowledge from a single source domain to a target domain with identical class sets. However, CD-FER faces two real-world challenges: 1) the need to leverage information from multiple sources, leading to multi-domain shift, and 2) the necessity to recognize unseen target classes, resulting in class shift. These issues give rise to a novel and challenging task, which we define as Multi-domain Open-set FER (MO-FER). In this paper, we propose PromptEmo, a novel CLIP-based framework that leverages bilateral textual prompts to address both shifts in the MO-FER task. Leveraging the generalizability of LLM, PromptEmo constructs trainable positive prompts with LLM-generated emotion descriptions for seen classes, as well as template-derived negative prompts to enhance the reasoning for unseen classes. Then, we introduce a modal-task optimization paradigm organized from two perspectives: textual semantics and visual domains, yielding Intra-modal Space-specific Optimization (ISO) and Cross-modal Emotion-aware Interaction (CEI) strategies. ISO refines the CLIP-based textual space to ensure semantic separation between bilateral prompts and improves the latent visual space by promoting inter-domain alignment. Founded on ISO, CEI facilitates effective vision-language interactions, resulting in four joint loss terms that improve emotion recognition by shaping a domain-invariant, discriminative feature space. PromptEmo surpasses the current SOTA method by 7.7% AUC on unseen classes across four FER datasets, serving as a strong baseline for the MO-FER task. Xinyi Zeng, Yuxiang Yang 0009, Pinxian Zeng, Wenxia Yin, Bo Liu 0113, Xi Wu 0004, Yan Wang 0015 |
AAAI | 6 |
| 2026 | SelCo: Efficient Distributed Multimodal LLM Training with Selective Co-location
Hao Yao, Xi Wu 0004, Jing Peng 0003, Jiqing Gu |
ICIC (16) | 2 |
| 2026 | Middle modality interactive feature attention learning for visible-infrared person re-identification
Haoyi Zhao, Shanmin Yang, Xiaojie Li 0001, Jing Peng 0003, Xi Wu 0004 |
Neurocomputing | 5 |
| 2026 | AMOS: Absent minority oversampling neural network for imbalanced data classification
Zhan ao Huang, Canghong Shi, Jia He 0003, Xiaojie Li 0001, Xi Wu 0004 |
Inf. Sci. | 5 |
| 2026 | RL-I2IT: Image-to-image translation with deep reinforcement learning
Jing Hu 0009, Ziwei Luo 0002, Chengming Feng, Shu Hu 0001, Bin B. Zhu, Xi Wu 0004, Xin Li 0005, Hongtu Zhu, Siwei Lyu, Xin Wang 0045 |
Neural Networks | 6 |
| 2026 | HSSN: Hierarchical Superpixel Segmentation Network guided by visual attention mechanism
Tingyu Zhao, Bo Peng 0006, Zhenguang Zhang, Daipeng Yang, Xi Wu 0004 |
Signal Process. | 5 |
| 2026 | CD-Former: A Cross-Modal Dual-Interaction Transformer With Whole-Slide Image Pyramids and Genomics for Survival PredictionabstractSurvival prediction is crucial for cancer patients as it provides essential early prognostic information for treatment planning and decision making. Despite impressive performance , current multi-modal survival prediction methods that integrate pathology and genomic data face two main challenges: (1) Whole-slide images (WSIs) generally exhibit hierarchical structures, but the interactions of phenotypes at different resolutions remain unexplored. More importantly, the potential semantic discrepancy arising from diverse resolutions is often ignored. (2) The absence of effective interactions between the inherent hierarchical structures of WSIs and genomic data. To address these challenges, in this paper, we propose Cross-modal Dual-interaction Transformer (CD-Former), a robust hierarchical framework for multi-modal survival prediction. Our CD-Former involves two key components: (1) an Multimodal Cross-Scale Calibration (MCSC) module for effectively capturing correlations across multiple resolutions and calibrating fine-grained features, thereby bridging the semantic discrepancy caused by different WSI resolutions; and (2) a hierarchical interaction module termed Multi-modal Dual-interaction (M2Di) for fully exploring multi-resolution cross-modal correlations and interactions, which comprises a Patch-level Cross-Attention Block (PCAB) and a Region-level Cross-Attention Block (RCAB) to investigate cross-modal associations between patch- or region-level features of WSI and genomic data. Additionally, we employ a scale-oriented WSI enhancer to capture the interactions among various components of WSIs. The experimental results demonstrate the effectiveness of our proposed framework, which achieves state-of-the-art performance compared to previous studies. Lifan Long, Xingchen Peng, Bo Liu 0113, Xi Wu 0004, Daoqiang Zhang, Yan Wang 0015 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | CDGR: Cross-Modal Dual Graph Reasoning for Weakly Supervised Semantic SegmentationabstractCurrent Convolutional Neural Networks (CNNs) for Weakly Supervised Semantic Segmentation (WSSS) often have difficulties in discovering distinctive feature locations for each category. Therefore, the pseudo-labels generated from the expanded seed regions are typically incomplete and contain a significant amount of noise. Without additional annotations, the numerous erroneous information will potentially propagate in the segmentation network’s training stage. In this work, we propose a Cross-Modal Dual Graph Reasoning (CDGR) framework to leverage both visual and language knowledge effectively. This framework can capture dependencies between the spatial and the semantic spaces, facilitating the discovery of discriminative feature locations. Specifically, we perform cross-modal graph reasoning between the visual and the language modal graphs to enhance global contextual relationships between pixels in the visual feature map. Additionally, we introduce a graph interaction attention network to thoroughly explore implicit relationships between visual and language graphs. We apply the CDGR network to generate more complete pseudo-labels for the classification network and utilize it in the segmentation network to unleash its self-correcting capabilities. Extensive experiments on the PASCAL VOC 2012 and MS COCO 2014 datasets demonstrate the effectiveness of CDGR compared to other state-of-the-art peers. Our code is provided at https://github.com/JIA-ZHANG666/CDGR. Jia Zhang 0027, Bo Peng 0006, Xi Wu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | EDVD: Cross-Modal Spatio-Temporal Fusion With Event and Diffusion for Video DeblurringabstractRestoring high-quality images from blurred videos is a highly challenging task, especially in severely blurred scenes. In recent years, event-based methods have achieved significant progress in video deblurring. However, the modal differences between the event and image increase the difficulty of feature fusion. Additionally, the sparsity of event makes it difficult to restore some local details. To address these issues, we propose a new video deblurring method. Firstly, we design a cross-modal collaborative attention mechanism to effectively fuse features from blurred frames and event frames, thereby deeply extracting motion information from event frames. Secondly, we utilize a diffusion model to generate spatial guiding prior feature, enhancing local details and textures. Furthermore, we propose an event-guided dynamic feature fusion module that adaptively integrates spatio-temporal information from neighboring frames. Experimental results on both synthetic and real datasets demonstrate that our method outperforms the current state-of-the-art approaches. The code is available at: https://github.com/Frank-Zhou-01/EDVD-main. Ying Fu 0003, Tao Wu 0010, Qing Li 0001, Xi Wu 0004, Wei Liu 0044 |
IEEE Trans. Image Process. | 6 |
| 2026 | MGTP: Multi-Granularity Textual Prompts for Low-Dose Brain PET Image Denoising via Adversarial Diffusion ModelabstractPositron emission tomography (PET) is an advanced nuclear imaging technique and has been widely applied in clinic. However, radiation risks associated with standard-dose PET imaging raise health concerns, whereas the quality of low-dose PET images fails to meet clinical requirements. To reduce the tracer dose while maintaining image quality, it is of great interest to estimate high-quality PET images from low-dose images. However, existing low-dose PET image denoising methods primarily focus on image data, overlooking crucial information in non-image textual data such as patients' clinical tabular and textual descriptions of general image quality. This neglect can lead to subpar denoising quality with inaccurate contexts and poor details. To address these problems, in this paper, we propose Multi-Granularity Textual Prompts, namely MGTP, to denoise low-dose PET images via an adversarial diffusion model. Different from prior methods that rely solely on image conditioning, our MGTP innovatively introduces textual prompts spanning diverse granularities to capture both high-level semantic-related contexts and low-level degradation-related details. To harmonize multi-granularity textual prompts with low-dose PET images, we design a Cross-Modality Selective Conditioning (CMSC) module, which prioritizes semantic- and detail-relevant information while eliminating irrelevant components. The resulting features are fed into diffusion model as conditions, enforcing a more controlled diffusion process. In addition, we develop a Masked Prompt Reconstruction Network (MPR-Net) to enhance the preservation of semantics and details in denoised images, mitigating distortions brought by the random noise in the diffusion process. Experiments on clinical PET data show that our method achieves the state-of-the-art performance. Xinyi Zeng, Pinxian Zeng, Bo Liu 0113, Xi Wu 0004, Deng Xiong, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | Data-Driven Robust Optimization Neural Network Method for Imbalanced Data Classification
Zhan ao Huang, Xiaojie Li 0001, Xi Wu 0004 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Content-Aware Dynamic Superpixel SegmentationabstractIn recent years, deep learning-based superpixel segmentation methods derived from SLIC have made significant progress by utilizing uniform grid-based seed initialization. However, due to the unequal pixel space variation rates in natural images, methods based on uniform grid initialization struggle to balance the compactness of superpixels in flat regions with the boundary adherence in non-flat regions. Inspired by the visual attention model based on saliency in the human visual system, we propose a content-aware dynamic superpixel segmentation network. Specifically, we propose a seed initialization strategy guided by geodesic distance transformation and design two segmentation heads for different scales, which are used for joint network training to encourage the network to focus more on areas with texture variations without causing unnecessary segmentation in flat regions. Extensive experiments on BSDS500 and NYUv2 datasets demonstrate that our method achieves state-of-the-art performance. Tingyu Zhao, Bo Peng 0006, Zhenguang Zhang, Daipeng Yang, Xi Wu 0004 |
ICASSP | 5 |
| 2025 | RLMiniStyler: Light-weight RL Style Agent for Arbitrary Sequential Neural Style GenerationabstractArbitrary style transfer aims to apply the style of any given artistic image to another content image. Still, existing deep learning-based methods often require significant computational costs to generate diverse stylized results. Motivated by this, we propose a novel reinforcement learning-based framework for arbitrary style transfer RLMiniStyler. This framework leverages a unified reinforcement learning policy to iteratively guide the style transfer process by exploring and exploiting stylization feedback, generating smooth sequences of stylized results while achieving model lightweight. Furthermore, we introduce an uncertainty-aware multi-task learning strategy that automatically adjusts loss weights to adapt to the content and style balance requirements at different training stages, thereby accelerating model convergence. Through a series of experiments across image various resolutions, we have validated the advantages of RLMiniStyler over other state-of-the-art methods in generating high-quality, diverse artistic image sequences at a lower cost. Codes are available at https://github.com/fengxiaoming520/RLMiniStyler. Jing Hu 0009, Chengming Feng, Shu Hu 0001, Ming-Ching Chang, Xin Li 0005, Xi Wu 0004, Xin Wang 0045 |
IJCAI | 6 |
| 2025 | Multi-Source Feature Fusion and Spatio-Temporal Unet for Precipitation NowcastingabstractPrecipitation nowcasting is an extremely critical task in the field of weather forecasting, as it facilitates advancements in meteorological observation. Nevertheless, accurate short-term precipitation forecasting remains a significant challenge at present. Traditional methods have relied on physical equations for predictions, which are often computationally consuming. Current deep learning approaches, using CNNs and RNNs, roughly extract the latent features of spatiotemporal data, but the feature extraction process usually overlooks the dynamic changes occurring between prediction image frames. Furthermore, most methods utilize a single precipitation variable as input for predicting future precipitation, neglecting that precipitation events are triggered by multiple meteorological factors. To tackle this issue, we propose a novel neural network model, Multi-Source Feature Fusion and Spatio-Temporal Unet (MFFST-Unet), which utilizes multi-source feature information to guide precipitation forecasting. Additionally, we introduce the Inter-Frame Difference Regularization(IFDR) Loss, which is combined with MSE Loss to optimize the frame stability of model predictions through adaptive weighting. We conducted training and testing on the SEVIR dataset, achieving high-resolution precipitation nowcasting for a one-hour forecast. Experimental results indicate that our MFFST-Unet model surpasses other deep learning methods, achieving a maximum improvement of 34.72% in the CSI precipitation metric, demonstrating its significant practical implications for weather forecasting applications. Dufu Liu, Xia Yuan, Xi Wu 0004, Jing Hu 0009 |
IJCNN | 5 |
| 2025 | LLM-MedQA: Enhancing Medical Question Answering through Case Studies in Large Language ModelsabstractAccurate and efficient question-answering systems are essential for high-quality patient care in the medical field. While Large Language Models (LLMs) have made remarkable strides across various domains, they still face challenges in medical question answering, particularly in understanding domain-specific terminology and performing complex reasoning, limiting their effectiveness in critical applications. To address this, we propose a multi-agent medical question-answering (MedQA) system incorporating similar case generation. We leverage the Llama3.1:70B model in a multi-agent architecture to enhance enhance zero-shot classification on the MedQA dataset, utilizing the model’s inherent medical knowledge and reasoning capabilities without additional training data. Experimental results show substantial gains over existing benchmark models, with improvements of 7% in both accuracy and F1-score across various medical QA tasks. Furthermore, we examine the model’s interpretability and reliability in addressing complex medical queries. This research not only offers a robust solution for medical question answering but also establishes a foundation for broader applications of LLMs in the medical domain. Yineng Chen, Chingsheng Lin, Shu Hu 0001, Jinrong Hu, Xi Wu 0004, Xin Wang 0045 |
IJCNN | 8 |
| 2025 | Wave Height Prediction: 3D Spatiotemporal FourCastNet Method with Multi-FactorabstractAccurate wave height prediction is essential for various marine operations. However, the complexity of the marine environment, influenced by numerous factors, underscores the importance of effectively leveraging available data. This paper introduces a novel spatiotemporal model, FourCastNet, which serves as a baseline for capturing wave height trends through spectral analysis. To refine spatiotemporal representations, we incorporate 3D convolution to extract local features from the data via nonlinear transformations. Employing Fourier transform, we convert the features into the frequency domain, filtering out high-frequency noise while enhancing the visibility of spatiotemporal patterns. Finally, we fuse various features and capture complex patterns via channel mixing, thereby enhancing the model’s predictive accuracy. Iterative forecasting techniques are applied to minimize prediction errors. Experiments using French wave reanalysis data, projected up to 24 hours at three-hour intervals, demonstrate that our method achieves approximately a 10% improvement in accuracy over existing state-of-the-art techniques, particularly for short-term forecasts within the first 6 hours. Chenchen He, Zhanao Huang, Canghong Shi, Xiaojie Li 0001, Xi Wu 0004 |
IJCNN | 8 |
| 2025 | Leveraging Visual Prompt with Diffusion Adversarial Network for Radiotherapy Dose Prediction
Zhenghao Feng, Lu Wen, Xi Wu 0004, Jianghong Xiao, Xingchen Peng, Dinggang Shen, Yan Wang 0015 |
MICCAI (15) | 4 |
| 2025 | PREMISE: Individual Preference-aware Multi-modal Cooperation for Survival PredictionabstractMulti-modal learning that combines whole-slide images (WSIs) and genomic data has recently emerged as a promising paradigm for improving cancer survival prediction. However, existing methods either utilize genomic data as guidance to integrate WSI features or treat both modalities as equally important across all patients, overlooking individual variations in modality importance. As critical survival-related features can reside in different modalities for different patients, prioritizing the modality with more discriminative information for each patient, referred to as individual modality preference, is crucial for enhancing prediction accuracy. In this paper, we propose a novel Individual PREference-aware Multi-modal CooperatIon framework for Survival PrEdiction (PREMISE), which collaborates with a uni-modal and a cross-modal preference learner to fully exploit individual modality preference. Specifically, the uni-modal preference learner adopts a task-aware preference estimator to dynamically assess the importance of each modality for each patient, thereby identifying the preferred modality for input individual. To promote cross-modal learning, the cross-modal preference learner embeds the obtained preferences as biases to construct a preference-aware mutual-attention module, enabling the individually adaptive focus and interactions between modalities. Meanwhile, inspired by clinical practice where doctors reference prior cases for survival evaluation, we introduce dual-level cross-modal alignment, incorporating both patient-level and group-level preferences. This alignment emphasizes the more discriminative modality and improves risk group separation during cross-modal knowledge transfer. Experiments have validated our superiority. Yilun Li, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
ACM Multimedia | 3 |
| 2025 | Rethinking Back Transformation in 2-stage Eigenvalue Decomposition on Heterogeneous ArchitecturesabstractThe 2-stage eigenvalue decomposition (EVD) method outperforms conventional 1-stage method on GPUs and heterogeneous architectures, especially when eigenvectors are not required. However, its performance advantage diminishes when performing back transformation to obtain eigenvectors. To address this, we propose two key solutions: 1) replacing BLAS3 operations with BLAS2 operations during the bulge-chasing back transformation for better performance, and 2) reordering the back transformation workflow from a backward pattern to a new parallelism-driven pattern to hide divide-and-conquer latency, at the cost of one additional GEMM computation. Experimentally, the proposed back transformation algorithm demonstrates significant performance improvements, outperforming the SOTA implementation in MAGMA by an average factor of 3.58x. For complete FP64 precision symmetric EVD with eigenvectors, the proposed algorithm, incorporating both solutions, surpasses the SOTA implementations in MAGMA and cuSOLVER by average factors of 2.62x and 2.21x, respectively. Dajun Huang, Gaoyuan Zou, Lu Shi 0006, Xu Jiang 0004, Xi Wu 0004, Hancong Duan, Shaoshuai Zhang |
SC | 6 |
| 2025 | Coupling importance sampling neural network for imbalanced data classification with multi-level learning bias
Zhan ao Huang, Xiaojie Li 0001, Xi Wu 0004 |
Neurocomputing | 5 |
| 2025 | A Method for Resolving Blockchain State Conflicts in IoT - Using Customized Smart Contract VariablesabstractSmart contracts are essential tools for enabling interaction between blockchain and Internet of Things (IoT) systems. For example, in cold chain logistics, the blockchain can obtain the states of the logistics system through smart contracts. However, direct interactions between smart contracts and these systems introduce uncertainties, potentially leading to network forks or state inconsistencies, which can compromise the security and reliability of the blockchain. To address these challenges, a novel smart contract variable, ExState, is proposed, specifically designed to track and store the dynamic states of IoT systems. Additionally, a corresponding operational logic is defined to organize these states into sequential records, ensuring that the state sequences obtained by each node remain consistent, effectively mitigating state conflicts. In addition, a formal model is developed, accompanied by a theoretical analysis of its determinacy. Experimental results demonstrate that, in cross-chain scenarios, this method achieves a performance improvement of up to 50.41% compared to the traditional Oracle method. Qing Fang, Hong Su, Xi Wu 0004 |
IEEE Internet Things J. | 3 |
| 2025 | A bio-inspired approach to line segment detection utilizing orientation-selective neurons
Daipeng Yang, Bo Peng 0006, Xi Wu 0004 |
Signal Process. | 3 |
| 2025 | Dual-Domain Classification-Aided High-Quality PET Synthesis With Shared Information MaximizationabstractPositron emission tomography (PET) is widely applied in clinic for providing crucial diagnosis information. However, its inherent radiation exposure inevitably brings potential health risk for patient. To reduce radiation risk while also obtaining high-quality PET image, we plan to synthesize standard-dose PET (SPET) from low-dose PET (LPET). Since PET images can be represented in both projection domain and image domain (dual domains) emphasizing different information, considering dual domains in PET synthesis could contribute to better performance. In this way, we propose a novel dual-domain model for high-quality PET synthesis, named DCBi-GAN, by introducing a denoising network for the projection domain and an enhancing network for the image domain to effectively exploit dual-domain information. Concretely, the denoising network takes the LPET sinogram converted from LPET image to suppress noise and artifacts in the projection domain. Then, the enhancing network in the image domain takes the denoised LPET image (transferred back from the denoised sinogram) to enhance image quality. Notably, as LPET and SPET images come from the same subject, the abundant shared information between LPET and SPET can be used for boosting synthesis performance. Specially, we design a bi-directional contrastive generative adversarial network (GAN) to encourage maximal preservation of the shared information. Besides, we introduce a mild cognitive impairment (MCI) classification task to enhance clinical applicability of the synthesized PET. Evaluation on both Real Human Brain dataset and Phantom Brain dataset demonstrates effectiveness and superiority of our proposed model.Note to Practitioners—Positron emission tomography (PET) is a primary nuclear imaging technique for tumor detection and brain disorder diagnosis in the early stage of diseases, while the inherent radiation exposure inevitably raises concerns about potential health risk. This article proposes a novel PET image synthesis model to obtain clinically accepted PET image at low dose, namely DCBi-GAN, by taking account of the complementary multi-domain information and the modality shared content information, with a mild cognitive impairment (MCI) classification task to further boost clinical applicability of synthesized PET images. We experimentally validate the effectiveness of proposed DCBi-GAN on two datasets. Our proposed method could facilitate diagnosis and treatment of disease, to be used in the existing computer-aided medical systems. Yuchen Fei, Chen Zu, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Multi-Modal Long-Short Distance Attention-Based Transformer-GAN for PET Reconstruction With Auxiliary MRIabstractTo obtain high-quality PET scans while minimizing potential radiation hazards for patients, various GAN-based methods have been developed to reconstruct high-quality standard-count PET (SPET) images from low-count PET (LPET) ones. While recent efforts try to integrate MRI or CT to enhance reconstruction in a multi-modal way, current architectures mainly face two limitations: 1) CNN backbones or simple Transformer bottleneck layers are insufficient for robust semantic understanding; and 2) the identical strategies for multi-modal feature extraction and fusion overlook each modality’s respective importance for the reconstruction task. In this work, we propose the Multi-modal Long-Short Distance Attention-based Transformer-GAN (MLSDA-GAN), a novel network combining 3D transformer and CNN architecture for PET image reconstruction. Specifically, to extract fine-grained features with a small number of parameters, our MLSDA-GAN integrates multi-scale convolution into the embedding part of the transformer. As for our multi-modal design, given the strong correlation between LPET and SPET in structural characteristics, we treat MRI as an auxiliary modality to LPET and achieve effective multi-modal extraction and fusion strategies. These strategies include 1) a PET-specific Self-attention Extraction (PSE) block for comprehensive feature extraction of the primary LPET and 2) a Multi-modality Cross-attention Fusion (MCF) block for effective multi-modal interaction and fusion, enabling us to more efficiently model both long- and short-range relationships in the corresponding feature extraction and fusion processes. Experiments demonstrate superiority of our method quantitatively and qualitatively. Code is available athttps://github.com/Aru321/MLSDA-GAN. Pinxian Zeng, Xinyi Zeng, Yan Wang 0015, Luping Zhou, Chen Zu, Xi Wu 0004, Jiliu Zhou, Dinggang Shen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Dual Graph Inference Network for Weakly Supervised Semantic SegmentationabstractEstablishing global contextual relationships between objects is crucial for weakly supervised semantic segmentation (WSSS) tasks that lack pixel-level labels. Limited by the efficiency of convolutional operations in capturing long-range dependencies with a limited receptive field and to bridge the gap between image-level annotations and pixel-level labels, we propose a Dual Graph Reasoning Mapping (DGRM) module. When integrated into a convolutional network, it conducts contextual graph reasoning on both spatial and interaction spaces of visual features. The first component of this graph reasoning module involves incorporating commonsense knowledge extracted from an external knowledge base into visual features to promote global contextual reasoning for visual graphs. The second component focuses on reasoning in the projected interaction space, utilizing abstracted object class attributes from high-level visual features to establish dependencies among channels in a potential low-dimensional space. Moreover, to capture correspondences at different semantic levels, we model the feature maps in a pyramid-like structure for graph reasoning at various levels. Extensive experiments on popular datasets, such as PASCAL VOC 2012 and MS COCO 2014, demonstrate the superiority of our approach. Our code is provided athttps://github.com/JIA-ZHANG666/DGRM. Jia Zhang 0027, Bo Peng 0006, Xi Wu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Spectral Spatial Window Attention Transformer for Hyperspectral Image Classification
Xi Wu 0004, Tahir Arshad, Bo Peng 0006 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | An Explanation Method Based on Interpretable Linear Model With Four Key CharacteristicsabstractFor the interpretability of deep neural networks (DNNs) in visual-related tasks, existing explanation methods commonly generate a saliency map based on the linear relation between output results and input features. However, when the explanation conflicts with a human visual examination, these methods do not provide further evidence to analyze the saliency explanation. Most may fail to provide feature attribution with identifiable semantics or produce misleading explanations due to their insufficient robustness. In this paper, we first propose four key characteristics (richness, adaptivity, exclusiveness, and fairness) to evaluate the existing linear relation-based explanation method, and then construct an interpretable linear model to satisfy them. We formalize the characteristics and develop a novel explanation method based on this. We extract and reconstruct key exclusive semantic features from the feature map using the Nonnegative Matrix Factorization (NMF) algorithm, utilize the information entropy model to determine the number of features adaptively and their richness, and then linearly combine each feature with fairly assigned weights using an approximate Shapley algorithm to generate the saliency map. Compared with the state-of-the-art methods, our explanations of different datasets and DNNs are more convincing and robust in terms of Average drop (AD), Average increase (AI), Deletions (Del), and Insertions (Ins). Our supplementary experiments provide sufficient evidence that the four characteristics guarantee the feasibility of feature attribution analysis and enhance the quality of the resulting explanations. Yuecan Yuan, Zhan ao Huang, Ying Fu 0003, Xuemin Zhao, Canghong Shi, Xiaojie Li 0001, Xi Wu 0004 |
IEEE Trans. Image Process. | 8 |
| 2025 | VB-KGN: Variational Bayesian Kernel Generation Networks for Motion Image DeblurringabstractMotion blur estimation is a critical and fundamental task in scene analysis and image restoration. While most state-of-the-art deep learning-based methods for single-image motion image deblurring focus on constructing deep networks or developing training strategies, the characterization of motion blur has received less attention. In this paper, we innovatively propose a non-parametric Variational Bayesian Kernel Generation Network (VB-KGN) for characterizing motion blur in a single image. To solve this model, we employ the variational inference framework to approximate the expected statistical distribution of motion blur images in a data-driven manner. The qualitative and quantitative evaluations of our experimental results demonstrate that our proposed model can generate highly accurate motion blur kernels, significantly improving motion image deblurring performance and substantially reducing the need for extensive training sample preprocessing for deblurring tasks. Ying Fu 0003, Xiaojie Li 0001, Xin Wang 0045, Xi Wu 0004, Shu Hu 0001, Siwei Lyu, Wei Liu 0044 |
IEEE Trans. Multim. | 5 |
| 2025 | A bio-inspired edge and segment detection method by modeling multiple visual regions
Daipeng Yang, Bo Peng 0006, Xi Wu 0004 |
Vis. Comput. | 3 |
| 2024 | Image2Points: A 3D Point-Based Context Clusters GAN for High-Quality Pet Image ReconstructionabstractTo obtain high-quality Positron emission tomography (PET) images while minimizing radiation exposure, numerous methods have been proposed to reconstruct standard-dose PET (SPET) images from the corresponding low-dose PET (LPET) images. However, these methods heavily rely on voxel-based representations, which fall short of adequately accounting for the precise structure and fine-grained context, leading to compromised reconstruction. In this paper, we propose a 3D point-based context clusters GAN, namely PCC-GAN, to reconstruct high-quality SPET images from LPET. Specifically, inspired by the geometric representation power of points, we resort to a point-based representation to enhance the explicit expression of the image structure, thus facilitating the reconstruction with finer details. Moreover, a context clustering strategy is applied to explore the contextual relationships among points, which mitigates the ambiguities of small structures in the reconstructed images. Experiments on both clinical and phantom datasets demonstrate that our PCC-GAN outperforms the state-of-the-art reconstruction methods qualitatively and quantitatively. Code is available at https://github.com/gluucose/PCCGAN. Yan Wang 0015, Lu Wen, Pinxian Zeng, Xi Wu 0004, Jiliu Zhou, Dinggang Shen |
ICASSP | 5 |
| 2024 | DCL-Net: Dual Contrastive Learning Network for Semi-Supervised Multi-Organ SegmentationabstractSemi-supervised learning (SSL) is a sound measure to relieve the strict demand of abundant annotated datasets, especially for challenging multi-organ segmentation (MoS). However, most existing SSL methods predict pixels in a single image independently, ignoring the relations among images and categories. In this paper, we propose a two-stage Dual Contrastive Learning Network (DCL-Net) for semi-supervised MoS, which utilizes global and local contrastive learning to strengthen the relations among images and classes. Concretely, in Stage I, we develop a similarity-guided global contrastive learning to explore the implicit continuity and similarity among images and learn global context. Then, in Stage II, we present an organ-aware local contrastive learning to further attract the class representations. To ease the computation burden, we introduce a mask center computation algorithm to compress the category representations for local contrastive learning. Experiments conducted on the public 2017 ACDC dataset and an in-house RC-OARs dataset has demonstrated the superior performance of our method. Lu Wen, Zhenghao Feng, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
ICASSP | 5 |
| 2024 | High fidelity medical image super-resolution based on Medical Multi-Feature Compensation Attention GANabstractMRI and CT medical images play a crucial role in clinical medicine, and high-resolution images enhance diagnostic quality. Improving image resolution through hardware upgrades is often expensive and may increase radiation exposure for patients. Common super-resolution methods used for natural images do not account for the unique relationships between colors and pixels in medical images. In this paper, we present a novel approach to medical image super-resolution using the Medical Multi-Feature Compensation Attention GAN (MMSRGAN). The proposed method integrates a feature compensation attention module, enhancing the reconstruction of high-resolution images by addressing the unique characteristics and correlations in medical data. Extensive experiments on IXI-T1 datasets demonstrate that MMSRGAN significantly outperforms existing deep learning-based methods, achieving state-of-the-art results in terms of PSNR and SSIM. This advancement highlights the potential for improved diagnostic accuracy and treatment planning in clinical settings through enhanced medical imaging. Qinrui Fan, Xi Wu 0004, Jing Hu 0009 |
IJCB | 3 |
| 2024 | Masked Conditional Diffusion Model for Enhancing Deepfake DetectionabstractRecent studies on deepfake detection have achieved promising results when training and testing faces are from the same dataset. However, their results severely degrade when confronted with forged samples that the model has not yet seen during training. In this paper, deepfake data to help detect deepfakes. this paper present we put a new insight into diffusion model-based data augmentation, and propose a Masked Conditional Diffusion Model (MCDM) for enhancing deepfake detection. It generates a variety of forged faces from a masked pristine one, encouraging the deepfake detection model to learn generic and robust representations without overfitting to special artifacts. Extensive experiments demonstrate that forgery images generated with our method are of high quality and helpful to improve the performance of deepfake detection models. Tiewen Chen, Shanmin Yang, Shu Hu 0001, Zhenghan Fang, Ying Fu 0003, Xi Wu 0004, Xin Wang 0045 |
IJCNN | 6 |
| 2024 | Efficient Image Super-Resolution via Symmetric Visual Attention NetworkabstractIn recent years, efficient super-resolution research has focused on reducing model complexity and improving efficiency by leveraging deep small-kernel convolution, but it has the problem of a small receptive field, which leads to a limited ability of the network to reconstruct details. Large kernel convolution can provide a large receptive field and lead to a substantial enhancement in the quality of image reconstruction, but its computational cost is too high. To minimize the model’s parameter count and achieve efficient super-resolution reconstruction, this study introduces a symmetric visual attention network. The network decomposes the large kernel convolution into three different lightweight and efficient convolutions. It then forms a bottleneck structure by leveraging the varied receptive field sizes of these convolutions in combination. The attention mechanism is integrated to create a bottleneck attention module, enhancing the network’s feature awareness. Furthermore, the bottleneck attention modules are symmetrically arranged to construct a symmetric large kernel attention block, thereby further enhancing the network’s capability to extract deep features. The experimental results demonstrate that the proposed model achieves competitive quantitative metrics when compared to other lightweight super-resolution methods, and the details of the reconstructed images are enhanced. With only 183K parameters, the model achieves a lightweight yet high-quality super-resolution model, offering a novel solution approach for efficient super-resolution. Qinrui Fan, Chengxu Wu, Shu Hu 0001, Xi Wu 0004, Xin Wang 0001, Jing Hu 0009 |
IJCNN | 4 |
| 2024 | Uncertainty-Aware Explainable Recommendation with Large Language ModelsabstractProviding explanations within the recommendation system would boost user satisfaction and foster trust, especially by elaborating on the reasons for selecting recommended items tailored to the user. The predominant approach in this domain revolves around generating text-based explanations, with a notable emphasis on applying large language models (LLMs). However, refining LLMs for explainable recommendations proves impractical due to time constraints and computing resource limitations. As an alternative, the current approach involves training the prompt rather than the LLM. In this study, we developed a model that utilizes the ID vectors of user and item inputs as prompts for GPT-2. We employed a joint training mechanism within a multi-task learning framework to optimize both the recommendation task and explanation task. This strategy enables a more effective exploration of users’ interests, improving recommendation effectiveness and user satisfaction. Through the experiments, our method achieving 1.59 DIV, 0.57 USR and 0.41 FCR on the Yelp, TripAdvisor and Amazon dataset respectively, demonstrates superior performance over four SOTA methods in terms of explainability evaluation metric. In addition, we identified that the proposed model is able to ensure stable textual quality on the three public datasets. Yicui Peng, Chingsheng Lin, Guo Huang, Jinrong Hu, Bin Kong 0001, Shu Hu 0001, Xi Wu 0004, Xin Wang 0045 |
IJCNN | 9 |
| 2024 | X-Transfer: A Transfer Learning-Based Framework for GAN-Generated Fake Image DetectionabstractGenerative adversarial networks (GANs) have remarkably advanced in diverse domains, especially image generation and editing. However, the misuse of GANs for generating deceptive images, such as face replacement, raises significant security concerns, which have gained widespread attention. Therefore, it is urgent to develop effective detection methods to distinguish between real and fake images. Current research centers around the application of transfer learning. Nevertheless, it encounters challenges such as knowledge forgetting from the original dataset and inadequate performance when dealing with imbalanced data during training. To alleviate this issue, this paper introduces a novel GAN-generated image detection algorithm called X-Transfer, which enhances transfer learning by utilizing two neural networks that employ interleaved parallel gradient transmission. In addition, we combine AUC loss and cross-entropy loss to improve the model’s performance. We carry out comprehensive experiments on multiple facial image datasets. The results show that our model outperforms the general transferring approach, and the best metric achieves 99.04%, which is increased by approximately 10%. Furthermore, we demonstrate excellent performance on non-face datasets, validating its generality and broader application prospects. Shu Hu 0001, Bin B. Zhu, Chingsheng Lin, Xi Wu 0004, Jinrong Hu, Xin Wang 0045 |
IJCNN | 6 |
| 2024 | MCAD: Multi-modal Conditioned Adversarial Diffusion Model for High-Quality PET Image Reconstruction
Xinyi Zeng, Pinxian Zeng, Bo Liu 0113, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
MICCAI (7) | 5 |
| 2024 | Robustly Optimized Deep Feature Decoupling Network for Fatty Liver Diseases Detection
Shu Hu 0001, Bo Peng 0006, Jiashu Zhang, Xi Wu 0004, Xin Wang 0045 |
MICCAI (1) | 5 |
| 2024 | Learning with Alignments: Tackling the Inter- and Intra-domain Shifts for Cross-multidomain Facial Expression RecognitionabstractFacial Expression Recognition (FER) holds significant importance in human-computer interactions. Existing cross-domain FER methods often transfer knowledge solely from a single labeled source domain to an unlabeled target domain, neglecting the comprehensive information across multiple sources. Nevertheless, cross-multidomain FER (CMFER) is very challenging for (i) the inherent inter-domain shifts across multiple domains and (ii) the intra-domain shifts stemming from the ambiguous expressions and low inter-class distinctions. In this paper, we propose a novel Learning with Alignments CMFER framework, named LA-CMFER, to handle both inter- and intra-domain shifts. Specifically, LA-CMFER is constructed with a global branch and a local branch to extract features from the full images and local subtle expressions, respectively. Based on this, LA-CMFER presents a dual-level inter-domain alignment method to force the model to prioritize hard-to-align samples in knowledge transfer at a sample level while gradually generating a well-clustered feature space with the guidance of class attributes at a cluster level, thus narrowing the inter-domain shifts. To address the intra-domain shifts, LA-CMFER introduces a multi-view intra-domain alignment method with a multi-view clustering consistency constraint where a prediction similarity matrix is built to pursue consistency between the global and local views, thus refining pseudo labels and eliminating latent noise. Extensive experiments on six benchmark datasets have validated the superiority of our LA-CMFER. Yuxiang Yang 0009, Lu Wen, Xinyi Zeng, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
ACM Multimedia | 5 |
| 2024 | Near-Surface Air Temperature Inversion Study Based on U-Net Family with Multi-source Data
Wanzhen Tang, Jing Peng 0003, Xuefei Hu, Xi Wu 0004, Xiaojie Li 0001, Shanmin Yang |
PRCV (4) | 4 |
| 2024 | Weakly supervised semantic segmentation by knowledge graph inference
Jia Zhang 0027, Bo Peng 0006, Xi Wu 0004, Jie Hu 0007 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | CAGAN: Classifier-augmented generative adversarial networks for weakly-supervised COVID-19 lung lesion localisationabstractAbstract The Coronavirus Disease 2019 (COVID‐19) epidemic has constituted a Public Health Emergency of International Concern. Chest computed tomography (CT) can help early reveal abnormalities indicative of lung disease. Thus, accurate and automatic localisation of lung lesions is particularly important to assist physicians in rapid diagnosis of COVID‐19 patients. The authors propose a classifier‐augmented generative adversarial network framework for weakly supervised COVID‐19 lung lesion localisation. It consists of an abnormality map generator, discriminator and classifier. The generator aims to produce the abnormality feature map M to locate lesion regions and then constructs images of the pseudo‐healthy subjects by adding M to the input patient images. Besides constraining the generated images of healthy subjects with real distribution by the discriminator, a pre‐trained classifier is introduced to enhance the generated images of healthy subjects to possess similar feature representations with real healthy people in terms of high‐level semantic features. Moreover, an attention gate is employed in the generator to reduce the noise effect in the irrelevant regions of M . Experimental results on the COVID‐19 CT dataset show that the method is effective in capturing more lesion areas and generating less noise in unrelated areas, and it has significant advantages in terms of quantitative and qualitative results over existing methods. Xiaojie Li 0001, Xin Fei, Hongping Ren, Canghong Shi, Xian Zhang 0008, Imran Mumtaz, Xi Wu 0004 |
IET Comput. Vis. | 9 |
| 2024 | Dual-granularity feature fusion in visible-infrared person re-identificationabstractAbstract Visible‐infrared person re‐identification (VI‐ReID) aims to recognize images of the same person captured in different modalities. Existing methods mainly focus on learning single‐granularity representations, which have limited discriminability and weak robustness. This paper proposes a novel dual‐granularity feature fusion network for VI‐ReID. Specifically, a dual‐branch module that extracts global and local features and then fuses them to enhance the representative ability is adopted. Furthermore, an identity‐aware modal discrepancy loss that promotes modality alignment by reducing the gap between features from visible and infrared modalities is proposed. Finally, considering the influence of non‐discriminative information in the modal‐shared features of RGB‐IR, a greyscale conversion is introduced to extract modality‐irrelevant discriminative features better. Extensive experiments on the SYSU‐MM01 and RegDB datasets demonstrate the effectiveness of the framework and superiority over state‐of‐the‐art methods. Shuang Cai, Shanmin Yang, Jing Hu 0009, Xi Wu 0004 |
IET Image Process. | 4 |
| 2024 | Cross-Chain Interoperability and Collaboration for Keyword-Based Embedded Smart Contracts in Internet of ThingsabstractIn recent years, blockchain technology has been widely applied in the Internet of Things (IoT) field, where devices located in different blockchains need to interact with each other cross-chain. Existing cross-chain models have high-implementation complexity, long response times, and are difficult to apply in IoT scenarios. In this article, we propose keyword-based embedded smart contract cross-chain interoperability and collaboration model. This model identifies the status of smart contracts through the keywords of smart contracts, completes cross-chain operations, and achieves smart contract collaboration and asset exchange among devices in different blockchains. To find matching collaborative contracts, contract-specific keywords will be set, including asset amounts and types. In the cross-chain asset exchange scenario of the model, the assets sent by the sender are frozen in the contract. When the contract is successfully executed, the assets will be sent to the receiver, otherwise returned to the sender. In the cross-chain contract collaboration scenario, we identify the number of contract interactions, and when the contract collaboration is completed, the data will be retained, otherwise it will be invalidated. We also use an embedded smart contract method to further optimize IoT application scenarios. This method embeds the steps of deploying smart contracts into the invoked transaction, enabling the deployment and invocation of smart contracts through a single transaction. The experiments show that the keyword-based embedded contract method is 47% faster and 49% less costly compared to notary schemes solutions in cross-chain scenarios. Hong Su, Xi Wu 0004, Yuliang Yang |
IEEE Internet Things J. | 3 |
| 2024 | 3D multi-modality Transformer-GAN for high-quality PET reconstruction
Yan Wang 0015, Yanmei Luo, Chen Zu, Bo Zhan, Zhengyang Jiao, Xi Wu 0004, Jiliu Zhou, Dinggang Shen, Luping Zhou |
Medical Image Anal. | 6 |
| 2024 | Source-free domain adaptation via dynamic pseudo labeling and Self-supervision
Qiankun Ma, Jie Zeng 0003, Jianjia Zhang, Chen Zu, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
Pattern Recognit. | 5 |
| 2024 | Semi-supervised medical image segmentation via hard positives oriented contrastive learning
Cheng Tang 0003, Xinyi Zeng, Luping Zhou, Qizheng Zhou, Xi Wu 0004, Hongping Ren, Jiliu Zhou, Yan Wang 0015 |
Pattern Recognit. | 6 |
| 2024 | Ma2SP: Missing-Aware Prompting With Modality-Adaptive Integration for Incomplete Multi-Modal Survival PredictionabstractSurvival prediction is crucial for head and neck (H&N) cancer patients. Recently, deep learning-based multi-modal models have achieved promising performance in accurate survival prediction. However, their clinical application is hindered by the difficulty of acquiring complete sets of multi-modal data. To tackle this limitation, in this paper, we propose a novel framework, namely Ma2SP, for incomplete multi-modal survival prediction in H&N cancer. Specifically, we develop missing-aware integration (MAI) modules to align heterogeneous multi-modal data and encourage dynamic interactions among available modalities, thereby achieving flexible multi-modal integration and enhancing robustness to incomplete data. Moreover, we employ missing-aware prompting (MAP) to provide explicit guidance on missing states during training, enabling effective training with incomplete data. In addition, we introduce tumor segmentation as an auxiliary task to capture tumor-related information, which further improves prediction accuracy. Experiments demonstrate our superior performance. Hanci Zheng, Yuanjun Liu 0002, Xi Wu 0004, Yan Wang 0015 |
IEEE Signal Process. Lett. | 4 |
| 2023 | LION: Label Disambiguation for Semi-supervised Facial Expression Recognition with Progressive Negative LearningabstractSemi-supervised deep facial expression recognition (SS-DFER) has recently attracted rising research interest due to its more practical setting of abundant unlabeled data. However, there are two main problems unconsidered in current SS-DFER methods: 1) label ambiguity, i.e., given labels mismatch with facial expressions; 2) inefficient utilization of unlabeled data with low-confidence. In this paper, we propose a novel SS-DFER method, including a Label DIsambiguation module and a PrOgressive Negative Learning module, namely LION, to simultaneously address both problems. Specifically, the label disambiguation module operates on labeled data, including data with accurate labels (clear data) and ambiguous labels (ambiguous data). It first uses clear data to calculate prototypes for all the expression classes, and then re-assign a candidate label set to all the ambiguous data. Based on the prototypes and the candidate label set, the ambiguous data can be relabeled more accurately. As for unlabeled data with low-confidence, the progressive negative learning module is developed to iteratively mine more complete complementary labels, which can guide the model to reduce the association between data and corresponding complementary labels. Experiments on three challenging datasets show that our method significantly outperforms the current state-of-the-art approaches in SS-DFER and surpasses fully-supervised baselines. Code will be available at https://github.com/NUM-7/LION. Zhongjing Du, Xu Jiang 0004, Qizheng Zhou, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
IJCAI | 5 |
| 2023 | Controlling Neural Style Transfer with Deep Reinforcement LearningabstractControlling the degree of stylization in the Neural Style Transfer (NST) is a little tricky since it usually needs hand-engineering on hyper-parameters. In this paper, we propose the first deep Reinforcement Learning (RL) based architecture that splits one-step style transfer into a step-wise process for the NST task. Our RL-based method tends to preserve more details and structures of the content image in early steps, and synthesize more style patterns in later steps. It is a user-easily-controlled style-transfer method. Additionally, as our RL-based model performs the stylization progressively, it is lightweight and has lower computational complexity than existing one-step Deep Learning (DL) based models. Experimental results demonstrate the effectiveness and robustness of our method. Chengming Feng, Jing Hu 0009, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Xi Wu 0004, Hongtu Zhu, Siwei Lyu |
IJCAI | 6 |
| 2023 | RMBench: Benchmarking Deep Reinforcement Learning for Robotic Manipulator ControlabstractReinforcement learning is used to tackle complex tasks with high-dimensional sensory inputs. Over the past decade, a wide range of reinforcement learning algorithms have been developed, with recent progress benefiting from deep learning for raw sensory signal representation. This raises a natural question: how well do these algorithms perform across different robotic manipulation tasks? To objectively compare algorithms, benchmarks use performance metrics. Benchmarks use objective performance metrics to offer a scientific way to compare algorithms. In this paper, we introduce RMBench, the first benchmark for robotic manipulations with high-dimensional continuous action and state spaces. We implement and evaluate reinforcement learning algorithms that take observed pixels as inputs and report their average performance and learning curves to demonstrate their performance and training stability. Our study concludes that none of the evaluated algorithms can handle all tasks well, with soft Actor-Critic outperforming most algorithms in terms of average reward and stability, and an algorithm combined with data augmentation potentially facilitating learning policies. Our code is publicly available at https://github.com/xiangyanfei212/RMBench-2022.git, including all benchmark tasks and studied algorithms. Yanfei Xiang, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Xiaomeng Huang, Xi Wu 0004, Siwei Lyu |
IROS | 6 |
| 2023 | TriDo-Former: A Triple-Domain Transformer for Direct PET Reconstruction from Low-Dose Sinograms
Pinxian Zeng, Xinyi Zeng, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
MICCAI (10) | 5 |
| 2023 | DiffDP: Radiotherapy Dose Prediction via a Diffusion Model
Zhenghao Feng, Lu Wen, Binyu Yan, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
MICCAI (6) | 5 |
| 2023 | Unsupervised Domain Adaptive Dose Prediction via Cross-Attention Transformer and Target-Specific Knowledge PreservationabstractRadiotherapy is one of the leading treatments for cancer. To accelerate the implementation of radiotherapy in clinic, various deep learning-based methods have been developed for automatic dose prediction. However, the effectiveness of these methods heavily relies on the availability of a substantial amount of data with labels, i.e. the dose distribution maps, which cost dosimetrists considerable time and effort to acquire. For cancers of low-incidence, such as cervical cancer, it is often a luxury to collect an adequate amount of labeled data to train a well-performing deep learning (DL) model. To mitigate this problem, in this paper, we resort to the unsupervised domain adaptation (UDA) strategy to achieve accurate dose prediction for cervical cancer (target domain) by leveraging the well-labeled high-incidence rectal cancer (source domain). Specifically, we introduce the cross-attention mechanism to learn the domain-invariant features and develop a cross-attention transformer-based encoder to align the two different cancer domains. Meanwhile, to preserve the target-specific knowledge, we employ multiple domain classifiers to enforce the network to extract more discriminative target features. In addition, we employ two independent convolutional neural network (CNN) decoders to compensate for the lack of spatial inductive bias in the pure transformer and generate accurate dose maps for both domains. Furthermore, to enhance the performance, two additional losses, i.e. a knowledge distillation loss (KDL) and a domain classification loss (DCL), are incorporated to transfer the domain-invariant features while preserving domain-specific information. Experimental results on a rectal cancer dataset and a cervical cancer dataset have demonstrated that our method achieves the best quantitative results with [Formula: see text], [Formula: see text], and HI of 1.446, 1.231, and 0.082, respectively, and outperforms other methods in terms of qualitative assessment. Jianghong Xiao, Xi Wu 0004, Jiliu Zhou, Xingchen Peng, Yan Wang 0015 |
Int. J. Neural Syst. | 4 |
| 2023 | A Transformer-Embedded Multi-Task Model for Dose Distribution PredictionabstractRadiation therapy is a fundamental cancer treatment in the clinic. However, to satisfy the clinical requirements, radiologists have to iteratively adjust the radiotherapy plan based on experience, causing it extremely subjective and time-consuming to obtain a clinically acceptable plan. To this end, we introduce a transformer-embedded multi-task dose prediction (TransMTDP) network to automatically predict the dose distribution in radiotherapy. Specifically, to achieve more stable and accurate dose predictions, three highly correlated tasks are included in our TransMTDP network, i.e. a main dose prediction task to provide each pixel with a fine-grained dose value, an auxiliary isodose lines prediction task to produce coarse-grained dose ranges, and an auxiliary gradient prediction task to learn subtle gradient information such as radiation patterns and edges in the dose maps. The three correlated tasks are integrated through a shared encoder, following the multi-task learning strategy. To strengthen the connection of the output layers for different tasks, we further use two additional constraints, i.e. isodose consistency loss and gradient consistency loss, to reinforce the match between the dose distribution features generated by the auxiliary tasks and the main task. Additionally, considering many organs in the human body are symmetrical and the dose maps present abundant global features, we embed the transformer into our framework to capture the long-range dependencies of the dose maps. Evaluated on an in-house rectum cancer dataset and a public head and neck cancer dataset, our method gains superior performance compared with the state-of-the-art ones. Code is available at https://github.com/luuuwen/TransMTDP. Lu Wen, Jianghong Xiao, Xi Wu 0004, Jiliu Zhou, Xingchen Peng, Yan Wang 0015 |
Int. J. Neural Syst. | 4 |
| 2023 | Facial Expression Recognition with Contrastive Learning and Uncertainty-Guided RelabelingabstractFacial expression recognition (FER) plays a vital role in the field of human-computer interaction. To achieve automatic FER, various approaches based on deep learning (DL) have been presented. However, most of them lack for the extraction of discriminative expression semantic information and suffer from the problem of annotation ambiguity. In this paper, we propose an elaborately designed end-to-end recognition network with contrastive learning and uncertainty-guided relabeling, to recognize facial expressions efficiently and accurately, as well as to alleviate the impact of annotation ambiguity. Specifically, a supervised contrastive loss (SCL) is introduced to promote inter-class separability and intra-class compactness, thus helping the network extract fine-grained discriminative expression features. As for the annotation ambiguity problem, we present an uncertainty estimation-based relabeling module (UERM) to estimate the uncertainty of each sample and relabel the unreliable ones. In addition, to deal with the padding erosion problem, we embed an amending representation module (ARM) into the recognition network. Experimental results on three public benchmarks demonstrate that our proposed method facilitates the recognition performance remarkably with 90.91% on RAF-DB, 88.59% on FERPlus and 61.00% on AffectNet, outperforming current state-of-the-art (SOTA) FER methods. Code will be available at http//github.com/xiaohu-run/fer_supCon. Chen Zu, Qizheng Zhou, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
Int. J. Neural Syst. | 5 |
| 2023 | Uncertainty-weighted and relation-driven consistency training for semi-supervised head-and-neck tumor segmentation
Yuang Shi, Chen Zu, Pinli Yang, Hongping Ren, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
Knowl. Based Syst. | 6 |
| 2023 | TransDose: Transformer-based radiotherapy dose prediction from CT images guided by super-pixel-level GCN classification
Zhengyang Jiao, Xingchen Peng, Yan Wang 0015, Jianghong Xiao, Dong Nie, Xi Wu 0004, Xin Wang 0045, Jiliu Zhou, Dinggang Shen |
Medical Image Anal. | 6 |
| 2023 | Automatic Head-and-Neck Tumor Segmentation in MRI via an End-to-End Adversarial Network
Pinli Yang, Xingchen Peng, Jianghong Xiao, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
Neural Process. Lett. | 4 |
| 2023 | Multi-level progressive transfer learning for cervical cancer dose prediction
Lu Wen, Jianghong Xiao, Jie Zeng 0003, Chen Zu, Xi Wu 0004, Jiliu Zhou, Xingchen Peng, Yan Wang 0015 |
Pattern Recognit. | 5 |
| 2023 | Pluralistic Face Inpainting With Transformation of Attribute InformationabstractMost face-inpainting methods perform well in face repair. However, these methods can only complete a single face image per input. Although existing various image-inpainting methods can achieve pluralistic image inpainting, they typically produce faces with distorted structures or the same texture. To resolve these shortcomings and achieve high-quality diverse face inpainting, we propose PFTANet, a two-stage pluralistic face-inpainting network that transforms attribute information. In the first stage, the face-parsing network is fine-tuned to obtain semantic facial region information. In the second stage, a generator consisting of SNBlock, CF_ShiftBlocks, and CF_MergeBlock, which ensures that high-quality pluralistic face results are generated, is used. Specifically, CF_ShiftBlocks completes pluralistic face generation by transforming the attribute information from the conditional face extracted by the attribute extractor and ensuring the consistency of the attribute information between the conditional and generated faces. CF_MergeBlock ensures structural consistency between the masked and background regions of the generated face using facial region semantic information. A multi-patch discriminator is used to enhance facial detail generation. Experimental results for the CelebA and CelebA-HQ datasets indicated that PFTANet achieved pluralistic and visually realistic face inpainting. Yang Zhang 0155, Xian Zhang 0008, Canghong Shi, Xi Wu 0004, Xiaojie Li 0001, Jing Peng 0003, Kunlin Cao, Jiancheng Lv 0001, Jiliu Zhou |
IEEE Trans. Multim. | 4 |
| 2022 | Stochastic Planner-Actor-Critic for Unsupervised Deformable Image RegistrationabstractLarge deformations of organs, caused by diverse shapes and nonlinear shape changes, pose a significant challenge for medical image registration. Traditional registration methods need to iteratively optimize an objective function via a specific deformation model along with meticulous parameter tuning, but which have limited capabilities in registering images with large deformations. While deep learning-based methods can learn the complex mapping from input images to their respective deformation field, it is regression-based and is prone to be stuck at local minima, particularly when large deformations are involved. To this end, we present Stochastic Planner-Actor-Critic (spac), a novel reinforcement learning-based framework that performs step-wise registration. The key notion is warping a moving image successively by each time step to finally align to a fixed image. Considering that it is challenging to handle high dimensional continuous action and state spaces in the conventional reinforcement learning (RL) framework, we introduce a new concept `Plan' to the standard Actor-Critic model, which is of low dimension and can facilitate the actor to generate a tractable high dimensional action. The entire framework is based on unsupervised training and operates in an end-to-end manner. We evaluate our method on several 2D and 3D medical image datasets, some of which contain large deformations. Our empirical results highlight that our work achieves consistent, significant gains and outperforms state-of-the-art methods. Ziwei Luo 0002, Jing Hu 0009, Xin Wang 0045, Shu Hu 0001, Bin Kong 0001, Youbing Yin, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
AAAI | 8 |
| 2022 | Contrastive Class-Specific Encoding for Few-Shot Object DetectionabstractIn this paper, we propose a new few-shot object detection (FSOD) framework that introduces a new contrastive branch to extract the class representation of images, which improves the generalization performance of the detection model for novel classes. Additionally, we investigate the effectiveness of both self-supervised and supervised contrastive losses for class-specific encoding in our framework. Experimental results on the benchmark datasets indicate that our proposed method archives the state-of-the-art performance compared with existing FSOD methods. Dizhong Lin, Ying Fu 0003, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
ICME | 7 |
| 2022 | Classification-Aided High-Quality PET Image Synthesis via Bidirectional Contrastive GAN with Shared Information Maximization
Yuchen Fei, Chen Zu, Zhengyang Jiao, Xi Wu 0004, Jiliu Zhou, Dinggang Shen, Yan Wang 0015 |
MICCAI (6) | 4 |
| 2022 | 3D CVT-GAN: A 3D Convolutional Vision Transformer-GAN for PET Reconstruction
Pinxian Zeng, Luping Zhou, Chen Zu, Xinyi Zeng, Zhengyang Jiao, Xi Wu 0004, Jiliu Zhou, Dinggang Shen, Yan Wang 0015 |
MICCAI (6) | 6 |
| 2022 | DDNet: 3D densely connected convolutional networks with feature pyramids for nasopharyngeal carcinoma segmentationabstractAbstract Radiation therapy is the standard treatment for early stage Nasopharyngeal cancer (NPC). Thus, accurate delineation of target volumes at risk in NPC is important. While manual delineation is time‐consuming and labour‐intensive process and also leads to significant inter‐ and intra‐practitioner variability. Thus, computer‐aided segmentation algorithm is required. However, segmentation task is not trivial due to large variations (e.g., shape and size) of nasopharynx structure across subjects. Moreover, extreme foreground and background class imbalance in NPC segmentation remains challenge. In this paper, we propose a threedimensional densely connected convolutional neural network with multi‐scale feature pyramids for NPC segmentation. We adapt the densely connected convolutional block into a new structure via adding feature pyramids. The concatenated pyramid feature carries multi‐scale and hierarchical semantic information which is effective for segmenting different size of tumors and perceiving hierarchical context information. To address the foreground and background imbalance problem, we propose an enhanced version of focal loss. It prevents the large number of negative voxels far from boundaries from overwhelming the segmentation algorithm. We validated the proposed method on 120 clinical subjects. Experimental results demonstrate that our approach out‐performed state‐of‐the‐art methods and human experts. Xiaojie Li 0001, Mingxuan Tang, Kunlin Cao, Qi Song 0001, Xi Wu 0004, Shanhui Sun, Jiliu Zhou |
IET Image Process. | 7 |
| 2022 | Multistage semantic-aware image inpainting with stacked generator networksabstractDeep learning has been widely applied into image inpainting. However, traditional image processing methods (i.e., patch-based and diffusion-based methods) generally fail to produce visually natural contents and semantically reasonable structures due to ineffectively processing the high-level semantic information of images. To solve the problem, we propose a stacked generator networks assisted by patch discriminator for image inpainting by multistage. In the proposed method, our generator network mainly consists of three-layer stacked encoder-decoder architecture, which could fuse different level feature information and achieve image inpainting via a coarse-to-fine hierarchical representation. Meanwhile, we split the masked image into different patches in each layer, which could effectively enlarge the receptive field and extract more useful features of images. Moreover, the patch discriminator is introduced to judge the patches of inpainting image are real or fake. In this way, our network can effectively utilize the semantic information to complete a fine result. Furthermore, both perceptual loss and style loss are used to improve the inpainting results in verse. Experimental results on Places2 and Paris StreetView illustrate that our approach could generate high-quality inpainting results, and our method is more effective than the existing image inpainting methods. Yongpeng Ren, Hongping Ren, Canghong Shi, Xian Zhang 0008, Xi Wu 0004, Xiaojie Li 0001, Jiancheng Lv 0001, Jiliu Zhou, Imran Mumtaz |
Int. J. Intell. Syst. | 5 |
| 2022 | An Efficient Semi-Supervised Framework with Multi-Task and Curriculum Learning for Medical Image SegmentationabstractA practical problem in supervised deep learning for medical image segmentation is the lack of labeled data which is expensive and time-consuming to acquire. In contrast, there is a considerable amount of unlabeled data available in the clinic. To make better use of the unlabeled data and improve the generalization on limited labeled data, in this paper, a novel semi-supervised segmentation method via multi-task curriculum learning is presented. Here, curriculum learning means that when training the network, simpler knowledge is preferentially learned to assist the learning of more difficult knowledge. Concretely, our framework consists of a main segmentation task and two auxiliary tasks, i.e. the feature regression task and target detection task. The two auxiliary tasks predict some relatively simpler image-level attributes and bounding boxes as the pseudo labels for the main segmentation task, enforcing the pixel-level segmentation result to match the distribution of these pseudo labels. In addition, to solve the problem of class imbalance in the images, a bounding-box-based attention (BBA) module is embedded, enabling the segmentation network to concern more about the target region rather than the background. Furthermore, to alleviate the adverse effects caused by the possible deviation of pseudo labels, error tolerance mechanisms are also adopted in the auxiliary tasks, including inequality constraint and bounding-box amplification. Our method is validated on ACDC2017 and PROMISE12 datasets. Experimental results demonstrate that compared with the full supervision method and state-of-the-art semi-supervised methods, our method yields a much better segmentation performance on a small labeled dataset. Code is available at https://github.com/DeepMedLab/MTCL. Kaiping Wang, Yan Wang 0015, Bo Zhan, Chen Zu, Xi Wu 0004, Jiliu Zhou, Dong Nie, Luping Zhou |
Int. J. Neural Syst. | 6 |
| 2022 | Semi-supervised NPC segmentation with uncertainty and attention guided consistency
Xingchen Peng, Jianghong Xiao, Bo Zhan, Chen Zu, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
Knowl. Based Syst. | 7 |
| 2022 | Explainable attention guided adversarial deep network for 3D radiotherapy dose distribution prediction
Huidong Li, Xingchen Peng, Jie Zeng 0003, Jianghong Xiao, Dong Nie, Chen Zu, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
Knowl. Based Syst. | 7 |
| 2022 | Unified medical image segmentation by learning from uncertainty in an end-to-end manner
Pin Tang, Pinli Yang, Dong Nie, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
Knowl. Based Syst. | 4 |
| 2022 | D2FE-GAN: Decoupled dual feature extraction based GAN for MRI image synthesis
Bo Zhan, Luping Zhou, Xi Wu 0004, Yi-Fei Pu, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
Knowl. Based Syst. | 4 |
| 2022 | Semi-supervised medical image segmentation via a tripled-uncertainty guided mean teacher model with contrastive learning
Kaiping Wang, Bo Zhan, Chen Zu, Xi Wu 0004, Jiliu Zhou, Luping Zhou, Yan Wang 0015 |
Medical Image Anal. | 4 |
| 2022 | Learning a deep dual-level network for robust DeepFake detection
Wenbo Pu, Jing Hu 0009, Xin Wang 0045, Yuezun Li, Shu Hu 0001, Bin B. Zhu, Rui Song 0006, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
Pattern Recognit. | 9 |
| 2022 | ASMFS: Adaptive-similarity-based multi-modality feature selection for classification of Alzheimer's disease
Yuang Shi, Chen Zu, Luping Zhou, Lei Wang 0001, Xi Wu 0004, Jiliu Zhou, Daoqiang Zhang, Yan Wang 0015 |
Pattern Recognit. | 6 |
| 2022 | DE-GAN: Domain Embedded GAN for High Quality Face Image Inpainting
Xian Zhang 0008, Xin Wang 0045, Canghong Shi, Xiaojie Li 0001, Bin Kong 0001, Siwei Lyu, Bin B. Zhu, Jiancheng Lv 0001, Youbing Yin, Qi Song 0001, Xi Wu 0004, Imran Mumtaz |
Pattern Recognit. | 12 |
| 2022 | Multi-Modal MRI Image Synthesis via GAN With Multi-Scale Gate MergenceabstractMulti-modal magnetic resonance imaging (MRI) plays a critical role in clinical diagnosis and treatment nowadays. Each modality of MRI presents its own specific anatomical features which serve as complementary information to other modalities and can provide rich diagnostic information. However, due to the limitations of time consuming and expensive cost, some image sequences of patients may be lost or corrupted, posing an obstacle for accurate diagnosis. Although current multi-modal image synthesis approaches are able to alleviate the issues to some extent, they are still far short of fusing modalities effectively. In light of this, we propose a multi-scale gate mergence based generative adversarial network model, namely MGM-GAN, to synthesize one modality of MRI from others. Notably, we have multiple down-sampling branches corresponding to input modalities to specifically extract their unique features. In contrast to the generic multi-modal fusion approach of averaging or maximizing operations, we introduce a gate mergence (GM) mechanism to automatically learn the weights of different modalities across locations, enhancing the task-related information while suppressing the irrelative information. As such, the feature maps of all the input modalities at each down-sampling level, i.e., multi-scale levels, are integrated via GM module. In addition, both the adversarial loss and the pixel-wise loss, as well as gradient difference loss (GDL) are applied to train the network to produce the desired modality accurately. Extensive experiments demonstrate that the proposed method outperforms the state-of-the-art multi-modal image synthesis methods. Bo Zhan, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and LocalizationabstractFor iris recognition in non-cooperative environments, iris segmentation has been regarded as the first most important challenge still open to the biometric community, affecting all downstream tasks from normalization to recognition. In recent years, deep learning technologies have gained significant popularity among various computer vision tasks and also been introduced in iris biometrics, especially iris segmentation. To investigate recent developments and attract more interest of researchers in the iris segmentation method, we organized the 2021 NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and Localization (NIR-ISL 2021) at the 2021 International Joint Conference on Biometrics (IJCB 2021). The challenge was used as a public platform to assess the performance of iris segmentation and localization methods on Asian and African NIR iris images captured in non-cooperative environments. The three best-performing entries achieved solid and satisfactory iris segmentation and localization results in most cases, and their code and models have been made publicly available for reproducibility research. Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Jawad Muhammad, Qi Zhang 0015, Qichuan Tian, Zhaofeng He 0001, Zhenan Sun, Tianbao Liu, Wei Yang 0006, Dongliang Wu, Yingfeng Liu, Ruiye Zhou, Huihai Wu, Junbao Wang, Wantong Xiong, Xueyu Shi, Shao Zeng, Peihua Li, Huijie Wu, Xinhui Zhang, Menghan Zhang, Fadi Boutros, Naser Damer, Arjan Kuijper, Juan E. Tapia, Andres Valenzuela, Christoph Busch 0001, Gourav Gupta, Kiran B. Raja, Xi Wu 0004, Xiaojie Li 0001, Jingfu Yang, Hongyan Jing, Xin Wang 0045, Bin Kong 0001, Youbing Yin, Qi Song 0001, Siwei Lyu, Shu Hu 0001, Leon Premk, Matej Vitek, Vitomir Struc, Peter Peer, Jalil Nourmohammadi-Khiarak, Farhang Jaryani, Samaneh Salehi Nasab, Seyed Naeim Moafinejad, Yasin Amini, Morteza Noshad |
IJCB | 44 |
| 2021 | Imperceptible Adversarial Examples For Fake Image DetectionabstractFooling people with highly realistic fake images generated with Deepfake or GANs brings a great social disturbance to our society. Many methods have been proposed to detect fake images, but they are vulnerable to adversarial perturbations – intentionally designed noises that can lead to the wrong prediction. Existing methods of attacking fake image detectors usually generate adversarial perturbations to perturb almost the entire image. This is redundant and increases the perceptibility of perturbations. In this paper, we propose a novel method to disrupt the fake image detection by determining key pixels to a fake image detector and attacking only the key pixels, which results in the L0and the L2norms of adversarial perturbations much less than those of existing works. Experiments on two public datasets with three fake image detectors indicate that our proposed method achieves state-of the-art performance in both white-box and black-box attacks. Quanyu Liao, Yuezun Li, Xin Wang 0045, Bin Kong 0001, Bin B. Zhu, Siwei Lyu, Youbing Yin, Qi Song 0001, Xi Wu 0004 |
ICIP | 9 |
| 2021 | Transferable Adversarial Examples for Anchor Free Object DetectionabstractDeep neural networks have been demonstrated to be vulnerable to adversarial attacks: subtle perturbation can completely change prediction result. The vulnerability has led to a surge of research in this direction, including adversarial attacks on object detection networks. However, previous studies are dedicated to attacking anchor-based object detectors. In this paper, we present the first adversarial attack on anchor-free object detectors. It conducts category-wise, instead of previously instance-wise, attacks on object detectors, and leverages high-level semantic information to efficiently generate transferable adversarial examples, which can also be transferred to attack other object detectors, even anchor-based detectors such as Faster R-CNN. Experimental results on two benchmark datasets demonstrate that our proposed method achieves state-of-the-art performance and transferability. Quanyu Liao, Xin Wang 0045, Bin Kong 0001, Siwei Lyu, Bin B. Zhu, Youbing Yin, Qi Song 0001, Xi Wu 0004 |
ICME | 8 |
| 2021 | Stochastic Actor-Executor-Critic for Image-to-Image TranslationabstractTraining a model-free deep reinforcement learning model to solve image-to-image translation is difficult since it involves high-dimensional continuous state and action spaces. In this paper, we draw inspiration from the recent success of the maximum entropy reinforcement learning framework designed for challenging continuous control problems to develop stochastic policies over high dimensional continuous spaces including image representation, generation, and control simultaneously. Central to this method is the Stochastic Actor-Executor-Critic (SAEC) which is an off-policy actor-critic model with an additional executor to generate realistic images. Specifically, the actor focuses on the high-level representation and control policy by a stochastic latent action, as well as explicitly directs the executor to generate low-level actions to manipulate the state. Experiments on several image-to-image translation tasks have demonstrated the effectiveness and robustness of the proposed SAEC when facing high-dimensional continuous space problems. Ziwei Luo 0002, Jing Hu 0009, Xin Wang 0045, Siwei Lyu, Bin Kong 0001, Youbing Yin, Qi Song 0001, Xi Wu 0004 |
IJCAI | 8 |
| 2021 | 3D Transformer-GAN for High-Quality PET Reconstruction
Yanmei Luo, Yan Wang 0015, Chen Zu, Bo Zhan, Xi Wu 0004, Jiliu Zhou, Dinggang Shen, Luping Zhou |
MICCAI (6) | 5 |
| 2021 | Coarse-To-Fine Segmentation of Organs at Risk in Nasopharyngeal Carcinoma Radiotherapy
Qiankun Ma, Chen Zu, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
MICCAI (1) | 3 |
| 2021 | Incorporating Isodose Lines and Gradient Information via Multi-task Learning for Dose Prediction in Radiotherapy
Pin Tang, Xingchen Peng, Jianghong Xiao, Chen Zu, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
MICCAI (7) | 6 |
| 2021 | Tripled-Uncertainty Guided Mean Teacher Model for Semi-supervised Medical Image Segmentation
Kaiping Wang, Bo Zhan, Chen Zu, Xi Wu 0004, Jiliu Zhou, Luping Zhou, Yan Wang 0015 |
MICCAI (2) | 4 |
| 2021 | Edge-preserving MRI image synthesis via adversarial network with iterative multi-scale fusion
Yanmei Luo, Dong Nie, Bo Zhan, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
Neurocomputing | 5 |
| 2021 | DA-DSUnet: Dual Attention-based Dense SU-net for automatic head-and-neck tumor segmentation in MRI images
Pin Tang, Chen Zu, Xingchen Peng, Jianghong Xiao, Xi Wu 0004, Jiliu Zhou, Luping Zhou, Yan Wang 0015 |
Neurocomputing | 7 |
| 2021 | End-to-end multimodal image registration via reinforcement learning
Jing Hu 0009, Ziwei Luo 0002, Xin Wang 0045, Shanhui Sun, Youbing Yin, Kunlin Cao, Qi Song 0001, Siwei Lyu, Xi Wu 0004 |
Medical Image Anal. | 9 |
| 2021 | Automatic vertebrae recognition from arbitrary spine MRI images by a category-Consistent self-calibration detection framework
Xi Wu 0004, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 2 |
| 2021 | Research on the Application of Visual SLAM in Embedded GPUabstractIn the automatic navigation robot field, robotic autonomous positioning is one of the most difficult challenges. Simultaneous localization and mapping (SLAM) technology can incrementally construct a map of the robot’s moving path in an unknown environment while estimating the position of the robot in the map, providing an effective solution for robots to fully navigate autonomously. The camera can obtain corresponding two‐dimensional digital images from the real three‐dimensional world. These images contain very rich colour, texture information, and highly recognizable features, which provide indispensable information for robots to understand and recognize the environment based on the ability to autonomously explore the unknown environment. Therefore, more and more researchers use cameras to solve SLAM problems, also known as visual SLAM. Visual SLAM needs to process a large number of image data collected by the camera, which has high performance requirements for computing hardware, and thus, its application on embedded mobile platforms is greatly limited. This paper presents a parallelization method based on embedded hardware equipped with embedded GPU. Use CUDA, a parallel computing platform, to accelerate the visual front‐end processing of the visual SLAM algorithm. Extensive experiments are done to verify the effectiveness of the method. The results show that the presented method effectively improves the operating efficiency of the visual SLAM algorithm and ensures the original accuracy of the algorithm. Tianji Ma, Nanyang Bai, Xi Wu 0004, Lutao Wang, Tao Wu 0010, Changming Zhao |
Wirel. Commun. Mob. Comput. | 4 |
| 2020 | Squeeze the Ball: Designing an Interactive Playground towards Aiding Social Activities of Children with Low-Function AutismabstractMost intervention methods used for social skills training in children with autism are dedicated to high-functioning autism (HFA). However, extensive neurological and developmental disorders of low-functioning autism (LFA) have hampered their adoption. In this study, we observed and interviewed children with LFA, and their teachers, from a local educational institution, to better understand the children's social needs and barriers. Then, with the aim of aiding the children with their social activities, we illustrate the design process of SqueeBall, an interactive playground equipment. We evaluated the design with 18 children (16 with LFA and 2 with HFA) between 2.5 and 7 years of age. Results showed that these children had a pleasant game experience when the group bonded, and the equipment had a positive effect on aiding them in various ways. Finally, we discuss the challenges and opportunities of multimedia interaction techniques in aiding children with LFA. Chenmei Yu, Jiayu Yao, Xi Wu 0004, Xiaolan Peng, Teng Han |
CHI | 5 |
| 2020 | Fast Local Attack: Generating Local Adversarial Examples for Object DetectorsabstractThe deep neural network is vulnerable to adversarial examples. Adding imperceptible adversarial perturbations to images is enough to make them fail. Most existing research focuses on attacking image classifiers or anchor-based object detectors, but they generate globally perturbation on the whole image, which is unnecessary. In our work, we leverage higher-level semantic information to generate high aggressive local perturbations for anchor-free object detectors. As a result, it is less computationally intensive and achieves a higher black-box attack as well as transferring attack performance. The adversarial examples generated by our method are not only capable of attacking anchor-free object detectors, but also able to be transferred to attack anchor-based object detector. Quanyu Liao, Xin Wang 0045, Bin Kong 0001, Siwei Lyu, Youbing Yin, Qi Song 0001, Xi Wu 0004 |
IJCNN | 7 |
| 2020 | Discriminative Dictionary-Embedded Network for Comprehensive Vertebrae Tumor Diagnosis
Heyou Chang, Xi Wu 0004, Shuo Li 0001 |
MICCAI (6) | 4 |
| 2020 | Conformal Feature-Selection Wrappers and ensembles for negative-transfer avoidance
Shuang Zhou 0001, Evgueni N. Smirnov, Gijs Schoenmakers, Ralf L. M. Peeters, Xi Wu 0004 |
Neurocomputing | 5 |
| 2020 | Multi-scale region composition of hierarchical image segmentation
Bo Peng 0006, Zaid Al-Huda, Zhuyang Xie, Xi Wu 0004 |
Multim. Tools Appl. | 4 |
| 2020 | Image segmentation of nasopharyngeal carcinoma using 3D CNN with long-range skip connection and multi-scale feature pyramid
Canghong Shi, Xiaojie Li 0001, Xi Wu 0004, Jiliu Zhou, Jiancheng Lv 0001 |
Soft Comput. | 4 |
| 2019 | Embeddings and Convolution, Is That the Best You can Do with Sentiment Features?abstractRapid growth of digital media motivates research on machine-assisted text analysis. Sentiment analysis, among one of the prevalent applications, has drawn great attention. In addition to the traditional bag-of-words models, embedding methods have become de facto standard for text representation, and various convolutional, recurrent and recursive neural networks are dominating leaderboards. Despite the large number of deep learning models in publication, the performance benchmarks in sentiment analysis are approaching a limit. If language-specific syntactic and semantic knowledge is excluded, is there still room for significant improvements? Over a general neural network that is based on word embedding, 2D convolution and max-pooling, we conduct extensive experiments on its various components, including convolutional kernels, pooling methods, recurrent layers, and attention mechanism. Certain combinations show moderate improvements in classification accuracy which are comparable to more sophisticated networks, but no sign of major breakthrough is in sight. We also extend the scope with potential game changers, covering context-aware representations, linguistic information, and large scale knowledge transfer in natural languages. Reported metrics show their great value in breaking the current performance bottleneck. Ao Feng, Shuang Zhou 0001, Xi Wu 0004 |
IJCNN | 4 |
| 2019 | A Multi-modality Network for Cardiomyopathy Death Risk Prediction with CMR Images and Clinical Information
Chaoyang Xia, Xiaojie Li 0001, Xin Wang 0045, Bin Kong 0001, Yucheng Chen 0003, Youbing Yin, Kunlin Cao, Qi Song 0001, Siwei Lyu, Xi Wu 0004 |
MICCAI (2) | 10 |
| 2019 | Automatic Vertebrae Recognition from Arbitrary Spine MRI Images by a Hierarchical Self-calibration Detection Framework
Xi Wu 0004, Bo Chen 0013, Shuo Li 0001 |
MICCAI (4) | 2 |
| 2019 | Generative Adversarial Networks with Enhanced Symmetric Residual Units for Single Image Super-Resolution
Xianyu Wu, Xiaojie Li 0001, Jia He 0003, Xi Wu 0004, Imran Mumtaz |
MMM (1) | 4 |
| 2019 | ACNET: Attention-based Convolution Network with Additional Discriminative Features for DCM Classification (S)abstractFor dilated cardiomyopathy (DCM) patients, immediate emergency diagnosis and treatment are critical for life saving and later recovery.T1 mapping is a non-invasive and effective diagnostic imaging approach to detect DCM.However, it is a demanding and time-consuming approach.In this paper, we propose an attention-based network structure, which can automatically identify DCM patients in a speedy manner to prioritize their treatment.In the proposed method, we adopt attention modules to generate attention-aware features.Inside each attention module, a bottom-up top-down feed-forward structure is used to unfold the feed-forward and feed-back attention processes into a single feed-forward process.It allows the network to focus more on determining useful information about the current output that is significant in the input data.Moreover, inspired by the residual network idea, we make full use of the characteristics of the original data.Combined residual block, we design down-residual modules for classification tasks.It consists of seven convolution layers and three layers of residual blocks.Our network achieves the most advanced recognition performance on cardiac datasets.We evaluated our approach on CMR(cardiac magnetic resonance) T1 mapping images with lower PSNR(peak signal to noise ratio), and the results demonstrate that our architecture outperforms previous approaches. Xin Wang 0045, Xiaojie Li 0001, Yucheng Chen 0003, Jiliu Zhou, Kunlin Cao, Qi Song 0001, Xi Wu 0004, Youbing Yin |
SEKE | 8 |
| 2019 | Automatic spondylolisthesis grading from MRIs across modalities using faster adversarial recognition network
Xi Wu 0004, Bo Chen 0013, Shuo Li 0001 |
Medical Image Anal. | 2 |
| 2019 | Patch-wise label propagation for MR brain segmentation based on multi-atlas images
Yan Wang 0015, Chen Zu, Zongqing Ma, Kun He 0007, Xi Wu 0004, Jiliu Zhou |
Multim. Syst. | 6 |
| 2019 | 3D Auto-Context-Based Locality Adaptive Multi-Modality GANs for PET SynthesisabstractPositron emission tomography (PET) has been substantially used recently. To minimize the potential health risk caused by the tracer radiation inherent to PET scans, it is of great interest to synthesize the high-quality PET image from the low-dose one to reduce the radiation exposure. In this paper, we propose a 3D auto-context-based locality adaptive multi-modality generative adversarial networks model (LA-GANs) to synthesize the high-quality FDG PET image from the low-dose one with the accompanying MRI images that provide anatomical information. Our work has four contributions. First, different from the traditional methods that treat each image modality as an input channel and apply the same kernel to convolve the whole image, we argue that the contributions of different modalities could vary at different image locations, and therefore a unified kernel for a whole image is not optimal. To address this issue, we propose a locality adaptive strategy for multi-modality fusion. Second, we utilize 1 ×1 ×1 kernel to learn this locality adaptive fusion so that the number of additional parameters incurred by our method is kept minimum. Third, the proposed locality adaptive fusion mechanism is learned jointly with the PET image synthesis in a 3D conditional GANs model, which generates high-quality PET images by employing large-sized image patches and hierarchical features. Fourth, we apply the auto-context strategy to our scheme and propose an auto-context LA-GANs model to further refine the quality of synthesized images. Experimental results show that our method outperforms the traditional multi-modality fusion methods used in deep networks, as well as the state-of-the-art PET estimation approaches. Yan Wang 0015, Luping Zhou, Biting Yu, Lei Wang 0001, Chen Zu, David S. Lalush, Weili Lin, Xi Wu 0004, Jiliu Zhou, Dinggang Shen |
IEEE Trans. Medical Imaging | 8 |
| 2018 | Robust Multimodal Image Registration Using Deep Recurrent Reinforcement Learning
Shanhui Sun, Jing Hu 0009, Mingqing Yao, Jinrong Hu, Qi Song 0001, Xi Wu 0004 |
ACCV (2) | 7 |
| 2018 | Deep Learning intra-image and inter-images features for Co-saliency detection
Shizhong Dong, Zhifan Gao, Xi Wu 0004, Heye Zhang, Guang Yang 0006, Shuo Li 0001 |
BMVC | 5 |
| 2018 | Noise Robust Single Image Super-Resolution Using a Multiscale Image PyramidabstractSingle image super-resolution (SR) generates a high-resolution (HR) image by estimating the mapping function between image patches of different resolutions. However, this kind of SR method cannot be directly applied to noisy images, since noise will be reinforced in the process of super-resolution. To this end, this paper presents a simultaneous super-resolution and denoising method by exploiting the noise decreasing property contained in the multiscale image pyramid. Experimental results confirm that our method is able to outperform other state-of-the-art super-resolution methods when super-resolving noisy images across differing noise levels. Jing Hu 0009, Xi Wu 0004, Jiliu Zhou |
ICIP | 3 |
| 2018 | Outlier Detection Based on the Data StructureabstractOutlier detection is one of the most frequently demanded task for optimizing results. Distance-based methods are a popular approach. They require no prior assumptions about the data generating distribution and are uncomplicated to implement. However, related methods have different parameters that are difficult to determine such that the identification results are generally unstable. Presenting related techniques without sacrificing stability is a challenging task. In this paper, we propose a new distance-based method that depends on the data structure to detect such points. In the proposed method, a global binary tree is constructed and the local distance score of a point is calculated to evaluate to what degree the observation is an outlier. The greater the value of the distance score, the more likely the point is an outlier point. Unlike typical distance-based methods, our algorithm has good scalability. Even when the dimension of the data points increases, the performance of our algorithm does not diminish. To reduce extra parameters, the top-p ranked points can be identified as outliers. Experimental results on synthetic and real-world datasets demonstrate the effectiveness and stability of our method. Canghong Shi, Xiaojie Li 0001, Jia He 0003, Xi Wu 0004 |
IJCNN | 5 |
| 2018 | Locality Adaptive Multi-modality GANs for High-Quality PET Image Synthesis
Yan Wang 0015, Luping Zhou, Lei Wang 0001, Biting Yu, Chen Zu, David S. Lalush, Weili Lin, Xi Wu 0004, Jiliu Zhou, Dinggang Shen |
MICCAI (1) | 8 |
| 2018 | Automatic Tumor Segmentation with Deep Convolutional Neural Networks for Radiotherapy Applications
Yan Wang 0015, Chen Zu, Guangliang Hu, Zongqing Ma, Kun He 0007, Xi Wu 0004, Jiliu Zhou |
Neural Process. Lett. | 7 |
| 2018 | Noise robust single image super-resolution using a multiscale image pyramid
Jing Hu 0009, Xi Wu 0004, Jiliu Zhou |
Signal Process. | 2 |
| 2018 | Automatic detection of boundary points based on local geometrical measures
Xiaojie Li 0001, Xi Wu 0004, Jiancheng Lv 0001, Jia He 0003, Jianping Gou, Mao Li 0001 |
Soft Comput. | 2 |
| 2018 | CUNet: A Compact Unsupervised Network For Image ClassificationabstractIn this paper, we propose a compact network called compact unsupervised network (CUNet) to address the image classification challenge. Contrasting the usual learning approach of convolutional neural networks, learning is achieved by the simple K-means on diverse image patches. This approach performs well even with scarcely labeled training images, greatly reducing the computational cost, while maintaining high discriminative power. Furthermore, we propose a new weighted pooling method in which different weighting values of adjacent neurons are considered. This strategy leads to improved classification since the network becomes more robust against small image distortions. In the output layer, CUNet integrates feature maps obtained in the last hidden layer, and straightforwardly computes histograms in nonoverlapped blocks. To reduce feature redundancy, we also implement the max-pooling operation on adjacent blocks to select the most competitive features. Comprehensive experiments on well-established databases are conducted to validate the classification performances of the introduced CUNet approach. Mengdie Mao, Gaipeng Kong, Xi Wu 0004, Qianni Zhang, Xiaochun Cao, Ebroul Izquierdo |
IEEE Trans. Multim. | 5 |
| 2017 | A Semi-supervised manifold alignment algorithm and an evaluation method based on local structure preservation
Xiaojie Li 0001, Jiancheng Lv 0001, Xi Wu 0004 |
Neurocomputing | 3 |
| 2016 | Analysis of micro-Doppler signatures of vibration targets using EMD and SPWVD
Yan Wang 0015, Xi Wu 0004, Wenzao Li, Yi Zhang 0018, Jiliu Zhou |
Neurocomputing | 2 |
| 2015 | A New Framework for Container Code Recognition by Using Segmentation-Based and HMM-Based ApproachesabstractTraditional methods for automatic recognition of container code in visual images are based on segmentation and recognition of isolated characters. However, when the segment fails to separate each character from the others, those methods will not function properly. Sometimes the container code characters are printed or arranged very closely, which makes it a challenge to isolate each character. To address this issue, a new framework for automatic container code recognition (ACCR) in visual images is proposed in this paper. In this framework, code-character regions are first located by applying a horizontal high-pass filter and scan line analysis. Then, character blocks are extracted from the code-character regions and further classified into two categories, i.e. single-character block and multi-character block. Finally, a segmentation-based approach is implemented for recognition of the characters in single-character blocks, and a hidden Markov model (HMM)-based method is proposed for the multi-character blocks. The experimental results demonstrate the effectiveness of the proposed method, which can successfully recognize the container code with closely arranged characters. Wei Wu 0002, Zheng Liu 0002, Zhiming Liu 0009, Xi Wu 0004, Xiaohai He |
Int. J. Pattern Recognit. Artif. Intell. | 5 |