VLDB 2026 Research / reviewers in the wild / expert
Bin Xiao 0002
dblp:43/5134-2
· DBLP profile ↗
145ranked-venue papers
23as first author
112since 2021 · last 2026
0000-0001-8469-5302ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 66 · 8 first-author · 58 since 2021Artificial intelligence and machine learning · 62 · 13 first-author · 45 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 18 since 2021Databases, data management, data science and information retrieval · 13 · 3 first-author · 6 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSR: Semantic and Spatial Rectification for CLIP-based Weakly Supervised SegmentationabstractIn recent years, Contrastive Language-Image Pretraining (CLIP) has been widely applied to Weakly Supervised Semantic Segmentation (WSSS) tasks due to its powerful cross-modal semantic understanding capabilities. This paper proposes a novel Semantic and Spatial Rectification (SSR) method to address the limitations of existing CLIP-based weakly supervised semantic segmentation approaches: over-activation in non-target foreground regions and background areas. Specifically, at the semantic level, the Cross-Modal Prototype Alignment (CMPA) establishes a contrastive learning mechanism to enforce feature space alignment across modalities, reducing inter-class overlap while enhancing semantic correlations, to rectify over-activation in non-target foreground regions effectively; at the spatial level, the Superpixel-Guided Correction (SGC) leverages superpixel-based spatial priors to precisely filter out interference from non-target regions during affinity propagation, significantly rectifying background over-activation. Extensive experiments on the PASCAL VOC and MS COCO datasets demonstrate that our method outperforms all single-stage approaches, as well as more complex multi-stage approaches, achieving mIoU scores of 79.5% and 50.6%, respectively. Xiuli Bi, Die Xiao, Junchao Fan, Bin Xiao 0002 |
AAAI | 4 |
| 2026 | Clear Nights Ahead: Towards Multi-Weather Nighttime Image RestorationabstractRestoring nighttime images affected by multiple adverse weather conditions is a practical yet under-explored research problem, as multiple weather degradations usually coexist in the real world alongside various lighting effects at night. This paper first explores the challenging multi-weather nighttime image restoration task, where various types of weather degradations are intertwined with flare effects. To support the research, we contribute the AllWeatherNight dataset, featuring large-scale nighttime images with diverse compositional degradations. By employing illumination-aware degradation generation, our dataset significantly enhances the realism of synthetic degradations in nighttime scenes, providing a more reliable benchmark for model training and evaluation. Additionally, we propose ClearNight, a unified nighttime image restoration framework, which effectively removes complex degradations in one go. Specifically, ClearNight extracts Retinex-based dual priors and explicitly guides the network to focus on uneven illumination regions and intrinsic texture contents respectively, thereby enhancing restoration effectiveness in nighttime scenarios. Moreover, to more effectively model the common and unique characteristics of multiple weather degradations, ClearNight performs weather-aware dynamic specificity and commonality collaboration that adaptively allocates optimal sub-networks associated with specific weather types. Comprehensive experiments on both synthetic and real-world images demonstrate the necessity of the AllWeatherNight dataset and the superior performance of ClearNight. Yuetong Liu, Yunqiu Xu, Yang Wei 0002, Xiuli Bi, Bin Xiao 0002 |
AAAI | 5 |
| 2026 | TGDD: Trajectory Guided Dataset Distillation with Balanced DistributionabstractDataset distillation compresses large datasets into compact synthetic ones to reduce storage and computational costs. Among various approaches, distribution matching (DM)-based methods have attracted attention for their high efficiency. However, they often overlook the evolution of feature representations during training, which limits the expressiveness of synthetic data and weakens downstream performance. To address this issue, we propose Trajectory Guided Dataset Distillation (TGDD), which reformulates distribution matching as a dynamic alignment process along the model’s training trajectory. At each training stage, TGDD captures evolving semantics by aligning the feature distribution between the synthetic and original dataset. Meanwhile, it introduces a distribution constraint regularization to reduce class overlap. This design helps synthetic data preserve both semantic diversity and representativeness, improving performance in downstream tasks. Without additional optimization overhead, TGDD achieves a favorable balance between performance and efficiency. Experiments on ten datasets demonstrate that TGDD achieves state-of-the-art performance, notably a 5.0% accuracy gain on high-resolution benchmarks. Fengli Ran, Xiao Pu 0002, Bo Liu 0047, Xiuli Bi, Bin Xiao 0002 |
AAAI | 5 |
| 2026 | Diff-AEPNet: Facial Aesthetic Enhancement and Prediction Network Based on Differential Average Aesthetic PerceptionsabstractWith the advent of the intelligent era, increasing attention has been given to facial aesthetics. While the academic community has achieved notable progress in facial aesthetic research, current efforts predominantly concentrate on two isolated subtasks: aesthetic evaluation and enhancement. Crucially, the intrinsic correlation between these tasks and their integration within a unified framework remain underexplored. To bridge this gap, this paper proposes a facial aesthetic enhancement and prediction network based on differential average aesthetic perceptions (Diff-AEPNet) that synergistically combines facial aesthetic enhancement with prediction. The proposed framework implements a four-stage architecture: (1) a transformer module learns latent code beautification trajectories to guide preenhancement feature modification; (2) a dual-stream encoder extracts and contrasts pre/postbeautification features to refine evaluation accuracy; (3) a lightweight network generates attention-guided image mask for image fusion; and (4) a deghosting block eliminates fusion artifacts through residual learning. The experimental results demonstrate that the model achieves a favorable beautification effect in the enhancement task and exhibits better generalization performance across datasets in the evaluation task than existing aesthetic evaluation models do. Weisheng Li 0001, Bin Xiao 0002, Yong Wang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Mask-Guided Proxy Mining Network for Few-Shot Medical Image SegmentationabstractFew-shot medical image segmentation (FSMIS) has attracted increasing attention as a promising technique for solving medical image segmentation tasks by relying on only a small amount of labeled data from new classes. Current FSMIS methods typically employ pixel-level semantic correlations between support-query image pairs to guide the segmentation of query images. However, the class information gap between support and query images may induce severe mismatches, leading to semantic ambiguity between foreground and background pixels. To address this issue, we propose a novel mask-guided proxy mining network (MPMNet), which mines a set of representative reference features (termed proxies) from support and query images to rectify foreground-background ambiguity. Specifically, to eliminate false pairwise matches caused by excessive intra-class variations, we design a mask-guided proxy mining module to adaptively learn representative proxies that can perceive visual differences between objects with different scales and shapes. Moreover, we integrate a hierarchical prior generation module and a context-aware feature enrichment module into MPMNet to obtain multi-scale information and enhance the discriminability of features. With these well-designed components and structures, our MPMNet can effectively overcome the adverse effects of false pixel matches by establishing proxy-level semantic correlations. Extensive experiments on three standard medical segmentation benchmarks demonstrate that our MPMNet significantly outperforms previous state-of-the-art methods, with a mean gain of 2.71% in DSC across all datasets. The code is available at: https://github.com/donglongzi/MPMNet. Wendong Huang, Jinwu Hu, Yongchao Wang 0004, Xiuli Bi, Yucheng Shu, Xuezong Yang, Bin Xiao 0002 |
IEEE Trans. Image Process. | 8 |
| 2026 | Dynamic Prompt Compression for Efficient Inference of Large Language ModelsabstractLarge language models (LLMs) have shown outstanding performance across a variety of tasks, partly due to advanced prompting techniques. However, these techniques often require lengthy prompts, which increase computational costs and can hinder performance because of the limited context windows of LLMs. While prompt compression is a straightforward solution, existing methods confront the challenges of retaining essential information, adapting to context changes, and remaining effective across different tasks. To tackle these issues, we propose a task-agnostic method called Dynamic Prompt Compression (LLM-DPC). Our method reduces the number of prompt tokens while minimizing any degradation in LLM performance. We model prompt compression as a Markov Decision Process (MDP), enabling the DPC-Agent to sequentially remove redundant tokens by adapting to dynamic contexts and retaining crucial content. We develop a reward function for training the DPC-Agent that balances the compression ratio, the quality of the LLM output, and the retention of key information. This allows for prompt token reduction without needing an external black-box LLM. Inspired by the progressive difficulty adjustment in curriculum learning, we introduce a Hierarchical Prompt Compression (HPC) training strategy that gradually increases the compression difficulty, enabling the DPC-Agent to learn an effective compression method that maintains information integrity. Experiments demonstrate that our method outperforms state-of-the-art techniques, especially at higher compression ratio. Jinwu Hu, Wei Zhang 0098, Yufeng Wang 0004, Yu Hu 0004, Bin Xiao 0002, Mingkui Tan |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2026 | Multi-Granularity Superpoint Graph Learning for Weakly Supervised 3D Semantic SegmentationabstractWeakly supervised 3D semantic segmentation has proven effective in alleviating the heavy dependence on dense annotations by generating high-quality pseudo-labels. However, due to the scene complexity and disorder of the point cloud, merely applying the model semantic prediction or hand-crafted feature similarity for pseudo labeling is inefficient and biased. This limitation inevitably results in incorrect pseudo labels. To tackle this challenge, we propose a new method called Multi-granularity Superpoint Graph Learning (MSGL) that leverages the multi-scale local features of point clouds to improve the quality of pseudo labels. We first design a multi-granularity local representation learning module on the superpoint graph to capture the neighboring structure information of each superpoint within complex scenes. Subsequently, the generated structural embedding is utilized to enhance the affinity matrix of label propagation, thereby yielding high-quality pseudo labels. To further enforce the generalization of the structural representation module under scenario changes or data fluctuations, we present a multi-granularity consistency loss in MSGL. This loss is applied across different views of the superpoint graph within each scene to ensure a robust and consistent learning process. Our experiments conducted on three benchmarks show that the proposed method outperforms existing weakly supervised methods under several sparse label settings, and improves the baseline by an average of 7.7% with only 1% extra computation cost. Moreover, our approach even compares favorably to some fully supervised methods with only one point labeled for each thing. Yan Fan 0002, Yu Wang 0106, Pengfei Zhu 0001, Le Hui, Jin Xie 0001, Bin Xiao 0002, Qinghua Hu |
IEEE Trans. Multim. | 6 |
| 2026 | Robust Image Stitching With Optimal PlaneabstractWe present RopStitch, an unsupervised deep image stitching framework with both robustness and naturalness. To ensure the robustness of RopStitch, we propose to incorporate the universal prior of content perception into the image stitching model by a dual-branch architecture. It separately captures coarse and fine features and integrates them to achieve highly generalizable performance across diverse unseen real-world scenes. Concretely, the dual-branch model consists of a pretrained branch to capture semantically invariant representations and a learnable branch to extract fine-grained discriminative features, which are then merged into a whole by a controllable factor at the correlation level. Besides, considering that content alignment and structural preservation are often contradictory to each other, we propose a concept of virtual optimal planes to relieve this conflict. To this end, we model this problem as a process of estimating homography decomposition coefficients, and design an iterative coefficient predictor and minimal semantic distortion constraint to identify the optimal plane. This scheme is finally incorporated into RopStitch by warping both views onto the optimal plane bidirectionally. Extensive experiments across various datasets demonstrate that RopStitch significantly outperforms existing methods, particularly in scene robustness and content naturalness. Lang Nie, Kang Liao, Yunqiu Xu, Chunyu Lin, Bin Xiao 0002 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | CustomTTT: Motion and Appearance Customized Video Generation via Test-Time TrainingabstractBenefiting from large-scale pre-training of text-video pairs, current text-to-video (T2V) diffusion models can generate high-quality videos from the text description. Besides, given some reference images or videos, the parameter-efficient fine-tuning method, i.e. LoRA, can generate high-quality customized concepts, e.g., the specific subject or the motions from a reference video. However, combining the trained multiple concepts from different references into a single network shows obvious artifacts. To this end, we propose CustomTTT, where we can joint custom the appearance and the motion of the given video easily. In detail, we first analyze the prompt influence in the current video diffusion model and find the LoRAs are only needed for the specific layers for appearance and motion customization. Besides, since each LoRA is trained individually, we propose a novel test-time training technique to update parameters after combination utilizing the trained customized models. We conduct detailed experiments to verify the effectiveness of the proposed methods. Our method outperforms several state-of-the-art works in both qualitative and quantitative evaluations. Xiuli Bi, Bo Liu 0047, Xiaodong Cun, Yong Zhang 0034, Weisheng Li 0001, Bin Xiao 0002 |
AAAI | 7 |
| 2025 | Power of Diversity: Enhancing Data-Free Black-Box Attack with Domain-Augmented LearningabstractSubstitute training-based data-free black-box attacks pose a significant threat to enterprise-deployed models. These attacks use a generator to synthesize data and query APIs, then train a substitute model to approximate the target model's decision boundary based on the returned results. However, existing attack methods often struggle to produce sufficiently diverse data, particularly for complex target models and extensive target data domains, severely limiting their practical application. To address this gap, we design domain-augmented learning to improve the quality of the synthetic data domain (SDD) generated by the generator from two perspectives. Specifically, (1) To broaden the SDD's coverage, we introduce textual semantic embeddings into the generator for the first time. (2) For enhancing the SDD's discretization, we propose a competitive optimization strategy that forces the generator to self-compete, along with heterogeneity excitation to overcome the constraints of information entropy on diversity. Comprehensive experiments demonstrate that our method is more effective. In non-targeted attacks on the CIFAR-10 and Tiny-ImageNet datasets, our method outperforms the state-of-the-art by 14% and 7% in attack success rate, respectively. Yang Wei 0002, Jingyu Tan, Guowen Xu, Zhuoran Ma 0002, Zhuo Ma 0001, Bin Xiao 0002 |
AAAI | 6 |
| 2025 | Towards Universal AI-Generated Image Detection by Variational Information Bottleneck NetworkabstractThe rapid advancement of generative models has significantly improved the quality of generated images. Mean-while, it challenges information authenticity and credibility. Current generated image detection methods based on large-scale pre-trained multimodal models have achieved impressive results. Although these models provide abundant features, the authentication task-related features are often submerged. Consequently, those authentication task-irrelated features cause models to learn superficial biases, thereby harming their generalization performance across different model genera (e.g., GANs and Diffusion Models). To this end, we proposed VIB-Net, which uses Variational Information Bottlenecks to enforce authentication task-related feature learning. We tested and analyzed the proposed method and existing methods on samples generated by 17 different generative models. Compared to SOTA methods, VIB-Net achieved a 5.55% improvement in mAP and a 9.33% increase in accuracy. Notably, in generalization tests on unseen generative models from different series, VIB-Net improved mAP by 12.48% and accuracy by 23.59% over SOTA methods. The code is available at https://github.com/oceanzhf/VIBAIGCDetect. Qinghui He, Xiuli Bi, Weisheng Li 0001, Bo Liu 0047, Bin Xiao 0002 |
CVPR | 6 |
| 2025 | Stacking U-Nets in U-shape: Redesigning the Information Flow in Model-based Networks for MRI ReconstructionabstractModel-based networks have shown convincing performance in MRI reconstruction. However, the unrolled cascades within the networks are constrained to solely obtain information from the preceding counterpart, resulting in potential error accumulation. Moreover, the linear structure fails to address the challenge of recovering fine-grained details. To tackle these problems, we propose to redesign the information flow in model-based networks. Our method features a large U-shaped network, where the nodes are built with unrolled cascades and U-Net-based regularizers. We design an input-level integration module to help the cascades acquire information from adjacent and skip-connected counterparts, building robust mappings to the target. We further design a coarse-to-fine feature-level integration module, aiming at guiding the network to progressively recover fine details. Intermediate reconstructions produced by subnetworks of different scales are integrated, enabling the extraction of complementary information to enhance the final performance. Compared with cutting-edge methods on different datasets, our method exhibits superior performances. Xiaoyu Qiao, Weisheng Li 0001, Bin Xiao 0002 |
ICASSP | 3 |
| 2025 | Subsampling Decomposition based k-Space Refinement for Accelerated MRI ReconstructionabstractIn accelerated MRI reconstruction problem, directly recovering all the missing k-space data from undersampled measurements is highly ill-posed and often leads to suboptimal performance. To address the problem, we propose a novel deep unfolding network (DUN) with subsampling decomposition (SD) based k-space refinement to mitigate the ill-posedness. Our method employs a parallel network architecture with a primary branch unfolded by gradient descent-inspired optimization process (GD-PB) for reconstruction. Additionally, we introduce an SD-based auxiliary branch (SD-AB) that decompose the inverse problem into moderately corrupted subproblems. We design a novel subsampling mask predictor that captures both global and local spatial correlations in k-space, ensuring the SD-AB effectively preserves the most well-reconstructed subsets as reliable region (RR). The RR in SD-AB is used to periodically refine the intermediate outputs of the GD-PB, achieving improved accuracy. Experimental results reveal that our method significantly outperforms conventional and SD-based DUN techniques, achieving superior PSNR and SSIM results compared with cutting-edge methods. Xiaoyu Qiao, Weisheng Li 0001, Bin Xiao 0002 |
ICASSP | 3 |
| 2025 | Learning Preconditioners in Gates-controlled Deep Unfolding Networks based on Quasi-Newton Methods For Accelerated MRI ReconstructionabstractDeep unfolding networks (DUNs) have made significant progress in MRI reconstruction, successfully tackling the problem of prolonged imaging time. However, the ill-conditioned nature of MRI reconstruction often causes slow convergence in iterative optimization, potentially compromising the performance of DUNs. In this study we propose a preconditioned and gates-controlled DUN (PGDUN) to address these challenges. Our approach starts with optimizing the step size of proximal gradient descent (PGD) through a preconditioner. To improve flexibility and adaptability, we relax the constrains on quasi-newton-based optimization procedure. We design ConvLSTM-based modules, where the gate units automatically preserve necessary long- and short-term information, facilitating the learning of optimized variables and their combinations. Furthermore, we design gate units to modulate the features fed to regularizers across different iterations, boosting their robustness against potential accumulated errors. Evaluations using PSNR and SSIM metrics reveal that our approach outperforms existing state-of-the-art methods, achieving superior reconstruction results across various sequences. Xiaoyu Qiao, Weisheng Li 0001, Bin Xiao 0002 |
ICASSP | 3 |
| 2025 | Covert and Potent: A Weather-Camouflaged Backdoor Attacks on Self-Supervised LearningabstractSelf-supervised learning is widely applied across various domains due to its advantage of learning data representations without the need for labels. However, recent research shows that backdoor attacks on self-supervised learning are achievable by coupling benign features with trigger features without manipulating labels. Existing methods, however, suffer from poor trigger disguise. When designing triggers, more emphasis is placed on attack strength rather than on disguising the triggers, which makes these triggers easily detectable through manual inspection or preprocessing methods. Therefore, we propose a camouflaged self-supervised backdoor attack method from the perspective of visual disguise. Specifically, we design triggers by embedding variable adverse weather information to achieve visual camouflage, which can bypass certain defence methods to some extent. Additionally, since our proposed camouflaged triggers have a global nature, they achieve more efficient backdoor attack capabilities. Experiments demonstrate that our method achieves attack success rates of 83.4% on the CIFAR-100 dataset and 44.8% on the ImageNet-100 dataset, surpassing existing state-of-the-art methods by 14.6% and 24.4%, respectively. At the same time, our method exhibits better stealthiness. Yang Wei 0002, Yonghao Yang, Bo Liu 0047, Bin Xiao 0002 |
ICASSP | 4 |
| 2025 | Who Controls the Authorization? Invertible Networks for Copyright Protection in Text-to-Image Synthesis
Baoyue Hu, Yang Wei 0002, Wendong Huang, Xiuli Bi, Bin Xiao 0002 |
ICCV | 6 |
| 2025 | Breaking Grid Constraints: Dynamic Graph Reconstruction Network for Multi-Organ Segmentation
Yang Wei 0002, Xiuli Bi, Bin Xiao 0002 |
ICCV | 6 |
| 2025 | Test-Time Learning for Large Language ModelsabstractWhile Large Language Models (LLMs) have exhibited remarkable emergent capabilities through extensive pre-training, they still face critical limitations in generalizing to specialized domains and handling diverse linguistic variations, known as distribution shifts. In this paper, we propose a Test-Time Learning (TTL) paradigm for LLMs, namely TLM, which dynamically adapts LLMs to target domains using only unlabeled test data during testing. Specifically, we first provide empirical evidence and theoretical insights to reveal that more accurate predictions from LLMs can be achieved by minimizing the input perplexity of the unlabeled test data. Based on this insight, we formulate the Test-Time Learning process of LLMs as input perplexity minimization, enabling self-supervised enhancement of LLM performance. Furthermore, we observe that high-perplexity samples tend to be more informative for model optimization. Accordingly, we introduce a Sample Efficient Learning Strategy that actively selects and emphasizes these high-perplexity samples for test-time updates. Lastly, to mitigate catastrophic forgetting and ensure adaptation stability, we adopt Low-Rank Adaptation (LoRA) instead of full-parameter optimization, which allows lightweight model updates while preserving more original knowledge from the model. We introduce the AdaptEval benchmark for TTL and demonstrate through experiments that TLM improves performance by at least 20% compared to original LLMs on domain knowledge adaptation. Jinwu Hu, Zitian Zhang, Xutao Wen, Chao Shuai, Wei Luo 0006, Bin Xiao 0002, Yuanqing Li 0001, Mingkui Tan |
ICML | 7 |
| 2025 | DGMIR: Dual-Guided Multimodal Medical Image Registration Based on Multi-view Augmentation and On-Site Modality Removal
Gao Le, Yucheng Shu, Lihong Qiao, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
MICCAI (1) | 5 |
| 2025 | Pathology-Aware Reconstruction with Discriminative Knowledge Boosting Alignment for Che-Xray Vision-Language Pre-trainingabstractCurrent medical vision-language pre-training models primarily follow two paradigms: report-supervised cross-modal alignment pre-training and reconstruction-based self-supervised pre-training. The former enhances the discriminative power of representations, while the latter facilitates fine-grained representation learning. However, naively combining these two paradigms inherits their inherent limitations: reconstruction-based methods treat all image patches equally during reconstruction, failing to effectively capture critical pathological details-since disease-related regions typically occupy only a small fraction of the image. Meanwhile, alignment-based methods suffer from suboptimal representations due to the presence of false negatives. To address these challenges, we propose a novel pre-training framework that integrates two key components: Pathology-Aware Reconstruction (PAR) and Discriminative Knowledge-Boosted Alignment (DKBA). Through a cascaded training strategy, our framework effectively combines the strengths of both paradigms while mitigating their inherent limitations. During the reconstruction pre-training stage, PAR incorporates pathology-aware priors to enhance the model's ability to capture fine-grained pathological details. In the alignment pre-training stage, DKBA leverages a medical knowledge graph as external supervision to improve cross-modal clustering alignment, thereby reducing the negative impact of false negatives. Extensive experiments on diverse downstream medical imaging tasks including image classification, object detection, and semantic segmentation, demonstrate the superior generalization capabilities of our method. Our code is publicly available at https://github.com/Felix1118/PADKB. Lihong Qiao, Shiyi Gao, Yucheng Shu, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
ACM Multimedia | 4 |
| 2025 | The Overlooked Matters: Revisiting Background, Prototype, and Activation in Few-Shot Medical Image Segmentation
Yucheng Shu, Lihong Qiao, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
ACM Multimedia | 5 |
| 2025 | Continual Knowledge Adaptation for Reinforcement LearningabstractReinforcement Learning enables agents to learn optimal behaviors through interactions with environments. However, real-world environments are typically non-stationary, requiring agents to continuously adapt to new tasks and changing conditions. Although Continual Reinforcement Learning facilitates learning across multiple tasks, existing methods often suffer from catastrophic forgetting and inefficient knowledge utilization. To address these challenges, we propose Continual Knowledge Adaptation for Reinforcement Learning (CKA-RL), which enables the accumulation and effective utilization of historical knowledge. Specifically, we introduce a Continual Knowledge Adaptation strategy, which involves maintaining a task-specific knowledge vector pool and dynamically using historical knowledge to adapt the agent to new tasks. This process mitigates catastrophic forgetting and enables efficient knowledge transfer across tasks by preserving and adapting critical model parameters. Additionally, we propose an Adaptive Knowledge Merging mechanism that combines similar knowledge vectors to address scalability challenges, reducing memory requirements while ensuring the retention of essential knowledge. Experiments on three benchmarks demonstrate that the proposed CKA-RL outperforms state-of-the-art methods, achieving an improvement of 4.20% in overall performance and 8.02% in forward transfer. The source code is available at https://github.com/Fhujinwu/CKA-RL. Jinwu Hu, Zihao Lian, Zhiquan Wen, Xutao Wen, Bin Xiao 0002, Mingkui Tan |
NeurIPS | 7 |
| 2025 | Rethinking the CNN and transformer for deformable image registration
Weisheng Li 0001, Yucheng Shu, Jian-Xun Mi, Guofen Wang, Bin Xiao 0002 |
Expert Syst. Appl. | 7 |
| 2025 | Transfer morphological features for segmentation with few labels on fluorescent mitochondria imagesabstractAbstract Automated segmentation of mitochondria is crucial for statistical analysis in biological research. Existing segmentation techniques often face challenges with fluorescence images. Handcrafted methods have poor segmentation results while deep learning‐based methods lack the labeled mitochondrial data. However, although the number of labeled mitochondrial images is limited, the unlabeled fluorescent data is easy to obtain. The authors aim to leverage a large amount of unlabeled data to learn mitochondrial morphological features. The approach begins with self‐supervised learning from a vast set of unlabeled images through masked image modeling. This technique involves presenting images with randomly masked patches, prompting the model to predict the content of these masked areas. By doing so, the model learns the distinctive features of mitochondria. In the subsequent phase, the trained encoder is transferred to the segmentation task, replacing the original reconstruction decoder with the Segformer segmentation decoder. The model is then fine‐tuned using a small labeled dataset. By reconstructing mitochondria in the masked regions, the model learns features more effectively on unlabeled samples, and improves segmentation performance even with limited labeled data. Empirical results validate the effectiveness of the approach, showing an 11.8% improvement in Intersection over Union metrics compared to existing fluorescence mitochondrial segmentation techniques. Junchao Fan, Xiuli Bi, Weisheng Li 0001, Bin Xiao 0002, Xiaoshuai Huang |
IET Image Process. | 5 |
| 2025 | CS-CoLBP: Cross-Scale Co-occurrence Local Binary Pattern for Image Classification
Bin Xiao 0002, Danyu Shi, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | Cardiac cavity segmentation review in the past decade: Methods and future perspectives
Weisheng Li 0001, Yucheng Shu, Yidong Peng, Bin Xiao 0002 |
Neurocomputing | 5 |
| 2025 | A multi-granularity facial aesthetic evaluation model based on image-text modality
Yong Wang 0009, Weisheng Li 0001, Bin Xiao 0002 |
Knowl. Based Syst. | 4 |
| 2025 | Improving the sparse coding model via hybrid Gaussian priors
Jian-Xun Mi, Weisheng Li 0001, Guofen Wang, Bin Xiao 0002 |
Pattern Recognit. | 5 |
| 2025 | Enhancing EEG-Based Cross-Subject Emotion Recognition via Adaptive Source Joint Domain AdaptationabstractEEG emotion recognition is crucial in both human-machine interaction and healthcare. However, recognizing emotions across different subjects remains challenging due to individual variability. While existing multi-source domain adaptation methods have been utilized for cross-subject EEG emotion decoding, they often struggle with irrelevant or weakly relevant source domains, leading to negative transfer. Additionally, variations within subdomains are often neglected in these studies. We propose a joint domain adaptation method, Adaptive Source Joint Domain Adaptation (ASJDA) to address these issues. ASJDA utilizes an unsupervised adaptive source selection strategy to select a subset of source domains by evaluating the Jensen-Shannon divergence between the source and target domains, choosing those most relevant to the target. Subsequently, it implements joint domain adaptation with these chosen sources at both the domain and category subdomain levels. Our proposed method outperforms existing state-of-the-art methods, achieving cross-subject accuracies of 96.81% in SEED, 89.69% in SEED-IV, and 69.31% in DEAP. This work significantly advances the state of the art in EEG emotion recognition by effectively addressing the challenges of cross-subject variability. Ke Liu 0008, Wenrui Zhu, Zhu Liang Yu, Hong Yu 0007, Bin Xiao 0002, Wei Wu 0022 |
IEEE Trans. Affect. Comput. | 6 |
| 2025 | Neurocognitive Insights: Cognitive Comprehension Attention in Multi-Organ SegmentationabstractIn multi-organ segmentation, attention mechanisms are frequently employed to enhance the focus on irregular organs, improving performance. However, current attention mechanisms exhibit notable limitations. On the one hand, their visual saliency-based attention bias results in incomplete region-of-interest coverage. On the other hand, their organ-specific cognitive deficiency exacerbates organ misclassification. Inspired by neurocognitive science, this paper proposes a Cognitive Comprehension Attention (CCA). Diverging from existing methods, CCA achieves refined attention allocation by decomposing visual representations into discrete visual stimuli. This fine-grained approach enables unbiased processing for each visual stimulus, preventing critical information omission and ensuring comprehensive organ region coverage. More importantly, CCA generates organ-specific attention representations by establishing distinct attention patterns across different organ regions, which empowers CCA with cognitive capacity, resolving organ misclassification. Extensive experiments across multiple datasets demonstrate that CCA significantly enhances backbone performance, achieving a max mDice improvement of 8.45% while surpassing state-of-the-art methods by 9% in Recall and 11.78% in Precision. Code is available at:https://github.com/robert1818118/CCA. Yang Wei 0002, Wendong Huang, Xiuli Bi, Xuezong Yang, Bin Xiao 0002 |
IEEE Trans. Big Data | 8 |
| 2025 | Let Images Speak More: An Efficient Method for Detecting Image Manipulation HistoryabstractDigital image forensics aims to verify the authenticity of digital images, which has emerged as a prominent research area. To reveal the manipulation history of an image, the existing methods can only detect specific image operations or are based on a general forensic feature with high dimensions. Moreover, these methods perform well only when the operation chain length is no greater than 2. However, their detection accuracy drops significantly for images with longer operation chains that are more representative of real-world scenarios. To break these limitations, we proposed a novel forensics frequency Feature based on Histogram and Detail Map (FHDM(79D)), which can distinguish various operation chains containing different numbers of operations. Specifically, compared to the traces left by image manipulation in the spatial domain, we have discovered that they are more distinct in the frequency domain. This observation has prompted us to extract features from the frequency domain of images by analyzing their histograms and detail maps to capture the manipulation traces of the images. Notably, the proposed feature extracted in the frequency domain has almost 90% fewer dimensions than the commonly used general forensic features, such as SRM(714D), which greatly reduces the computational complexity. Meanwhile, compared to deep learning-based methods, the experiments show that the proposed method achieves a detection accuracy of over 95% for image operations across multiple datasets, while other deep learning-based methods do not exceed 90% accuracy. Extensive experimental results show that the proposed method is more versatile and effective, showing good performance in complex operation chain detection and local forgery detection. The code is available at https://github.com/CherishL-J/Op-detection. Yang Wei 0002, Xiaochen Yuan, Xiuli Bi, Bin Xiao 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Jointly RS Image Deblurring and Super-Resolution With Adjustable-Kernel and Multi-Domain AttentionabstractRemote sensing (RS) image deblurring and super-resolution (SR) are common tasks in computer vision that aim at restoring RS image detail and spatial scale, respectively. However, real-world RS images often suffer from a complex combination of global low-resolution (LR) degeneration and local blurring degeneration. Although carefully designed deblurring and SR models perform well on these two tasks individually, a unified model that performs jointly RS image deblurring and SR (JRSIDSR) task is still challenging due to the vital dilemma of reconstructing the global and local degeneration simultaneously. In addition, existing methods struggle to capture the interrelationship between deblurring and SR processes, leading to suboptimal results. To tackle these issues, we give a unified theoretical analysis of RS images’ spatial and blur degeneration processes and propose a dual-branch parallel network named adjustable-kernel and multi-domain network (AKMD-Net) for the JRSIDSR task. AKMD-Net consists of two main branches: deblurring and SR branches. In the deblurring branch, we design a pixel-adjustable kernel block (PAKB) to estimate the local and spatial-varying blur kernels. In the SR branch, a multi-domain attention block (MDAB) is proposed to capture the global contextual information enhanced with high-frequency details. Furthermore, we develop an adaptive feature fusion (AFF) module to model the contextual relationships between the deblurring and SR branches. Finally, we design an adaptive Wiener loss (AW Loss) to depress the prior noise in the reconstructed images. Extensive experiments demonstrate that the proposed AKMD-Net achieves state-of-the-art (SOTA) quantitative and qualitative performance on commonly used RS image datasets. The source code is publicly available at:https://github.com/zpc456/AKMD-Net. Yan Zhang 0108, Chengxiao Zeng, Bin Xiao 0002, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | HEOI: Human Attention Prediction in Natural Daily Life With Fine-Grained Human-Environment-Object Interaction ModelabstractThis paper handles the problem of human attention prediction in natural daily life from the third-person view. Due to the significance of this topic in various applications, researchers in the computer vision community have proposed many excellent models in the past few decades, and many models have begun to focus on natural daily life scenarios in recent years. However, existing mainstream models usually ignore a basic fact that human attention is a typical interdisciplinary concept. Specifically, the mainstream definition is direction-level or pixel-level, while many interdisciplinary studies argue the object-level definition. Additionally, the mainstream model structure converges to the dual-pathway architecture or its variants, while the majority of interdisciplinary studies claim attention is involved in the human-environment interaction procedure. Grounded on solid theories and studies in interdisciplinary fields including computer vision, cognition, neuroscience, psychology, and philosophy, this paper proposes a fine-grained Human-Environment-Object Interaction (HEOI) model, which for the first time integrates multi-granularity human cues to predict human attention. Our model is explainable and lightweight, and validated to be effective by a wide range of comparison, ablation, and visualization experiments on two public datasets. Zhixiong Nan, Leiyu Jia, Bin Xiao 0002 |
IEEE Trans. Image Process. | 3 |
| 2025 | Contrastive Learning Guided Fusion Network for Brain CT and MRIabstractMedical image fusion technology provides professionals with more detailed and precise diagnostic information. This paper introduces a new efficient CT and MRI fusion network, CLGFusion, based on a contrastive learning-guided network. CLGFusion includes two encoding branches at the feature encoding stage, enabling them to interact and learn from each other. The approach begins with training a single-view encoder to predict the feature representation of an image from varied augmented views. Simultaneously, the multi-view encoder is improved using the exponential moving average of the single-view encoder. Contrastive learning is integrated into medical image fusion by creating a feature contrast space without constructing negative samples. This feature contrast space cleverly uses the information of the difference in the feature product of the source image and its corresponding augmented image. It continuously guides the network to constantly optimize its fusion effect by combining the method of structural similarity loss, to achieve more accurate and efficient image fusion. This approach represents an end-to-end unsupervised fusion model. Experimental validation shows that our proposed method demonstrates performance comparable to state-of-the-art techniques in both subjective evaluation and objective metrics. Weisheng Li 0001, Bin Xiao 0002, Guofen Wang, Dan He 0010, Xiaoyu Qiao |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | ADMM-ESINet: A Deep Unrolling Network for EEG Extended Source ImagingabstractElectroencephalography (EEG) source imaging (ESI) methods aim to reconstruct cortical sources from scalp EEG signals, a crucial task for understanding the normal brain as well as brain disorders. Traditional model-driven ESI methods face challenges in real-time reconstruction, while deep neural network (DNN)-based ESI methods often struggle with generalization to new data. To address these issues, we propose ADMM-ESINet, a novel deep unfolding neural network for robust and efficient reconstruction of EEG extended sources. ADMM-ESINet leverages a structured sparsity constraint within a regularization framework and employs the Alternating Direction Method of Multipliers (ADMM) to achieve iterative solutions. By unrolling the ADMM algorithm into a cascaded network architecture, ADMM-ESINet effectively integrates prior knowledge, enabling end-to-end, real-time ESI. Crucially, both the regularization parameters and the spatial transform operator are learned directly from the training data. Numerical results demonstrate that ADMM-ESINet surpasses traditional DNN-based methods in generalization ability and accurately reconstructs the location, extent, and temporal dynamics of extended sources, establishing ADMM-ESINet as a promising method for real-time ESI. Ke Liu 0008, Jun Zhang 0026, Zhenghui Gu, Zhu Liang Yu, Yu Zhang 0009, Bin Xiao 0002, Wei Wu 0022 |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | DMSACNN: Deep Multiscale Attentional Convolutional Neural Network for EEG-Based Motor DecodingabstractOBJECTIVE: Accurate decoding of electroencephalogram (EEG) signals has become more significant for the brain-computer interface (BCI). Specifically, motor imagery and motor execution (MI/ME) tasks enable the control of external devices by decoding EEG signals during imagined or real movements. However, accurately decoding MI/ME signals remains a challenge due to the limited utilization of temporal information and ineffective feature selection methods. METHODS: This paper introduces DMSACNN, an end-to-end deep multiscale attention convolutional neural network for MI/ME-EEG decoding. DMSACNN incorporates a deep multiscale temporal feature extraction module to capture temporal features at various levels. These features are then processed by a spatial convolutional module to extract spatial features. Finally, a local and global feature fusion attention module is utilized to combine local and global information and extract the most discriminative spatiotemporal features. MAIN RESULTS: DMSACNN achieves impressive accuracies of 78.20%, 96.34% and 70.90% for hold-out analysis on the BCI-IV-2a, High Gamma and OpenBMI datasets, respectively, outperforming most of the state-of-the-art methods. CONCLUSION AND SIGNIFICANCE: These results highlight the potential of DMSACNN in robust BCI applications. Our proposed method provides a valuable solution to improve the accuracy of the MI/ME-EEG decoding, which can pave the way for more efficient and reliable BCI systems. Ke Liu 0008, Zhu Liang Yu, Bin Xiao 0002, Guoyin Wang 0001, Wei Wu 0022 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | CryptIF: Toward Cloud-Based IoT Anomaly Detection Over Encrypted Feature Streams
Teng Li 0003, Zejian Lin, Yebo Feng, Chong Wang 0013, Zhuo Ma 0001, Bin Xiao 0002, Jianfeng Ma 0001, Yang Liu 0003 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2025 | Prototype-Guided Graph Reasoning Network for Few-Shot Medical Image SegmentationabstractFew-shot semantic segmentation (FSS) is of tremendous potential for data-scarce scenarios, particularly in medical segmentation tasks with merely a few labeled data. Most of the existing FSS methods typically distinguish query objects with the guidance of support prototypes. However, the variances in appearance and scale between support and query objects from the same anatomical class are often exceedingly considerable in practical clinical scenarios, thus resulting in undesirable query segmentation masks. To tackle the aforementioned challenge, we propose a novel prototype-guided graph reasoning network (PGRNet) to explicitly explore potential contextual relationships in structured query images. Specifically, a prototype-guided graph reasoning module is proposed to perform information interaction on the query graph under the guidance of support prototypes to fully exploit the structural properties of query images to overcome intra-class variances. Moreover, instead of fixed support prototypes, a dynamic prototype generation mechanism is devised to yield a collection of dynamic support prototypes by mining rich contextual information from support images to further boost the efficiency of information interaction between support and query branches. Equipped with the proposed two components, PGRNet can learn abundant contextual representations for query images and is therefore more resilient to object variations. We validate our method on three publicly available medical segmentation datasets, namely CHAOS-T2, MS-CMRSeg, and Synapse. Experiments indicate that the proposed PGRNet outperforms previous FSS methods by a considerable margin and establishes a new state-of-the-art performance. Wendong Huang, Jinwu Hu, Yang Wei 0002, Xiuli Bi, Bin Xiao 0002 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | MDFA: A Quantitative Framework for the Analysis of Multimodal Facial EstheticsabstractIn the era of big data, the problem of facial beauty prediction (FBP) has been addressed using a combination of deep learning and esthetics based on data and models. Most existing methods are based on 2-D unimodal information processing. Owing to the high cost of 3-D data acquisition equipment, studies on the use of multimodal features of 2-D and 3-D for esthetic evaluation are scarce. Moreover, most existing methods are based on self-built 3-D datasets, which are limited to practical application scenarios of 2-D facial images. This study proposed a label distribution-based multimodal facial esthetic analysis framework (LDMFE). The LDMFE performed facial esthetic evaluation by combining 2-D and 3-D information following the process used by the human brain to conduct the 3-D esthetic evaluation. FBP was performed by extracting facial depth structure information using a depth information extraction network, DIENet, which comprises a facial structure perception layer (FSP-Layer) and an attention decision block (AD-Block). Furthermore, to ensure a high degree of agreement between the predicted label distribution of the network and the true distribution, a simple and efficient distribution measurement loss function called ${\mathcal {L}}_{\text {WD}}$ was proposed. Compared with the label distribution-based FBP loss and the latest FBP loss, ${\mathcal {L}}_{\text {WD}}$ was more stable and effective. The performance of LDMFE was evaluated using three datasets. The experimental results demonstrate that the LDMFE exhibits state-of-the-art performance. Weisheng Li 0001, Bin Xiao 0002, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Subgraph Propagation and Contrastive Calibration for Incomplete Multiview Data ClusteringabstractThe success of multiview raw data mining relies on the integrity of attributes. However, each view faces various noises and collection failures, which leads to a condition that attributes are only partially available. To make matters worse, the attributes in multiview raw data are composed of multiple forms, which makes it more difficult to explore the structure of the data especially in multiview clustering task. Due to the missing data in some views, the clustering task on incomplete multiview data confronts the following challenges, namely: 1) mining the topology of missing data in multiview is an urgent problem to be solved; 2) most approaches do not calibrate the complemented representations with common information of multiple views; and 3) we discover that the cluster distributions obtained from incomplete views have a cluster distribution unaligned problem (CDUP) in the latent space. To solve the above issues, we propose a deep clustering framework based on subgraph propagation and contrastive calibration (SPCC) for incomplete multiview raw data. First, the global structural graph is reconstructed by propagating the subgraphs generated by the complete data of each view. Then, the missing views are completed and calibrated under the guidance of the global structural graph and contrast learning between views. In the latent space, we assume that different views have a common cluster representation in the same dimension. However, in the unsupervised condition, the fact that the cluster distributions of different views do not correspond affects the information completion process to use information from other views. Finally, the complemented cluster distributions for different views are aligned by contrastive learning (CL), thus solving the CDUP in the latent space. Our method achieves advanced performance on six benchmarks, which validates the effectiveness and superiority of our SPCC. Zhibin Dong, Jiaqi Jin, Yuyang Xiao, Bin Xiao 0002, Siwei Wang 0001, Xinwang Liu 0002, En Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | SAMCL: Subgraph-Aligned Multiview Contrastive Learning for Graph Anomaly DetectionabstractGraph anomaly detection (GAD) has gained increasing attention in various attribute graph applications, i.e., social communication and financial fraud transaction networks. Recently, graph contrastive learning (GCL)-based methods have been widely adopted as the mainstream for GAD with remarkable success. However, existing GCL strategies in GAD mainly focus on node-node and node-subgraph contrast and fail to explore subgraph-subgraph level comparison. Furthermore, the different sizes or component node indices of the sampled subgraph pairs may cause the "nonaligned" issue, making it difficult to accurately measure the similarity of subgraph pairs. In this article, we propose a novel subgraph-aligned multiview contrastive approach for graph anomaly detection, named SAMCL, which fills the subgraph-subgraph contrastive-level blank for GAD tasks. Specifically, we first generate the multiview augmented subgraphs by capturing different neighbors of target nodes forming contrasting subgraph pairs. Then, to fulfill the nonaligned subgraph pair contrast, we propose a subgraph-aligned strategy that estimates similarities with the Earth mover's distance (EMD) of both considering the node embedding distributions and typology awareness. With the newly established similarity measure for subgraphs, we conduct the interview subgraph-aligned contrastive learning module to better detect changes for nodes with different local subgraphs. Moreover, we conduct intraview node-subgraph contrastive learning to supplement richer information on abnormalities. Finally, we also employ the node reconstruction task for the masked subgraph to measure the local change of the target node. Finally, the anomaly score for each node is jointly calculated by these three modules. Extensive experiments conducted on benchmark datasets verify the effectiveness of our approach compared to existing state-of-the-art (SOTA) methods with significant performance gains (up to 6.36% improvement on ACM). Our code can be verified at https://github.com/hujingtao/SAMCL. Jingtao Hu, Bin Xiao 0002, Hu Jin 0005, Jingcan Duan, Siwei Wang 0001, Zhao Lv, Siqi Wang 0001, Xinwang Liu 0002, En Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | SARF: Aliasing Relation-Assisted Self-Supervised Learning for Few-Shot Relation ReasoningabstractFew-shot relation reasoning on knowledge graphs (FS-KGR) is an important and practical problem that aims to infer long-tail relations and has drawn increasing attention these years. Among all the proposed methods, self-supervised learning (SSL) methods, which effectively extract the hidden essential inductive patterns relying only on the support sets, have achieved promising performance. However, the existing SSL methods simply cut down connections between high-frequency and long-tail relations, which ignores the fact, i.e., the two kinds of information could be highly related to each other. Specifically, we observe that relations with similar contextual meanings, called aliasing relations (ARs), may have similar attributes. In other words, the ARs of the target long-tail relation could be in high-frequency, and leveraging such attributes can largely improve the reasoning performance. Based on the interesting observation above, we proposed a novel Self-supervised learning model by leveraging Aliasing Relations to assist FS-KGR, termed SARF. Specifically, we propose a graph neural network (GNN)-based AR-assist module to encode the ARs. Besides, we further provide two fusion strategies, i.e., simple summation and learnable fusion, to fuse the generated representations, which contain extra abundant information underlying the ARs, into the self-supervised reasoning backbone for performance enhancement. Extensive experiments on three few-shot benchmarks demonstrate that SARF achieves state-of-the-art (SOTA) performance compared with other methods in most cases. Lingyuan Meng, Ke Liang 0006, Bin Xiao 0002, Sihang Zhou 0001, Yue Liu 0008, Meng Liu 0014, Xihong Yang, Xinwang Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Revisiting Initializing Then Refining: An Incomplete and Missing Graph Imputation NetworkabstractWith the development of various applications, such as recommendation systems and social network analysis, graph data have been ubiquitous in the real world. However, graphs usually suffer from being absent during data collection due to copyright restrictions or privacy-protecting policies. The graph absence could be roughly grouped into attribute-incomplete and attribute-missing cases. Specifically, attribute-incomplete indicates that a portion of the attribute vectors of all nodes are incomplete, while attribute-missing indicates that all attribute vectors of partial nodes are missing. Although various graph imputation methods have been proposed, none of them is custom-designed for a common situation where both types of graph absence exist simultaneously. To fill this gap, we develop a novel graph imputation network termed revisiting initializing then refining (RITR), where both attribute-incomplete and attribute-missing samples are completed under the guidance of a novel initializing-then-refining imputation criterion. Specifically, to complete attribute-incomplete samples, we first initialize the incomplete attributes using Gaussian noise before network learning, and then introduce a structure-attribute consistency constraint to refine incomplete values by approximating a structure-attribute correlation matrix to a high-order structure matrix. To complete attribute-missing samples, we first adopt structure embeddings of attribute-missing samples as the embedding initialization, and then refine these initial values by adaptively aggregating the reliable information of attribute-incomplete samples according to a dynamic affinity structure. To the best of our knowledge, this newly designed method is the first end-to-end unsupervised framework dedicated to handling hybrid-absent graphs. Extensive experiments on six datasets have verified that our methods consistently outperform the existing state-of-the-art competitors. Our source code is available at https://github.com/WxTu/RITR. Wenxuan Tu, Bin Xiao 0002, Xinwang Liu 0002, Sihang Zhou 0001, Zhiping Cai, Jieren Cheng |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Boosting Pseudo-Labeling With Curriculum Self-Reflection for Attributed Graph ClusteringabstractAttributed graph clustering is an unsupervised learning task that aims to partition various nodes of a graph into distinct groups. Existing approaches focus on devising diverse pretext tasks to obtain suitable supervised information for representation learning, among which the predictive methods show great potential. However, these methods 1) generate auxiliary task bias toward the clustering target and 2) introduce label noise due to static thresholds. To address this issue, we propose a new self-supervised learning method, namely, pseudo-labeling with curriculum self-reflection (PLCSR), that learns reliable pseudo-labels by mining its information to achieve progressive processing of nodes in a self-reflection manner. First, a self-auxiliary encoder is constructed using the exponential moving average (EMA) of the original encoder's parameters to replace the auxiliary tasks, which provides an additional perspective of finding highly confident pseudo-labels. Second, a curriculum selection strategy using dynamic thresholds is designed to take full advantage of graph nodes more accurately. Besides simple nodes with high confidence at the initial stage, nodes that yield consistent predictions from both encoders are then assigned pseudo-labels to avoid the under-learning problem. For the rest difficult nodes that are highly uncertain, we abstain from making judgments to minimize their adverse impact on the model. Extensive experiments have shown that PLCSR significantly outperforms the state-of-the-art predictive method CDRS, achieving more than 6% improvements in terms of clustering accuracy. The code is available at: https://github.com/Jillian555/PLCSR. Pengfei Zhu 0001, Yu Wang 0106, Bin Xiao 0002, Jinglin Zhang 0001, Wanyu Lin, Qinghua Hu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Focus Stacking with High Fidelity and Superior Visual EffectsabstractFocus stacking is a technique in computational photography, and it synthesizes a single all-in-focus image from different focal plane images. It is difficult for previous works to produce a high-quality all-in-focus image that meets two goals: high-fidelity to its source images and good visual effects without defects or abnormalities. This paper proposes a novel method based on optical imaging process analysis and modeling. Based on a foreground segmentation - diffusion elimination architecture, the foreground segmentation makes most of the areas in full-focus images heritage information from the source images to achieve high fidelity; diffusion elimination models the physical imaging process and is specially used to solve the transition region (TR) problem that is a long-term neglected issue and degrades visual effects of synthesized images. Based on extensive experiments on simulated dataset, existing realistic dataset and our proposed BetaFusion dataset, the results show that our proposed method can generate high-quality all-in-focus images by achieving two goals simultaneously, especially can successfully solve the TR problem and eliminate the visual effect degradation of synthesized images caused by the TR problem. Bo Liu 0047, Xiuli Bi, Weisheng Li 0001, Bin Xiao 0002 |
AAAI | 5 |
| 2024 | Learning from Inside: Self-driven Intra-modality Siamese Knowledge Generation and Inter-modality Alignment for Chest X-rays Vision-Language Pre-trainingabstractSince pathology occupies only a small portion of an X-ray, which means that a large portion of the information may be irrelevant to the paired radiology report, the Chest X-rays Report Understanding (CRU) task focuses on how to utilize small regions of the case to improve the performance of medical VLP. However, existing studies have neglected the fine-grained false negative samples of medical visual representations, resulting in their poor performance in CRU scenarios, which we attribute this to the fine-grained feature collapse problem. To address this issue, we propose an intra-modality siamese knowledge generation and inter-modality alignment framework, termed Chest X-rays Report Understanding Framework(CRUF). CRUF leverages the siamese knowledge in image-text pairs as guiding signals to distinguish fine-grained false negative and negative samples within the modality, and further narrows the distance between false negative and positive samples between modalities, accurately aligning the case regions of each image with the corresponding medical terms. Experimental results on multiple downstream medical image datasets covering tasks such as image classification, object detection, and semantic segmentation demonstrate the stability and outstanding performance of our framework. Code is available at https://github.com/cl-red/CRUF. Lihong Qiao, Yucheng Shu, Xiao Luan, Bin Xiao 0002 |
BIBM | 5 |
| 2024 | Using My Artistic Style? You Must Obtain My Authorization
Xiuli Bi, Weisheng Li 0001, Bo Liu 0047, Bin Xiao 0002 |
ECCV (86) | 5 |
| 2024 | Facial Aesthetic Enhancement Network for Asian Faces Based on Differential Facial Aesthetic ActivationsabstractIn this paper, we addressed facial aesthetic enhancement (FAE). Although existing methods have made great progress, the beautified images generated by them are highly prone to poor beautification, which limits their application to real-world scenes. To tackle this problem, we proposed a new method called the facial aesthetic enhancement network for Asian faces based on differential facial aesthetic activations (Diff-FANet), which comprises three important modules: aesthetic average difference perception block (ADP), aesthetic difference evaluation block (ADE), and aesthetic fusion optimization block (AFO). ADP learns the transformation of the latent code of an image before and after beautification. The ADE learns the features of an enhanced image, which guides image fusion. The AFO was used to eliminate ghosting. To evaluate the effectiveness of Diff-FANet, we utilized the wedding dataset for training and the SCUT-FBP5500 and Asian face datasets for testing. The results of experiments revealed that Diff-FANet achieved excellent results. Weisheng Li 0001, Xinbo Gao 0001, Bin Xiao 0002 |
ICASSP | 4 |
| 2024 | Focal-Guided Multi-Consistency for Unsupervised Partial-to-Partial Point Cloud RegistrationabstractPoint Cloud Registration (PCR) is fundamental for the automatic perception of our space. With the rapid development of deep neural network, the community has swiftly adapted to this data-driven technique, and achieved promising performances. However, most existing learning-based methods attempt to conduct PCR within specific ideal experimental settings, in which the ground truth transformations are accessible and most of the data points have one-to-one correspondences. But in real-world scenarios, the GT transformations are often unknown, and point clouds may only share partially overlapped regions. It leads us to a challenging yet practical issue: How to perform Partial-to-Partial (PtP) Point Cloud Registration without pre-acquired supervisions? In this paper, we aim to tackle both challenges under a unified framework. To achieve this, we propose a novel Focal Anchor Generator to emulate the human perceptual process, particularly focusing on the mutual cloud parts. On top of it, a set of Multi-Consistency constraints are introduced to equip our model with the unsupervised learning ability, which is highly applicable. Extensive experiments have demonstrated the distinctive quality of our proposed framework. We believe this work will broaden the scope of PCR research and enhance the applicative potential of PCR algorithms. (The project code has been released on github.com/chengxiaojin/FGMC-UPCR). Yucheng Shu, Longjin Cheng, Bin Xiao 0002, Lihong Qiao, Weisheng Li 0001, Xinbo Gao 0001 |
ICME | 3 |
| 2024 | C3T: Contrastive Consistency Cross-Network Learning for Semi-Supervised Semantic SegmentationabstractSemi-supervised image semantic segmentation, a vital but challenging task in multimedia applications, aims to accurately classify pixels with limited labeled data. Traditional approaches in this domain often grapple with the confirmation bias problem, where models, influenced by their own predictions, become prone to replicating errors. To address this critical issue, our research introduces a cross-network-crossview consistency learning framework. This novel paradigm significantly reduce the confirmation bias through diversifying the learning perspectives. Integral to our approach are two components: a pseudo-label validation and filtering mechanism, and a cross-contrastive learning module within the feature domain. These elements work in synergy to not only amplify the accuracy of the model but also its robustness against varied data scenarios. Extensive evaluations, conducted across multiple datasets, clearly demonstrate the effectiveness of our method. In comparison to existing state-of-the-art models, our approach exhibits marked improvements, especially in the challenging contexts of semisupervised image semantic segmentation. The code is available at https://github.com/Sstar2orchid/C3T. Yucheng Shu, Jiaxin Xie, Lihong Qiao, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
ICME | 4 |
| 2024 | PriFU: Capturing Task-Relevant Information Without Adversarial LearningabstractAs machine learning advances, machine learning as a service (MLaaS) in the cloud brings convenience to human lives but also privacy risks, as powerful neural networks used for generation, classification or other tasks can also become privacy snoopers. This motivates privacy preservation in the inference phase. Many approaches for preserving privacy in the inference phase introduce multi-objective functions, training models to remove specific private information from users' uploaded data. Although effective, these adversarial learning-based approaches suffer not only from convergence difficulties, but also from limited generalization beyond the specific privacy for which they are trained. To address these issues, we propose a method for privacy preservation in the inference phase by removing task-irrelevant information, which requires no knowledge of the privacy attacks nor introduction of adversarial learning. Specifically, we introduce a metric to distinguish task-irrelevant information from task-relevant information, and achieve more efficient metric estimation to remove task-irrelevant features. The experiments demonstrate the potential of our method in several tasks. Our code will be available at: https://github.com/iwhoyoung/PriFU. Xiuli Bi, Bo Liu 0047, Weisheng Li 0001, Pamela C. Cosman, Bin Xiao 0002 |
ACM Multimedia | 6 |
| 2024 | Anatomical Prior Guided Spatial Contrastive Learning for Few-Shot Medical Image Segmentation
Wendong Huang, Jinwu Hu, Xiuli Bi, Bin Xiao 0002 |
ACM Multimedia | 4 |
| 2024 | ShiftMorph: A Fast and Robust Convolutional Neural Network for 3D Deformable Medical Image Registration
Weisheng Li 0001, Yucheng Shu, Jian-Xun Mi, Bin Xiao 0002 |
ACM Multimedia | 6 |
| 2024 | CMRVAE: Contrastive margin-restrained variational auto-encoder for class-separated domain adaptation in cardiac segmentation
Lihong Qiao, Rui Wang 0173, Yucheng Shu, Bin Xiao 0002, Xidong Xu, Baobin Li, Weisheng Li 0001, Xinbo Gao 0001, Bai Ying Lei |
Knowl. Based Syst. | 4 |
| 2024 | D-Net: A dual-encoder network for image splicing forgery detection and localization
Bo Liu 0047, Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001, Guoyin Wang 0001, Xinbo Gao 0001 |
Pattern Recognit. | 4 |
| 2024 | Boosting Robust Multi-Focus Image Fusion With Frequency Mask and Hyperdimensional ComputingabstractMulti-focus image fusion (MFIF) creates an image from different source images with various sensors or optical settings as the devices can’t focus all objects at different distances. Most of the MFIF methods have several limitations in encoder enough features from the images and the result are not robust. To overcome the primary issue, we present a robust fusion algorithm based on the Frequency mask and the Hyperdimensional computing. We propose the Frequency Mask Filter (FMF) to get the narrow-band signals by encoding the frequency domain vector through the mask filter in the frequency domain. The Hyperdimensional encoder uses monogenic mapping, in which the multi-modulation features (MMF) such as the frequency, phase and amplitude are dynamically selected to obtain robust focus maps. Generated by multiscale monogenic representations of each image, the narrow-band image are mapped to hypervector encoding. Hyperdimensional encoder shows the energetic and structural information and leads to robust fusion results. Our proposed method is far superior to the existing MFIF method in terms of both objective evaluation metrics and visual effects on three publicly available datasets.Additionally, our proposed method requires only 0.88 seconds and has a parameter count of 0.13 million for multi-focus image fusion. Lihong Qiao, Shixin Wu, Bin Xiao 0002, Yucheng Shu, Xiao Luan, Sicheng Lu, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Learning Discriminative Representations From Cross-Scale Features for Camouflaged Object DetectionabstractThe key that hinders the performance improvement of current camouflaged object detection (COD) models is the lack of discriminability of features at fine granularity. We solve this problem from two complementary perspectives. Firstly, complex scenes result in the discriminative feature representations of camouflaged objects being present at different scales and semantic abstraction levels. Therefore, a mechanism is needed to increase the diversity of features to integrate more information potentially beneficial for COD. Second, appearance similarity between objects and environments will inevitably lead to similarity in features. Enhancing feature diversity alone is not enough to solve the above problems. Therefore, it is necessary to give the model semantic perception capabilities to expand the subtle discrepancies between objects and environments in feature embedding. Inspired by the first point, we propose a cross-scale interaction module (CSIM) that utilizes cross-attention between different scales to enhance the diversity of feature representations. Regarding the second point, the semantic guided feature learning (SGFL) is proposed to promote the model to expand feature discrepancies through explicit supervision. Experiments on four popular COD datasets show that our method outperforms recent SOTA methods. In addition, polyp segmentation experiments show that it is also effective for other COD-like tasks. Yongchao Wang 0004, Xiuli Bi, Bo Liu 0047, Yang Wei 0002, Weisheng Li 0001, Bin Xiao 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | CTNet: Contrastive Transformer Network for Polyp SegmentationabstractSegmenting polyps from colonoscopy images is very important in clinical practice since it provides valuable information for colorectal cancer. However, polyp segmentation remains a challenging task as polyps have camouflage properties and vary greatly in size. Although many polyp segmentation methods have been recently proposed and produced remarkable results, most of them cannot yield stable results due to the lack of features with distinguishing properties and those with high-level semantic details. Therefore, we proposed a novel polyp segmentation framework called contrastive Transformer network (CTNet), with three key components of contrastive Transformer backbone, self-multiscale interaction module (SMIM), and collection information module (CIM), which has excellent learning and generalization abilities. The long-range dependence and highly structured feature map space obtained by CTNet through contrastive Transformer can effectively localize polyps with camouflage properties. CTNet benefits from the multiscale information and high-resolution feature maps with high-level semantic obtained by SMIM and CIM, respectively, and thus can obtain accurate segmentation results for polyps of different sizes. Without bells and whistles, CTNet yields significant gains of 2.3%, 3.7%, 3.7%, 18.2%, and 10.1% over classical method PraNet on Kvasir-SEG, CVC-ClinicDB, Endoscene, ETIS-LaribPolypDB, and CVC-ColonDB respectively. In addition, CTNet has advantages in camouflaged object detection and defect detection. The code is available at https://github.com/Fhujinwu/CTNet. Bin Xiao 0002, Jinwu Hu, Weisheng Li 0001, Chi-Man Pun, Xiuli Bi |
IEEE Trans. Cybern. | 1 |
| 2024 | Effectively Improving Data Diversity of Substitute Training for Data-Free Black-Box AttackabstractRecent substitute training methods have utilized the concept of Generative Adversarial Networks (GANs) to implement data-free black-box attacks. Specifically, in designing the generators, the substitute training methods use a similar structure to the generators in GANs. However, this design approach ignores the potential situation that the generators in GANs operate under real data supervision, while the generators in substitute training methods lack such supervision. This difference in data-supervised conditions constrain the diversity of data generated by the substitute training methods, resulting in inadequate data to support effective training of the substitute model. This impacts the substitute model's ability to attack the target model further. Consequently, to solve the above issues, we propose three strategies to improve the attack success rates. For the generator, we first propose a dense projection space that projects the input noise into various latent feature spaces to diversify feature information. Then, we introduce a novel disguised natural color mode. This mode improves information exchange between the generator's output layer and previous layers, allowing for more diverse generated data. Besides, we present a regularization method for the substitute model, called noise-based balanced learning, to prevent the potential risk of overfitting due to the lack of diversity of the generated data. In the experimental analysis, extensive experiments are conducted to validate the effectiveness of these proposed strategies. Yang Wei 0002, Zhuo Ma 0001, Zhuoran Ma 0002, Zhan Qin, Yang Liu 0118, Bin Xiao 0002, Xiuli Bi, Jianfeng Ma 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2024 | Fast Continual Multi-View Clustering With Incomplete ViewsabstractMulti-view clustering (MVC) has attracted broad attention due to its capacity to exploit consistent and complementary information across views. This paper focuses on a challenging issue in MVC called the incomplete continual data problem (ICDP). Specifically, most existing algorithms assume that views are available in advance and overlook the scenarios where data observations of views are accumulated over time. Due to privacy considerations or memory limitations, previous views cannot be stored in these situations. Some works have proposed ways to handle this problem, but all of them fail to address incomplete views. Such an incomplete continual data problem (ICDP) in MVC is difficult to solve since incomplete information with continual data increases the difficulty of extracting consistent and complementary knowledge among views. We propose Fast Continual Multi-View Clustering with Incomplete Views (FCMVC-IV) to address this issue. Specifically, the method maintains a scalable consensus coefficient matrix and updates its knowledge with the incoming incomplete view rather than storing and recomputing all the data matrices. Considering that the given views are incomplete, the newly collected view might contain samples that have yet to appear; two indicator matrices and a rotation matrix are developed to match matrices with different dimensions. In addition, we design a three-step iterative algorithm to solve the resultant problem with linear complexity and proven convergence. Comprehensive experiments conducted on various datasets demonstrate the superiority of FCMVC-IV over the competing approaches. The code is publicly available at https://github.com/wanxinhang/FCMVC-IV. Xinhang Wan, Bin Xiao 0002, Xinwang Liu 0002, Jiyuan Liu 0003, Weixuan Liang, En Zhu |
IEEE Trans. Image Process. | 2 |
| 2024 | Boundary-Aware Prototype in Semi-Supervised Medical Image SegmentationabstractThe true label plays an important role in semi-supervised medical image segmentation (SSMIS) because it can provide the most accurate supervision information when the label is limited. The popular SSMIS method trains labeled and unlabeled data separately, and the unlabeled data cannot be directly supervised by the true label. This limits the contribution of labels to model training. Is there an interactive mechanism that can break the separation between two types of data training to maximize the utilization of true labels? Inspired by this, we propose a novel consistency learning framework based on the non-parametric distance metric of boundary-aware prototypes to alleviate this problem. This method combines CNN-based linear classification and nearest neighbor-based non-parametric classification into one framework, encouraging the two segmentation paradigms to have similar predictions for the same input. More importantly, the prototype can be clustered from both labeled and unlabeled data features so that it can be seen as a bridge for interactive training between labeled and unlabeled data. When the prototype-based prediction is supervised by the true label, the supervisory signal can simultaneously affect the feature extraction process of both data. In addition, boundary-aware prototypes can explicitly model the differences in boundaries and centers of adjacent categories, so pixel-prototype contrastive learning is introduced to further improve the discriminability of features and make them more suitable for non-parametric distance measurement. Experiments show that although our method uses a modified lightweight UNet as the backbone, it outperforms the comparison method using a 3D VNet with more parameters. Yongchao Wang 0004, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | ARISE: Graph Anomaly Detection on Attributed Networks via Substructure AwarenessabstractRecently, graph anomaly detection on attributed networks has attracted growing attention in data mining and machine learning communities. Apart from attribute anomalies, graph anomaly detection also aims at suspicious topological-abnormal nodes that exhibit collective anomalous behavior. Closely connected uncorrelated node groups form uncommonly dense substructures in the network. However, existing methods overlook that the topology anomaly detection performance can be improved by recognizing such a collective pattern. To this end, we propose a new graph anomaly detection framework on attributed networks via substructure awareness (ARISE). Unlike previous algorithms, we focus on the substructures in the graph to discern abnormalities. Specifically, we establish a region proposal module to discover high-density substructures in the network as suspicious regions. The average node-pair similarity can be regarded as the topology anomaly degree of nodes within substructures. Generally, the lower the similarity, the higher the probability that internal nodes are topology anomalies. To distill better embeddings of node attributes, we further introduce a graph contrastive learning scheme, which observes attribute anomalies in the meantime. In this way, ARISE can detect both topology and attribute anomalies. Ultimately, extensive experiments on benchmark datasets show that ARISE greatly improves detection performance (up to 7.30% AUC and 17.46% AUPRC gains) compared to state-of-the-art attributed networks anomaly detection (ANAD) algorithms. Jingcan Duan, Bin Xiao 0002, Siwei Wang 0001, Haifang Zhou, Xinwang Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Self-Supervised Image Local Forgery Detection by JPEG Compression TraceabstractFor image local forgery detection, the existing methods require a large amount of labeled data for training, and most of them cannot detect multiple types of forgery simultaneously. In this paper, we firstly analyzed the JPEG compression traces which are mainly caused by different JPEG compression chains, and designed a trace extractor to learn such traces. Then, we utilized the trace extractor as the backbone and trained self-supervised to strengthen the discrimination ability of learned traces. With its benefits, regions with different JPEG compression chains can easily be distinguished within a forged image. Furthermore, our method does not rely on a large amount of training data, and even does not require any forged images for training. Experiments show that the proposed method can detect image local forgery on different datasets without re-training, and keep stable performance over various types of image local forgery. Xiuli Bi, Wuqing Yan, Bo Liu 0047, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
AAAI | 4 |
| 2023 | Location-Aware Transformer Network for Few-Shot Medical Image SegmentationabstractAutomatic and precise organ segmentation plays a significant role in promoting the development of the diagnosis and treatment of the disease. Despite making enormous strides in medical image segmentation, conventional deep neural network-based methods are inherently massive data-driven techniques and are challenging to adapt to novel classes with a small number of labeled samples. Few-shot learning is a promising solution through learning novel classes from extremely limited annotated examples. However, existing few-shot segmentation methods focus excessively on targets in individual images while neglecting to model the global spatial correlation across images, which may cause severe performance degradation. To solve this issue, we propose a new Transformer-based few-shot segmentation framework for medical imaging, namely location-aware transformer network (LATNet), which establishes the spatial correlation between support and query objects, yielding location-aware prototypes, and then performs segmentation by computing the semantic similarity between query features and location-aware prototypes. Additionally, to further enhance the representativeness of the obtained location-aware prototypes in low-data regimes, we design a prediction iterative refinement module, which can iteratively exploit the query predictions output by each iteration to update the location-aware prototypes and progressively refine the query predictions. Extensive experiments on three challenging medical image datasets, i.e., Abd-MRI, Card-MRI, and Abd-CT, show that the proposed LATNet achieves remarkable improvements over current state-of-the-art methods by an average of 4.17%, 1.50%, and 4.63% in terms of the Dice Score, respectively. Wendong Huang, Bin Xiao 0002, Jinwu Hu, Xiuli Bi |
BIBM | 2 |
| 2023 | Non-rigid Medical Image Registration Based on Unsupervised Self-driven Prior FusionabstractDeformable image registration is a basic building block in intelligent bioinformatical analysis and biomedicine systems. With the rapid development of deep learning, the community has witnessed a great leap via this effective data-driven technique. Recently, the Vision Transformer, famous by its long-range modeling ability, has been successfully used in the field of medical image registration. However, the existing ViT based techniques are deemed to have certain limitations. Firstly, these methods often transplanted the transformer module directly into the networks, while did not dive deeper to explore its compatibility to practical registration tasks. Moreover, self-attention’s relatively rigid all-to-all patching strategy may cause undesirable discontinuity effect to the spatial calculation. To address these issues, we propose a novel medical image registration framework based on an efficient image prior learning and fusion mechanism. Unlike the existing prior-based registration methods, our model is capable of learning task-specific saliency priors, without the need of hand-crafted features, or heavy-loaded auxiliary tasks, or pre-acquired expensive annotations. Then, followed by a multi-scale patch embedding module, the self-driven saliency prior is integrated into a ViT block with an active feature fusion mechanism, to further expand our network’s structural learning abilities. Extensive experiments on multiple data sets have demonstrated the superior quality of the proposed framework. We believe this plug-and-play model will bring about more application potentials to the community (Project webpage: https://github.com/raincity212/SPF-Net). Yucheng Shu, Xuxuan Guan, Bin Xiao 0002, Lihong Qiao, Weisheng Li 0001, Xinbo Gao 0001 |
BIBM | 3 |
| 2023 | MCF: Mutual Correction Framework for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning is a promising method for medical image segmentation under limited annotation. However, the model cognitive bias impairs the segmentation performance, especially for edge regions. Furthermore, current mainstream semi-supervised medical image segmentation (SSMIS) methods lack designs to handle model bias. The neural network has a strong learning ability, but the cognitive bias will gradually deepen during the training, and it is difficult to correct itself. We propose a novel mutual correction framework (MCF) to explore network bias correction and improve the performance of SSMIS. Inspired by the plain contrast idea, MCF introduces two different subnets to explore and utilize the discrepancies between subnets to correct cognitive bias of the model. More concretely, a contrastive difference review (CDR) module is proposed to find out inconsistent prediction regions and perform a review training. Additionally, a dynamic competitive pseudo-label generation (DCPLG) module is proposed to evaluate the performance of subnets in real-time, dynamically selecting more reliable pseudo-labels. Experimental results on two medical image databases with different modalities (CT and MRI) show that our method achieves superior performance compared to several state-of-the-art methods. The code will be available at https://github.com/WYC-321/MCF. Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
CVPR | 2 |
| 2023 | DLBD: A Self-Supervised Direct-Learned Binary DescriptorabstractFor learning-based binary descriptors, the binarization process has not been well addressed. The reason is that the binarization blocks gradient back-propagation. Existing learning-based binary descriptors learn real-valued output, and then it is converted to binary descriptors by their proposed binarization processes. Since their binarizaiion processes are not a component of the network, the learning-based binary descriptor cannot fully utilize the advances of deep learning. To solve this issue, we propose a model-agnostic plugin binary transformation layer (BTL), making the network directly generate binary descriptors. Then, we present the first self-supervised, direct-learned binary descriptor, dubbed DLBD. Furthermore, we propose ultra-wide temperature-scaled crossentropy loss to adjust the distribution of learned descriptors in a larger range. Experiments demonstrate that the proposed BTL can substitute the previous binarization process. Our proposed DLBD outperforms SOTA on different tasks such as image retrieval and classification11Our code is available at: https://github.com/CQUPT-CV/DLBD. Bin Xiao 0002, Bo Liu 0047, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
CVPR | 1 |
| 2023 | Cross-slice Context Consistency for Semi-supervised 3D Left Atrium SegmentationabstractSemi-supervised learning is a promising approach in reducing the requirement to collect large amounts of dense annotations, especially in medical image segmentation. However, most existing semi-supervised 3D medical image segmentation methods tend to ignore the cross-slice context that contains extensive structural information. We believe cross-slice context can help the model capture semantic information complementary to slice context and achieve robust and more accurate segmentation. Therefore, in this paper, we propose a novel cross-slice context consistency framework for 3D left atrium segmentation named CSC2-Net. Our method can effectively utilize unlabeled data by encouraging consistent results between slice segmentation and cross-slice inference segmentation. To achieve this, we design a bidirectional gated context inference module (Bi-GCM) to model cross-slice context and predict slice segmentation without direct slice features. Experiments on a public left atrium (LA) databases show that our method achieves higher performance and outperforms state-of-the-art methods by imposing cross-slice consistency constraint. Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
ICME | 2 |
| 2023 | BMI-Net: A Brain-inspired Multimodal Interaction Network for Image Aesthetic AssessmentabstractImage aesthetic assessment (IAA) has drawn wide attention in recent years as more and more users post images and texts on the Internet to share their views. The intense subjectivity and complexity of IAA make it extremely challenging. Text triggers the subjective expression of human aesthetic experience based on human implicit memory, so incorporating the textual information and identifying the relationship with the image is of great importance for IAA. However, IAA with the image as input fails to fully consider subjectivity, while existing multimodal IAA ignores the interrelationship among modalities. To this end, we propose a brain-inspired multimodal interaction network (BMI-Net) that simulates how the association area of the cerebral cortex processes sensory stimuli. In particular, the knowledge integration LSTM (KI-LSTM) is proposed to learn the image-text interaction relation. The proposed scalable multimodal fusion (SMF) based on low-rank decomposition fuses image, text and interaction modalities to predict the aesthetic distribution. Extensive experiments show that the proposed BMI-Net outperforms existing state-of-the-art methods on three IAA tasks. Xixi Nie, Bo Hu 0008, Xinbo Gao 0001, Leida Li, Xiaodan Zhang 0005, Bin Xiao 0002 |
ACM Multimedia | 6 |
| 2023 | Secondary Labeling: A Novel Labeling Strategy for Image Manipulation DetectionabstractImage manipulation detection methods typically rely on a binary annotation called Primary Labeling (PrLa) to identify tampered and authentic regions in a tampered image. However, PrLa only focuses on the difference between authentic and tampered regions, ignoring the distinctions among tampered regions in different images. This transforms the task of image manipulation detection into salient object detection, with the goal shifting towards identifying the most attention-grabbing objects in images. To address this issue, this paper proposes a novel labeling strategy called Secondary Labeling (SeLa). SeLa generates a query table containing multiple tampered categories and randomly reassigns these tampered classes to different types of tampered data, effectively improving the detection performance of models by refocusing the differences among the various data. Additionally, to further improve the detection performance, this paper introduces an Adaptive Label Smoothing (ALS) regularization method. This method addresses the loss of correlation among tampered classes in SeLa caused by the one-hot encoding method. Experimental results show that compared with PrLa, SeLa not only improves the performance of detection models by up to 17%, but also enhances the robustness and convergence rate. Yang Wei 0002, Bin Xiao 0002, Xiuli Bi, Zhuoran Ma 0002, Yang Liu 0118, Zhuo Ma 0001 |
ACM Multimedia | 2 |
| 2023 | AEP-GAN: Aesthetic Enhanced Perception Generative Adversarial Network for Asian facial beauty synthesis
Weisheng Li 0001, Xinbo Gao 0001, Bin Xiao 0002 |
Appl. Intell. | 4 |
| 2023 | QDRJL: Quaternion dynamic representation with joint learning neural network for heart sound signal abnormal detectionabstractAt present, deep learning based heart sound diagnosis algorithms are mostly complex and large models for high accuracy, which are difficult to deploy on mobile devices due to the high number of parameters and large computational cost. The current mainstream approach for processing heart sound signals involves utilizing their Mel-frequency cepstral coefficients (MFCC) features. However, most existing methods have overlooked the multi-channel characteristics of MFCC. To address this issue, we propose a Quaternion Dynamic Representation with Joint Learning (QDRJL) neural network for learning MFCC multi-channel features. Our proposed approach combines quaternion dynamic convolution with dynamic weighting and the Quaternion Interior Learning Block (QILB). Finally, we present a global and energy joint learning branch for jointly learning MFCC features. The success of the proposed quaternion network depends on its ability to utilize the internal relations between quaternion-valued input features and the definition of the dynamic weight variables in the augmented quaternion domain. We assessed various state-of-the-art classification algorithms for detecting heart sounds and found that our proposed classifier achieved an accuracy of up to 97.2%, outperforming existing models. Our experimental evaluation, using the 2016 PhysioNet/CinC Challenge dataset, revealed that our model could reduce the number of network parameters to 25% due to quaternion properties. Lihong Qiao, Bin Xiao 0002, Yucheng Shu, Yuhang Shi, Weisheng Li 0001, Xinbo Gao 0001 |
Neurocomputing | 3 |
| 2023 | A Dual Self-Calibrating Framework for Noninvasive Fetal ECG R-Peak DetectionabstractFetal heart rate (fHR) is critical for assessing fetal health and diagnosing disorders, such as fetal distress, congenital heart disease, and intrauterine growth retardation. With the rapid development of the Internet of Medical Things (IoMT), fetal R-peak detection plays an important role in diagnosing heart defects during pregnancy. However, due to the nonlinear mixing of multiple sources in the noninvasive signals and the low signal-to-noise ratio (SNR), it is difficult to obtain accurate R-peak detection result. This article presents a dual self-calibrating system based on a spectral attention kernel independent component analysis (SA-KICA) module and a self-calibrating fetal R-peak detection (SC-FRD) module. SA-KICA is an ICA-based calibration module constructed by the spectral attention mechanism, which was sought from short-time Fourier transform (STFT) and was shipped back to original signal with convolution to achieve perfect maternal electrocardiogram (MECG) separation in high-dimensional linear separable space. Then, a periodic and morphological-based channel selector is designed to select the optimal MECG. After MECG removal, to further improve the performance of fetal R-peak detection, the SC-FRD module is introduced to utilize the interior peak information and self-calibrating strategy, which includes variance-based fetal R-peak seed selection, time-varying coarse prediction, and adaptive probability mask calibration. The proposed framework is a primary attempt to concurrently introduce the nonlinear feature, spectral information, and self-calibrating strategy in the field of fetal ECG processing. The framework achieved excellent performance in fetal R-peak detection accuracy on a simulated data set and two public data sets with varying divergence and richness of resources. The experimental results show that our framework is superior to existing methods and can be used as a potential fetal monitoring method in the application of IoMT. The code is released inhttps://github.com/bfyjr/NI-FECG-Extraction. Lihong Qiao, Shuai Hu, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Internet Things J. | 3 |
| 2023 | Learning an Invariant and Equivariant Network for Weakly Supervised Object DetectionabstractWeakly Supervised Object Detection (WSOD) is of increasing importance in the community of computer vision as its extensive applications and low manual cost. Most of the advanced WSOD approaches build upon an indefinite and quality-agnostic framework, leading to unstable and incomplete object detectors. This paper attributes these issues to the process of inconsistent learning for object variations and the unawareness of localization quality and constructs a novel end-to-end Invariant and Equivariant Network (IENet). It is implemented with a flexible multi-branch online refinement, to be naturally more comprehensive-perceptive against various objects. Specifically, IENet first performs label propagation from the predicted instances to their transformed ones in a progressive manner, achieving affine-invariant learning. Meanwhile, IENet also naturally utilizes rotation-equivariant learning as a pretext task and derives an instance-level rotation-equivariant branch to be aware of the localization quality. With affine-invariance learning and rotation-equivariant learning, IENet urges consistent and holistic feature learning for WSOD without additional annotations. On the challenging datasets of both natural scenes and aerial scenes, we substantially boost WSOD to new state-of-the-art performance. The codes have been released at: https://github.com/XiaoxFeng/IENet. Xiaoxu Feng, Xiwen Yao, Hui Shen 0005, Gong Cheng 0003, Bin Xiao 0002, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Sniffer: A Novel Model Type Detection System against Machine-Learning-as-a-Service PlatformsabstractRecent works explore several attacks against Machine-Learning-as-a-Service (MLaaS) platforms (e.g., the model stealing attack), allegedly posing potential real-world threats beyond viability in laboratories. However, hampered by model-type-sensitive , most of the attacks can hardly break mainstream real-world MLaaS platforms. That is, many MLaaS attacks are designed against only one certain type of model, such as tree models or neural networks. As the black-box MLaaS interface hides model type info, the attacker cannot choose a proper attack method with confidence, limiting the attack performance. In this paper, we demonstrate a system, named Sniffer, that is capable of making model-type-sensitive attacks "great again" in real-world applications. Specifically, Sniffer consists of four components: Generator, Querier, Probe, and Arsenal. The first two components work for preparing attack samples. Probe, as the most characteristic component in Sniffer, implements a series of self-designed algorithms to determine the type of models hidden behind the black-box MLaaS interfaces. With model type info unraveled, an optimum method can be selected from Arsenal (containing multiple attack methods) to accomplish its attack. Our demonstration shows how the audience can interact with Sniffer in a web-based interface against five mainstream MLaaS platforms. Zhuo Ma 0001, Yilong Yang 0004, Bin Xiao 0002, Yang Liu 0118, Xinjing Liu, Zhuoran Ma 0002, Tong Yang 0003 |
Proc. VLDB Endow. | 3 |
| 2023 | IEMask R-CNN: Information-Enhanced Mask R-CNNabstractThe instance segmentation task is relatively difficult in computer vision, which requires not only high-quality masks but also high-accuracy instance category classification. Mask R-CNN has been proven to be a feasible method. However, due to the Feature Pyramid Network (FPN) structure lack useful channel information, global information and low-level texture information, and mask branch cannot obtain useful local-global information, Mask R-CNN is prevented from obtaining high-quality masks and high-accuracy instance category classification. Therefore, we proposed the Information-enhanced Mask R-CNN, called IEMask R-CNN. In the FPN structure of IEMask R-CNN, the information-enhanced FPN will enhance the useful channel information and the global information of the feature maps to solve the issues that the high-level feature map loses useful channel information and inaccurate of instance category classification, meanwhile the bottom-up path enhancement with adaptive feature fusion will ultilize the precise positioning signal in the lower layer to enhance the feature pyramid. In the mask branch of IEMask R-CNN, an encoding-decoding mask head will strength local-global information to gain a high-quality mask. Without bells and whistles, IEMask R-CNN gains significant gains of about 2.60%, 4.00%, 3.17% over Mask R-CNN on MS COCO2017, Cityscapes and LVIS1.0 benchmarks respectively. Xiuli Bi, Jinwu Hu, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Big Data | 3 |
| 2023 | Novel Multi-Feature Fusion Facial Aesthetic Analysis FrameworkabstractMachine learning has been used in facial beauty prediction studies. However, the integrity of facial geometric information is not considered in facial aesthetic feature extraction, and the impact of other facial attributes (expression) on aesthetics. We propose a novel multi-feature fusion facial aesthetic analysis framework (NMFA) to overcome this problem. First, we designed a facial shape feature, which is an intuitive, visual quantitative description, based on B-spline. Second, we designed a representative low-dimensional facial structural feature to establish the theoretical basis of the facial structure, based on facial aesthetic structure and expression recognition theory. Next, we designed texture and holistic features based on Gabor and VGG-face network. Finally, we used a multi-feature fusion strategy to fuse them for aesthetic evaluation. Experiments were conducted on four databases. The results revealed that the proposed method realizes the visualization of facial shape features, enriches geometric information, solves the problem of lack of facial geometric information and difficulty to understand, and achieves excellent performance with fewer parameters. Weisheng Li 0001, Xinbo Gao 0001, Bin Xiao 0002 |
IEEE Trans. Big Data | 4 |
| 2023 | Outsourced Privacy-Preserving Data Alignment on Vertically Partitioned DatabaseabstractIn the context of real-world secure outsourced computations, private data alignment has been always the essential preprocessing step. However, current private data alignment schemes, mainly circuit-based, suffer from high communication overhead and often need to transfer potentially gigabytes of data. In this paper, we propose a lightweight private data alignment protocol (called SC-PSI) that can overcome the bottleneck of communication. Specifically, SC-PSI involves four phases of computations, including data preprocessing, data outsourcing, private set member (PSM) evaluation and circuit computation (CC). Like prior works, the major overhead of SC-PSI mainly lies in the latter two phases. The improvement is SC-PSI utilizes the function secret sharing technique to develop the PSM protocol, which avoids the multiple rounds of communication to compute intersection set members. Moreover, benefited from our specially designed PSM protocol, SC-PSI does not to execute complex secure comparison circuits in the CC phase. Experimentally, we validate that compared to prior works, SC-PSI can save around 61.39% running time and 89.61% communication overhead. Cui Hu, Bin Xiao 0002, Yang Liu 0118, Teng Li 0003, Zhuo Ma 0001, Jianfeng Ma 0001 |
IEEE Trans. Big Data | 3 |
| 2023 | A Versatile Detection Method for Various Contrast Enhancement ManipulationsabstractContrast enhancement manipulation is a common method to improve the visual effect of an image. Meanwhile, it can also be considered a type of global image forgery because it changes the image’s visual appearance without alerting its semantics. Moreover, for local image forgery, a tampered image may be composited by images with different contrast enhancement manipulations or post-processed by a contrast enhancement manipulation to conceal the trails of tampering. Therefore, contrast enhancement manipulation detection is critical to global image forgery detection. The existing methods can only detect a particular type of contrast enhancement manipulation, such as gamma correction or histogram equalization. To break this limitation, we propose the zero-gap spans (ZGS) as the fingerprint to explore the traces of contrast enhancement manipulations. Based on ZGS, various contrast enhancement manipulations can be distinguished by a simple classification method at image-level and patch-level; different gamma corrections can be identified, and their gamma value can be estimated. Experimental results indicate that the proposed ZGS-based classification method can achieve and maintain good classification performance under different cases (gamma correction, simple histogram equalization, modified histogram equalization techniques). Meanwhile, ZGS can estimate the gamma value with the mean squared error (MSE) below 0.1156. For the local forgery images, the proposed ZGS also can be utilized to locate the regions with different contrast enhancement manipulations. Xiuli Bi, Yixuan Shang, Bo Liu 0047, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Target-Aware Transformer TrackingabstractObject tracking is aimed at locating a specific object in the image sequence, such as pedestrians, vehicles, and so on. The existing algorithms based on siamese neural network predict the target through similarity matching. Although these algorithms have achieved satisfactory performance, in the process of similarity calculation between template image and search image, only local information is often concerned, which makes the algorithms difficult to obtain the optimal solution. To deal with the abovementioned problems, we propose a model based on Transformer, named TaTrack. Specifically, we first use the encoders to enhance the features. Then, the dependency between template features and search features is established through the target-aware module. Finally, we utilize the classification regression network to locate the target, and use the classification score to adapt to update the template image. Experiments show that our model can achieve great performance on GOT-10k, LaSOT, and TrackingNet datasets. Yuhui Zheng, Yan Zhang 0108, Bin Xiao 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | HS-Vectors: Heart Sound Embeddings for Abnormal Heart Sound Detection Based on Time-Compressed and Frequency-Expanded TDNN With Dynamic Mask EncoderabstractIn recent years, auxiliary diagnosis technology for cardiovascular disease based on abnormal heart sound detection has become a research hotspot. Heart sound signals are promising in the preliminary diagnosis of cardiovascular diseases. Previous studies have focused on capturing the local characteristics of heart sounds. In this paper, we investigate a method for mapping heart sound signals with complex patterns to fixed-length feature embedding called HS-Vectors for abnormal heart sound detection. To get the full embedding of the complex heart sound, HS-Vectors are obtained through the Time-Compressed and Frequency-Expanded Time-Delay Neural Network(TCFE-TDNN) and the Dynamic Masked-Attention (DMA) module. HS-Vectors extract and utilize the global and critical heart sound characteristics by masking out irreverent information. Based on the TCFE-TDNN module, the heart sound signal within a certain time is projected into fixed-length embedding. Then, with a learnable mask attention matrix, DMA stats pooling aggregates multi-scale hidden features from different TCFE-TDNN layers and masks out irrelevant frame-level features. Experimental evaluations are performed on a 10-fold cross-validation task using the 2016 PhysioNet/CinC Challenge dataset and the new publicly available pediatric heart sound dataset we collected. Experimental results demonstrate that the proposed method excels the state-of-the-art models in abnormality detection. Lihong Qiao, Yonghao Gao, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Reveal Your Images: Gradient Leakage Attack Against Unbiased Sampling-Based Secure AggregationabstractRecently, some Unbiased Gradient Sampling-based (UGS) methods have been proposed to enhance the security and efficiency of federated learning through crafted unbiased random transformation and sampling, such as MinMax Sampling in SIGMOD ’22. In this paper, we propose a novel attack, GLAUS, to show that UGS is not as secure as claimed in these works and is still vulnerable to the gradient leakage attack (GLA). Specifically, we demonstrate an idea to approximately infer the gradient for GLA in the context of the UGS scenario where the real gradient is not available. Once the gradient is approximately obtained, the security of the UGS frameworks is downgraded to that of the original federated learning. The approximate gradient is refined by the following steps: 1)narrow the gradient searching rangeto the finite set; 2)obtain the magnitudeof each gradient value approximately; 3)revise the gradient signs. Versus the failure of existing attacks, extensive experiments on six datasets show that our attack is effective in reconstructing private datapoints with pixel-wise accuracy on four network sizes and three image resolutions. Finally, we show how to defend against GLAUS while maintaining the high efficiency of UGS and only introducing an additional step to hide the sampled gradient indices. Yilong Yang 0004, Zhuo Ma 0001, Bin Xiao 0002, Yang Liu 0118, Teng Li 0003, Junwei Zhang 0008 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Cross-Mix Monitoring for Medical Image Segmentation With Limited SupervisionabstractImage segmentation is a fundamental building block of automatic medical applications. It has been greatly improved since the emergence of deep neural networks. However, deep-learning based models often require a large number of manual annotations, which has seriously hindered its practical usage. To alleviate this problem, numerous works were proposed by utilizing unlabeled data based on semi-supervised frameworks. Recently, the Mean-Teacher (MT) model has been successfully applied in many scenarios due to its effective learning strategy. Nevertheless, the existing MT model still have certain limitations. Firstly, various sorts of perturbations are often added to the training data to gain extra generalization ability through consistency training. However, if the variation is too weak, it may cause the Lazy Student Phenomenon, and bring large fluctuations to the learning model. On the contrary, large image perturbations may enlarge the performance gap between the teacher and student. In this case, the student may lose its learning momentum, and more seriously, drag down the overall performance of the whole system. In order to address these issues, we introduce a novel semi-supervised medical image segmentation framework, in which a Cross-Mix Teaching paradigm is proposed to provide extra data flexibility, thus effectively avoid Lazy Student Phenomenon. Moreover, a lightweight Transductive Monitor is applied to server as the bridge that connect the teacher and student for active knowledge distillation. In the light of this cross-network information mixing and transfer mechanism, our method is able to continuously explore the discriminative information contained in unlabeled data. Extensive experiments on challenging medical image data sets demonstrate that our method is able to outperform current state-of-the-art semi-supervised segmentation methods under severe lack of supervision. Yucheng Shu, Hengbo Li, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Collaborative Decision-Reinforced Self-Supervision for Attributed Graph ClusteringabstractAttributed graph clustering aims to partition nodes of a graph structure into different groups. Recent works usually use variational graph autoencoder (VGAE) to make the node representations obey a specific distribution. Although they have shown promising results, how to introduce supervised information to guide the representation learning of graph nodes and improve clustering performance is still an open problem. In this article, we propose a Collaborative Decision-Reinforced Self-Supervision (CDRS) method to solve the problem, in which a pseudo node classification task collaborates with the clustering task to enhance the representation learning of graph nodes. First, a transformation module is used to enable end-to-end training of existing methods based on VGAE. Second, the pseudo node classification task is introduced into the network through multitask learning to make classification decisions for graph nodes. The graph nodes that have consistent decisions on clustering and pseudo node classification are added to a pseudo-label set, which can provide fruitful self-supervision for subsequent training. This pseudo-label set is gradually augmented during training, thus reinforcing the generalization capability of the network. Finally, we investigate different sorting strategies to further improve the quality of the pseudo-label set. Extensive experiments on multiple datasets show that the proposed method achieves outstanding performance compared with state-of-the-art methods. Our code is available at https://github.com/Jillian555/TNNLS_CDRS. Pengfei Zhu 0001, Yu Wang 0106, Bin Xiao 0002, Qinghua Hu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | MU-TEIR: Traceable Encrypted Image Retrieval in the Multi-User SettingabstractThe encrypted image retrieval technique allows users to retrieve images in an encrypted manner without decrypting images. However, most of the existing schemes still are vulnerable to security threats and inefficiency, caused by malicious users and inefficient feature extraction methods, respectively. To this end, we propose a traceable encrypted image retrieval in the multi-user setting in this article, termed as MU-TEIR. First, MU-TEIR employs a convolutional neural network VGG16 to extract image feature vectors and calculate the mean and variance of the feature vectors to construct the index, then encrypts index with the distributed two trapdoors public-key cryptosystem. After that, MU-TEIR protects image content by encrypting each image pixel with a standard stream cipher. Furthermore, MU-TEIR utilizes a watermark-based mechanism to prevent malicious query users from maliciously distributing images. Detailed security analysis shows that MU-TEIR protects the outsourced images and indexes security as well as query privacy, and can track malicious users. Experimental results verify effectiveness of MU-TEIR. Jianfeng Ma 0001, Yinbin Miao, Yue Wang 0063, Ximeng Liu, Kim-Kwang Raymond Choo, Bin Xiao 0002 |
IEEE Trans. Serv. Comput. | 7 |
| 2022 | FF-Net: An End-to-end Feature-Fusion Network for Double JPEG Detection and Localization
Bo Liu 0047, Ranglei Wu, Xiuli Bi, Bin Xiao 0002 |
ACML | 4 |
| 2022 | Detecting Generated Images by Real Images
Bo Liu 0047, Fan Yang 0159, Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
ECCV (14) | 4 |
| 2022 | Context Correlation Aware Network for Cardiac SegmentationabstractAutomatically segmenting the anatomical structure of the heart from the cardiac magnetic resonance (CMR) images offers a great potential to augment the traditional healthcare strategy for the quantitative analysis of cardiac contractile function. Most of the existing CNN-based methods for cardiac segmentation tend to ignore the misalignment issues during the feature aggregation process and not fully use multi-scale context and contour information, which may lead to the unexpected misclassification caused by the falsely aligned contextual features and the discontinuity in the edge of segmentation maps. To resolve these issues, we proposed a context correlation aware network (CCA-Net). In CCA-Net, a volume correlation flow module was designed to align contour features and semantic features from adjacent levels, which offered the guidance to wrap low-resolution semantic features into high-resolution features. Besides, a residual gated squeeze module was utilized to explicitly model the boundaries and enhance the representations. Extensive experiments on the multi-sequence cardiac magnetic resonance segmentation challenge (MS-CMRSeg 2019) dataset and MICCAI challenge 2017 automatic cardiac diagnosis challenge (ACDC) dataset demonstrated that CCA-Net was superior to other state-of-the-art methods. Junchao Fan, Jiawei Pei, Xiuli Bi, Bin Xiao 0002, Pietro Liò |
ICME | 4 |
| 2022 | Partial-to-Partial Point Cloud Registration Based on Multi-Level Semantic-Structural Cognitionabstract3D point cloud registration attempt to establish spatial correspondences between the source point cloud and the target point cloud. It is a fundamental task in computer vision and multimedia applications. Recently, many learning-based methods have been proposed and achieved promising performance. However, in the Partial-to-Partial (PtP) registration problem, the existence of a large number of external points may greatly handicap the effectiveness of these methods. In this paper, we propose to address the PtP issue under a novel multi-task cognition framework. At the global semantic level, we introduce a multi-scale feature exchanging network to actively evaluate the matching credibility. For the local structural learning, an inner-inter attention fusion branch is applied to generate discriminative features. Moreover, we integrate a novel alternating correspondences searching mechanism with a flexible bi-directional dislocation loss to perform robust learning, and a simple yet effective SVD weighting scheme is introduced at the inference stage. Experiment results on two challenging PtP 3D point cloud registration data sets show that our proposed method outperforms all the SOTA methods with higher precision and robustness. Yucheng Shu, Zongzhuang Hou, Bin Xiao 0002, Xiuli Bi, Xiao Luan, Weisheng Li 0001 |
ICME | 3 |
| 2022 | Mixed Color Channels (MCC): A Universal Module for Mixed Sample Data Augmentation MethodsabstractColor invariance is critical for computer vision systems since it significantly increases the robustness and effectiveness of the system. MSDA approaches (e.g., FMix and CutMix) have attracted considerable attention in recent years since they are simple, effective, and do not require extra computation con-sumption. By mixing samples, these approaches extend the distribution of training samples. However, the color information of these mixed samples is not changed, which makes it still difficult for trained models to achieve color invariance. To address this issue, we propose a universal module called Mixed Color Channels (MCC) that implements color changes by mixing the sample and its color variants, which enables trained models to achieve color invariance. In the experimen-tal section, we insert MCC into four state-of-the-art MSDA approaches, evaluate its effectiveness, and embed MCC into a non-MSDA method to demonstrate its extensibility. Yang Wei 0002, Jianfeng Ma 0001, Zhongyuan Jiang, Bin Xiao 0002 |
ICME | 4 |
| 2022 | HessHist: A Hessian-matrix weighted histogram for image contrast enhancementabstractAbstract For image contrast enhancement operation, it is a keypoint to obtain more natural enhanced results and keep more details without distortion. In this paper, a novel image Hessian‐matrix weighted histogram for image contrast enhancement is proposed, which can improve the contrast of smooth regions and simultaneously restrain the contrast of texture regions. In the proposed method, the multi‐scale fractional‐order Hessian‐matrix is firstly utilized to detect and quantify the texture information of the input image, which explores the regions that should be contrasted or should be restrained. Then, the strong texture regions are suppressed by a designed suppress function. Finally, the information on unsuppressed regions and suppressed texture regions will be count by a histogram, which is termed as Hessian‐matrix weighted Histogram (HessHist) in this paper. According to HessHist, the corresponding cumulative distribution function will realize the contrast enhancement operation on the input image. For real‐time application, the integral images are introduced for fast computation of the HessHist. Experimental results show that the proposed HessHist‐based image enhancement algorithm preserves more details of input image without distortion, and is competitive with state‐of‐the‐art image enhancement algorithms in both subjective visual perception and objective evaluation metrics. Junchao Fan, Xuyang Zong, Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001 |
IET Image Process. | 5 |
| 2022 | Image splicing forgery detection by combining synthetic adversarial networks and hybrid dense U-net based on multiple spacesabstractWith the popularity of image editing tools, the originality and information security of images are facing serious threats. The most common threat is splicing forgery that copies a part of the area from one donor image to the acceptor one. Some research works were proposed to protect the image originality, whereas they are still difficult to apply in practice. There are two main reasons: (a) very limited data for learning models; (b) huge attribute differences between the donor and acceptor images. We propose two novel tasks to conquer the above challenges: Synthetic Adversarial Networks (SANs) and Hybrid Dense U-Net (HDU-Net). SAN finds the most secluded position for inserting tampered areas in an image by learning the association between scenes and objects, and can enlarge the original small data set by more than 40 times. We call the data set SF-Data generated by SAN. We combine the dense U-Net that detects the differences of the essential attributes of image with four spaces containing more available feature information to propose HDU-Net. Then, the synthetic data set SF-Data are used to train HDU-Net. We perform various attack experiments on several public data sets to demonstrate the effectiveness and robustness of our method. Yang Wei 0002, Jianfeng Ma 0001, Bin Xiao 0002, Wenying Zheng |
Int. J. Intell. Syst. | 4 |
| 2022 | Multimodal medical image fusion based on multichannel coupled neural P systems and max-cloud models in spectral total variation domain
Guofen Wang, Weisheng Li 0001, Xinbo Gao 0001, Bin Xiao 0002, Jiao Du |
Neurocomputing | 4 |
| 2022 | DCNet: Diversity convolutional network for ventricle segmentation on short-axis cardiac magnetic resonance images
Weisheng Li 0001, Xinbo Gao 0001, Bin Xiao 0002 |
Knowl. Based Syst. | 5 |
| 2022 | Recent advancement in haze removal approaches
Hira Khan, Bin Xiao 0002, Weisheng Li 0001, Muhammad Nazeer |
Multim. Syst. | 2 |
| 2022 | Multi-Task Convolution Operators With Object Detection for Visual TrackingabstractRecently, multi-task correlation filters has drawn much attention in the object tracking field, which utilizes the multi-task learning (MTL) approach to explore the interdependencies among deep features for object tracking. However, the existing multi-task correlation filters based method fails to consider the relations between the correlation filters. To address this problem, a novel correlation filters based visual tracking method is proposed in this paper, with the integration of multi-task convolution operators and object detection. In our method, convolution and correction filters are jointly learnt through using the MTL technique, with the purpose of exploring not only the interdependencies of deep features but also the internal relevance of the convolution filters. In addition, object detection is introduced into our algorithm to handle the problem of object missing to ensure a better performance of our tracking method. Experiments on five benchmark datasets demonstrate that the proposed visual tracking method outperforms existing state-of-the-art approaches. Yuhui Zheng, Xinyan Liu 0002, Bin Xiao 0002, Xu Cheng 0003, Yi Wu 0001, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | PAM-DenseNet: A Deep Convolutional Neural Network for Computer-Aided COVID-19 DiagnosisabstractCurrently, several convolutional neural network (CNN)-based methods have been proposed for computer-aided COVID-19 diagnosis based on lung computed tomography (CT) scans. However, the lesions of pneumonia in CT scans have wide variations in appearances, sizes, and locations in the lung regions, and the manifestations of COVID-19 in CT scans are also similar to other types of viral pneumonia, which hinders the further improvement of CNN-based methods. Delineating infection regions manually is a solution to this issue, while excessive workload of physicians during the epidemic makes it difficult for manual delineation. In this article, we propose a CNN called dense connectivity network with parallel attention module (PAM-DenseNet), which can perform well on coarse labels without manually delineated infection regions. The parallel attention module automatically learns to strengthen informative features from both channelwise and spatialwise simultaneously, which can make the network pay more attention to the infection regions without any manual delineation. The dense connectivity structure performs feature maps reuse by introducing direct connections from previous layers to all subsequent layers, which can extract representative features from fewer CT slices. The proposed network is first trained on 3530 lung CT slices selected from 382 COVID-19 lung CT scans, 372 lung CT scans infected by other pneumonia, and 200 normal lung CT scans to obtain a pretrained model for slicewise prediction. We then apply this pretrained model to a CT scans dataset containing 94 COVID-19 CT scans, 93 other pneumonia CT scans, and 93 normal lung scans, and achieve patientwise prediction through a voting mechanism. The experimental results show that the proposed network achieves promising results with an accuracy of 94.29%, a precision of 93.75%, a sensitivity of 95.74%, and a specificity of 96.77%, which is comparable to the methods that are based on manually delineated infection regions. Bin Xiao 0002, Xiaoming Qiu, Guoyin Wang 0001, Wenbing Zeng, Weisheng Li 0001, Yongjian Nian |
IEEE Trans. Cybern. | 1 |
| 2022 | Registration-Is-Evaluation: Robust Point Set Matching With Multigranular Prior AssessmentabstractPoint set registration is one of the challenging tasks in remote sensing image processing and analysis. Its critical step is to find the corresponding relationships between the fixed scene point set and the moving model point set that undergo different sorts of transformations. Existing algorithms primarily utilize different types of prior information to improve the registration performance, such as spatial consistency, local similarity, and uniform outliers. However, due to the lack of active evaluation on the prior and intermediate information during the registration process, these strategies are susceptible to large data transformations. In order to enhance the robustness and accuracy for point set registration, we propose in this article a novel framework, namely Registration-is-Evaluation (RisE). Based on a multigranular probability model, our method exploits and utilizes both prior and posterior information to dynamically evaluate the matching status. What is more, instead of adding an extra uniform prior, we unified the outliers, noise, missing points, and heavily warped points into our registration evaluation model and address them simultaneously. We also apply a novel point set descriptor, called local polar relative geometry (LPRG), to have a more robust local similarity measurement. It adopts the local polar coordinate to perform multiscale pooling and relative geometric computation. Based on our proposed method, the matching relationships and the spatial transformations can be actively evaluated to provide useful contextual guidance for the registration process. Experimental results on multiple data sets show that our algorithm outperforms the state-of-the-art methods, in terms of both accuracy and robustness under large data degradations. Yucheng Shu, Zhenlong Liao, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Privacy-Preserving Color Image Feature Extraction by Quaternion Discrete Orthogonal MomentsabstractTo implement image storage and computation in cloud servers without violating users’ privacy, privacy-preserving feature extraction has been a new research interest. The existing works are mainly designed for grayscale images. For color images, they tend to convert them to grayscale images or obtain the results of the combination of single-channel processes. While the capabilities of features extracted from the encrypted color images will be affected if color information and inter-relationship between color channels are ignored. To fully preserve features of color images, we introduce quaternion theory to encode each color image and propose an improved vector homomorphic encryption scheme (IVHE) to encrypt quaternion-based color images. IVHE helps protect image content and keep vector characteristics of color images. Based on IVHE, the framework for feature extraction of privacy-preserving Quaternion Discrete Orthogonal Moments (PPQDOMs) is presented. Theoretical analyses prove that Quaternion Discrete Orthogonal Moments (QDOMs) can be extracted from the encrypted color images by PPQDOMs. Furthermore, we apply three common Discrete Orthogonal Moments to the proposed framework to evaluate its performance. Experimental results demonstrate that the proposed framework can protect color image content and perform well compared to QDOMs in image reconstruction and image recognition. Xiuli Bi, Chao Shuai, Bo Liu 0047, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | A Novel Framework With Weighted Decision Map Based on Convolutional Neural Network for Cardiac MR SegmentationabstractFor diagnosing cardiovascular disease, an accurate segmentation method is needed. There are several unresolved issues in the complex field of cardiac magnetic resonance imaging, some of which have been partially addressed by using deep neural networks. To solve two problems of over-segmentation and under-segmentation of anatomical shapes in the short-axis view from different cardiac magnetic resonance sequences, we propose a novel two-stage framework with a weighted decision map based on convolutional neural networks to segment the myocardium (Myo), left ventricle (LV), and right ventricle (RV) simultaneously. The framework comprises a decision map extractor and a cardiac segmenter. A cascaded U-Net++ is used as a decision map extractor to acquire the decision map that decides the category of each pixel. Cardiac segmenter is a multiscale dual-path feature aggregation network (MDFA-Net) which consists of a densely connected network and an asymmetric encoding and decoding network. The input to the cardiac segmenter is derived from processed original images weighted by the output of the decision map extractor. We conducted experiments on two datasets of multi-sequence cardiac magnetic resonance segmentation challenge 2019 (MS-CMRSeg 2019) and myocardial pathology segmentation challenge 2020 (MyoPS 2020). Test results obtained on MyoPS 2020 show that the average Dice coefficients of the proposed method on the segmentation tasks of Myo, LV and RV are 84.70%, 86.00%, and 86.31%, respectively. Weisheng Li 0001, Xinbo Gao 0001, Bin Xiao 0002 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Medical Image Fusion and Denoising Algorithm Based on a Decomposition Model of Hybrid Variation-Sparse RepresentationabstractMedical image fusion technology integrates the contents of medical images of different modalities, thereby assisting users of medical images to better understand their meaning. However, the fusion of medical images corrupted by noise remains a challenge. To solve the existing problems in medical image fusion and denoising algorithms related to excessive blur, unclean denoising, gradient information loss, and color distortion, a novel medical image fusion and denoising algorithm is proposed. First, a new image layer decomposition model based on hybrid variation-sparse representation and weighted Schatten p-norm is proposed. The alternating direction method of multipliers is used to update the structure, detail layer dictionary, and detail layer coefficient map of the input image while denoising. Subsequently, appropriate fusion rules are employed for the structure layers and detail layer coefficient maps. Finally, the fused image is restored using the fused structure layer, detail layer dictionary, and detail layer coefficient maps. A large number of experiments confirm the superiority of the proposed algorithm over other algorithms. The proposed medical image fusion and denoising algorithm can effectively remove noise while retaining the gradient information without color distortion. Guofen Wang, Weisheng Li 0001, Jiao Du, Bin Xiao 0002, Xinbo Gao 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | DTMNet: A Discrete Tchebichef Moments-based Deep Neural Network for Multi-focus Image FusionabstractCompared with traditional methods, the deep learning-based multi-focus image fusion methods can effectively improve the performance of image fusion tasks. However, the existing deep learning-based methods encounter a common issue of a large number of parameters, which leads to the deep learning models with high time complexity and low fusion efficiency. To address this issue, we propose a novel discrete Tchebichef moment-based Deep neural network, termed as DTMNet, for multi-focus image fusion. The proposed DTMNet is an end-to-end deep neural network with only one convolutional layer and three fully connected layers. The convolutional layer is fixed with DTM co-efficients (DTMConv) to extract high/low-frequency information without learning parameters effectively. The three fully connected layers have learnable parameters for feature classification. Therefore, the proposed DTMNet for multi-focus image fusion has a small number of parameters (0.01M paras vs. 4.93M paras of regular CNN) and high computational efficiency (0.32s vs. 79.09s by regular CNN to fuse an image). In addition, a large-scale multi-focus image dataset is synthesized for training and verifying the deep learning model. Experimental results on three public datasets demonstrate that the proposed method is competitive with or even outperforms the state-of-the-art multi-focus image fusion methods in terms of subjective visual perception and objective evaluation metrics. Bin Xiao 0002, Haifeng Wu, Xiuli Bi |
ICCV | 1 |
| 2021 | Reality Transform Adversarial Generators for Image Splicing Forgery Detection and LocalizationabstractWhen many forgery images become more and more realistic with help of image editing tools and convolutional neural networks (CNNs), authenticators need to improve their ability to verify these forgery images. The process of generating and detecting forgery images is the same as the principle of Generative Adversarial Networks (GANs). In this paper, since the retouching progress of forgery images requires to suppress the tampering artifacts and to keep the structural information, we consider this retouching progress as an image style transform, and then propose a fake-to-realistic transform generator GT. For detecting the tampered regions, a localization generator GMis proposed too, which is based on a multi-decoder-single-task strategy. By adversarial training two generators, the proposed α-learnable whitening and coloring transform α-learnable WCT) block in GTautomatically suppress the tampering artifacts in the forgery images. Meanwhile, the detection and localization abilities of GMwill be improved by learning the forgery images retouched by GT. The experiment results demonstrate that the proposed two generators in GAN can simulate confrontation between the faker and the authenticator well; the localization generator GMoutperforms the state-of-the-art methods in splicing forgery detection and localization on four public datasets. Xiuli Bi, Bin Xiao 0002 |
ICCV | 3 |
| 2021 | Multi-Task Wavelet Corrected Network for Image Splicing Forgery Detection and LocalizationabstractAlthough the existing image splicing forgery detection networks can achieve a promising performance, most of these networks utilize regular pooling operations (max-pooling and mean-pooling) and a single task strategy, which limits the comprehensiveness and representativeness of the features learned by the networks. In this paper, we propose a multi-task wavelet corrected network (MWC-Net) that can learn more comprehensive and representative features for image splicing forgery detection and localization. MWC-Net exploits wavelet-pooling and wavelet un-pooling to compress and reconstruct the features of splicing forgery images, which can reduce information loss during learning features. Mean-while, MWC-Net implements a multi-task strategy to improve its ability to learn and utilize more comprehensive and representative features. The experimental results demonstrate that MWC-Net outperforms the state-of-the-art methods in splicing forgery detection and localization on four public datasets. Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001 |
ICME | 4 |
| 2021 | Medical Image Registration Based on Uncoupled Learning and Accumulative Enhancement
Yucheng Shu, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001 |
MICCAI (4) | 3 |
| 2021 | A focus measure in discrete cosine transform domain for multi-focus image fast fusion
Xixi Nie, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2021 | Locally GAN-generated face detection based on an improved Xception
Beijing Chen, Xingwang Ju, Bin Xiao 0002, Weiping Ding 0001, Yuhui Zheng, Victor Hugo C. de Albuquerque |
Inf. Sci. | 3 |
| 2021 | Medical image segmentation based on active fusion-transduction of multi-stream features
Yucheng Shu, Bin Xiao 0002, Weisheng Li 0001 |
Knowl. Based Syst. | 3 |
| 2021 | 2D-LCoLBP: A Learning Two-Dimensional Co-Occurrence Local Binary Pattern for Image RecognitionabstractThe rotation, scale and translation invariance of extracted features have a high significance in image recognition. Local binary pattern (LBP) and LBP-based descriptors have been widely used in image recognition due to feature discrimination and computational efficiency. However, most of the existing LBP-based descriptors have been designed to achieve rotation invariance while fail to achieve scale invariance. Moreover, it is usually difficult to achieve a good trade-off between the feature discrimination and the feature dimension. In this work, a learning 2D co-occurrence LBP termed 2D-LCoLBP is proposed to address these issues. Firstly, a weighted joint histogram is constructed in different neighborhoods and scales of an image to represent the multi-neighborhood and multi-scale LBP (2D-MLBP) and achieve the rotation invariance. A feature learning strategy is then designed to learn the compact and robust descriptor (2D-LCoLBP) from LBP pattern pairs across different scales in the extracted 2D-MLBP to characterize the most stable local structures and achieve the scale invariance, as well as decrease the feature dimension and improve the noise robustness. Finally, a linear SVM classifier is employed for recognition. We applied the proposed 2D-LCoLBP on four image recognition tasks-texture, object, face and food recognition with ten image databases. Experimental results show that 2D-LCoLBP has obviously low feature dimension but outperforms the state-of-the-art LBP-based descriptors in terms of recognition accuracy under noise-free, Gaussian noise and JPEG compression conditions. Xiuli Bi, Yuan Yuan 0015, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Global-Feature Encoding U-Net (GEU-Net) for Multi-Focus Image FusionabstractThe convolutional neural network (CNN)-based multi-focus image fusion methods which learn the focus map from the source images have greatly enhanced fusion performance compared with the traditional methods. However, these methods have not yet reached a satisfactory fusion result, since the convolution operation pays too much attention on the local region and generating the focus map as a local classification (classify each pixel into focus or de-focus classes) problem. In this article, a global-feature encoding U-Net (GEU-Net) is proposed for multi-focus image fusion. In the proposed GEU-Net, the U-Net network is employed for treating the generation of focus map as a global two-class segmentation task, which segments the focused and defocused regions from a global view. For improving the global feature encoding capabilities of U-Net, the global feature pyramid extraction module (GFPE) and global attention connection upsample module (GACU) are introduced to effectively extract and utilize the global semantic and edge information. The perceptual loss is added to the loss function, and a large-scale dataset is constructed for boosting the performance of GEU-Net. Experimental results show that the proposed GEU-Net can achieve superior fusion performance than some state-of-the-art methods in both human visual quality, objective assessment and network complexity. Bin Xiao 0002, Bocheng Xu, Xiuli Bi, Weisheng Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Pyramidal Multiple Instance Detection Network With Mask Guided Self-Correction for Weakly Supervised Object DetectionabstractWeakly supervised object detection has attracted more and more attention as it only needs image-level annotations for training object detectors. A popular solution to this task is to train a multiple instance detection network (MIDN) which integrates multiple instance learning into a deep convolutional neural network. One major issue of the MIDN is that it is prone to be stuck at local discriminative regions. To address this local optimum issue, we propose a pyramidal MIDN (P-MIDN) comprised of a sequence of multiple MIDNs. In particular, one MIDN performs proposal removal for its subsequent MIDN to reduce the exposure of local discriminative proposal regions to the latter during training. In this manner, it allows our MIDNs to focus on proposals which cover objects more completely. Furthermore, we integrate the P-MIDN into an online instance classifier refinement (OICR) framework. Combined with the P-MIDN, a mask guided self-correction (MGSC) method is proposed to generate high-quality pseudo ground-truths for training the OICR. Experimental results on PASCAL VOC 2007, PASCAL VOC 2010, PASCAL VOC 2012, ILSVRC 2013 DET and MS-COCO benchmarks demonstrate that our approach achieves state-of-the-art performance. Yunqiu Xu, Chunluan Zhou, Xin Yu 0002, Bin Xiao 0002, Yi Yang 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Controlling Neural Learning Network with Multiple Scales for Image Splicing Forgery DetectionabstractThe guarantee of social stability comes from many aspects of life, and image information security as one of them is being subjected to various malicious attacks. As a means of information attack, image splicing forgery refers to copying some areas of an image to another image to hide the traces of the original information and leads to grave consequences. Image splicing forgery is extremely complex since the attributes of the two images subjected to the pasting and copying operations are greatly different. In order to solve the issue mentioned above, we propose a method by applying a neural learning network controlled by multiple scales (MCNL-Net) based on U-Net to identify whether an image has been tampered and to locate the tampered regions. Firstly, the learning capacity of MCNL-Net is enhanced by the combination of a residual propagation module and a residual feedback module. An ingenious strategy is designed to control the size of local receptive field in each building block of MCNL-Net. The strategy makes MCNL-Net able to achieve properties and superiorities of multi-scale structure and learn specified features. For further improving the detection performance of MCNL-Net, a block attention mechanism is proposed to control the advanced degree of the input information in each building block. In addition, a MaxBlurPool method is applied into image splicing forgery detection for the first time, preserving the shift-equivariance of a convolutional neural network. Through experiments, we demonstrate that MCNL-Net can achieve more promising results and offer stronger robustness than the state-of-the-art splicing forgery detection methods. Yang Wei 0002, Bin Xiao 0002, Ximeng Liu, Zheng Yan 0002, Jianfeng Ma 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | AFT-Net: Active Fusion-Transduction for Multi-stream Medical Image SegmentationabstractAs an important building block in automatic medical applications, image segmentation has made a great progress due to the data-driving mechanism of deep architecture. Recently, numerous methods have been proposed to boost the segmentation performance based on U-shape network. However, they often built feature encoders with only one data routine, which have limited the representation ability of the networks. Although some methods applied multiple learning paths to fix this problem, the deep supervision techniques are required to monitor the training status at individual path, which may bring extra burden to practical usage of the algorithm. Additionally, under these frameworks, the semantic gap between different paths may interfere with model's learning performance, and the potential transduction ability of skip connections still needs further investigation. To address these issues, we introduce a novel medical image segmentation framework, namely AFT-Net, in which an attention-based data fusion model is proposed to effectively cooperate with multi-stream encoder. By progressively accumulating the features from different paths, our method can establish meaningful connections between structural and semantic features, while keeping an integral and flexible layout without deeply customized supervisions. Extensive experiments on two medical image data sets demonstrate that our method is able to acquire image features with both diversity and quality, thereby outperforms current state-of-the-art segmentation methods. Yucheng Shu, Bin Xiao 0002, Xiao Luan, Linghui Liu, Chunlong Hu |
ICTAI | 3 |
| 2020 | Computer aided Alzheimer's disease diagnosis by an unsupervised deep learning technology
Xiuli Bi, Shutong Li, Bin Xiao 0002, Yu Li 0018, Guoyin Wang 0001 |
Neurocomputing | 3 |
| 2020 | Heart sounds classification using a novel 1-D convolutional neural network with extremely low parameter consumption
Bin Xiao 0002, Yunqiu Xu, Xiuli Bi |
Neurocomputing | 1 |
| 2020 | Follow the Sound of Children's Heart: A Deep-Learning-Based Computer-Aided Pediatric CHDs Diagnosis SystemabstractAuscultation of heart sounds is a noninvasive and less costly way for congenital heart disease (CHD) diagnosis, especially for pediatric individuals. The deep-learning-based computer-aided heart sound analysis has been widely studied and developed in recent years. In this article, we develop a deep-learning-based computer-aided system for pediatric CHDs diagnosis using two novel lightweight convolution neural networks (CNNs). One key issue of most existing deep-learning-based systems is the scarcity of large-scale data sets for CNN learning. To this end, we collect heart sounds from newborns and children with physicians' annotations to construct a pediatric heart sound data set that contains 528 high-quality recordings (nearly 4 h in total) from 137 subjects. With the constructed data set, deep CNN models can be easily trained as classifiers in computer-aided CHDs diagnosis systems. The experimental results demonstrate the superiority of our proposed methods in terms of diagnosis performance and parameter consumption in the application of Internet of Things. Bin Xiao 0002, Yunqiu Xu, Xiuli Bi, Weisheng Li 0001, Zhuo Ma 0001 |
IEEE Internet Things J. | 1 |
| 2020 | Fractional discrete Tchebyshev moments and their applications in image encryption and watermarking
Bin Xiao 0002, Jiangxia Luo, Xiuli Bi, Weisheng Li 0001, Beijing Chen |
Inf. Sci. | 1 |
| 2020 | Image splicing forgery detection combining coarse to refined convolutional neural network and adaptive clustering
Bin Xiao 0002, Yang Wei 0002, Xiuli Bi, Weisheng Li 0001, Jianfeng Ma 0001 |
Inf. Sci. | 1 |
| 2020 | Privacy-Preserving Krawtchouk Moment feature extraction over encrypted image data
Jianfeng Ma 0001, Yinbin Miao, Ximeng Liu, Xuan Wang 0006, Bin Xiao 0002 |
Inf. Sci. | 6 |
| 2020 | Image splicing localization using residual image and residual-based fully convolutional network
Beijing Chen, Xiaoming Qi, Guanyu Yang 0001, Yuhui Zheng, Bin Xiao 0002 |
J. Vis. Commun. Image Represent. | 6 |
| 2020 | Multi-Focus Image Fusion by Hessian Matrix Based DecompositionabstractIn this paper, a Hessian matrix based multi-focus image fusion method is proposed. First, the integral map is introduced for fast compute the Hessian matrix of source images at different scales, and the multi-scale Hessian matrix of source image is obtained. Second, the multi-scale Hessian matrix is used to decompose each source image into two kinds of regions: the feature and background regions. In order to improve the fusion performance, two new focus measures based on the multi-scale Hessian matrix and two different fusion strategies for both feature and background regions are utilized to obtain the initial decision maps, respectively. Finally, the final decision map for image fusion is achieved by post-processing on the results of the previous step. The proposed method is a primary attempt to introduce image feature and background regions decomposition strategies in the field of multi-focus image fusion. The experimental results also show that our method outperforms the existing image fusion methods in both visual perception and objective evaluations. Bin Xiao 0002, Ge Ou, Xiuli Bi, Weisheng Li 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Quaternion weighted spherical Bessel-Fourier moment and its invariant for color image reconstruction and object recognition
Jianfeng Ma 0001, Yinbin Miao, Xuan Wang 0006, Bin Xiao 0002 |
Inf. Sci. | 5 |
| 2019 | 2D-LBP: An Enhanced Local Binary Feature for Texture Image ClassificationabstractThe local binary pattern (LBP) and its variants have shown the effectiveness in texture images classification, face recognition, and other applications. However, most of these LBP methods only focus on the histogram of LBP patterns and ignore the spatial contextual information between LBP patterns. In this paper, we propose a 2D-LBP method which uses a sliding window to count the weighted occurrence number of the rotation invariant uniform LBP pattern pairs to obtain the spatial contextual information. The multi-resolution 2D-LBP features can also be obtained when the radius of 2D-LBP is changed. At last, a two-stage classifier which acts as an ensemble learning step is followed to achieve an accurate classification by combining the predictions on each 2D-LBP with single resolution. Theoretical proof shows that the proposed 2D-LBP is a general framework and can be integrated on other LBP variants to derive new feature extraction methods. Experimental results show that, the proposed method achieves 99.71%, 97.09%, 98.48%, and 49.00% classification accuracy on the public “Brodatz,” “CUReT,” “UIUC,” and “FMD” texture image databases, respectively. Compared with the original LBP and its variants, the proposed method obtains higher classification accuracy under different cases, and simultaneously owns shorter time complexity. Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Brightness and contrast controllable image enhancement based on histogram specification
Bin Xiao 0002, Yanjun Jiang, Weisheng Li 0001, Guoyin Wang 0001 |
Neurocomputing | 1 |
| 2018 | Fusion of anatomical and functional images using parallel saliency features
Jiao Du, Weisheng Li 0001, Bin Xiao 0002 |
Inf. Sci. | 3 |
| 2018 | Pixel convolutional neural network for multi-focus image fusion
Bin Xiao 0002, Weisheng Li 0001, Guoyin Wang 0001 |
Inf. Sci. | 2 |
| 2017 | A 3D polar-radius-moment invariant as a shape circularity measure
Ziping Ma 0001, Jinlin Ma, Bin Xiao 0002, Ke Lu 0002 |
Neurocomputing | 3 |
| 2017 | Image analysis by fractional-order orthogonal moments
Bin Xiao 0002, Linping Li, Yu Li 0018, Weisheng Li 0001, Guoyin Wang 0001 |
Inf. Sci. | 1 |
| 2017 | Anatomical-Functional Image Fusion by Information of Interest in Local Laplacian Filtering DomainabstractA novel method for performing anatomical (MRI)-functional (PET or SPECT) image fusion is presented. The method merges specific feature information from input image signals of a single or multiple medical imaging modalities into a single fused image while preserving more information and generating less distortion. The proposed method uses a local Laplacian filtering based technique realized through a novel multi-scale system architecture. Firstly, the input images are generated in a multi-scale image representation and are processed using local Laplacian filtering. Secondly, at each scale, the decomposed images are combined to produce fused approximate images using a local energy maximum scheme and produce the fused residual images using an information of interest-based scheme. Finally, a fused image is obtained using a reconstruction process that is analogous to that of conventional Laplacian pyramid transform. Experimental results computed using individual multi-scale analysis-based decomposition schemes or fusion rules clearly demonstrate the superiority of the proposed method through subjective observation as well as objective metrics. Furthermore, the proposed method can obtain better performance, compared to the state-of-the-art fusion methods. Jiao Du, Weisheng Li 0001, Bin Xiao 0002 |
IEEE Trans. Image Process. | 3 |
| 2016 | An overview of multi-modal medical image fusion
Jiao Du, Weisheng Li 0001, Ke Lu 0002, Bin Xiao 0002 |
Neurocomputing | 4 |
| 2016 | Union Laplacian pyramid with multiple features for medical image fusion
Jiao Du, Weisheng Li 0001, Bin Xiao 0002, Qamar Nawaz |
Neurocomputing | 3 |
| 2016 | Object recognition based on the Region of Interest and optimal Bag of Words model
Weisheng Li 0001, Bin Xiao 0002, Lifang Zhou |
Neurocomputing | 3 |
| 2016 | Lossless image compression based on integer Discrete Tchebichef Transform
Bin Xiao 0002, Yanhong Zhang, Weisheng Li 0001, Guoyin Wang 0001 |
Neurocomputing | 1 |
| 2016 | Medical image fusion by combining parallel features on multi-scale local extrema scheme
Jiao Du, Weisheng Li 0001, Bin Xiao 0002, Qamar Nawaz |
Knowl. Based Syst. | 3 |
| 2015 | Compact and discriminative representation of Bag-of-Features
Jiangtao Cui, Miaomiao Cui, Bin Xiao 0002, Guangxin Li |
Neurocomputing | 3 |
| 2015 | Moments and moment invariants in the Radon space
Bin Xiao 0002, Jiangtao Cui, Hongxing Qin, Weisheng Li 0001, Guoyin Wang 0001 |
Pattern Recognit. | 1 |
| 2015 | Errata and comments on "Orthogonal moments based on exponent functions: Exponent-Fourier moments"
Bin Xiao 0002, Weisheng Li 0001, Guoyin Wang 0001 |
Pattern Recognit. | 1 |
| 2014 | Radial shifted Legendre moments for image analysis and invariant image recognition
Bin Xiao 0002, Guoyin Wang 0001, Weisheng Li 0001 |
Image Vis. Comput. | 1 |
| 2013 | Generic radial orthogonal moment invariants for invariant image recognition
Bin Xiao 0002, Guoyin Wang 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2012 | Radial Tchebichef moment invariants for image recognition
Bin Xiao 0002, Jianfeng Ma 0001, Jiangtao Cui |
J. Vis. Commun. Image Represent. | 1 |
| 2012 | Combined blur, translation, scale and rotation invariant image recognition by Radon and pseudo-Fourier-Mellin transforms
Bin Xiao 0002, Jianfeng Ma 0001, Jiangtao Cui |
Pattern Recognit. | 1 |
| 2011 | Invariant Image Recognition Using Radial Jacobi Moment InvariantsabstractAs orthogonal moments in the polar coordinate, radial orthogonal moments such as Zernike, pseudo-Zernike and orthogonal Fourier-Mellin moments have been successfully used in the field of pattern recognition. However, the scale and rotation invariant property of these moments has not been studied. In this paper, we present a generic approach based on Jacobi-Fourier moments for scale and rotation invariant analysis of radial orthogonal moments. It provides a fundamental mathematical tool for invariant analysis of the radial orthogonal moments since Jacobi-Fourier moments are the generic expressions of radial orthogonal moments. Experimental results show the efficiency and the robustness to noise of the proposed method for recognition tasks. Bin Xiao 0002, Jianfeng Ma 0001, Jiangtao Cui |
ICIG | 1 |
| 2010 | Rotation invariant analysis and orientation estimation method for texture classification based on Radon transform and correlation analysis
Xuan Wang 0006, Fang-xia Guo, Bin Xiao 0002, Jianfeng Ma 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2010 | Image analysis by Bessel-Fourier moments
Bin Xiao 0002, Jianfeng Ma 0001, Xuan Wang 0006 |
Pattern Recognit. | 1 |
| 2007 | Scaling and rotation invariant analysis approach to object recognition based on Radon and Fourier-Mellin transforms
Xuan Wang 0006, Bin Xiao 0002, Jianfeng Ma 0001, Xiuli Bi |
Pattern Recognit. | 2 |