EDBT 2026 Demo / reviewers in the wild / expert
Xiuli Bi
dblp:92/860 · also Xiu-Li Bi
· DBLP profile ↗
64ranked-venue papers
16as first author
51since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 10 first-author · 35 since 2021Artificial intelligence and machine learning · 27 · 7 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorSecurity and privacy · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSR: Semantic and Spatial Rectification for CLIP-based Weakly Supervised SegmentationabstractIn recent years, Contrastive Language-Image Pretraining (CLIP) has been widely applied to Weakly Supervised Semantic Segmentation (WSSS) tasks due to its powerful cross-modal semantic understanding capabilities. This paper proposes a novel Semantic and Spatial Rectification (SSR) method to address the limitations of existing CLIP-based weakly supervised semantic segmentation approaches: over-activation in non-target foreground regions and background areas. Specifically, at the semantic level, the Cross-Modal Prototype Alignment (CMPA) establishes a contrastive learning mechanism to enforce feature space alignment across modalities, reducing inter-class overlap while enhancing semantic correlations, to rectify over-activation in non-target foreground regions effectively; at the spatial level, the Superpixel-Guided Correction (SGC) leverages superpixel-based spatial priors to precisely filter out interference from non-target regions during affinity propagation, significantly rectifying background over-activation. Extensive experiments on the PASCAL VOC and MS COCO datasets demonstrate that our method outperforms all single-stage approaches, as well as more complex multi-stage approaches, achieving mIoU scores of 79.5% and 50.6%, respectively. Xiuli Bi, Die Xiao, Junchao Fan, Bin Xiao 0002 |
AAAI | 1 |
| 2026 | Clear Nights Ahead: Towards Multi-Weather Nighttime Image RestorationabstractRestoring nighttime images affected by multiple adverse weather conditions is a practical yet under-explored research problem, as multiple weather degradations usually coexist in the real world alongside various lighting effects at night. This paper first explores the challenging multi-weather nighttime image restoration task, where various types of weather degradations are intertwined with flare effects. To support the research, we contribute the AllWeatherNight dataset, featuring large-scale nighttime images with diverse compositional degradations. By employing illumination-aware degradation generation, our dataset significantly enhances the realism of synthetic degradations in nighttime scenes, providing a more reliable benchmark for model training and evaluation. Additionally, we propose ClearNight, a unified nighttime image restoration framework, which effectively removes complex degradations in one go. Specifically, ClearNight extracts Retinex-based dual priors and explicitly guides the network to focus on uneven illumination regions and intrinsic texture contents respectively, thereby enhancing restoration effectiveness in nighttime scenarios. Moreover, to more effectively model the common and unique characteristics of multiple weather degradations, ClearNight performs weather-aware dynamic specificity and commonality collaboration that adaptively allocates optimal sub-networks associated with specific weather types. Comprehensive experiments on both synthetic and real-world images demonstrate the necessity of the AllWeatherNight dataset and the superior performance of ClearNight. Yuetong Liu, Yunqiu Xu, Yang Wei 0002, Xiuli Bi, Bin Xiao 0002 |
AAAI | 4 |
| 2026 | TGDD: Trajectory Guided Dataset Distillation with Balanced DistributionabstractDataset distillation compresses large datasets into compact synthetic ones to reduce storage and computational costs. Among various approaches, distribution matching (DM)-based methods have attracted attention for their high efficiency. However, they often overlook the evolution of feature representations during training, which limits the expressiveness of synthetic data and weakens downstream performance. To address this issue, we propose Trajectory Guided Dataset Distillation (TGDD), which reformulates distribution matching as a dynamic alignment process along the model’s training trajectory. At each training stage, TGDD captures evolving semantics by aligning the feature distribution between the synthetic and original dataset. Meanwhile, it introduces a distribution constraint regularization to reduce class overlap. This design helps synthetic data preserve both semantic diversity and representativeness, improving performance in downstream tasks. Without additional optimization overhead, TGDD achieves a favorable balance between performance and efficiency. Experiments on ten datasets demonstrate that TGDD achieves state-of-the-art performance, notably a 5.0% accuracy gain on high-resolution benchmarks. Fengli Ran, Xiao Pu 0002, Bo Liu 0047, Xiuli Bi, Bin Xiao 0002 |
AAAI | 4 |
| 2026 | Breaking the Generator Barrier: Disentangled Representation for Generalizable AI-Text DetectionabstractAs large language models (LLMs) generate text that increasingly resembles human writing, the subtle cues that distinguish AI-generated content from human-written content become increasingly challenging to capture. Reliance on generator-specific artifacts is inherently unstable, since new models emerge rapidly and reduce the robustness of such shortcuts. This generalizes unseen generators as a central and challenging problem for AI-text detection. To tackle this challenge, we propose a progressively structured framework that disentangles AI-detection semantics from generator-aware artifacts. This is achieved through a compact latent encoding that encourages semantic minimality, followed by perturbation-based regularization to reduce residual entanglement, and finally a discriminative adaptation stage that aligns representations with task objectives. Experiments on MAGE benchmark, covering 20 representative LLMs across 7 categories, demonstrate consistent improvements over state-of-the-art methods, achieving up to 24.2% accuracy gain and 26.2% F_1 improvement. Notably, performance continues to improve as the diversity of training generators increases, confirming strong scalability and generalization in open-set scenarios. Our source code will be publicly available at https://github.com/PuXiao06/DRGD. Xiao Pu 0002, Zepeng Cheng, Lin Yuan 0002, Yu Wu 0001, Xiuli Bi |
ACL (1) | 5 |
| 2026 | Mask-Guided Proxy Mining Network for Few-Shot Medical Image SegmentationabstractFew-shot medical image segmentation (FSMIS) has attracted increasing attention as a promising technique for solving medical image segmentation tasks by relying on only a small amount of labeled data from new classes. Current FSMIS methods typically employ pixel-level semantic correlations between support-query image pairs to guide the segmentation of query images. However, the class information gap between support and query images may induce severe mismatches, leading to semantic ambiguity between foreground and background pixels. To address this issue, we propose a novel mask-guided proxy mining network (MPMNet), which mines a set of representative reference features (termed proxies) from support and query images to rectify foreground-background ambiguity. Specifically, to eliminate false pairwise matches caused by excessive intra-class variations, we design a mask-guided proxy mining module to adaptively learn representative proxies that can perceive visual differences between objects with different scales and shapes. Moreover, we integrate a hierarchical prior generation module and a context-aware feature enrichment module into MPMNet to obtain multi-scale information and enhance the discriminability of features. With these well-designed components and structures, our MPMNet can effectively overcome the adverse effects of false pixel matches by establishing proxy-level semantic correlations. Extensive experiments on three standard medical segmentation benchmarks demonstrate that our MPMNet significantly outperforms previous state-of-the-art methods, with a mean gain of 2.71% in DSC across all datasets. The code is available at: https://github.com/donglongzi/MPMNet. Wendong Huang, Jinwu Hu, Yongchao Wang 0004, Xiuli Bi, Yucheng Shu, Xuezong Yang, Bin Xiao 0002 |
IEEE Trans. Image Process. | 4 |
| 2025 | CustomTTT: Motion and Appearance Customized Video Generation via Test-Time TrainingabstractBenefiting from large-scale pre-training of text-video pairs, current text-to-video (T2V) diffusion models can generate high-quality videos from the text description. Besides, given some reference images or videos, the parameter-efficient fine-tuning method, i.e. LoRA, can generate high-quality customized concepts, e.g., the specific subject or the motions from a reference video. However, combining the trained multiple concepts from different references into a single network shows obvious artifacts. To this end, we propose CustomTTT, where we can joint custom the appearance and the motion of the given video easily. In detail, we first analyze the prompt influence in the current video diffusion model and find the LoRAs are only needed for the specific layers for appearance and motion customization. Besides, since each LoRA is trained individually, we propose a novel test-time training technique to update parameters after combination utilizing the trained customized models. We conduct detailed experiments to verify the effectiveness of the proposed methods. Our method outperforms several state-of-the-art works in both qualitative and quantitative evaluations. Xiuli Bi, Bo Liu 0047, Xiaodong Cun, Yong Zhang 0034, Weisheng Li 0001, Bin Xiao 0002 |
AAAI | 1 |
| 2025 | Towards Universal AI-Generated Image Detection by Variational Information Bottleneck NetworkabstractThe rapid advancement of generative models has significantly improved the quality of generated images. Mean-while, it challenges information authenticity and credibility. Current generated image detection methods based on large-scale pre-trained multimodal models have achieved impressive results. Although these models provide abundant features, the authentication task-related features are often submerged. Consequently, those authentication task-irrelated features cause models to learn superficial biases, thereby harming their generalization performance across different model genera (e.g., GANs and Diffusion Models). To this end, we proposed VIB-Net, which uses Variational Information Bottlenecks to enforce authentication task-related feature learning. We tested and analyzed the proposed method and existing methods on samples generated by 17 different generative models. Compared to SOTA methods, VIB-Net achieved a 5.55% improvement in mAP and a 9.33% increase in accuracy. Notably, in generalization tests on unseen generative models from different series, VIB-Net improved mAP by 12.48% and accuracy by 23.59% over SOTA methods. The code is available at https://github.com/oceanzhf/VIBAIGCDetect. Qinghui He, Xiuli Bi, Weisheng Li 0001, Bo Liu 0047, Bin Xiao 0002 |
CVPR | 3 |
| 2025 | Who Controls the Authorization? Invertible Networks for Copyright Protection in Text-to-Image Synthesis
Baoyue Hu, Yang Wei 0002, Wendong Huang, Xiuli Bi, Bin Xiao 0002 |
ICCV | 5 |
| 2025 | Breaking Grid Constraints: Dynamic Graph Reconstruction Network for Multi-Organ Segmentation
Yang Wei 0002, Xiuli Bi, Bin Xiao 0002 |
ICCV | 5 |
| 2025 | Transfer morphological features for segmentation with few labels on fluorescent mitochondria imagesabstractAbstract Automated segmentation of mitochondria is crucial for statistical analysis in biological research. Existing segmentation techniques often face challenges with fluorescence images. Handcrafted methods have poor segmentation results while deep learning‐based methods lack the labeled mitochondrial data. However, although the number of labeled mitochondrial images is limited, the unlabeled fluorescent data is easy to obtain. The authors aim to leverage a large amount of unlabeled data to learn mitochondrial morphological features. The approach begins with self‐supervised learning from a vast set of unlabeled images through masked image modeling. This technique involves presenting images with randomly masked patches, prompting the model to predict the content of these masked areas. By doing so, the model learns the distinctive features of mitochondria. In the subsequent phase, the trained encoder is transferred to the segmentation task, replacing the original reconstruction decoder with the Segformer segmentation decoder. The model is then fine‐tuned using a small labeled dataset. By reconstructing mitochondria in the masked regions, the model learns features more effectively on unlabeled samples, and improves segmentation performance even with limited labeled data. Empirical results validate the effectiveness of the approach, showing an 11.8% improvement in Intersection over Union metrics compared to existing fluorescence mitochondrial segmentation techniques. Junchao Fan, Xiuli Bi, Weisheng Li 0001, Bin Xiao 0002, Xiaoshuai Huang |
IET Image Process. | 3 |
| 2025 | CS-CoLBP: Cross-Scale Co-occurrence Local Binary Pattern for Image Classification
Bin Xiao 0002, Danyu Shi, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | DEAR: Disentangled Event-Agnostic Representation Learning for Early Fake News DetectionabstractAbstract Detecting fake news early is challenging due to the absence of labeled articles for emerging events in training data. To address this, we propose a Disentangled Event-Agnostic Representation (DEAR) learning approach. Our method begins with a BERT-based adaptive multi-grained semantic encoder that captures hierarchical and comprehensive textual representations of the input news content. To effectively separate latent authenticity-related and event-specific knowledge within the news content, we employ a disentanglement architecture. To further enhance the decoupling effect, we introduce a cross-perturbation mechanism that perturbs authenticity-related representation with the event-specific one, and vice versa, deriving a robust and discerning authenticity-related signal. Additionally, we implement a refinement learning scheme to minimize potential interactions between two decoupled representations, ensuring that the authenticity signal remains strong and unaffected by event-specific details. Experimental results demonstrate that our approach effectively mitigates the impact of event-specific influence, outperforming state-of-the-art methods. In particular, it achieves a 6.0% improvement in accuracy on the PHEME dataset over MDDA, a similar approach that decouples latent content and style knowledge, in scenarios involving articles from unseen events different from the topics of the training set. Xiao Pu 0002, Xiuli Bi, Yu Wu 0001, Xinbo Gao 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | Neurocognitive Insights: Cognitive Comprehension Attention in Multi-Organ SegmentationabstractIn multi-organ segmentation, attention mechanisms are frequently employed to enhance the focus on irregular organs, improving performance. However, current attention mechanisms exhibit notable limitations. On the one hand, their visual saliency-based attention bias results in incomplete region-of-interest coverage. On the other hand, their organ-specific cognitive deficiency exacerbates organ misclassification. Inspired by neurocognitive science, this paper proposes a Cognitive Comprehension Attention (CCA). Diverging from existing methods, CCA achieves refined attention allocation by decomposing visual representations into discrete visual stimuli. This fine-grained approach enables unbiased processing for each visual stimulus, preventing critical information omission and ensuring comprehensive organ region coverage. More importantly, CCA generates organ-specific attention representations by establishing distinct attention patterns across different organ regions, which empowers CCA with cognitive capacity, resolving organ misclassification. Extensive experiments across multiple datasets demonstrate that CCA significantly enhances backbone performance, achieving a max mDice improvement of 8.45% while surpassing state-of-the-art methods by 9% in Recall and 11.78% in Precision. Code is available at:https://github.com/robert1818118/CCA. Yang Wei 0002, Wendong Huang, Xiuli Bi, Xuezong Yang, Bin Xiao 0002 |
IEEE Trans. Big Data | 5 |
| 2025 | TransHFC: Joints Hypergraph Filtering Convolution and Transformer Framework for TemporalForgery LocalizationabstractThe authenticity of audio-visual content is being challenged by advanced multimedia editing technologies inspired by Artificial Intelligence-Generated Content (AIGC). Temporal forgery localization aims to detect suspicious contents by locating forged segments. So far, most of the existing methods are based on Convolutional Neural Networks (CNNs) or Transformers, yet neither of them has fully considered the complex relationships within forged audio-visual content. To address this issue, in this paper, we propose a novel method, named TransHFC, which innovatively introduces hypergraphs to model group relationships among segments while considering point-to-point relationships through Transformers. Through its dual hypergraph filtering convolution branch, TransHFC captures both temporal and spatial level group relationships, enhancing the representation of forged segment features. Furthermore, we propose a new hypergraph filtering convolution Auto-Encoder that uses a multi-frequency filter bank for adaptive signal capture. This design compensates for the limitation of a single hypergraph filter. Our extensive experiments on Lav-DF, TVIL, Psynd, and HAD datasets demonstrate that TransHFC achieves state-of-the-art performance. Xiaochen Yuan, Chan-Tong Lam, Sio Kei Im, Fangyuan Lei, Xiuli Bi |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Let Images Speak More: An Efficient Method for Detecting Image Manipulation HistoryabstractDigital image forensics aims to verify the authenticity of digital images, which has emerged as a prominent research area. To reveal the manipulation history of an image, the existing methods can only detect specific image operations or are based on a general forensic feature with high dimensions. Moreover, these methods perform well only when the operation chain length is no greater than 2. However, their detection accuracy drops significantly for images with longer operation chains that are more representative of real-world scenarios. To break these limitations, we proposed a novel forensics frequency Feature based on Histogram and Detail Map (FHDM(79D)), which can distinguish various operation chains containing different numbers of operations. Specifically, compared to the traces left by image manipulation in the spatial domain, we have discovered that they are more distinct in the frequency domain. This observation has prompted us to extract features from the frequency domain of images by analyzing their histograms and detail maps to capture the manipulation traces of the images. Notably, the proposed feature extracted in the frequency domain has almost 90% fewer dimensions than the commonly used general forensic features, such as SRM(714D), which greatly reduces the computational complexity. Meanwhile, compared to deep learning-based methods, the experiments show that the proposed method achieves a detection accuracy of over 95% for image operations across multiple datasets, while other deep learning-based methods do not exceed 90% accuracy. Extensive experimental results show that the proposed method is more versatile and effective, showing good performance in complex operation chain detection and local forgery detection. The code is available at https://github.com/CherishL-J/Op-detection. Yang Wei 0002, Xiaochen Yuan, Xiuli Bi, Bin Xiao 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Prototype-Guided Graph Reasoning Network for Few-Shot Medical Image SegmentationabstractFew-shot semantic segmentation (FSS) is of tremendous potential for data-scarce scenarios, particularly in medical segmentation tasks with merely a few labeled data. Most of the existing FSS methods typically distinguish query objects with the guidance of support prototypes. However, the variances in appearance and scale between support and query objects from the same anatomical class are often exceedingly considerable in practical clinical scenarios, thus resulting in undesirable query segmentation masks. To tackle the aforementioned challenge, we propose a novel prototype-guided graph reasoning network (PGRNet) to explicitly explore potential contextual relationships in structured query images. Specifically, a prototype-guided graph reasoning module is proposed to perform information interaction on the query graph under the guidance of support prototypes to fully exploit the structural properties of query images to overcome intra-class variances. Moreover, instead of fixed support prototypes, a dynamic prototype generation mechanism is devised to yield a collection of dynamic support prototypes by mining rich contextual information from support images to further boost the efficiency of information interaction between support and query branches. Equipped with the proposed two components, PGRNet can learn abundant contextual representations for query images and is therefore more resilient to object variations. We validate our method on three publicly available medical segmentation datasets, namely CHAOS-T2, MS-CMRSeg, and Synapse. Experiments indicate that the proposed PGRNet outperforms previous FSS methods by a considerable margin and establishes a new state-of-the-art performance. Wendong Huang, Jinwu Hu, Yang Wei 0002, Xiuli Bi, Bin Xiao 0002 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Focus Stacking with High Fidelity and Superior Visual EffectsabstractFocus stacking is a technique in computational photography, and it synthesizes a single all-in-focus image from different focal plane images. It is difficult for previous works to produce a high-quality all-in-focus image that meets two goals: high-fidelity to its source images and good visual effects without defects or abnormalities. This paper proposes a novel method based on optical imaging process analysis and modeling. Based on a foreground segmentation - diffusion elimination architecture, the foreground segmentation makes most of the areas in full-focus images heritage information from the source images to achieve high fidelity; diffusion elimination models the physical imaging process and is specially used to solve the transition region (TR) problem that is a long-term neglected issue and degrades visual effects of synthesized images. Based on extensive experiments on simulated dataset, existing realistic dataset and our proposed BetaFusion dataset, the results show that our proposed method can generate high-quality all-in-focus images by achieving two goals simultaneously, especially can successfully solve the TR problem and eliminate the visual effect degradation of synthesized images caused by the TR problem. Bo Liu 0047, Xiuli Bi, Weisheng Li 0001, Bin Xiao 0002 |
AAAI | 3 |
| 2024 | Depth-Aware Test-Time Training for Zero-Shot Video Object SegmentationabstractZero-shot Video Object Segmentation (ZSVOS) aims at segmenting the primary moving object without any human annotations. Mainstream solutions mainly focus on learning a single model on large-scale video datasets, which struggle to generalize to unseen videos. In this work, we introduce a test-time training (TTT) strategy to address the problem. Our key insight is to enforce the model to predict consistent depth during the TTT process. In detail, we first train a single network to perform both segmentation and depth prediction tasks. This can be effectively learned with our specifically designed depth modulation layer. Then, for the TTT process, the model is updated by predicting consistent depth maps for the same frame under different data augmentations. In addition, we explore different TTT weight updating strategies. Our empirical results suggest that the momentum-based weight initialization and looping-based training scheme lead to more stable improvements. Experiments show that the proposed method achieves clear improvements on ZSVOS. Our proposed video TTT strategy provides significant superiority over state-of-the-art TTT methods. Our code is available at: https://nifangbaage.github.io/DATTT/. Weihuang Liu, Xi Shen 0001, Haolun Li 0001, Xiuli Bi, Bo Liu 0047, Chi-Man Pun, Xiaodong Cun |
CVPR | 4 |
| 2024 | Using My Artistic Style? You Must Obtain My Authorization
Xiuli Bi, Weisheng Li 0001, Bo Liu 0047, Bin Xiao 0002 |
ECCV (86) | 1 |
| 2024 | PriFU: Capturing Task-Relevant Information Without Adversarial LearningabstractAs machine learning advances, machine learning as a service (MLaaS) in the cloud brings convenience to human lives but also privacy risks, as powerful neural networks used for generation, classification or other tasks can also become privacy snoopers. This motivates privacy preservation in the inference phase. Many approaches for preserving privacy in the inference phase introduce multi-objective functions, training models to remove specific private information from users' uploaded data. Although effective, these adversarial learning-based approaches suffer not only from convergence difficulties, but also from limited generalization beyond the specific privacy for which they are trained. To address these issues, we propose a method for privacy preservation in the inference phase by removing task-irrelevant information, which requires no knowledge of the privacy attacks nor introduction of adversarial learning. Specifically, we introduce a metric to distinguish task-irrelevant information from task-relevant information, and achieve more efficient metric estimation to remove task-irrelevant features. The experiments demonstrate the potential of our method in several tasks. Our code will be available at: https://github.com/iwhoyoung/PriFU. Xiuli Bi, Bo Liu 0047, Weisheng Li 0001, Pamela C. Cosman, Bin Xiao 0002 |
ACM Multimedia | 1 |
| 2024 | Anatomical Prior Guided Spatial Contrastive Learning for Few-Shot Medical Image Segmentation
Wendong Huang, Jinwu Hu, Xiuli Bi, Bin Xiao 0002 |
ACM Multimedia | 3 |
| 2024 | D-Net: A dual-encoder network for image splicing forgery detection and localization
Bo Liu 0047, Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001, Guoyin Wang 0001, Xinbo Gao 0001 |
Pattern Recognit. | 3 |
| 2024 | Learning Discriminative Representations From Cross-Scale Features for Camouflaged Object DetectionabstractThe key that hinders the performance improvement of current camouflaged object detection (COD) models is the lack of discriminability of features at fine granularity. We solve this problem from two complementary perspectives. Firstly, complex scenes result in the discriminative feature representations of camouflaged objects being present at different scales and semantic abstraction levels. Therefore, a mechanism is needed to increase the diversity of features to integrate more information potentially beneficial for COD. Second, appearance similarity between objects and environments will inevitably lead to similarity in features. Enhancing feature diversity alone is not enough to solve the above problems. Therefore, it is necessary to give the model semantic perception capabilities to expand the subtle discrepancies between objects and environments in feature embedding. Inspired by the first point, we propose a cross-scale interaction module (CSIM) that utilizes cross-attention between different scales to enhance the diversity of feature representations. Regarding the second point, the semantic guided feature learning (SGFL) is proposed to promote the model to expand feature discrepancies through explicit supervision. Experiments on four popular COD datasets show that our method outperforms recent SOTA methods. In addition, polyp segmentation experiments show that it is also effective for other COD-like tasks. Yongchao Wang 0004, Xiuli Bi, Bo Liu 0047, Yang Wei 0002, Weisheng Li 0001, Bin Xiao 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | CTNet: Contrastive Transformer Network for Polyp SegmentationabstractSegmenting polyps from colonoscopy images is very important in clinical practice since it provides valuable information for colorectal cancer. However, polyp segmentation remains a challenging task as polyps have camouflage properties and vary greatly in size. Although many polyp segmentation methods have been recently proposed and produced remarkable results, most of them cannot yield stable results due to the lack of features with distinguishing properties and those with high-level semantic details. Therefore, we proposed a novel polyp segmentation framework called contrastive Transformer network (CTNet), with three key components of contrastive Transformer backbone, self-multiscale interaction module (SMIM), and collection information module (CIM), which has excellent learning and generalization abilities. The long-range dependence and highly structured feature map space obtained by CTNet through contrastive Transformer can effectively localize polyps with camouflage properties. CTNet benefits from the multiscale information and high-resolution feature maps with high-level semantic obtained by SMIM and CIM, respectively, and thus can obtain accurate segmentation results for polyps of different sizes. Without bells and whistles, CTNet yields significant gains of 2.3%, 3.7%, 3.7%, 18.2%, and 10.1% over classical method PraNet on Kvasir-SEG, CVC-ClinicDB, Endoscene, ETIS-LaribPolypDB, and CVC-ColonDB respectively. In addition, CTNet has advantages in camouflaged object detection and defect detection. The code is available at https://github.com/Fhujinwu/CTNet. Bin Xiao 0002, Jinwu Hu, Weisheng Li 0001, Chi-Man Pun, Xiuli Bi |
IEEE Trans. Cybern. | 5 |
| 2024 | Effectively Improving Data Diversity of Substitute Training for Data-Free Black-Box AttackabstractRecent substitute training methods have utilized the concept of Generative Adversarial Networks (GANs) to implement data-free black-box attacks. Specifically, in designing the generators, the substitute training methods use a similar structure to the generators in GANs. However, this design approach ignores the potential situation that the generators in GANs operate under real data supervision, while the generators in substitute training methods lack such supervision. This difference in data-supervised conditions constrain the diversity of data generated by the substitute training methods, resulting in inadequate data to support effective training of the substitute model. This impacts the substitute model's ability to attack the target model further. Consequently, to solve the above issues, we propose three strategies to improve the attack success rates. For the generator, we first propose a dense projection space that projects the input noise into various latent feature spaces to diversify feature information. Then, we introduce a novel disguised natural color mode. This mode improves information exchange between the generator's output layer and previous layers, allowing for more diverse generated data. Besides, we present a regularization method for the substitute model, called noise-based balanced learning, to prevent the potential risk of overfitting due to the lack of diversity of the generated data. In the experimental analysis, extensive experiments are conducted to validate the effectiveness of these proposed strategies. Yang Wei 0002, Zhuo Ma 0001, Zhuoran Ma 0002, Zhan Qin, Yang Liu 0118, Bin Xiao 0002, Xiuli Bi, Jianfeng Ma 0001 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2024 | A Creative Weak Supervised Semantic Segmentation for Remote Sensing ImagesabstractIn weakly supervised semantic segmentation (WSSS) tasks on remote sensing images, it is a common practice to train a classification network from scratch using a large batch of images with a limited number of classes. Subsequently, class activation maps are extracted from the model based on predefined class indices, and these maps are then optimized to obtain pseudolabels. To make this strategy effective when introducing a new class, a substantial amount of data needs to be provided to the model. In this article, we present an innovative framework, RS-TextWS-Seg, designed to efficiently generate high-quality segmentation results for a wide range of remote sensing objects using concise descriptions. Our proposed framework comprises three sequential stages: initially, we undertake parameter fine-tuning of the contrastive language-image pretraining (CLIP) model to swiftly strengthen its capacity for zero-shot detection of a limited number of remote sensing features. Subsequently, we introduce a text-driven background suppression mechanism aimed at deriving class activation maps from the refined CLIP model based on textual cues, while concurrently mitigating background noises. Finally, we use the segment anything model (SAM) to refine the edges of the extracted class activation map. We widely researched the leading-edge methodologies in WSSS and conducted a range of comparative experiments and ablation studies to prove the efficacy of our proposed framework. The research findings underscore that RS-TextWS-Seg outperforms other state-of-the-art methods on renowned datasets such as DLRSD and Potsdam, as well as on bespoke datasets specifically curated for overground petroleum pipelines and oil well fields. Zhibao Wang, Lu Bai 0006, Liangfu Chen, Xiuli Bi |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Boundary-Aware Prototype in Semi-Supervised Medical Image SegmentationabstractThe true label plays an important role in semi-supervised medical image segmentation (SSMIS) because it can provide the most accurate supervision information when the label is limited. The popular SSMIS method trains labeled and unlabeled data separately, and the unlabeled data cannot be directly supervised by the true label. This limits the contribution of labels to model training. Is there an interactive mechanism that can break the separation between two types of data training to maximize the utilization of true labels? Inspired by this, we propose a novel consistency learning framework based on the non-parametric distance metric of boundary-aware prototypes to alleviate this problem. This method combines CNN-based linear classification and nearest neighbor-based non-parametric classification into one framework, encouraging the two segmentation paradigms to have similar predictions for the same input. More importantly, the prototype can be clustered from both labeled and unlabeled data features so that it can be seen as a bridge for interactive training between labeled and unlabeled data. When the prototype-based prediction is supervised by the true label, the supervisory signal can simultaneously affect the feature extraction process of both data. In addition, boundary-aware prototypes can explicitly model the differences in boundaries and centers of adjacent categories, so pixel-prototype contrastive learning is introduced to further improve the discriminability of features and make them more suitable for non-parametric distance measurement. Experiments show that although our method uses a modified lightweight UNet as the backbone, it outperforms the comparison method using a 3D VNet with more parameters. Yongchao Wang 0004, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Self-Supervised Image Local Forgery Detection by JPEG Compression TraceabstractFor image local forgery detection, the existing methods require a large amount of labeled data for training, and most of them cannot detect multiple types of forgery simultaneously. In this paper, we firstly analyzed the JPEG compression traces which are mainly caused by different JPEG compression chains, and designed a trace extractor to learn such traces. Then, we utilized the trace extractor as the backbone and trained self-supervised to strengthen the discrimination ability of learned traces. With its benefits, regions with different JPEG compression chains can easily be distinguished within a forged image. Furthermore, our method does not rely on a large amount of training data, and even does not require any forged images for training. Experiments show that the proposed method can detect image local forgery on different datasets without re-training, and keep stable performance over various types of image local forgery. Xiuli Bi, Wuqing Yan, Bo Liu 0047, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
AAAI | 1 |
| 2023 | Location-Aware Transformer Network for Few-Shot Medical Image SegmentationabstractAutomatic and precise organ segmentation plays a significant role in promoting the development of the diagnosis and treatment of the disease. Despite making enormous strides in medical image segmentation, conventional deep neural network-based methods are inherently massive data-driven techniques and are challenging to adapt to novel classes with a small number of labeled samples. Few-shot learning is a promising solution through learning novel classes from extremely limited annotated examples. However, existing few-shot segmentation methods focus excessively on targets in individual images while neglecting to model the global spatial correlation across images, which may cause severe performance degradation. To solve this issue, we propose a new Transformer-based few-shot segmentation framework for medical imaging, namely location-aware transformer network (LATNet), which establishes the spatial correlation between support and query objects, yielding location-aware prototypes, and then performs segmentation by computing the semantic similarity between query features and location-aware prototypes. Additionally, to further enhance the representativeness of the obtained location-aware prototypes in low-data regimes, we design a prediction iterative refinement module, which can iteratively exploit the query predictions output by each iteration to update the location-aware prototypes and progressively refine the query predictions. Extensive experiments on three challenging medical image datasets, i.e., Abd-MRI, Card-MRI, and Abd-CT, show that the proposed LATNet achieves remarkable improvements over current state-of-the-art methods by an average of 4.17%, 1.50%, and 4.63% in terms of the Dice Score, respectively. Wendong Huang, Bin Xiao 0002, Jinwu Hu, Xiuli Bi |
BIBM | 4 |
| 2023 | MCF: Mutual Correction Framework for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning is a promising method for medical image segmentation under limited annotation. However, the model cognitive bias impairs the segmentation performance, especially for edge regions. Furthermore, current mainstream semi-supervised medical image segmentation (SSMIS) methods lack designs to handle model bias. The neural network has a strong learning ability, but the cognitive bias will gradually deepen during the training, and it is difficult to correct itself. We propose a novel mutual correction framework (MCF) to explore network bias correction and improve the performance of SSMIS. Inspired by the plain contrast idea, MCF introduces two different subnets to explore and utilize the discrepancies between subnets to correct cognitive bias of the model. More concretely, a contrastive difference review (CDR) module is proposed to find out inconsistent prediction regions and perform a review training. Additionally, a dynamic competitive pseudo-label generation (DCPLG) module is proposed to evaluate the performance of subnets in real-time, dynamically selecting more reliable pseudo-labels. Experimental results on two medical image databases with different modalities (CT and MRI) show that our method achieves superior performance compared to several state-of-the-art methods. The code will be available at https://github.com/WYC-321/MCF. Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
CVPR | 3 |
| 2023 | DLBD: A Self-Supervised Direct-Learned Binary DescriptorabstractFor learning-based binary descriptors, the binarization process has not been well addressed. The reason is that the binarization blocks gradient back-propagation. Existing learning-based binary descriptors learn real-valued output, and then it is converted to binary descriptors by their proposed binarization processes. Since their binarizaiion processes are not a component of the network, the learning-based binary descriptor cannot fully utilize the advances of deep learning. To solve this issue, we propose a model-agnostic plugin binary transformation layer (BTL), making the network directly generate binary descriptors. Then, we present the first self-supervised, direct-learned binary descriptor, dubbed DLBD. Furthermore, we propose ultra-wide temperature-scaled crossentropy loss to adjust the distribution of learned descriptors in a larger range. Experiments demonstrate that the proposed BTL can substitute the previous binarization process. Our proposed DLBD outperforms SOTA on different tasks such as image retrieval and classification11Our code is available at: https://github.com/CQUPT-CV/DLBD. Bin Xiao 0002, Bo Liu 0047, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
CVPR | 4 |
| 2023 | Cross-slice Context Consistency for Semi-supervised 3D Left Atrium SegmentationabstractSemi-supervised learning is a promising approach in reducing the requirement to collect large amounts of dense annotations, especially in medical image segmentation. However, most existing semi-supervised 3D medical image segmentation methods tend to ignore the cross-slice context that contains extensive structural information. We believe cross-slice context can help the model capture semantic information complementary to slice context and achieve robust and more accurate segmentation. Therefore, in this paper, we propose a novel cross-slice context consistency framework for 3D left atrium segmentation named CSC2-Net. Our method can effectively utilize unlabeled data by encouraging consistent results between slice segmentation and cross-slice inference segmentation. To achieve this, we design a bidirectional gated context inference module (Bi-GCM) to model cross-slice context and predict slice segmentation without direct slice features. Experiments on a public left atrium (LA) databases show that our method achieves higher performance and outperforms state-of-the-art methods by imposing cross-slice consistency constraint. Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
ICME | 3 |
| 2023 | Secondary Labeling: A Novel Labeling Strategy for Image Manipulation DetectionabstractImage manipulation detection methods typically rely on a binary annotation called Primary Labeling (PrLa) to identify tampered and authentic regions in a tampered image. However, PrLa only focuses on the difference between authentic and tampered regions, ignoring the distinctions among tampered regions in different images. This transforms the task of image manipulation detection into salient object detection, with the goal shifting towards identifying the most attention-grabbing objects in images. To address this issue, this paper proposes a novel labeling strategy called Secondary Labeling (SeLa). SeLa generates a query table containing multiple tampered categories and randomly reassigns these tampered classes to different types of tampered data, effectively improving the detection performance of models by refocusing the differences among the various data. Additionally, to further improve the detection performance, this paper introduces an Adaptive Label Smoothing (ALS) regularization method. This method addresses the loss of correlation among tampered classes in SeLa caused by the one-hot encoding method. Experimental results show that compared with PrLa, SeLa not only improves the performance of detection models by up to 17%, but also enhances the robustness and convergence rate. Yang Wei 0002, Bin Xiao 0002, Xiuli Bi, Zhuoran Ma 0002, Yang Liu 0118, Zhuo Ma 0001 |
ACM Multimedia | 3 |
| 2023 | A Dual Self-Calibrating Framework for Noninvasive Fetal ECG R-Peak DetectionabstractFetal heart rate (fHR) is critical for assessing fetal health and diagnosing disorders, such as fetal distress, congenital heart disease, and intrauterine growth retardation. With the rapid development of the Internet of Medical Things (IoMT), fetal R-peak detection plays an important role in diagnosing heart defects during pregnancy. However, due to the nonlinear mixing of multiple sources in the noninvasive signals and the low signal-to-noise ratio (SNR), it is difficult to obtain accurate R-peak detection result. This article presents a dual self-calibrating system based on a spectral attention kernel independent component analysis (SA-KICA) module and a self-calibrating fetal R-peak detection (SC-FRD) module. SA-KICA is an ICA-based calibration module constructed by the spectral attention mechanism, which was sought from short-time Fourier transform (STFT) and was shipped back to original signal with convolution to achieve perfect maternal electrocardiogram (MECG) separation in high-dimensional linear separable space. Then, a periodic and morphological-based channel selector is designed to select the optimal MECG. After MECG removal, to further improve the performance of fetal R-peak detection, the SC-FRD module is introduced to utilize the interior peak information and self-calibrating strategy, which includes variance-based fetal R-peak seed selection, time-varying coarse prediction, and adaptive probability mask calibration. The proposed framework is a primary attempt to concurrently introduce the nonlinear feature, spectral information, and self-calibrating strategy in the field of fetal ECG processing. The framework achieved excellent performance in fetal R-peak detection accuracy on a simulated data set and two public data sets with varying divergence and richness of resources. The experimental results show that our framework is superior to existing methods and can be used as a potential fetal monitoring method in the application of IoMT. The code is released inhttps://github.com/bfyjr/NI-FECG-Extraction. Lihong Qiao, Shuai Hu, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Internet Things J. | 4 |
| 2023 | IEMask R-CNN: Information-Enhanced Mask R-CNNabstractThe instance segmentation task is relatively difficult in computer vision, which requires not only high-quality masks but also high-accuracy instance category classification. Mask R-CNN has been proven to be a feasible method. However, due to the Feature Pyramid Network (FPN) structure lack useful channel information, global information and low-level texture information, and mask branch cannot obtain useful local-global information, Mask R-CNN is prevented from obtaining high-quality masks and high-accuracy instance category classification. Therefore, we proposed the Information-enhanced Mask R-CNN, called IEMask R-CNN. In the FPN structure of IEMask R-CNN, the information-enhanced FPN will enhance the useful channel information and the global information of the feature maps to solve the issues that the high-level feature map loses useful channel information and inaccurate of instance category classification, meanwhile the bottom-up path enhancement with adaptive feature fusion will ultilize the precise positioning signal in the lower layer to enhance the feature pyramid. In the mask branch of IEMask R-CNN, an encoding-decoding mask head will strength local-global information to gain a high-quality mask. Without bells and whistles, IEMask R-CNN gains significant gains of about 2.60%, 4.00%, 3.17% over Mask R-CNN on MS COCO2017, Cityscapes and LVIS1.0 benchmarks respectively. Xiuli Bi, Jinwu Hu, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Big Data | 1 |
| 2023 | A Versatile Detection Method for Various Contrast Enhancement ManipulationsabstractContrast enhancement manipulation is a common method to improve the visual effect of an image. Meanwhile, it can also be considered a type of global image forgery because it changes the image’s visual appearance without alerting its semantics. Moreover, for local image forgery, a tampered image may be composited by images with different contrast enhancement manipulations or post-processed by a contrast enhancement manipulation to conceal the trails of tampering. Therefore, contrast enhancement manipulation detection is critical to global image forgery detection. The existing methods can only detect a particular type of contrast enhancement manipulation, such as gamma correction or histogram equalization. To break this limitation, we propose the zero-gap spans (ZGS) as the fingerprint to explore the traces of contrast enhancement manipulations. Based on ZGS, various contrast enhancement manipulations can be distinguished by a simple classification method at image-level and patch-level; different gamma corrections can be identified, and their gamma value can be estimated. Experimental results indicate that the proposed ZGS-based classification method can achieve and maintain good classification performance under different cases (gamma correction, simple histogram equalization, modified histogram equalization techniques). Meanwhile, ZGS can estimate the gamma value with the mean squared error (MSE) below 0.1156. For the local forgery images, the proposed ZGS also can be utilized to locate the regions with different contrast enhancement manipulations. Xiuli Bi, Yixuan Shang, Bo Liu 0047, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | HS-Vectors: Heart Sound Embeddings for Abnormal Heart Sound Detection Based on Time-Compressed and Frequency-Expanded TDNN With Dynamic Mask EncoderabstractIn recent years, auxiliary diagnosis technology for cardiovascular disease based on abnormal heart sound detection has become a research hotspot. Heart sound signals are promising in the preliminary diagnosis of cardiovascular diseases. Previous studies have focused on capturing the local characteristics of heart sounds. In this paper, we investigate a method for mapping heart sound signals with complex patterns to fixed-length feature embedding called HS-Vectors for abnormal heart sound detection. To get the full embedding of the complex heart sound, HS-Vectors are obtained through the Time-Compressed and Frequency-Expanded Time-Delay Neural Network(TCFE-TDNN) and the Dynamic Masked-Attention (DMA) module. HS-Vectors extract and utilize the global and critical heart sound characteristics by masking out irreverent information. Based on the TCFE-TDNN module, the heart sound signal within a certain time is projected into fixed-length embedding. Then, with a learnable mask attention matrix, DMA stats pooling aggregates multi-scale hidden features from different TCFE-TDNN layers and masks out irrelevant frame-level features. Experimental evaluations are performed on a 10-fold cross-validation task using the 2016 PhysioNet/CinC Challenge dataset and the new publicly available pediatric heart sound dataset we collected. Experimental results demonstrate that the proposed method excels the state-of-the-art models in abnormality detection. Lihong Qiao, Yonghao Gao, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Cross-Mix Monitoring for Medical Image Segmentation With Limited SupervisionabstractImage segmentation is a fundamental building block of automatic medical applications. It has been greatly improved since the emergence of deep neural networks. However, deep-learning based models often require a large number of manual annotations, which has seriously hindered its practical usage. To alleviate this problem, numerous works were proposed by utilizing unlabeled data based on semi-supervised frameworks. Recently, the Mean-Teacher (MT) model has been successfully applied in many scenarios due to its effective learning strategy. Nevertheless, the existing MT model still have certain limitations. Firstly, various sorts of perturbations are often added to the training data to gain extra generalization ability through consistency training. However, if the variation is too weak, it may cause the Lazy Student Phenomenon, and bring large fluctuations to the learning model. On the contrary, large image perturbations may enlarge the performance gap between the teacher and student. In this case, the student may lose its learning momentum, and more seriously, drag down the overall performance of the whole system. In order to address these issues, we introduce a novel semi-supervised medical image segmentation framework, in which a Cross-Mix Teaching paradigm is proposed to provide extra data flexibility, thus effectively avoid Lazy Student Phenomenon. Moreover, a lightweight Transductive Monitor is applied to server as the bridge that connect the teacher and student for active knowledge distillation. In the light of this cross-network information mixing and transfer mechanism, our method is able to continuously explore the discriminative information contained in unlabeled data. Extensive experiments on challenging medical image data sets demonstrate that our method is able to outperform current state-of-the-art semi-supervised segmentation methods under severe lack of supervision. Yucheng Shu, Hengbo Li, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | FF-Net: An End-to-end Feature-Fusion Network for Double JPEG Detection and Localization
Bo Liu 0047, Ranglei Wu, Xiuli Bi, Bin Xiao 0002 |
ACML | 3 |
| 2022 | Detecting Generated Images by Real Images
Bo Liu 0047, Fan Yang 0159, Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
ECCV (14) | 3 |
| 2022 | Context Correlation Aware Network for Cardiac SegmentationabstractAutomatically segmenting the anatomical structure of the heart from the cardiac magnetic resonance (CMR) images offers a great potential to augment the traditional healthcare strategy for the quantitative analysis of cardiac contractile function. Most of the existing CNN-based methods for cardiac segmentation tend to ignore the misalignment issues during the feature aggregation process and not fully use multi-scale context and contour information, which may lead to the unexpected misclassification caused by the falsely aligned contextual features and the discontinuity in the edge of segmentation maps. To resolve these issues, we proposed a context correlation aware network (CCA-Net). In CCA-Net, a volume correlation flow module was designed to align contour features and semantic features from adjacent levels, which offered the guidance to wrap low-resolution semantic features into high-resolution features. Besides, a residual gated squeeze module was utilized to explicitly model the boundaries and enhance the representations. Extensive experiments on the multi-sequence cardiac magnetic resonance segmentation challenge (MS-CMRSeg 2019) dataset and MICCAI challenge 2017 automatic cardiac diagnosis challenge (ACDC) dataset demonstrated that CCA-Net was superior to other state-of-the-art methods. Junchao Fan, Jiawei Pei, Xiuli Bi, Bin Xiao 0002, Pietro Liò |
ICME | 3 |
| 2022 | Partial-to-Partial Point Cloud Registration Based on Multi-Level Semantic-Structural Cognitionabstract3D point cloud registration attempt to establish spatial correspondences between the source point cloud and the target point cloud. It is a fundamental task in computer vision and multimedia applications. Recently, many learning-based methods have been proposed and achieved promising performance. However, in the Partial-to-Partial (PtP) registration problem, the existence of a large number of external points may greatly handicap the effectiveness of these methods. In this paper, we propose to address the PtP issue under a novel multi-task cognition framework. At the global semantic level, we introduce a multi-scale feature exchanging network to actively evaluate the matching credibility. For the local structural learning, an inner-inter attention fusion branch is applied to generate discriminative features. Moreover, we integrate a novel alternating correspondences searching mechanism with a flexible bi-directional dislocation loss to perform robust learning, and a simple yet effective SVD weighting scheme is introduced at the inference stage. Experiment results on two challenging PtP 3D point cloud registration data sets show that our proposed method outperforms all the SOTA methods with higher precision and robustness. Yucheng Shu, Zongzhuang Hou, Bin Xiao 0002, Xiuli Bi, Xiao Luan, Weisheng Li 0001 |
ICME | 4 |
| 2022 | HessHist: A Hessian-matrix weighted histogram for image contrast enhancementabstractAbstract For image contrast enhancement operation, it is a keypoint to obtain more natural enhanced results and keep more details without distortion. In this paper, a novel image Hessian‐matrix weighted histogram for image contrast enhancement is proposed, which can improve the contrast of smooth regions and simultaneously restrain the contrast of texture regions. In the proposed method, the multi‐scale fractional‐order Hessian‐matrix is firstly utilized to detect and quantify the texture information of the input image, which explores the regions that should be contrasted or should be restrained. Then, the strong texture regions are suppressed by a designed suppress function. Finally, the information on unsuppressed regions and suppressed texture regions will be count by a histogram, which is termed as Hessian‐matrix weighted Histogram (HessHist) in this paper. According to HessHist, the corresponding cumulative distribution function will realize the contrast enhancement operation on the input image. For real‐time application, the integral images are introduced for fast computation of the HessHist. Experimental results show that the proposed HessHist‐based image enhancement algorithm preserves more details of input image without distortion, and is competitive with state‐of‐the‐art image enhancement algorithms in both subjective visual perception and objective evaluation metrics. Junchao Fan, Xuyang Zong, Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001 |
IET Image Process. | 4 |
| 2022 | Privacy-Preserving Color Image Feature Extraction by Quaternion Discrete Orthogonal MomentsabstractTo implement image storage and computation in cloud servers without violating users’ privacy, privacy-preserving feature extraction has been a new research interest. The existing works are mainly designed for grayscale images. For color images, they tend to convert them to grayscale images or obtain the results of the combination of single-channel processes. While the capabilities of features extracted from the encrypted color images will be affected if color information and inter-relationship between color channels are ignored. To fully preserve features of color images, we introduce quaternion theory to encode each color image and propose an improved vector homomorphic encryption scheme (IVHE) to encrypt quaternion-based color images. IVHE helps protect image content and keep vector characteristics of color images. Based on IVHE, the framework for feature extraction of privacy-preserving Quaternion Discrete Orthogonal Moments (PPQDOMs) is presented. Theoretical analyses prove that Quaternion Discrete Orthogonal Moments (QDOMs) can be extracted from the encrypted color images by PPQDOMs. Furthermore, we apply three common Discrete Orthogonal Moments to the proposed framework to evaluate its performance. Experimental results demonstrate that the proposed framework can protect color image content and perform well compared to QDOMs in image reconstruction and image recognition. Xiuli Bi, Chao Shuai, Bo Liu 0047, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | DTMNet: A Discrete Tchebichef Moments-based Deep Neural Network for Multi-focus Image FusionabstractCompared with traditional methods, the deep learning-based multi-focus image fusion methods can effectively improve the performance of image fusion tasks. However, the existing deep learning-based methods encounter a common issue of a large number of parameters, which leads to the deep learning models with high time complexity and low fusion efficiency. To address this issue, we propose a novel discrete Tchebichef moment-based Deep neural network, termed as DTMNet, for multi-focus image fusion. The proposed DTMNet is an end-to-end deep neural network with only one convolutional layer and three fully connected layers. The convolutional layer is fixed with DTM co-efficients (DTMConv) to extract high/low-frequency information without learning parameters effectively. The three fully connected layers have learnable parameters for feature classification. Therefore, the proposed DTMNet for multi-focus image fusion has a small number of parameters (0.01M paras vs. 4.93M paras of regular CNN) and high computational efficiency (0.32s vs. 79.09s by regular CNN to fuse an image). In addition, a large-scale multi-focus image dataset is synthesized for training and verifying the deep learning model. Experimental results on three public datasets demonstrate that the proposed method is competitive with or even outperforms the state-of-the-art multi-focus image fusion methods in terms of subjective visual perception and objective evaluation metrics. Bin Xiao 0002, Haifeng Wu, Xiuli Bi |
ICCV | 3 |
| 2021 | Reality Transform Adversarial Generators for Image Splicing Forgery Detection and LocalizationabstractWhen many forgery images become more and more realistic with help of image editing tools and convolutional neural networks (CNNs), authenticators need to improve their ability to verify these forgery images. The process of generating and detecting forgery images is the same as the principle of Generative Adversarial Networks (GANs). In this paper, since the retouching progress of forgery images requires to suppress the tampering artifacts and to keep the structural information, we consider this retouching progress as an image style transform, and then propose a fake-to-realistic transform generator GT. For detecting the tampered regions, a localization generator GMis proposed too, which is based on a multi-decoder-single-task strategy. By adversarial training two generators, the proposed α-learnable whitening and coloring transform α-learnable WCT) block in GTautomatically suppress the tampering artifacts in the forgery images. Meanwhile, the detection and localization abilities of GMwill be improved by learning the forgery images retouched by GT. The experiment results demonstrate that the proposed two generators in GAN can simulate confrontation between the faker and the authenticator well; the localization generator GMoutperforms the state-of-the-art methods in splicing forgery detection and localization on four public datasets. Xiuli Bi, Bin Xiao 0002 |
ICCV | 1 |
| 2021 | Multi-Task Wavelet Corrected Network for Image Splicing Forgery Detection and LocalizationabstractAlthough the existing image splicing forgery detection networks can achieve a promising performance, most of these networks utilize regular pooling operations (max-pooling and mean-pooling) and a single task strategy, which limits the comprehensiveness and representativeness of the features learned by the networks. In this paper, we propose a multi-task wavelet corrected network (MWC-Net) that can learn more comprehensive and representative features for image splicing forgery detection and localization. MWC-Net exploits wavelet-pooling and wavelet un-pooling to compress and reconstruct the features of splicing forgery images, which can reduce information loss during learning features. Mean-while, MWC-Net implements a multi-task strategy to improve its ability to learn and utilize more comprehensive and representative features. The experimental results demonstrate that MWC-Net outperforms the state-of-the-art methods in splicing forgery detection and localization on four public datasets. Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001 |
ICME | 1 |
| 2021 | Medical Image Registration Based on Uncoupled Learning and Accumulative Enhancement
Yucheng Shu, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001 |
MICCAI (4) | 4 |
| 2021 | A focus measure in discrete cosine transform domain for multi-focus image fast fusion
Xixi Nie, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
Neurocomputing | 3 |
| 2021 | 2D-LCoLBP: A Learning Two-Dimensional Co-Occurrence Local Binary Pattern for Image RecognitionabstractThe rotation, scale and translation invariance of extracted features have a high significance in image recognition. Local binary pattern (LBP) and LBP-based descriptors have been widely used in image recognition due to feature discrimination and computational efficiency. However, most of the existing LBP-based descriptors have been designed to achieve rotation invariance while fail to achieve scale invariance. Moreover, it is usually difficult to achieve a good trade-off between the feature discrimination and the feature dimension. In this work, a learning 2D co-occurrence LBP termed 2D-LCoLBP is proposed to address these issues. Firstly, a weighted joint histogram is constructed in different neighborhoods and scales of an image to represent the multi-neighborhood and multi-scale LBP (2D-MLBP) and achieve the rotation invariance. A feature learning strategy is then designed to learn the compact and robust descriptor (2D-LCoLBP) from LBP pattern pairs across different scales in the extracted 2D-MLBP to characterize the most stable local structures and achieve the scale invariance, as well as decrease the feature dimension and improve the noise robustness. Finally, a linear SVM classifier is employed for recognition. We applied the proposed 2D-LCoLBP on four image recognition tasks-texture, object, face and food recognition with ten image databases. Experimental results show that 2D-LCoLBP has obviously low feature dimension but outperforms the state-of-the-art LBP-based descriptors in terms of recognition accuracy under noise-free, Gaussian noise and JPEG compression conditions. Xiuli Bi, Yuan Yuan 0015, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Global-Feature Encoding U-Net (GEU-Net) for Multi-Focus Image FusionabstractThe convolutional neural network (CNN)-based multi-focus image fusion methods which learn the focus map from the source images have greatly enhanced fusion performance compared with the traditional methods. However, these methods have not yet reached a satisfactory fusion result, since the convolution operation pays too much attention on the local region and generating the focus map as a local classification (classify each pixel into focus or de-focus classes) problem. In this article, a global-feature encoding U-Net (GEU-Net) is proposed for multi-focus image fusion. In the proposed GEU-Net, the U-Net network is employed for treating the generation of focus map as a global two-class segmentation task, which segments the focused and defocused regions from a global view. For improving the global feature encoding capabilities of U-Net, the global feature pyramid extraction module (GFPE) and global attention connection upsample module (GACU) are introduced to effectively extract and utilize the global semantic and edge information. The perceptual loss is added to the loss function, and a large-scale dataset is constructed for boosting the performance of GEU-Net. Experimental results show that the proposed GEU-Net can achieve superior fusion performance than some state-of-the-art methods in both human visual quality, objective assessment and network complexity. Bin Xiao 0002, Bocheng Xu, Xiuli Bi, Weisheng Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Computer aided Alzheimer's disease diagnosis by an unsupervised deep learning technology
Xiuli Bi, Shutong Li, Bin Xiao 0002, Yu Li 0018, Guoyin Wang 0001 |
Neurocomputing | 1 |
| 2020 | Heart sounds classification using a novel 1-D convolutional neural network with extremely low parameter consumption
Bin Xiao 0002, Yunqiu Xu, Xiuli Bi |
Neurocomputing | 3 |
| 2020 | Follow the Sound of Children's Heart: A Deep-Learning-Based Computer-Aided Pediatric CHDs Diagnosis SystemabstractAuscultation of heart sounds is a noninvasive and less costly way for congenital heart disease (CHD) diagnosis, especially for pediatric individuals. The deep-learning-based computer-aided heart sound analysis has been widely studied and developed in recent years. In this article, we develop a deep-learning-based computer-aided system for pediatric CHDs diagnosis using two novel lightweight convolution neural networks (CNNs). One key issue of most existing deep-learning-based systems is the scarcity of large-scale data sets for CNN learning. To this end, we collect heart sounds from newborns and children with physicians' annotations to construct a pediatric heart sound data set that contains 528 high-quality recordings (nearly 4 h in total) from 137 subjects. With the constructed data set, deep CNN models can be easily trained as classifiers in computer-aided CHDs diagnosis systems. The experimental results demonstrate the superiority of our proposed methods in terms of diagnosis performance and parameter consumption in the application of Internet of Things. Bin Xiao 0002, Yunqiu Xu, Xiuli Bi, Weisheng Li 0001, Zhuo Ma 0001 |
IEEE Internet Things J. | 3 |
| 2020 | Fractional discrete Tchebyshev moments and their applications in image encryption and watermarking
Bin Xiao 0002, Jiangxia Luo, Xiuli Bi, Weisheng Li 0001, Beijing Chen |
Inf. Sci. | 3 |
| 2020 | Image splicing forgery detection combining coarse to refined convolutional neural network and adaptive clustering
Bin Xiao 0002, Yang Wei 0002, Xiuli Bi, Weisheng Li 0001, Jianfeng Ma 0001 |
Inf. Sci. | 3 |
| 2020 | Multi-Focus Image Fusion by Hessian Matrix Based DecompositionabstractIn this paper, a Hessian matrix based multi-focus image fusion method is proposed. First, the integral map is introduced for fast compute the Hessian matrix of source images at different scales, and the multi-scale Hessian matrix of source image is obtained. Second, the multi-scale Hessian matrix is used to decompose each source image into two kinds of regions: the feature and background regions. In order to improve the fusion performance, two new focus measures based on the multi-scale Hessian matrix and two different fusion strategies for both feature and background regions are utilized to obtain the initial decision maps, respectively. Finally, the final decision map for image fusion is achieved by post-processing on the results of the previous step. The proposed method is a primary attempt to introduce image feature and background regions decomposition strategies in the field of multi-focus image fusion. The experimental results also show that our method outperforms the existing image fusion methods in both visual perception and objective evaluations. Bin Xiao 0002, Ge Ou, Xiuli Bi, Weisheng Li 0001 |
IEEE Trans. Multim. | 4 |
| 2019 | 2D-LBP: An Enhanced Local Binary Feature for Texture Image ClassificationabstractThe local binary pattern (LBP) and its variants have shown the effectiveness in texture images classification, face recognition, and other applications. However, most of these LBP methods only focus on the histogram of LBP patterns and ignore the spatial contextual information between LBP patterns. In this paper, we propose a 2D-LBP method which uses a sliding window to count the weighted occurrence number of the rotation invariant uniform LBP pattern pairs to obtain the spatial contextual information. The multi-resolution 2D-LBP features can also be obtained when the radius of 2D-LBP is changed. At last, a two-stage classifier which acts as an ensemble learning step is followed to achieve an accurate classification by combining the predictions on each 2D-LBP with single resolution. Theoretical proof shows that the proposed 2D-LBP is a general framework and can be integrated on other LBP variants to derive new feature extraction methods. Experimental results show that, the proposed method achieves 99.71%, 97.09%, 98.48%, and 49.00% classification accuracy on the public “Brodatz,” “CUReT,” “UIUC,” and “FMD” texture image databases, respectively. Compared with the original LBP and its variants, the proposed method obtains higher classification accuracy under different cases, and simultaneously owns shorter time complexity. Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Multi-scale feature extraction and adaptive matching for copy-move forgery detection
Xiuli Bi, Chi-Man Pun, Xiaochen Yuan |
Multim. Tools Appl. | 1 |
| 2018 | Fast copy-move forgery detection using local bidirectional coherency error refinement
Xiuli Bi, Chi-Man Pun |
Pattern Recognit. | 1 |
| 2017 | Fast reflective offset-guided searching method for copy-move forgery detection
Xiuli Bi, Chi-Man Pun |
Inf. Sci. | 1 |
| 2016 | Multi-Level Dense Descriptor and Hierarchical Feature Matching for Copy-Move Forgery Detection
Xiuli Bi, Chi-Man Pun, Xiaochen Yuan |
Inf. Sci. | 1 |
| 2015 | Image Forgery Detection Using Adaptive Oversegmentation and Feature Point MatchingabstractA novel copy-move forgery detection scheme using adaptive oversegmentation and feature point matching is proposed in this paper. The proposed scheme integrates both block-based and keypoint-based forgery detection methods. First, the proposed adaptive oversegmentation algorithm segments the host image into nonoverlapping and irregular blocks adaptively. Then, the feature points are extracted from each block as block features, and the block features are matched with one another to locate the labeled feature points; this procedure can approximately indicate the suspected forgery regions. To detect the forgery regions more accurately, we propose the forgery region extraction algorithm, which replaces the feature points with small superpixels as feature blocks and then merges the neighboring blocks that have similar local color features into the feature blocks to generate the merged regions. Finally, it applies the morphological operation to the merged regions to generate the detected forgery regions. The experimental results indicate that the proposed copy-move forgery detection scheme can achieve much better detection results even under various challenging conditions compared with the existing state-of-the-art copy-move forgery detection methods. Chi-Man Pun, Xiaochen Yuan, Xiuli Bi |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2007 | Scaling and rotation invariant analysis approach to object recognition based on Radon and Fourier-Mellin transforms
Xuan Wang 0006, Bin Xiao 0002, Jianfeng Ma 0001, Xiuli Bi |
Pattern Recognit. | 4 |