Yongchao Xu

dblp:86/11269 · DBLP profile ↗
← Back
91ranked-venue papers
16as first author
64since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 46 · 9 first-author · 28 since 2021Artificial intelligence and machine learning · 44 · 7 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 29 · 1 first-author · 23 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Decoupling What to Count and Where to See for Referring Expression Counting
abstract
Referring Expression Counting (REC) extends class-level object counting to the fine-grained subclass-level, aiming to enumerate objects matching a textual expression that specifies both the class and distinguishing attribute. A fundamental challenge, however, has been overlooked: annotation points are typically placed on class-representative locations (e.g., heads), forcing models to focus on class-level features while neglecting attribute information from other visual regions (e.g., legs for ''walking''). To address this, we propose W2-Net, a novel framework that explicitly decouples the problem into ''what to count'' and ''where to see'' via a dual-query mechanism. Specifically, alongside the standard what-to-count (w2c) queries that localize the object, we introduce dedicated where-to-see (w2s) queries. The w2s queries are guided to seek and extract features from attribute-specific visual regions, enabling precise subclass discrimination. Furthermore, we introduce Subclass Separable Matching (SSM), a novel matching strategy that incorporates a repulsive force to enhance inter-subclass separability during label assignment. W2-Net significantly outperforms the state-of-the-art on the REC-8K dataset, reducing counting error by 22.5% (validation) and 18.0% (test), and improving localization F1 by 7% and 8%, respectively.
Yuda Zou, Yongchao Xu
AAAI3
2026 CAFL: Conditional Attention Federated Learning for Image Emotion Analysis
abstract
Abstract The rapid proliferation of images on online platforms has made emotion analysis a task of paramount significance. However, these images are often privacy-sensitive, making Federated Learning (FL) a compelling paradigm over traditional centralized methods. A critical yet largely unaddressed challenge in applying FL to this domain is the severe concept drift stemming from the subjective and culturally diverse nature of emotional expression, which causes conventional FL algorithms to fail. In this paper, we propose CAFL (Conditional Attention Federated Learning) to fill this gap. CAFL empowers clients to learn collaboratively yet personally. It intelligently routes information through an adaptive gate that separates features into a personalized stream and a global stream. These streams are then processed by dedicated local and global prediction heads. Crucially, collaboration is guided by a conditional attention mechanism, where the server computes a personalized reference model for each client based on an attention-weighted aggregation of peer models, promoting knowledge sharing among kindred clients. Extensive experiments on various lightweight foundation models show that CAFL consistently outperforms existing FL methods, demonstrating its robustness and superior performance as a solution for distributed, privacy-sensitive image emotion analysis.
Chang Liu 0046, Zengmao Wang, Yongchao Xu, Bo Du 0001
Data Sci. Eng.3
2026 Mamba-Driven Comprehensive Context Learning for Zero-Shot HOI Detection
Jiawei Liu 0001, Yongchao Xu, Sen Tao, Yuexuan Qi, Zhengjun Zha
Int. J. Comput. Vis.2
2026 Boosting Active Prompt Learning via Discriminative Self-Training Dual-Curriculum Learning
Sen Tao, Jiawei Liu 0001, Yongchao Xu, Bingyu Hu, Zhengjun Zha
Int. J. Comput. Vis.4
2026 Understanding and manipulating color representations in CLIP
Peilin Zhou, Jiongyi Wang, Yanmin Dai, Hanyu Du, Yuliang Gu, Yongchao Xu
Neurocomputing6
2026 Incorporating modality-specific intensity prior as text prompt for multimodal myocardial pathology segmentation
Donggen Fang, Yuliang Gu, Lingyi Yu, Bo Du 0001, Yongchao Xu, Lianming Wu
Medical Image Anal.5
2026 Vicinal Gaussian Transform: Rethinking Source-Free Domain Adaptation Through Source-Informed Label Consistency
abstract
A central challenge in source-free domain adaptation (SFDA) is the lack of a theoretical framework for explicitly analyzing domain shifts, as the absence of source data prevents direct domain comparisons. In this paper, we introduce the Vicinal Gaussian Transform (VGT), an analytical operator that models source-informed latent vicinities as Gaussians and shows that vicinal prediction divergence is bounded by their covariance. By this formulation, SFDA can be reframed as shrinking covariance to reinforce label consistency. To operationalize this idea, we introduce the Energy-based VGT (EBVGT), a novel SDE that realizes the Gaussian transform by contracting covariance through a denoising mechanism. A recovery-likelihood with a Schrödinger-Bridge smoothness penalty denoises perturbed states, while a BYOL-derived energy function, directly obtained from model predictions, provides the score to guide label-consistent trajectories within the vicinity. This design not only yields noise-suppressed vicinal features for adaptation without source data, but also eliminates the need for additional learnable parameters for score estimation, in contrast to conventional deep SDEs. Our EBVGT is model- and modality-agnostic, efficient for classification, and improves state-of-the-art SFDA methods by 1.3-3.0% (2.0% on average) across both 2D image and 3D point cloud benchmarks.
Jing Wang 0112, Yongchao Xu, Zeyu Gong, Bo Tao 0001, Clarence W. de Silva, Xiang Bai
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Prompt-guided Modality Completion for cardiac pathology segmentation
Donggen Fang, Yajie Chen, Yuliang Gu, Lingyi Yu, Zhongyuan Wang 0001, Bo Du 0001, Lianming Wu, Yongchao Xu
Pattern Recognit.9
2026 Coarse-to-fine crack cue for robust crack detection
Zelong Liu, Yuliang Gu, Zhichao Sun 0004, Huachao Zhu, Xin Xiao 0010, Bo Du 0001, Laurent Najman, Yongchao Xu
Pattern Recognit.8
2026 Leveraging Textual Anatomical Knowledge for Class-Imbalanced Semi-Supervised Multi-Organ Segmentation
abstract
Imbalanced class distributions among different organs pose significant challenges in real-world semi-supervised multi-organ segmentation. Integrating anatomical priors offers a promising research direction to mitigate these imbalances. In this paper, we explore the capabilities of Multimodal Large Language Models (MLLM) to extract robust, generic textual anatomical insights serving as prior knowledge for segmentation model. Specifically, we employ GPT-4o to generate detailed textual descriptions of anatomical priors-including both inter-organ relative positional relationships and organ shape characteristics. These priors generated only once for the whole training and testing are then seamlessly integrated into the segmentation model as parameters within the segmentation head. Furthermore, we align the textual priors with visual features using contrastive learning. The inter-organ positional priors guide the model in localizing smaller organs relative to larger ones, while the organ shape priors help ensure that the learned morphological structures are more anatomically plausible. Extensive experiments demonstrate that our method significantly outperforms some state-of-the-art approaches. The source code is available at: https://github.com/Lunn88/TAK-Semi.
Yuliang Gu, Weilun Tsao, Yepeng Liu 0002, Lianming Wu, Thierry Géraud, Bo Du 0001, Yongchao Xu
IEEE Trans. Medical Imaging7
2026 Anomaly or Characteristic: Memory-Based Coarse-to-Fine Feature Fusion for Industrial Anomaly Detection
abstract
Unsupervised anomaly detection methods primarily focus on modeling the distribution of normal samples at image/feature level. Significant deviation from the modeled distribution is then considered as anomaly. Yet, each normal sample may have its own unique characteristic drifting from the idea distribution, making it difficult to distinguish between anomaly and characteristic. In this paper, we propose a memory-based Coarse-to-Fine Feature Fusion (C3F) module to tackle this challenge. Specifically, we construct a group of memory banks that model the feature distribution of normal samples at various levels of granularity. The memory-based C3F is applied to each skip-connection between the encoder and the decoder, and progressively removes anomaly while maintaining characteristic. This helps to reconstruct defect-free image with characteristic preserved, encouraging large (respsmall) deviation from the modeled distribution for anomaly (respcharacteristic). Besides, we also introduce a novel rough anomaly score map Guided Segmentation (GS) module to achieve precise anomaly localization. Extensive experiments on widely used VisA and MVTec-AD benchmarks demonstrate the wide-ranging applicability of the proposed method termed C3FGS on industrial components of various forms. The implementation code is publicly available athttps://github.com/LZL501/c3f_industrial_anomaly_detection.
Huachao Zhu, Zelong Liu, Zhichao Sun 0004, Xin Xiao 0010, Yongchao Xu
IEEE Trans. Multim.7
2026 DayPQ: Dynamic Layerwise Pruning and Quantization for LLM Inference Acceleration
abstract
The deployment of the prevailing large language models (LLMs) on resource-constrained hardware faces critical challenges due to their massive computational and memory requirements. Pruning and quantization (PQ), as two predominant techniques for LLM acceleration, face the challenge of fully preserving model capabilities while achieving significant acceleration in practice. Static PQ irreversibly discards potentially critical parameters and degrades precision across diverse inputs, which must be compensated through fine-tuning. Dynamic PQ preserves accuracy but incurs prohibitive hardware overhead and additional latency from runtime adaptations. Achieving low-latency and overhead-efficient dynamic PQ requires the co-design of adaptive compression tightly integrated with hardware. In this work, we present DayPQ, an algorithm-hardware co-optimization technique for efficient LLM inference: self-adaptive norm-based pruning and bias-shifting FP8 quantization at the algorithm level, synergized with a multiplication-free shift–add PE array and switchable dataflow at the hardware level, achieving dynamic acceleration with the negligible accuracy loss and zero additional latency. DayPQ enables adaptive PQ by intrinsically exploiting normalization operations from the model’s inference process instead of extra operations. Simultaneously, it establishes dynamic orthogonality between sparsity and quantization, where these dual compression mechanisms synergistically enhance the redundancy reduction beyond mere additive effects of independent operations. Our evaluation shows that DayPQ obtains an average$3.13\times $higher energy efficiency and$2.03\times $performance boost than three state-of-the-art (SOTA) accelerators.
Yongchao Xu, Ziying Zhuo, Qinzhe Zhi
IEEE Trans. Very Large Scale Integr. Syst.1
2025 HOIMamba: Efficient Mamba-based Disentangled Progressive Learning for HOI Detection
abstract
Human-object interaction (HOI) detection aims to detect the spatial positions of human-object pairs and recognize their interactions. Existing single-branch, two-branch, and three-branch methods are challenging to make an appropriate trade-off on efficiency, multi-task decoupling, and collaborative learning, while they fail to identify rare and complex interaction categories effectively as well. In this work, we propose a novel Efficient Mamba-based Disentangled Progressive Learning (HOIMamba) for HOI Detection to absorb the advantages of the existing three approaches and adaptively aggregate multi-level interaction semantics guided by cross-task bidirectional information contexts. Specifically, HOIMamba builds an efficient and effective decoder through cascaded Low-Rank Adaptations (LoRAs), with high efficiency, thorough decoupling of tasks, and good multi-task collaborative learning. Furthermore, to alleviate the recognition problem of interactions in difficult HOI samples, a novel Mamba-based comprehensive progressive learning strategy with Cross-enhance Mamba (CEM) blocks and Detection Context Propagation (DCP) blocks is designed to gradually excavate interaction-related discriminative cues from four levels. CEM blocks automatically aggregate context to generate diverse task-shared semantics and simultaneously realize the cross-task interaction between human and object branches, guiding the interaction branch to extract more expressive HOI representation. DCP blocks further transfer the comprehensive interaction context to human and object branches to achieve rich and effective information exchange, facilitating the model to discover more HOI instances. Extensive experimental results on two standard benchmarks demonstrate the effectiveness of our HOIMamba.
Yongchao Xu, Jiawei Liu 0001, Sen Tao, Qiang Zhang 0051, Zhengjun Zha
AAAI1
2025 Hierarchical Knowledge Prompt Tuning for Multi-task Test-Time Adaptation
abstract
Test-time adaptation using vision- language models (such as CLIP) to quickly adjust to distributional shifts of downstream tasks has shown great potential. Despite significant progress, existing methods are still limited to single- task test- time adaptation scenarios and have not effectively explored the issue of multi- task adaptation. To address this practical problem, we propose a novel Hierarchical Knowledge Prompt Tuning (HKPT) method, which achieves joint adaptation to multiple target domains by mining more comprehensive source domain discriminative knowledge and hierarchically modeling task- specific and task- shared knowledge. Specifically, HKPT constructs a CLIP prompt distillation framework that utilizes the broader source domain knowledge of large teacher CLIP to guide prompt tuning for lightweight student CLIP from multiple views during testing. Meanwhile, HKPT establishes task- specific dual dynamic knowledge graph to capture fine- grained contextual knowledge from continuous test data. To fully exploit the complementarity among multiple target tasks, HKPT employs an adaptive task grouping strategy for achieving intertask knowledge sharing. Furthermore, HKPT can seamlessly transfer to basic single- task test- time adaptation scenarios while maintaining robust performance. Extensive experimental results in both multi- task and single- task testtime adaptation settings demonstrate that our HKPT significantly outperforms state- of- the- art methods.
Qiang Zhang 0051, Mengsheng Zhao, Jiawei Liu 0001, Fanrui Zhang, Yongchao Xu, Zhengjun Zha
CVPR5
2025 Pathological Prior-Guided Multiple Instance Learning For Mitigating Catastrophic Forgetting in Breast Cancer Whole Slide Image Classification
abstract
In histopathology, intelligent diagnosis of Whole Slide Images (WSIs) is essential for automating and objectifying diagnoses, reducing the workload of pathologists. However, diagnostic models often face the challenge of forgetting previously learned data during incremental training on datasets from different sources. To address this issue, we propose a new framework PaGMIL to mitigate catastrophic forgetting in breast cancer WSI classification. Our framework introduces two key components into the common MIL model architecture. First, it leverages microscopic pathological prior to select more accurate and diverse representative patches for MIL. Secondly, it trains separate classification heads for each task and uses macroscopic pathological prior knowledge, treating the thumbnail as a prompt guide (PG) to select the appropriate classification head. We evaluate the continual learning performance of PaGMIL across several public breast cancer datasets. PaGMIL achieves a better balance between the performance of the current task and the retention of previous tasks, outperforming other continual learning methods. Our code will be open-sourced upon acceptance.
Weixi Zheng, Aoling Huang, Jingping Yuan, Yongchao Xu, Thierry Géraud
ICASSP6
2025 Beyond Pixel Uncertainty: Bounding the OoD Objects in Road Scenes
Huachao Zhu, Zelong Liu, Zhichao Sun 0004, Yuda Zou, Gui-Song Xia, Yongchao Xu
ICCV6
2025 LiftFeat: 3D Geometry-Aware Local Feature Matching
abstract
Robust and efficient local feature matching plays a crucial role in applications such as SLAM and visual localization for robotics. Despite great progress, it is still very challenging to extract robust and discriminative visual features in scenarios with drastic lighting changes, low texture areas, or repetitive patterns. In this paper, we propose a new lightweight network called LiftFeat, which lifts the robustness of raw descriptor by aggregating 3D geometric feature. Specifically, we first adopt a pre-trained monocular depth estimation model to generate pseudo surface normal label, supervising the extraction of 3D geometric feature in terms of predicted surface normal. We then design a 3D geometry-aware feature lifting module to fuse surface normal feature with raw 2D descriptor feature. Integrating such 3D geometric feature enhances the discriminative ability of 2D feature description in extreme conditions. Extensive experimental results on relative pose estimation, homography estimation, and visual localization tasks, demonstrate that our LiftFeat outperforms some lightweight state-of-the-art methods. Code will be released at: https://github.com/lyp-deeplearning/LiftFeat.
Yepeng Liu 0002, Wenpeng Lai, Yuxuan Xiong, Jinchi Zhu, Jun Cheng 0003, Yongchao Xu
ICRA7
2025 Pixel-wise Divide and Conquer for Federated Vessel Segmentation
abstract
Accurate vessel segmentation is essential for diagnosing and managing vascular and ophthalmic diseases. Traditional learning-based vessel segmentation methods heavily rely on high-quality, pixel-level annotated datasets. However, segmentation performance suffers significantly when applied in federated learning settings due to vessel morphology inconsistency and vessel-background imbalance. The former limits the ability of models to capture fine-grained vessels, while the latter overemphasizes background pixels and biases the model towards them. To address these challenges, we propose a novel method named Federated Vessel-Aware Calibration (FVAC), which leverages global uncertainty to provide differentiated guidance for clients, focusing on pixels of various morphologies that are difficult to distinguish. Furthermore, we introduce a foreground-background decoupling alignment strategy that utilizes more stable and balanced global features to mitigate semantic drift caused by vessel-background imbalance in local clients. Comprehensive experiments confirm the effectiveness of our method
Wenke Huang 0003, Zhihao Wang 0002, Zekun Shi, He Li 0054, Mang Ye, Bo Du 0001, Yongchao Xu
IJCAI9
2025 Test-Time Training with Local Contrast-Preserving Copy-Pasted Image for Domain Generalization in Retinal Vessel Segmentation
Yuliang Gu, Zhichao Sun 0004, Zelong Liu, Yongchao Xu
MICCAI (7)4
2025 Neighborhood-Consistent Binary Transformation for Domain-Invariant Chest X-Ray Diagnosis
Zelong Liu, Huachao Zhu, Zhichao Sun 0004, Yuda Zou, Yuliang Gu, Bo Du 0001, Yongchao Xu
MICCAI (5)7
2025 Noise-Robust Tuning of SAM for Domain Generalized Ultrasound Image Segmentation
Zhikai Wei, Hanyu Du, Rui Yu 0002, Bo Du 0001, Yongchao Xu
MICCAI (5)6
2025 CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection
abstract
With the exponential growth of data, traditional object detection methods are increasingly struggling to handle vast vocabulary object detection tasks effectively. We analyze two key limitations of classification-based detectors: positive gradient dilution, where rare positive categories receive insufficient learning signals, and hard negative gradient dilution, where discriminative gradients are overwhelmed by numerous easy negatives. To address these challenges, we propose CQ-DINO, a category query-based object detection framework that reformulates classification as a contrastive task between object queries and learnable category queries. Our method introduces image-guided query selection, which reduces the negative space by adaptively retrieving top-K relevant categories per image via cross-attention, thereby rebalancing gradient distributions and facilitating implicit hard example mining. Furthermore, CQ-DINO flexibly integrates explicit hierarchical category relationships in structured datasets (e.g., V3Det) or learns implicit category correlations via self-attention in generic datasets (e.g., COCO). Experiments demonstrate that CQ-DINO achieves superior performance on the challenging V3Det benchmark (surpassing previous methods by 2.1% AP) while maintaining competitiveness in COCO. Our work provides a scalable solution for real-world detection systems requiring wide category coverage.
Huazhang Hu, Yidong Ma, Xu Tang 0007, Yao Hu 0002, Yongchao Xu
NeurIPS8
2025 Development of residual learning in deep neural networks for computer vision: A survey
Guoping Xu, Xiaxia Wang 0005, Xuesong Leng, Yongchao Xu
Eng. Appl. Artif. Intell.5
2025 Shape-intensity knowledge distillation for robust medical image segmentation
Bo Du 0001, Yongchao Xu
Frontiers Comput. Sci.3
2025 Not All Pixels are Equal: Learning Pixel Hardness for Semantic Segmentation
Xin Xiao 0010, Daiguo Zhou, Jiagao Hu, Yongchao Xu
Int. J. Comput. Vis.5
2025 Dual structure-aware image filterings for semi-supervised medical image segmentation
Yuliang Gu, Zhichao Sun 0004, Xin Xiao 0010, Yepeng Liu 0002, Yongchao Xu, Laurent Najman
Medical Image Anal.6
2025 Calibration Matters: Prototype-Aware Diffusion for OCT Cervical Classification With Calibration
abstract
Cervical optical coherence tomography (OCT) imaging serves as an effective diagnostic tool, and the development of deep learning classification models for OCT has the potential to enhance diagnosis. However, the complex imaging patterns of OCT data, significant noise, and the substantial domain gap from multi-center data result in high uncertainty and low accuracy in classification networks. To address these challenges, we propose a Multi-scale Prototype-Guided Diffusion learning method (MPGD), which is constructed with theMulti-scale Feature Condition (MFC),Diffusion-based Classification Calibrator (DCC), andMulti-scale Prototype Bank (MPB)modules. Specifically, MFC provides initial classification based on multi-scale features, DCC calibrates MFC's classification results through a diffusion model, and MPB refines DCC's visual guidance using prototypes obtained from clustering. Extensive experiments demonstrate that MPGD outperforms widely-used competitors for cervical OCT image classification, showing excellent generalization performance.
Yuxuan Xiong, Yongchao Xu, Yan Zhang 0123, Bo Du 0001
IEEE Signal Process. Lett.3
2025 Learning Modality-Invariant Feature for Multimodal Image Matching via Knowledge Distillation
abstract
Multimodal remote sensing image matching is essential for multi-source information fusion. Recently, learning-based feature matching networks have significantly enhanced the performance of unimodal image matching tasks through data-driven approaches. However, progress in applying these learning-based methods to multimodal image matching has been slower. A major obstacle is the substantial nonlinear radiometric differences between modalities, which require networks to learn modality-invariant features from large amounts of paired data. To address this, we propose EMINet, an efficient method for learning modality-invariant features from limited data to improve matching performance. Our approach constructs a high-performance teacher network by combining the DINOv2 foundational model, the keypoint and descriptor extraction network SuperPoint, and the feature matching network SuperGlue. Leveraging the strong semantic representation capability of DINOv2, the teacher network achieves excellent cross-modality matching ability. To meet low-latency requirements in practical applications, we introduce two novel knowledge distillation strategies: Semantic Window Relation Distillation (SWRD) and Cross-Triplet Descriptor Distillation (CTDD). SWRD improves the discriminative power of the student network’s descriptors by learning patch-level distributions from DINOv2, while CTDD enforces cross-modality triplet constraints to enhance modality invariance of the student network. Experimental results demonstrate that EMINet outperforms several state-of-the-art methods on various datasets, including Optical-SAR, Optical-NIR, and Optical-IR datasets.
Yepeng Liu 0002, Wenpeng Lai, Yuliang Gu, Gui-Song Xia, Bo Du 0001, Yongchao Xu
IEEE Trans. Geosci. Remote. Sens.6
2025 MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching
abstract
Many keypoint detection and description methods have been proposed for image matching or registration. While these methods demonstrate promising performance for single-modality image matching, they often struggle with multimodal data because the descriptors trained on single-modality data tend to lack robustness against the non-linear variations present in multimodal data. Extending such methods to multimodal image matching often requires well-aligned multimodal data to learn modality-invariant descriptors. However, acquiring such data is often costly and impractical in many real-world scenarios. To address this challenge, we propose a modality-invariant feature learning network (MIFNet) to compute modality-invariant features for keypoint descriptions in multimodal image matching using only single-modality training data. Specifically, we propose a novel latent feature aggregation module and a cumulative hybrid aggregation module to enhance the base keypoint descriptors trained on single-modality data by leveraging pre-trained features from Stable Diffusion models. We validate our method with recent keypoint detection and description methods in three multimodal retinal image datasets (CF-FA, CF-OCT, EMA-OCTA) and two remote sensing datasets (Optical-SAR and Optical-NIR). Extensive experiments demonstrate that the proposed MIFNet is able to learn modality-invariant feature for multimodal image matching without accessing the targeted modality and has good zero-shot generalization ability. The code will be released at https://github.com/lyp-deeplearning/MIFNet.
Yepeng Liu 0002, Zhichao Sun 0004, Baosheng Yu, Yitian Zhao, Bo Du 0001, Yongchao Xu, Jun Cheng 0003
IEEE Trans. Image Process.6
2024 Shape Transformation Driven by Active Contour for Class-Imbalanced Semi-Supervised Medical Image Segmentation
abstract
Annotating 3D medical images demands expert knowledge and is time-consuming. As a result, semi-supervised learning (SSL) approaches have gained significant interest in 3D medical image segmentation. The significant size differences among various organs in the human body lead to imbalanced class distribution, which is a major challenge in the real-world application of these SSL approaches. To address this issue, we develop a novel Shape Transformation driven by Active Contour (STAC), that enlarges smaller organs to alleviate imbalanced class distribution across different organs. Inspired by curve evolution theory in active contour methods, STAC employs a signed distance function (SDF) as the level set function, to implicitly represent the shape of organs, and deforms voxels in the direction of the steepest descent of SDF (i.e., the normal vector). To ensure that the voxels far from expansion organs remain unchanged, we design an SDF-based weight function to control the degree of deformation for each voxel. We then use STAC as a data-augmentation process during the training stage. Experimental results on two benchmark datasets demonstrate that the proposed method significantly outperforms some state-of-the-art methods. Source code is publicly available at https://github.com/GuGuLL123/STAC.
Yuliang Gu, Yepeng Liu 0002, Zhichao Sun 0004, Jinchi Zhu, Yongchao Xu, Laurent Najman
BIBM5
2024 Progressive Retinal Image Registration via Global and Local Deformable Transformations
abstract
Retinal image registration plays an important role in the ophthalmological diagnosis process. Since there exist variances in viewing angles and anatomical structures across different retinal images, keypoint-based approaches become the mainstream methods for retinal image registration thanks to their robustness and low latency. These methods typically assume the retinal surfaces are planar, and adopt feature matching to obtain the homography matrix that represents the global transformation between images. Yet, such a planar hypothesis inevitably introduces registration errors since retinal surface is approximately curved. This limitation is more prominent when registering image pairs with significant differences in viewing angles. To address this problem, we propose a hybrid registration framework called HybridRetina, which progressively registers retinal images with global and local deformable transformations. For that, we use a keypoint detector and a deformation network called GAMorph to estimate the global transformation and local deformable transformation, respectively. Specifically, we integrate multi-level pixel relation knowledge to guide the training of GAMorph. Additionally, we utilize an edge attention module that includes the geometric priors of the images, ensuring the deformation field focuses more on the vascular regions of clinical interest. Experiments on two widely-used datasets, FIRE and FLoRI21, show that our proposed HybridRetina significantly outperforms some state-of-the-art methods. The code is available at https://github.com/lyp-deeplearning/awesome-retinal-registration.
Yepeng Liu 0002, Baosheng Yu, Yuliang Gu, Bo Du 0001, Yongchao Xu, Jun Cheng 0003
BIBM6
2024 Uncertainty-aware Condensation for OCT Cervical Dataset Combining Coreset Selection and Distillation
abstract
Optical Coherence Tomography (OCT) cervical images offer a sensitive and non-invasive solution for screening cervical cancer. Currently, utilizing deep learning techniques to interpret these images has emerged as a promising approach for computer-aided diagnosis. However, challenges in computational efficiency and privacy protection persist in this field. To address these issues, we propose an uncertainty-aware filtering before distillation (UFD) framework to compress the OCT cervical dataset in a trustworthy and efficient manner. Our approach involves pre-training an uncertainty proxy to evaluate the importance of each sample, allowing us to extract a valuable subset (i.e., coreset) from the original dataset. We then apply importance-weighted distribution matching for dataset distillation. The distribution matching loss is calculated as the distance between the class centers of the uncertainty-weighted coreset and the synthetic distilled dataset. Extensive experiments validate the feasibility and effectiveness of our proposed UFD framework. It is worth noting that after selecting the most representative samples from the original OCT dataset through coreset selection, the performance of the trained network is comparable to, or sometimes even slightly higher than, that of the network trained with all the original data. This indicates that automated dataset compression and cleaning techniques are crucial for improving image classification problems in medical imaging datasets, such as OCT and ultrasound, which are rich in noise and erroneous samples.
Yuxuan Xiong, Yongchao Xu, Yan Zhang 0123, Kai Gui, Bo Du 0001
BIBM2
2024 HFGS: 4D Gaussian Splatting with Emphasis on Spatial and Temporal High-Frequency Components for Endoscopic Scene Reconstruction
Xingyue Zhao, Lingting Zhu, Weixi Zheng, Yongchao Xu
BMVC5
2024 Shifted Autoencoders for Point Annotation Restoration in Object Counting
Yuda Zou, Xin Xiao 0010, Peilin Zhou, Zhichao Sun 0004, Bo Du 0001, Yongchao Xu
ECCV (25)6
2024 Auxiliary Information Guided Segmentation for the Clinical Target Volume of Cervical Cancer
Yongchao Xu
ICPR (27)2
2024 P2A: Transforming Proposals to Anomaly Masks
Huachao Zhu, Zhichao Sun 0004, Zelong Liu, Yongchao Xu
ICPR (33)4
2024 Position-Guided Prompt Learning for Anomaly Detection in Chest X-Rays
Zhichao Sun 0004, Yuliang Gu, Yepeng Liu 0002, Yongchao Xu
MICCAI (1)6
2024 Prompting Segment Anything Model with Domain-Adaptive Prototype for Generalizable Medical Image Segmentation
Zhikai Wei, Peilin Zhou, Yuliang Gu, Yongchao Xu
MICCAI (8)6
2024 Spatial-Aware Attention Generative Adversarial Network for Semi-supervised Anomaly Detection in Medical Image
Zhichao Sun 0004, Zelong Liu, Rui Yu 0002, Bo Du 0001, Yongchao Xu
MICCAI (5)7
2024 MoreStyle: Relax Low-Frequency Constraint of Fourier-Based Image Reconstruction in Generalizable Medical Image Segmentation
Rui Yu 0002, Bo Du 0001, Yongchao Xu
MICCAI (8)6
2024 WIA-LD2ND: Wavelet-Based Image Alignment for Self-supervised Low-Dose CT Denoising
Yuliang Gu, Bo Du 0001, Yongchao Xu, Rui Yu 0002
MICCAI (7)5
2024 A Generalized Contrast-Adjustment Guided Growth Method for Medical Image Segmentation
Qikui Zhu, Yongchao Xu, Bo Du 0001
PRCV (15)3
2024 ESCNet: Entity-enhanced and Stance Checking Network for Multi-modal Fact-Checking
abstract
Recently, misinformation incorporating both texts and images has been disseminated more effectively than those containing text alone on social media, raising significant concerns for multi-modal fact-checking. Existing research makes contributions to multi-modal feature extraction and interaction, but fails to fully enhance the valuable semantic representations or excavate the intricate entity information. Besides, existing multi-modal fact-checking datasets are primarily focused on English and merely concentrate on a single type of misinformation, thereby neglecting a comprehensive summary and coverage of various types of misinformation. Taking these factors into account, we construct the first large-scale Chinese Multi-modal Fact-Checking (CMFC) dataset which encompasses 46,000 claims. The CMFC covers all types of misinformation for fact-checking and is divided into two sub-datasets, Collected Chinese Multi-modal Fact-Checking (CCMF) and Synthetic Chinese Multi-modal Fact-Checking (SCMF). To establish baseline performance, we propose a novel Entity-enhanced and Stance Checking Network (ESCNet), which includes Multi-modal Feature Extraction Module, Stance Transformer, and Entity-enhanced Encoder. The ESCNet jointly models stance semantic reasoning features and knowledge-enhanced entity pair features, in order to simultaneously learn effective semantic-level and knowledge-level claim representations. Our work offers the first step and establishes a benchmark for evidence-based, multi-type, multi-modal fact-checking.
Fanrui Zhang, Jiawei Liu 0001, Qiang Zhang 0051, Yongchao Xu, Zhengjun Zha
WWW5
2024 Shape-intensity-guided U-net for medical image segmentation
Bo Du 0001, Yongchao Xu
Neurocomputing3
2024 Source domain prior-assisted segment anything model for single domain generalization in medical image segmentation
Bo Du 0001, Yongchao Xu
Image Vis. Comput.3
2024 Nighttime image semantic segmentation with retinex theory
Zhichao Sun 0004, Huachao Zhu, Xin Xiao 0010, Yuliang Gu, Yongchao Xu
Image Vis. Comput.5
2024 Distilling OCT cervical dataset with evidential uncertainty proxy
Yuxuan Xiong, Yongchao Xu, Yan Zhang 0123, Bo Du 0001
Image Vis. Comput.2
2024 HPFL: hyper-network guided personalized federated learning for multi-center tuberculosis chest x-ray diagnosis
Chang Liu 0046, Yong Luo 0002, Yongchao Xu, Bo Du 0001
Multim. Tools Appl.3
2024 Query-guided generalizable medical image segmentation
Zhiyi Yang, Yuliang Gu, Yongchao Xu
Pattern Recognit. Lett.4
2024 Edge-and-Mask Integration-Driven Diffusion Models for Medical Image Segmentation
abstract
Denoising diffusion probabilistic models (DDPMs) exhibit significant potential in the realm of medical image segmentation. Nevertheless, current DDPM implementations rely on original image features as conditional information, thus lacking the ability to specifically emphasize edge information, a critical aspect in addressing the primary challenge of segmentation. Furthermore, the necessary semantic features for conditioning the diffusion process lack effective alignment with the noise embedding. To address the above issues, we propose a novel edge-and-mask integration-driven diffusion model (EMidDiff). Specifically, 1) an edge-and-mask condition strategy is proposed for the segmentation diffusion model to effectively leverage rich semantic features, particularly the edge feature. 2) A novel co-attention guidance block is designed to align the segmentation map and condition features. The experimental results on brain tumor segmentation and optic-cup segmentation underscore the effectiveness of our approach, surpassing the performance of some state-of-the-art segmentation diffusion models.
Qikui Zhu, Yuxuan Xiong, Yongchao Xu, Bo Du 0001
IEEE Signal Process. Lett.4
2024 Enhancing Building Footprint Extraction With Partial Occlusion by Exploring Building Integrity
abstract
Building footprint extraction (BFE) is essential for applications like land use management, urban planning, and database updates. However, many existing methods struggle with occluded buildings in remote sensing images, causing performance degradation. In addition, methods designed for occluded buildings often depend on external data, such as LiDAR or multiview images, increasing extraction costs. To address this challenge, we present a robust building extraction framework that effectively targets occluded buildings. Specifically, we design a low-cost image enhancement strategy. This strategy leverages the segmented labels to introduce object information about the buildings, aiding in generating enhanced images with strong building integrity, no occlusions, and high contrast with the background. Then, a building integrity enhancement branch (BIEB) is proposed. This branch incorporates enhanced building images as supervision, in order to leverage building integrity to mitigate the impact of occlusion. Finally, we introduce the idea of residual learning. This approach facilitates the model’s ability to leverage building integrity to mitigate the impact of occlusion. Extensive experiments on the WHU, Massachusetts, Inria, and Aerial Imagery for Roof Segmentation (AIRS) datasets have demonstrated the effectiveness of our approach in extracting occluded buildings. The experimental results show that, compared with the latest HD-Net, our BIENet achieved IoU score improvements of 1.49%, 1.16%, and 2.26%, and$F1$score improvements of 0.82%, 0.70%, and 1.22%, respectively, on the WHU, Inria, and AIRS datasets. The source code of the proposed BIENet is available athttps://github.com/tqwhdx19/bienet.
Yongchao Xu, Bo Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 WASPSYN: A Challenge for Domain Adaptive Synapse Detection in Microwasp Brain Connectomes
abstract
The size of image volumes in connectomics studies now reaches terabyte and often petabyte scales with a great diversity of appearance due to different sample preparation procedures. However, manual annotation of neuronal structures (e.g., synapses) in these huge image volumes is time-consuming, leading to limited labeled training data often smaller than 0.001% of the large-scale image volumes in application. Methods that can utilize in-domain labeled data and generalize to out-of-domain unlabeled data are in urgent need. Although many domain adaptation approaches are proposed to address such issues in the natural image domain, few of them have been evaluated on connectomics data due to a lack of domain adaptation benchmarks. Therefore, to enable developments of domain adaptive synapse detection methods for large-scale connectomics applications, we annotated 14 image volumes from a biologically diverse set of Megaphragma viggianii brain regions originating from three different whole-brain datasets and organized the WASPSYN challenge at ISBI 2023. The annotations include coordinates of pre-synapses and post-synapses in the 3D space, together with their one-to-many connectivity information. This paper describes the dataset, the tasks, the proposed baseline, the evaluation method, and the results of the challenge. Limitations of the challenge and the impact on neuroscience research are also discussed. The challenge is and will continue to be available at https://codalab.lisn.upsaclay.fr/competitions/9169. Successful algorithms that emerge from our challenge may potentially revolutionize real-world connectomics research and further the cause that aims to unravel the complexity of brain structure and function.
Yicong Li 0002, Wanhua Li 0001, Qi Chen 0014, Wei Huang 0036, Yuda Zou, Kazunori Shinomiya, Pat Gunn, Nishika Gupta, Alexey Polilov, Yongchao Xu, Yueyi Zhang 0001, Zhiwei Xiong, Hanspeter Pfister, Donglai Wei 0001, Jingpeng Wu
IEEE Trans. Medical Imaging11
2024 Foundation models matter: federated learning for multi-center tuberculosis diagnosis via adaptive regularization and model-contrastive learning
Chang Liu 0046, Yong Luo 0002, Yongchao Xu, Bo Du 0001
World Wide Web (WWW)3
2023 FedARC: Federated Learning for Multi-Center Tuberculosis Chest X-ray Diagnosis with Adaptive Regularizing Contrastive Representation
abstract
Tuberculosis (TB) poses a significant global health threat and leads to millions of deaths annually. While early diagnosis and treatment can substantially enhance survival prospects, it continues to present a major challenge, particularly in developing countries. In recent years, machine learning has emerged as a valuable tool for tuberculosis diagnosis. However, the training of a dependable diagnostic model necessitates a large volume of data, typically distributed across multiple medical centers. To safeguard data privacy across various centers, we have incorporated federated learning (FL) into TB diagnosis. However, conventional FL methods suffer from substantial performance degradation due to the considerable variation in TB data distribution across different centers. Consequently, we introduce a novel personalized FL approach, FedARC, to address this issue. To mitigate data heterogeneity across centers, we guide the objective function for each center with adaptive regularization to align it with the stationary point of the global loss, thereby enabling the model to converge towards the global optimum. Simultaneously, model-contrastive learning enables the exploration of the specific attributes of each client, enabling the local model to learn more generalizable features. Extensive experimental results on five publicly available chest X-ray image datasets demonstrate the significant outperformance of our proposed method over state-of-the-art methods in diverse settings.
Chang Liu 0046, Yong Luo 0002, Yongchao Xu, Bo Du 0001
BIBM3
2023 Scratch Each Other's Back: Incomplete Multi-modal Brain Tumor Segmentation Via Category Aware Group Self-Support Learning
abstract
Although Magnetic Resonance Imaging (MRI) is very helpful for brain tumor segmentation and discovery, it often lacks some modalities in clinical practice. As a result, degradation of prediction performance is inevitable. According to current implementations, different modalities are considered to be independent and non-interfering with each other during the training process of modal feature extraction, however they are complementary. In this paper, considering the sensitivity of different modalities to diverse tumor regions, we propose a Category Aware Group Self-Support Learning framework, called GSS, to make up for the information deficit among the modalities in the individual modal feature extraction phase. Precisely, within each prediction category, predictions of all modalities form a group, where the prediction with the most extraordinary sensitivity is selected as the group leader. Collaborative efforts between group leaders and members identify the communal learning target with high consistency and certainty. As our minor contribution, we introduce a random mask to reduce the possible biases. GSS adopts the standard training strategy without specific architectural choices and thus can be easily plugged into existing incomplete multi-modal brain tumor segmentation. Remarkably, extensive experiments on BraTS2020, BraTS2018, and BraTS2015 datasets demonstrate that GSS can improve the performance of existing SOTA algorithms by 1.27-3.20% in Dice on average. The code is released at https://github.com/qysgithubopen/GSS.
Yansheng Qiu, Delin Chen, Hongdou Yao, Yongchao Xu, Zheng Wang 0007
ICCV4
2023 Scale-Aware Test-Time Click Adaptation for Pulmonary Nodule and Mass Segmentation
Jiancheng Yang, Yongchao Xu, Li Zhang 0085, Bo Du 0001
MICCAI (3)3
2022 Self-supervised Learning Based on Max-tree Representation for Medical Image Segmentation
abstract
In recent years, the convolutional neural network(CNN) based deep learning architectures have achieved great success in medical image segmentation. However, CNN usually relies on abundant labeled data for training. At the same time, collecting labeled training data is time-consuming and expensive. Therefore, in addition to the common unsupervised learning methods, a series of self-supervised learning(SSL) methods have been proposed in medical image analysis using a large amount of unlabeled data. These SSL strategies usually extract potential supervised signals by pretext tasks and help the networks learn a way of feature representation. However, the learned feature representations in the pretext tasks are not commonly directly related to downstream tasks like segmentation. We assume that the more pretext tasks help the model learn the structural features of the image, the better the model's performance on the downstream segmentation task will be. In this paper, we propose an SSL strategy based on max-tree representation to extract the image structure information, and the CNN learns this max-tree representation in the pretext task. To the best of our knowledge, we are the first to take the structure information into account in the SSL pretext task. Extensive experiments show that our SSL strategy based on max-tree representation can help the CNN to learn abundant structural information, which is significantly useful for the downstream segmentation task.
Bo Du 0001, Yongchao Xu
IJCNN3
2022 Pulmonary Nodule Classification with Multi-View Convolutional Vision Transformer
abstract
Pulmonary nodule classification from computerized tomography(CT) Scans is a vital task for the early screening of Lung cancers. The algorithm is aiming at distinguishing malignant pulmonary nodules, benign nodules and the ones with their subtypes. In this paper, we defined a detailed pulmonary nodule classification task considering 5 semantic labels. We are facing with a series of non-trival problems dealing with such a task. First, the available medical image data for training is quite limited. We enlarged the training dataset by cropping out three-dimension(3D) volume of each pulmonary nodule and generating 15 planes with different orientations from these volumes. Secondly, the global modeling ability of the existing convolutional neural network(CNN) based architectures can not meet the need of medical image analysis well. To learn discriminative abstract information, we down-sample feature maps between successive stages and adopt the BotNet-50 backbone which is a combination of ResNet backbone and self-attention modules. Such an architecture can extract local and non-local information in low-level and high-level layers, respectively. Last but not the least, the data distribution of training data and testing data don't share similar distribution in real-world multi-center medical image classification scenes. We assigned the samples with modified wights while calculating the loss value for optimization. The proposed method can eliminate the spurious correlation between features and labels. Experiments demonstrate the effectiveness of each component.
Yuxuan Xiong, Bo Du 0001, Yongchao Xu, Jiajun Deng, Yunlang She
IJCNN3
2022 AutoScale: Learning to Scale for Crowd Counting
Chenfeng Xu, Dingkang Liang, Yongchao Xu, Song Bai 0001, Xiang Bai, Masayoshi Tomizuka
Int. J. Comput. Vis.3
2022 Local Intensity Order Transformation for Robust Curvilinear Object Segmentation
abstract
Segmentation of curvilinear structures is important in many applications, such as retinal blood vessel segmentation for early detection of vessel diseases and pavement crack segmentation for road condition evaluation and maintenance. Currently, deep learning-based methods have achieved impressive performance on these tasks. Yet, most of them mainly focus on finding powerful deep architectures but ignore capturing the inherent curvilinear structure feature (e.g., the curvilinear structure is darker than the context) for a more robust representation. In consequence, the performance usually drops a lot on cross-datasets, which poses great challenges in practice. In this paper, we aim to improve the generalizability by introducing a novel local intensity order transformation (LIOT). Specifically, we transfer a gray-scale image into a contrast-invariant four-channel image based on the intensity order between each pixel and its nearby pixels along with the four (horizontal and vertical) directions. This results in a representation that preserves the inherent characteristic of the curvilinear structure while being robust to contrast changes. Cross-dataset evaluation on three retinal blood vessel segmentation datasets demonstrates that LIOT improves the generalizability of some state-of-the-art methods. Additionally, the cross-dataset evaluation between retinal blood vessel segmentation and pavement crack segmentation shows that LIOT is able to preserve the inherent characteristic of curvilinear structure with large appearance gaps. An implementation of the proposed method is available at https://github.com/TY-Shi/LIOT.
Nicolas Boutry, Yongchao Xu, Thierry Géraud
IEEE Trans. Image Process.3
2022 Cell Localization and Counting Using Direction Field Map
abstract
Automatic cell counting in pathology images is challenging due to blurred boundaries, low-contrast, and overlapping between cells. In this paper, we train a convolutional neural network (CNN) to predict a two-dimensional direction field map and then use it to localize cell individuals for counting. Specifically, we define a direction field on each pixel in the cell regions (obtained by dilating the original annotation in terms of cell centers) as a two-dimensional unit vector pointing from the pixel to its corresponding cell center. Direction field for adjacent pixels in different cells have opposite directions departing from each other, while those in the same cell region have directions pointing to the same center. Such unique property is used to partition overlapped cells for localization and counting. To deal with those blurred boundaries or low contrast cells, we set the direction field of the background pixels to be zeros in the ground-truth generation. Thus, adjacent pixels belonging to cells and background will have an obvious difference in the predicted direction field. To further deal with cells of varying density and overlapping issues, we adopt geometry adaptive (varying) radius for cells of different densities in the generation of ground-truth direction field map, which guides the CNN model to separate cells of different densities and overlapping cells. Extensive experimental results on three widely used datasets (i.e., VGG Cell, CRCHistoPhenotype2016, and MBM datasets) demonstrate the effectiveness of the proposed approach.
Yajie Chen, Dingkang Liang, Xiang Bai, Yongchao Xu, Xin Yang 0008
IEEE J. Biomed. Health Informatics4
2021 DeepFlux for Skeleton Detection in the Wild
Yongchao Xu, Yukang Wang, Stavros Tsogkas, Jianqiang Wan, Xiang Bai, Sven J. Dickinson, Kaleem Siddiqi
Int. J. Comput. Vis.1
2021 Gliding Vertex on the Horizontal Bounding Box for Multi-Oriented Object Detection
abstract
Object detection has recently experienced substantial progress. Yet, the widely adopted horizontal bounding box representation is not appropriate for ubiquitous oriented objects such as objects in aerial images and scene texts. In this paper, we propose a simple yet effective framework to detect multi-oriented objects. Instead of directly regressing the four vertices, we glide the vertex of the horizontal bounding box on each corresponding side to accurately describe a multi-oriented object. Specifically, We regress four length ratios characterizing the relative gliding offset on each corresponding side. This may facilitate the offset learning and avoid the confusion issue of sequential label points for oriented objects. To further remedy the confusion issue for nearly horizontal objects, we also introduce an obliquity factor based on area ratio between the object and its horizontal bounding box, guiding the selection of horizontal or oriented detection for each object. We add these five extra target variables to the regression head of faster R-CNN, which requires ignorable extra computation time. Extensive experimental results demonstrate that without bells and whistles, the proposed method achieves superior performances on multiple multi-oriented object detection benchmarks including object detection in aerial images, scene text detection, pedestrian detection in fisheye images.
Yongchao Xu, Mingtao Fu, Qimeng Wang, Yukang Wang, Kai Chen 0006, Gui-Song Xia, Xiang Bai
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Affinity Space Adaptation for Semantic Segmentation Across Domains
abstract
Semantic segmentation with dense pixel-wise annotation has achieved excellent performance thanks to deep learning. However, the generalization of semantic segmentation in the wild remains challenging. In this paper, we address the problem of unsupervised domain adaptation (UDA) in semantic segmentation. Motivated by the fact that source and target domain have invariant semantic structures, we propose to exploit such invariance across domains by leveraging co-occurring patterns between pairwise pixels in the output of structured semantic segmentation. This is different from most existing approaches that attempt to adapt domains based on individual pixel-wise information in image, feature, or output level. Specifically, we perform domain adaptation on the affinity relationship between adjacent pixels termed affinity space of source and target domain. To this end, we develop two affinity space adaptation strategies: affinity space cleaning and adversarial affinity space alignment. Extensive experiments demonstrate that the proposed method achieves superior performance against some state-of-the-art methods on several challenging benchmarks for semantic segmentation across domains. The code is available at https://github.com/idealwei/ASANet.
Wei Zhou 0068, Yukang Wang, Jiajia Chu, Jiehua Yang, Xiang Bai, Yongchao Xu
IEEE Trans. Image Process.6
2020 All You Need Is Boundary: Toward Arbitrary-Shaped Text Spotting
abstract
Recently, end-to-end text spotting that aims to detect and recognize text from cluttered images simultaneously has received particularly growing interest in computer vision. Different from the existing approaches that formulate text detection as bounding box extraction or instance segmentation, we localize a set of points on the boundary of each text instance. With the representation of such boundary points, we establish a simple yet effective scheme for end-to-end text spotting, which can read the text of arbitrary shapes. Experiments on three challenging datasets, including ICDAR2015, TotalText and COCO-Text demonstrate that the proposed method consistently surpasses the state-of-the-art in both scene text detection and end-to-end text recognition tasks.
Hao Wang 0207, Pu Lu, Hui Zhang 0085, Xiang Bai, Yongchao Xu, Mengchao He, Yongpan Wang, Wenyu Liu 0001
AAAI6
2020 Super-BPD: Super Boundary-to-Pixel Direction for Fast Image Segmentation
abstract
Image segmentation is a fundamental vision task and still remains a crucial step for many applications. In this paper, we propose a fast image segmentation method based on a novel super boundary-to-pixel direction (super-BPD) and a customized segmentation algorithm with super-BPD. Precisely, we define BPD on each pixel as a two-dimensional unit vector pointing from its nearest boundary to the pixel. In the BPD, nearby pixels from different regions have opposite directions departing from each other, and nearby pixels in the same region have directions pointing to the other or each other (i.e., around medial points). We make use of such property to partition image into super-BPDs, which are novel informative superpixels with robust direction similarity for fast grouping into segmentation regions. Extensive experimental results on BSDS500 and Pascal Context demonstrate the accuracy and efficiency of the proposed super-BPD in segmenting images. Specifically, we achieve comparable or superior performance with MCG while running at ~25fps vs 0.07fps. Super-BPD also exhibits a noteworthy transferability to unseen scenes.
Jianqiang Wan, Yang Liu 0271, Donglai Wei 0001, Xiang Bai, Yongchao Xu
CVPR5
2020 Intra-class Feature Variation Distillation for Semantic Segmentation
Yukang Wang, Tao Jiang 0002, Xiang Bai, Yongchao Xu
ECCV (7)5
2020 AutoSTR: Efficient Backbone Search for Scene Text Recognition
Hui Zhang 0085, Quanming Yao, Yongchao Xu, Xiang Bai
ECCV (24)4
2020 Learning Directional Feature Maps for Cardiac MRI Segmentation
Yukang Wang, Heshui Shi, Yukun Cao, Dandan Tu, Changzheng Zhang, Yongchao Xu
MICCAI (4)8
2020 Pay More Attention to Discontinuity for Medical Image Segmentation
Jiajia Chu, Yajie Chen, Heshui Shi, Yukun Cao, Dandan Tu, Richu Jin, Yongchao Xu
MICCAI (4)8
2020 Deep-Person: Learning discriminative deep features for person Re-Identification
Xiang Bai, Tengteng Huang, Zhiyong Dou, Rui Yu 0002, Yongchao Xu
Pattern Recognit.6
2019 DeepFlux for Skeletons in the Wild
abstract
Computing object skeletons in natural images is challenging, owing to large variations in object appearance and scale, and the complexity of handling background clutter. Many recent methods frame object skeleton detection as a binary pixel classification problem, which is similar in spirit to learning-based edge detection, as well as to semantic segmentation methods. In the present article, we depart from this strategy by training a CNN to predict a two-dimensional vector field, which maps each scene point to a candidate skeleton pixel, in the spirit of flux-based skeletonization algorithms. This ``image context flux'' representation has two major advantages over previous approaches. First, it explicitly encodes the relative position of skeletal pixels to semantically meaningful entities, such as the image points in their spatial context, and hence also the implied object boundaries. Second, since the skeleton detection context is a region-based vector field, it is better able to cope with object parts of large width. We evaluate the proposed method on three benchmark datasets for skeleton detection and two for symmetry detection, achieving consistently superior performance over state-of-the-art methods.
Yukang Wang, Yongchao Xu, Stavros Tsogkas, Xiang Bai, Sven J. Dickinson, Kaleem Siddiqi
CVPR2
2019 Learn to Scale: Generating Multipolar Normalized Density Maps for Crowd Counting
abstract
Dense crowd counting aims to predict thousands of human instances from an image, by calculating integrals of a density map over image pixels. Existing approaches mainly suffer from the extreme density variations. Such density pattern shift poses challenges even for multi-scale model ensembling. In this paper, we propose a simple yet effective approach to tackle this problem. First, a patch-level density map is extracted by a density estimation model and further grouped into several density levels which are determined over full datasets. Second, each patch density map is automatically normalized by an online center learning strategy with a multipolar center loss. Such a design can significantly condense the density distribution into several clusters, and enable that the density variance can be learned by a single model. Extensive experiments demonstrate the superiority of the proposed method. Our work outperforms the state-of-the-art by 4.2%, 14.3%, 27.1% and 20.1% in MAE, on the ShanghaiTech Part A, ShanghaiTech Part B, UCF_CC_50 and UCF-QNRF datasets, respectively.
Chenfeng Xu, Kai Qiu 0001, Jianlong Fu, Song Bai 0001, Yongchao Xu, Xiang Bai
ICCV5
2019 Visual Urban Perception with Deep Semantic-Aware Network
Yongchao Xu, Qizheng Yang, Chaoran Cui, Guangle Song, Xiaohui Han, Yilong Yin
MMM (2)1
2019 Feature context learning for human parsing
Tengteng Huang, Yongchao Xu, Song Bai 0001, Yongpan Wang, Xiang Bai
Sci. China Inf. Sci.2
2019 SegLink++: Detecting Dense and Arbitrary-shaped Scene Text by Instance-aware Component Grouping
Jun Tang 0008, Zhibo Yang 0003, Yongpan Wang, Qi Zheng 0002, Yongchao Xu, Xiang Bai
Pattern Recognit.5
2019 TextField: Learning a Deep Direction Field for Irregular Scene Text Detection
abstract
Scene text detection is an important step in the scene text reading system. The main challenges lie in significantly varied sizes and aspect ratios, arbitrary orientations, and shapes. Driven by the recent progress in deep learning, impressive performances have been achieved for multi-oriented text detection. Yet, the performance drops dramatically in detecting the curved texts due to the limited text representation (e.g., horizontal bounding boxes, rotated rectangles, or quadrilaterals). It is of great interest to detect the curved texts, which are actually very common in natural scenes. In this paper, we present a novel text detector named TextField for detecting irregular scene texts. Specifically, we learn a direction field pointing away from the nearest text boundary to each text point. This direction field is represented by an image of 2D vectors and learned via a fully convolutional neural network. It encodes both binary text mask and direction information used to separate adjacent text instances, which is challenging for the classical segmentation-based approaches. Based on the learned direction field, we apply a simple yet effective morphological-based post-processing to achieve the final detection. The experimental results show that the proposed TextField outperforms the state-of-the-art methods by a large margin (28% and 8%) on two curved text datasets: Total-Text and SCUT-CTW1500, respectively; TextField also achieves very competitive performance on multi-oriented datasets: ICDAR 2015 and MSRA-TD500. Furthermore, TextField is robust in generalizing unseen datasets.
Yongchao Xu, Yukang Wang, Wei Zhou 0068, Yongpan Wang, Zhibo Yang 0003, Xiang Bai
IEEE Trans. Image Process.1
2019 Standardized Assessment of Automatic Segmentation of White Matter Hyperintensities and Results of the WMH Segmentation Challenge
abstract
Quantification of cerebral white matter hyperintensities (WMH) of presumed vascular origin is of key importance in many neurological research studies. Currently, measurements are often still obtained from manual segmentations on brain MR images, which is a laborious procedure. The automatic WMH segmentation methods exist, but a standardized comparison of the performance of such methods is lacking. We organized a scientific challenge, in which developers could evaluate their methods on a standardized multi-center/-scanner image dataset, giving an objective comparison: the WMH Segmentation Challenge. Sixty T1 + FLAIR images from three MR scanners were released with the manual WMH segmentations for training. A test set of 110 images from five MR scanners was used for evaluation. The segmentation methods had to be containerized and submitted to the challenge organizers. Five evaluation metrics were used to rank the methods: 1) Dice similarity coefficient; 2) modified Hausdorff distance (95th percentile); 3) absolute log-transformed volume difference; 4) sensitivity for detecting individual lesions; and 5) F1-score for individual lesions. In addition, the methods were ranked on their inter-scanner robustness; 20 participants submitted their methods for evaluation. This paper provides a detailed analysis of the results. In brief, there is a cluster of four methods that rank significantly better than the other methods, with one clear winner. The inter-scanner robustness ranking shows that not all the methods generalize to unseen scanners. The challenge remains open for future submissions and provides a public platform for method evaluation.
Hugo J. Kuijf, Adrià Casamitjana, D. Louis Collins, Mahsa Dadar, Achilleas Georgiou, Mohsen Ghafoorian, Dakai Jin, April Khademi, Jesse Knight, Hongwei Li 0004, Xavier Lladó, J. Matthijs Biesbroek, Miguel Luna, Qaiser Mahmood, Richard McKinley, Alireza Mehrtash, Sébastien Ourselin, Bo-yong Park, Hyunjin Park, Simon Pezold, Élodie Puybareau, Jeroen de Bresser, Letícia Rittner, Carole H. Sudre, Sergi Valverde, Verónica Vilaplana, Roland Wiest, Yongchao Xu, Ziyue Xu 0004, Guodong Zeng, Jianguo Zhang 0001, Guoyan Zheng, Rutger Heinen, Christopher Li Hsian Chen, Wiesje M. van der Flier, Frederik Barkhof, Max A. Viergever, Geert Jan Biessels, Simon Andermatt, Mariana P. Bento, Matt Berseth, Mikhail Belyaev, Manuel Jorge Cardoso
IEEE Trans. Medical Imaging29
2019 Benchmark on Automatic Six-Month-Old Infant Brain Segmentation Algorithms: The iSeg-2017 Challenge
abstract
Accurate segmentation of infant brain magnetic resonance (MR) images into white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF) is an indispensable foundation for early studying of brain growth patterns and morphological changes in neurodevelopmental disorders. Nevertheless, in the isointense phase (approximately 6-9 months of age), due to inherent myelination and maturation process, WM and GM exhibit similar levels of intensity in both T1-weighted (T1w) and T2-weighted (T2w) MR images, making tissue segmentation very challenging. Despite many efforts were devoted to brain segmentation, only few studies have focused on the segmentation of 6-month infant brain images. With the idea of boosting methodological development in the community, iSeg-2017 challenge (http://iseg2017.web.unc.edu) provides a set of 6-month infant subjects with manual labels for training and testing the participating methods. Among the 21 automatic segmentation methods participating in iSeg-2017, we review the 8 top-ranked teams, in terms of Dice ratio, modified Hausdorff distance and average surface distance, and introduce their pipelines, implementations, as well as source codes. We further discuss limitations and possible future directions. We hope the dataset in iSeg-2017 and this review article could provide insights into methodological development for the community.
Li Wang 0026, Dong Nie, Élodie Puybareau, Jose Dolz, Qian Zhang 0066, Fan Wang 0023, Zhengwang Wu, Jiawei Chen 0001, Kim-Han Thung, Toan Duc Bui, Jitae Shin, Guodong Zeng, Guoyan Zheng, Vladimir S. Fonov, Andrew Doyle, Yongchao Xu, Pim Moeskops, Josien P. W. Pluim, Christian Desrosiers, Ismail Ben Ayed, Gerard Sanroma, Oualid M. Benkarim, Adrià Casamitjana, Verónica Vilaplana, Weili Lin, Gang Li 0001, Dinggang Shen
IEEE Trans. Medical Imaging18
2018 Hard-Aware Point-to-Set Deep Metric for Person Re-identification
Rui Yu 0002, Zhiyong Dou, Song Bai 0001, Zhaoxiang Zhang 0001, Yongchao Xu, Xiang Bai
ECCV (16)5
2018 The challenge of cerebral magnetic resonance imaging in neonates: A new method using mathematical morphology for the segmentation of structures including diffuse excessive high signal intensities
Yongchao Xu, Baptiste Morel, Sonia Dahdouh, Élodie Puybareau, Alessio Virzi, Hélène Urien, Thierry Géraud, Catherine Adamsbaum, Isabelle Bloch
Medical Image Anal.1
2017 From neonatal to adult brain MR image segmentation in a few seconds using 3D-like fully convolutional network and transfer learning
abstract
Brain magnetic resonance imaging (MRI) is widely used to assess brain development in neonates and to diagnose a wide range of neurological diseases in adults. Such studies are usually based on quantitative analysis of different brain tissues, so it is essential to be able to classify them accurately. In this paper, we propose a fast automatic method that segments 3D brain MR images into different tissues using fully convolutional network (FCN) and transfer learning. As compared to existing deep learning-based approaches that rely either on 2D patches or on fully 3D FCN, our method is way much faster: it only takes a few seconds, and only a single modality (T1 or T2) is required. In order to take the 3D information into account, all 3 successive 2D slices are stacked to form a set of 2D “color” images, which serve as input for the FCN pre-trained on ImageNet for natural image classification. To the best of our knowledge, this is the first method that applies transfer learning to segment both neonatal and adult brain 3D MR images. Our experiments on two public datasets show that our method achieves state-of-the-art results.
Yongchao Xu, Thierry Géraud, Isabelle Bloch
ICIP1
2017 Hierarchical Segmentation Using Tree-Based Shape Spaces
abstract
Current trends in image segmentation are to compute a hierarchy of image segmentations from fine to coarse. A classical approach to obtain a single meaningful image partition from a given hierarchy is to cut it in an optimal way, following the seminal approach of the scale-set theory. While interesting in many cases, the resulting segmentation, being a non-horizontal cut, is limited by the structure of the hierarchy. In this paper, we propose a novel approach that acts by transforming an input hierarchy into a new saliency map. It relies on the notion of shape space: a graph representation of a set of regions extracted from the image. Each region is characterized with an attribute describing it. We weigh the boundaries of a subset of meaningful regions (local minima) in the shape space by extinction values based on the attribute. This extinction-based saliency map represents a new hierarchy of segmentations highlighting regions having some specific characteristics. Each threshold of this map represents a segmentation which is generally different from any cut of the original hierarchy. This new approach thus enlarges the set of possible partition results that can be extracted from a given hierarchy. Qualitative and quantitative illustrations demonstrate the usefulness of the proposed method.
Yongchao Xu, Edwin Carlinet, Thierry Géraud, Laurent Najman
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 Morphology-based hierarchical representation with application to text segmentation in natural images
abstract
Many text segmentation methods are elaborate and thus are not suitable to real-time implementation on mobile devices. Having an efficient and effective method, robust to noise, blur, or uneven illumination, is interesting due to the increasing number of mobile applications needing text extraction. We propose a hierarchical image representation, based on the morphological Laplace operator, which is used to give a robust text segmentation. This representation relies on several very sound theoretical tools; its computation eventually translates to a simple labeling algorithm, and for text segmentation and grouping, to an easy tree-based processing. We also show that this method can also be applied to document binarization, with the interesting feature of getting also reverse-video text.
Lê Duy Huynh, Yongchao Xu, Thierry Géraud
ICPR2
2016 Connected Filtering on Tree-Based Shape-Spaces
abstract
Connected filters are well-known for their good contour preservation property. A popular implementation strategy relies on tree-based image representations: for example, one can compute an attribute characterizing the connected component represented by each node of the tree and keep only the nodes for which the attribute is sufficiently high. This operation can be seen as a thresholding of the tree, seen as a graph whose nodes are weighted by the attribute. Rather than being satisfied with a mere thresholding, we propose to expand on this idea, and to apply connected filters on this latest graph. Consequently, the filtering is performed not in the space of the image, but in the space of shapes built from the image. Such a processing of shape-space filtering is a generalization of the existing tree-based connected operators. Indeed, the framework includes the classical existing connected operators by attributes. It also allows us to propose a class of novel connected operators from the leveling family, based on non-increasing attributes. Finally, we also propose a new class of connected operators that we call morphological shapings. Some illustrations and quantitative evaluations demonstrate the usefulness and robustness of the proposed shape-space filters.
Yongchao Xu, Thierry Géraud, Laurent Najman
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 Hierarchical image simplification and segmentation based on Mumford-Shah-salient level line selection
Yongchao Xu, Thierry Géraud, Laurent Najman
Pattern Recognit. Lett.1
2014 Meaningful disjoint level lines selection
abstract
Many methods based on the morphological notion of shapes (i.e., connected components of level sets) have been proved to be very efficient in shape recognition and shape analysis. The inclusion relationship of the level lines (boundaries of level sets) forms the tree of shapes, a tree-based image representation with a high potential. Numerous applications using this tree representation have been proposed. In this article, we propose an efficient algorithm that extracts a set of disjoint level lines in the image. These selected level lines yield a simplified image with clean contours, which provides an intuitive idea about the main structure of the tree of shapes. Besides, we obtain a saliency map without transition problems around the contours by weighting level lines with their significance. Experimental results demonstrate the efficiency and usefulness of our method.
Yongchao Xu, Edwin Carlinet, Thierry Géraud, Laurent Najman
ICIP1
2014 Tree-Based Morse Regions: A Topological Approach to Local Feature Detection
abstract
This paper introduces a topological approach to local invariant feature detection motivated by Morse theory. We use the critical points of the graph of the intensity image, revealing directly the topology information as initial interest points. Critical points are selected from what we call a tree-based shape-space. In particular, they are selected from both the connected components of the upper level sets of the image (the Max-tree) and those of the lower level sets (the Min-tree). They correspond to specific nodes on those two trees: 1) to the leaves (extrema) and 2) to the nodes having bifurcation (saddle points). We then associate to each critical point the largest region that contains it and is topologically equivalent in its tree. We call such largest regions the tree-based Morse regions (TBMRs). The TBMR can be seen as a variant of maximally stable extremal region (MSER), which are contrasted regions. Contrarily to MSER, TBMR relies only on topological information and thus fully inherit the invariance properties of the space of shapes (e.g., invariance to affine contrast changes and covariance to continuous transformations). In particular, TBMR extracts the regions independently of the contrast, which makes it truly contrast invariant. Furthermore, it is quasi-parameter free. TBMR extraction is fast, having the same complexity as MSER. Experimentally, TBMR achieves a repeatability on par with state-of-the-art methods, but obtains a significantly higher number of features. Both the accuracy and robustness of TBMR are demonstrated by applications to image registration and 3D reconstruction.
Yongchao Xu, Pascal Monasse, Thierry Géraud, Laurent Najman
IEEE Trans. Image Process.1
2013 Salient level lines selection using the Mumford-Shah functional
abstract
Many methods relying on the morphological notion of shapes, (i.e., connected components of level sets) have been proved to be very useful for pattern analysis and recognition. Selecting meaningful level lines (boundaries of level sets) yields to simplify images while preserving salient structures. Many image simplification and/or segmentation methods are driven by the optimization of an energy functional, for instance the Mumford-Shah functional. In this article, we propose an efficient shape-based morphological filtering that very quickly compute to a locally (subordinated to the tree of shapes) optimal solution of the piecewise-constant Mumford-Shah functional. Experimental results demonstrate the efficiency, usefulness, and robustness of our method, when applied to image simplification, pre-segmentation, and detection of affine regions with viewpoint changes.
Yongchao Xu, Thierry Géraud, Laurent Najman
ICIP1
2012 Context-based energy estimator: Application to object segmentation on the tree of shapes
abstract
Image segmentation can be defined as the detection of closed contours surrounding objects of interest. Given a family of closed curves obtained by some means, a difficulty is to extract the relevant ones. A classical approach is to define an energy minimization framework, where interesting contours correspond to local minima of this energy. Active contours, graph cuts or minimum ratio cuts are instances of such approaches. In this article, we propose a novel efficient ratio-cut estimator which is both context-based and can be interpreted as an active contour. As a first example of the effectiveness of our formulation, we consider the tree of shapes, which provides a family of level lines organized in a tree hierarchy through an inclusion relationship. Thanks to the tree structure, the estimator can be computed incrementally in an efficient fashion. Experimental results on synthetic and real images demonstrate the robustness and usefulness of our method.
Yongchao Xu, Thierry Géraud, Laurent Najman
ICIP1
2012 Morphological filtering in shape spaces: Applications using tree-based image representations
Yongchao Xu, Thierry Géraud, Laurent Najman
ICPR1