Dongnan Liu

dblp:226/2662 · DBLP profile ↗
← Back
46ranked-venue papers
7as first author
39since 2021 · last 2026
0000-0001-8102-3949ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 5 first-author · 27 since 2021Artificial intelligence and machine learning · 18 · 3 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training
abstract
Qiuyi Qi, Tian Liang, Mutian Bao, Jinjian Zhang, Dongnan Liu, Wei Zhou, Linjian Mo, Ming Kong, Jie Liu, Feng Zhang, Qiang Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qiuyi Qi, Mutian Bao, Jinjian Zhang, Dongnan Liu, Linjian Mo, Ming Kong 0001
ACL (1)5
2026 Monotonic Rank Knowledge Distillation via Kendall Correlation
abstract
The computational and memory demands of deep neural networks for vision tasks remain a critical barrier to their deployment on resource-constrained edge devices. Although knowledge distillation (KD) effectively transfers over-parameterized models’ knowledge into compact students, its efficacy diminishes substantially when a significant capacity gap exists between them. Current approaches often impose linear mapping constraints between output distributions, an assumption that becomes prohibitively restrictive under such capacity gaps. This paper proposes a fundamental relaxation of alignment requirements. Specifically, rather than enforcing strict parametric relationships, we experimentally validate that preserving monotonic rank correlation between teacher and student outputs suffices for effective knowledge transfer. To operationalize this insight, we introduceMonotonic Rank Knowledge Distillation, a novel framework that leverages differentiable approximations of Kendall’s rank correlation coefficient to measure and optimize rank-order consistency. Our methodology further decomposes rank correlation into inter-class and intra-class components, ensuring the student network retains both global discriminative patterns and fine-grained categorical distinctions inherent to the teacher’s outputs. Extensive experiments across CIFAR-100 and ImageNet-1K benchmarks validate the effectiveness of our approach, demonstrating consistent performance gains over state-of-the-art distillation methods. The proposed framework achieves superior generalization across diverse architectures, including CNN-based, MLP-based, and ViT-based, with particular efficacy in various compression scenarios.
Xuewan He, Jielei Wang, Yuchen Su 0001, Dongnan Liu, Guoming Lu
IEEE Trans. Circuits Syst. Video Technol.4
2026 MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
abstract
Histopathology and transcriptomics are fundamental modalities in cancer diagnostics, encapsulating the morphological and molecular characteristics of the disease. Multi-modal self-supervised learning has demonstrated remarkable potential in learning pathological representations by integrating diverse data sources. Conventional multi-modal integration methods primarily emphasize modality alignment, while paying insufficient attention to retaining the modality-specific intrinsic structures. However, unlike conventional scenarios where multi-modal inputs often share highly overlapping features, histopathology and transcriptomics exhibit pronounced heterogeneity, offering orthogonal yet complementary insights. Histopathology data provides morphological and spatial context, elucidating tissue architecture and cellular topology, whereas transcriptomics data delineates molecular signatures through quantifying gene expression patterns. This inherent disparity introduces a major challenge in aligning these modalities while maintaining modality-specific fidelity. To address these challenges, we present MIRROR, a novel multi-modal representation learning framework designed to foster both modality alignment and retention. MIRROR employs dedicated encoders to extract comprehensive feature representations for each modality, which is further complemented by a modality alignment module to achieve seamless integration between phenotype patterns and molecular profiles. Furthermore, a modality retention module safeguards unique attributes from each modality, while a style clustering module mitigates redundancy and enhances disease-relevant information by modeling and aligning consistent pathological signatures within a clustering space. Extensive evaluations on The Cancer Genome Atlas (TCGA) cohorts for cancer subtyping and survival analysis highlight MIRROR's superior performance, demonstrating its effectiveness in constructing comprehensive oncological feature representations and benefiting the cancer diagnosis. Code is available at https://github.com/TianyiFranklinWang/MIRROR.
Jianan Fan, Dingxin Zhang 0001, Dongnan Liu, Yong Xia 0001, Heng Huang 0001, Tom Weidong Cai
IEEE Trans. Medical Imaging4
2025 ScSAM: Debiasing Morphology and Distributional Variability in Subcellular Semantic Segmentation
abstract
The significant morphological and distributional variability among subcellular components poses a long-standing challenge for learning-based organelle segmentation models, significantly increasing the risk of biased feature learning. Existing methods often rely on single mapping relationships, overlooking feature diversity and thereby inducing biased training. Although the Segment Anything Model (SAM) provides rich feature representations, its application to subcellular scenarios is hindered by two key challenges: (1) The variability in subcellular morphology and distribution creates gaps in the label space, leading the model to learn spurious or biased features. (2) SAM focuses on global contextual understanding and often ignores fine-grained spatial details, making it challenging to capture subtle structural alterations and cope with skewed data distributions. To address these challenges, we introduce ScSAM, a method that enhances feature robustness by fusing pre-trained SAM with Masked Autoencoder (MAE)-guided cellular prior knowledge to alleviate training bias from data imbalance. Specifically, we design a feature alignment and fusion module to align pre-trained embeddings to the same feature space and efficiently combine different representations. Moreover, we present a cosine similarity matrix-based class prompt encoder to activate class-specific features to recognize subcellular categories. Extensive experiments on diverse subcellular image datasets demonstrate that ScSAM outperforms state-of-the-art methods.
Jianan Fan, Dongnan Liu, Hang Chang, Gerald J. Shami, Filip Braet, Tom Weidong Cai
ECAI3
2025 DEQuant: Distribution-Enhanced Reconstruction for Post-Training Quantization
abstract
Post-training quantization (PTQ) has emerged as a promising approach for converting full-precision models into compact, low-precision models with minimal computational overhead, making them ideal for deployment in resource-constrained edge scenarios. While most existing PTQ techniques focus on minimizing the numerical discrepancy between model activations before and after quantization, such methods often overlook the inherent noise and distributional shifts caused by quantization, which can lead to severe performance degradation. To address this, we propose Distribution-Enhanced Reconstruction for PTQ (DEQuant), a novel approach that enhances the performance of quantized models by introducing a module that further enhances the alignment of activation pre- and post-quantization during model reconstruction. Extensive experiments demonstrate the effectiveness of DEQuant in several low-bit settings, achieving superior performance compared to existing methods. For instance, DEQuant achieves 14.18% accuracy on MobileNetV2 under the W2A2 configuration, representing a 5.72% improvement over the baseline QDrop and surpassing other baselines by 1–3%.
Guoming Lu, Guodong Zou, Dongnan Liu, Jielei Wang, Guangchun Luo
ICME3
2025 RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
abstract
After pre-training on extensive image-text pairs, Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of multimodal interleaved documents remains underutilized for contrastive vision-language representation learning. To fully leverage these unpaired documents, we initially establish a Real-World Data Extraction pipeline to extract high-quality images and texts. Then we design a hierarchical retrieval method to efficiently associate each image with multiple semantically relevant realistic texts. To further enhance fine-grained visual information, we propose an image semantic augmented generation module for synthetic text production. Furthermore, we employ a semantic balance sampling strategy to improve dataset diversity, enabling better learning of long-tail concepts. Based on these innovations, we construct RealSyn, a dataset combining realistic and synthetic texts, available in three scales: 15M, 30M, and 100M. We compare our dataset with other widely used datasets of equivalent scale for CLIP training. Models pre-trained on RealSyn consistently achieve state-of-the-art performance across various downstream tasks, including linear probe, zero-shot transfer, zero-shot robustness, and zero-shot retrieval. Furthermore, extensive experiments confirm that RealSyn significantly enhances contrastive vision-language representation learning and demonstrates robust scalability. The code will be released in https://garygutc.github.io/RealSyn.
Tiancheng Gu, Kaicheng Yang 0002, Chaoyi Zhang, Yin Xie, Xiang An, Ziyong Feng, Dongnan Liu, Tom Weidong Cai, Jiankang Deng
ACM Multimedia7
2025 KeyRegionPose: Region-Aware Feature Interaction and Multi-Scale Token Pruning for Efficient Human Pose Estimation
abstract
The primary challenge in deploying Human Pose Estimation (HPE) methods in real-world applications lies in balancing computational speed, model compactness, and prediction accuracy. Existing methods achieve strong performance in one or two aspects, but usually at the expense of the remaining one. To overcome this trade-off, we propose KeyRegionPose, a novel framework that achieves high accuracy while reducing model size and computational cost. Central to our design is the Region Focus Mechanism, which enables the model to concentrate on keypoint-relevant regions rather than the entire image. During training, we generate intermediate keypoint proposals to estimate keypoint-specific areas, from which the model learns region-focused features and refines predictions. To ensure accurate keypoint localization and enhance final pose estimation performance, we introduce a Cross-Representation Consistency Loss (CRC Loss) that enforces alignment between the predicted heatmaps and the regressed keypoint coordinates. Additionally, we propose Progressive Multi-Scale Token Pruning (PMTP), a strategy that prunes irrelevant tokens across multiple feature scales to accelerate inference. KeyRegionPose achieves 76.0 AP on the COCO validation set and 75.4 AP on the test-dev set, with only 20.0 million parameters and 8.6 GFLOPs—representing a 27.3% reduction in parameter count, 21.8% decrease in GFLOPs, and a competitive result (+0.2%) over state-of-the-art lightweight HPE models.
Xuanchen Wang, Heng Wang 0007, Dongnan Liu, Tom Weidong Cai
MMAsia3
2025 ORID: Organ-Regional Information Driven Framework for Radiology Report Generation
abstract
The objective of Radiology Report Generation (RRG) is to automatically generate coherent textual analyses of diseases based on radiological images, thereby alleviating the workload of radiologists. Current AI-based methods for RRG primarily focus on modifications to the encoder-decoder model architecture. To advance these approaches, this paper introduces an Organ-Regional Information Driven (ORID) framework which can effectively integrate multi-modal information and reduce the influence of noise from unrelated organs. Specifically, based on the LLaVA-Med, we first construct an RRG-related instruction dataset to improve organ-regional diagnosis description ability and get the LLaVA-Med-RRG. After that, we propose an organ-based cross-modal fusion module to effectively combine the information from the organ-regional diagnosis description and radiology image. To further reduce the influence of noise from unrelated organs on the radiology report generation, we introduce an organ importance coefficient analysis module, which leverages Graph Neural Network (GNN) to examine the interconnections of the cross-modal information of each organ region. Extensive experiments and comparisons with state-of-the-art methods across various evaluation metrics demonstrate the superior performance of our proposed method.
Tiancheng Gu, Kaicheng Yang 0002, Xiang An, Ziyong Feng, Dongnan Liu, Tom Weidong Cai
WACV5
2025 AMNCutter: Affinity-Attention-Guided Multi-View Normalized Cutter for Unsupervised Surgical Instrument Segmentation
abstract
Surgical instrument segmentation (SIS) is pivotal for robotic-assisted minimally invasive surgery, assisting surgeons by identifying surgical instruments in endoscopic video frames. Recent unsupervised surgical instrument segmentation (USIS) methods primarily rely on pseudo-labels derived from low-level features such as color and optical flow, but these methods show limited effective-ness and generalizability in complex and unseen endo-scopic scenarios. In this work, we propose a label-free unsupervised model featuring a novel module named Multi-View Normalized Cutter (m-NCutter). Different from previous USIS works, our model is trained using a graph-cutting loss function that leverages patch affini-ties for supervision, eliminating the need for pseudo-labels. The framework adaptively determines which affini-ties from which levels should be prioritized. Therefore, the low- and high-level features and their affinities are effectively integrated to train a label-free unsupervised model, showing superior effectiveness and generalization abil-ity. We conduct comprehensive experiments across mul-tiple SIS datasets to validate our approach's state-of-the-art (SOTA) performance, robustness, and exceptional potential as a pre-trained model. Our code is released at https://github.com/MingyuShengSMYIAMNCutter.
Mingyu Sheng, Jianan Fan, Dongnan Liu, Ron Kikinis, Tom Weidong Cai
WACV3
2025 Dance any Beat: Blending Beats with Visuals in Dance Video Generation
abstract
Generating dance from music is crucial for advancing automated choreography. Current methods typically produce skeleton keypoint sequences instead of dance videos and lack the capability to make specific individuals dance, which reduces their real-world applicability. These methods also require precise keypoint annotations, complicating data collection and limiting the use of self-collected video datasets. To overcome these challenges, we introduce a novel task: generating dance videos directly from images of individuals guided by music. This task enables the dance generation of specific individuals without requiring keypoint annotations, making it more versatile and applicable to various situations. Our solution, the Dance Any Beat Diffusion model (DabFusion), utilizes a reference image and a music piece to generate dance videos featuring various dance types and choreographies. The music is analyzed by our specially designed music encoder, which identifies essential features including dance style, movement, and rhythm. DabFusion excels in generating dance videos not only for individuals in the training dataset but also for any previously unseen person. This versatility stems from its approach of generating latent optical flow, which contains all necessary motion information to animate any person in the image. We evaluate DabFusion's performance using the AIST + + dataset, focusing on video quality, audio-video synchronization, and motion-music alignment. We propose a 2D Motion-Music Alignment Score (2D-MM Align), which builds on the Beat Alignment Score to more effectively evaluate motion-music alignment for this new task. Experiments show that our DabFusion establishes a solid baseline for this innovative task. Video results can be found on our project page: https://DabFusion.github.io.
Xuanchen Wang, Heng Wang 0007, Dongnan Liu, Tom Weidong Cai
WACV3
2025 On Structuring Hyperspherical Manifold for Probing Novel Biomedical Entities
abstract
The insufficient high-throughput modeling capability for high-dimensional, multiscale, and nonlinear real-world observations and measurements stands as one of the major impediments for modern science advancements. In this regard, machine learning holds tremendous promise for transforming the fundamental practice of scientific discovery by virtue of its data-driven disposition. With the ever-increasing stream of research data collection, it would be appealing to automate the exploration of patterns and insights from observational data for discovering novel classes of phenotypes and entities. However, in the discipline of biomedical investigation, the cumulative data is intrinsically subjected to non-i.i.d. distribution and severe biases amongst different clusters, inducing disorganization and ambiguity in the learned representation space. To contend with the inherent challenges, in this paper, we present a geometry- constrained probabilistic modeling treatment on hyperspherical manifolds. It firstly parameterizes the approximated posterior of instance-wise embedding as a marginal von MisesFisher distribution to account for the interference of distributional latent shift, and thereafter incorporates a suite of critical inductive biases to organically shape the layout of tailored embedding space. Together, these advancements offer a systematic solution to regularize the uncontrollable risk for unseen class learning and prospecting. Furthermore, we propose a spectral graph-theoretic method to efficiently estimate the number of potential novel classes and endow the prediction with adorable taxonomy adaptability. Through extensive experiments under various settings, we demonstrate the effectiveness and general applicability of the proposed methods in recognizing and structurally phenotyping novel visual concepts.
Jianan Fan, Dongnan Liu, Hang Chang, Heng Huang 0001, Tom Weidong Cai
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Contrastive Neuron Pruning for Backdoor Defense
abstract
Recent studies have revealed that deep neural networks (DNNs) are susceptible to backdoor attacks, in which attackers insert a pre-defined backdoor into a DNN model by poisoning a few training samples. A small subset of neurons in DNN is responsible for activating this backdoor and pruning these backdoor-associated neurons has been shown to mitigate the impact of such attacks. Current neuron pruning techniques often face challenges in accurately identifying these critical neurons, and they typically depend on the availability of labeled clean data, which is not always feasible. To address these challenges, we propose a novel defense strategy called Contrastive Neuron Pruning (CNP). This approach is based on the observation that poisoned samples tend to cluster together and are distinguishable from benign samples in the feature space of a backdoored model. Given a backdoored model, we initially apply a reversed trigger to benign samples, generating multiple positive (benign-benign) and negative (benign-poisoned) feature pairs from the backdoored model. We then employ contrastive learning on these pairs to improve the separation between benign and poisoned features. Subsequently, we identify and prune neurons in the Batch Normalization layers that show significant response differences to the generated pairs. By removing these backdoor-associated neurons, CNP effectively defends against backdoor attacks while requiring the pruning of only about 1% of the total neurons. Comprehensive experiments conducted on various benchmarks validate the efficacy of CNP, demonstrating its robustness and effectiveness in mitigating backdoor attacks compared to existing methods.
Benteng Ma, Dongnan Liu, Yanning Zhang 0001, Tom Weidong Cai, Yong Xia 0001
IEEE Trans. Image Process.3
2024 Enhancing Robustness to Noise Corruption for Point Cloud Recognition via Spatial Sorting and Set-Mixing Aggregation Module
Dingxin Zhang 0001, Jianhui Yu, Tengfei Xue, Chaoyi Zhang, Dongnan Liu, Tom Weidong Cai
ACCV (9)5
2024 Unsupervised Domain Adaptation for Tubular Structure Segmentation Across Different Anatomical Sources
Yuxiang An, Dongnan Liu, Tom Weidong Cai
BMVC2
2024 Rethinking Domain Adaptive Optic Disc and Cup Segmentation in Fundus Image through Dynamic Diffusion Flow
Canran Li, Dongnan Liu, Tom Weidong Cai
BMVC2
2024 Seeing Unseen: Discover Novel Biomedical Concepts via Geometry-Constrained Probabilistic Modeling
abstract
Machine learning holds tremendous promise for trans-forming the fundamental practice of scientific discovery by virtue of its data-driven nature. With the ever-increasing stream of research data collection, it would be appealing to autonomously explore patterns and insights from obser-vational data for discovering novel classes of phenotypes and concepts. However, in the biomedical domain, there are several challenges inherently presented in the cumu-lated data which hamper the progress of novel class dis-covery. The non-i.i.d. data distribution accompanied by the severe imbalance among different groups of classes es-sentially leads to ambiguous and biased semantic represen-tations. In this work, we present a geometry-constrained probabilistic modeling treatment to resolve the identified is-sues. First, we propose to parameterize the approximated posterior of instance embedding as a marginal von Mises-Fisher distribution to account for the interference of distri-butional latent bias. Then, we incorporate a suite of critical geometric properties to impose proper constraints on the layout of constructed embedding space, which in turn min-imizes the uncontrollable risk for unknown class learning and structuring. Furthermore, a spectral graph-theoretic method is devised to estimate the number of potential novel classes. It inherits two intriguing merits compared to exis-tent approaches, namely high computational efficiency and flexibility for taxonomy-adaptive estimation. Extensive ex-periments across various biomedical scenarios substantiate the effectiveness and general applicability of our method.
Jianan Fan, Dongnan Liu, Hang Chang, Heng Huang 0001, Tom Weidong Cai
CVPR2
2024 Revisiting Adaptive Cellular Recognition Under Domain Shifts: A Contextual Correspondence View
Jianan Fan, Dongnan Liu, Canran Li, Hang Chang, Heng Huang 0001, Filip Braet, Tom Weidong Cai
ECCV (73)2
2024 RWKV-CLIP: A Robust Vision-Language Representation Learner
abstract
Contrastive Language-Image Pre-training (CLIP) has significantly improved performance in various vision-language tasks by expanding the dataset with image-text pairs obtained from the web.This paper further explores CLIP from the perspectives of data and model architecture.To mitigate the impact of the noise data and enhance the quality of large-scale image-text data crawled from the internet, we introduce a diverse description generation framework that can leverage Large Language Models (LLMs) to combine and refine information from web-based image-text pairs, synthetic captions, and detection tags.Additionally, we propose RWKV-CLIP, the first RWKV-driven vision-language representation learning model that combines the effective parallel training of transformers with the efficient inference of RNNs.Extensive experiments across different model scales and pre-training datasets demonstrate that RWKV-CLIP is a robust vision-language representation learner and it achieves state-of-the-art performance across multiple downstream tasks, including linear probing, zero-shot classification, and zero-shot image-text retrieval.To facilitate future research, the code and pre-trained models are released at https: //github.com/deepglint/RWKV-CLIP.
Tiancheng Gu, Kaicheng Yang 0002, Xiang An, Ziyong Feng, Dongnan Liu, Tom Weidong Cai, Jiankang Deng
EMNLP5
2024 Enhancing Advanced Visual Reasoning Ability of Large Language Models
abstract
Recent advancements in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models’ advanced reasoning ability. Traditional Vision-Language models (VLMs) perform well in visual perception tasks while struggling with complex reasoning scenarios. Conversely, Large Language Models (LLMs) demonstrate robust text reasoning capabilities; however, they lack visual acuity. To bridge this gap, we propose Complex Visual Reasoning Large Language Models (CVR-LLM), capitalizing on VLMs’ visual perception proficiency and LLMs’ extensive reasoning capability. Unlike recent multimodal large language models (MLLMs) that require a projection layer, our approach transforms images into detailed, context-aware descriptions using an iterative self-refinement loop and leverages LLMs’ text knowledge for accurate predictions without extra training. We also introduce a novel multi-modal in-context learning (ICL) methodology to enhance LLMs’ contextual understanding and reasoning. Additionally, we introduce Chain-of-Comparison (CoC), a step-by-step comparison technique enabling contrasting various aspects of predictions. Our CVR-LLM presents the first comprehensive study across a wide array of complex visual reasoning tasks and achieves SOTA performance among all.
Dongnan Liu, Chaoyi Zhang, Heng Wang 0007, Tengfei Xue, Tom Weidong Cai
EMNLP2
2024 Cross-View Consistency Regularisation for Knowledge Distillation
abstract
Knowledge distillation (KD) is an established paradigm for transferring privileged knowledge from a cumbersome model to a more lightweight and efficient one. In recent years, logit-based KD methods are quickly catching up in performance with their feature-based counterparts. However, previous research has pointed out that logit-based methods are still fundamentally limited by two major issues in their training process, namely overconfident teacher and confirmation bias. Inspired by the success of cross-view learning in fields such as semi-supervised learning, in this work we introduce within-view and cross-view regularisations to standard logit-based distillation frameworks to combat the above cruxes. We also perform confidence-based soft label mining to improve the quality of distilling signals from the teacher, which further mitigates the confirmation bias problem. Despite its apparent simplicity, the proposed Consistency-Regularisation-based Logit Distillation (CRLD) significantly boosts student learning, setting new state-of-the-art results on the standard CIFAR-100, Tiny-ImageNet, and ImageNet datasets across a diversity of teacher and student architectures, whilst introducing no extra network parameters. Orthogonal to on-going logit-based distillation research, our method enjoys excellent generalisation properties and, without bells and whistles, boosts the performance of various existing approaches by considerable margins.
Dongnan Liu, Tom Weidong Cai, Chao Ma 0004
ACM Multimedia2
2024 Exploring Annotation-free Image Captioning with Retrieval-augmented Pseudo Sentence Generation
Dongnan Liu, Heng Wang 0007, Chaoyi Zhang, Tom Weidong Cai
MMAsia2
2024 Complex Organ Mask Guided Radiology Report Generation
abstract
The goal of automatic report generation is to generate a clinically accurate and coherent phrase from a single given X-ray image, which could alleviate the workload of traditional radiology reporting. However, in a real-world scenario, radiologists frequently face the challenge of producing extensive reports derived from numerous medical images, thereby medical report generation from multi-image perspective is needed. In this paper, we propose the Complex Organ Mask Guided (termed as COMG) report generation model, which incorporates masks from multiple organs (e.g., bones, lungs, heart, and mediastinum), to pro-vide more detailed information and guide the model’s attention to these crucial body regions. Specifically, we leverage prior knowledge of the disease corresponding to each organ in the fusion process to enhance the disease identification phase during the report generation process. Additionally, cosine similarity loss is introduced as target function to ensure the convergence of cross-modal consistency and facilitate model optimization. Experimental results on two public datasets show that COMG achieves a 11.4% and 9.7% improvement in terms of BLEU@4 scores over the SOTA model KiUT on IU-Xray and MIMIC, respectively. The code is publicly available at https://github.com/GaryGuTC/COMG_model.
Tiancheng Gu, Dongnan Liu, Tom Weidong Cai
WACV2
2024 Alleviating Foreground Sparsity for Semi-Supervised Monocular 3D Object Detection
abstract
Monocular 3D object detection (M3OD) is a significant yet inherently challenging task in autonomous driving due to absence of explicit depth cues in a single RGB image. In this paper, we strive to boost currently underperforming monocular 3D object detectors by leveraging an abundance of unlabelled data via semi-supervised learning. Our proposed ODM3D framework entails cross-modal knowledge distillation at various levels to inject LiDAR-domain knowledge into a monocular detector during training. By identifying foreground sparsity as a main culprit behind existing methods’ suboptimal training, we exploit the precise localisation information embedded in LiDAR points to enable more foreground-attentive and efficient distillation via the proposed BEV occupancy guidance mask, leading to notably improved knowledge transfer and M3OD performance. Besides, motivated by insights into why existing cross-modal GT-sampling techniques fail on our task at hand, we further design a novel cross-modal object-wise data augmentation strategy for effective RGB-LiDAR joint learning. Our method ranks 1stin both KITTI validation and test benchmarks, significantly surpassing all existing monocular methods, supervised or semi-supervised, on both BEV and 3D detection metrics. Code will be released at https://github.com/arcaninez/odm3d.
Dongnan Liu, Chao Ma 0004, Tom Weidong Cai
WACV2
2024 Improving multiple sclerosis lesion segmentation across clinical sites: A federated learning approach with noise-resilient training
abstract
Accurately measuring the evolution of Multiple Sclerosis (MS) with magnetic resonance imaging (MRI) critically informs understanding of disease progression and helps to direct therapeutic strategy. Deep learning models have shown promise for automatically segmenting MS lesions, but the scarcity of accurately annotated data hinders progress in this area. Obtaining sufficient data from a single clinical site is challenging and does not address the heterogeneous need for model robustness. Conversely, the collection of data from multiple sites introduces data privacy concerns and potential label noise due to varying annotation standards. To address this dilemma, we explore the use of the federated learning framework while considering label noise. Our approach enables collaboration among multiple clinical sites without compromising data privacy under a federated learning paradigm that incorporates a noise-robust training strategy based on label correction. Specifically, we introduce a Decoupled Hard Label Correction (DHLC) strategy that considers the imbalanced distribution and fuzzy boundaries of MS lesions, enabling the correction of false annotations based on prediction confidence. We also introduce a Centrally Enhanced Label Correction (CELC) strategy, which leverages the aggregated central model as a correction teacher for all sites, enhancing the reliability of the correction process. Extensive experiments conducted on two multi-site datasets demonstrate the effectiveness and robustness of our proposed methods, indicating their potential for clinical applications in multi-site collaborations to train better deep learning models with lower cost in data collection and annotation.
Lei Bai 0001, Dongang Wang, Hengrui Wang, Michael Barnett 0006, Mariano Cabezas, Tom Weidong Cai, Fernando Calamante, Kain Kyle, Dongnan Liu, Linda Ly, Aria Nguyen, Chun-Chien Shieh, Ryan Sullivan, Geng Zhan, Wanli Ouyang, Chenyu Wang 0001
Artif. Intell. Medicine9
2024 Learning to Generalize over Subpartitions for Heterogeneity-Aware Domain Adaptive Nuclei Segmentation
abstract
Abstract Annotation scarcity and cross-modality/stain data distribution shifts are two major obstacles hindering the application of deep learning models for nuclei analysis, which holds a broad spectrum of potential applications in digital pathology. Recently, unsupervised domain adaptation (UDA) methods have been proposed to mitigate the distributional gap between different imaging modalities for unsupervised nuclei segmentation in histopathology images. However, existing UDA methods are built upon the assumption that data distributions within each domain should be uniform. Based on the over-simplified supposition, they propose to align the histopathology target domain with the source domain integrally, neglecting severe intra-domain discrepancy over subpartitions incurred by mixed cancer types and sampling organs. In this paper, for the first time, we propose to explicitly consider the heterogeneity within the histopathology domain and introduce open compound domain adaptation (OCDA) to resolve the crux. In specific, a two-stage disentanglement framework is proposed to acquire domain-invariant feature representations at both image and instance levels. The holistic design addresses the limitations of existing OCDA approaches which struggle to capture instance-wise variations. Two regularization strategies are specifically devised herein to leverage the rich subpartition-specific characteristics in histopathology images and facilitate subdomain decomposition. Moreover, we propose a dual-branch nucleus shape and structure preserving module to prevent nucleus over-generation and deformation in the synthesized images. Experimental results on both cross-modality and cross-stain scenarios over a broad range of diverse datasets demonstrate the superiority of our method compared with state-of-the-art UDA and OCDA methods. Graphical abstract
Jianan Fan, Dongnan Liu, Hang Chang, Tom Weidong Cai
Int. J. Comput. Vis.2
2023 Taxonomy Adaptive Cross-Domain Adaptation in Medical Imaging via Optimization Trajectory Distillation
abstract
The success of automated medical image analysis depends on large-scale and expert-annotated training sets. Unsupervised domain adaptation (UDA) has been raised as a promising approach to alleviate the burden of labeled data collection. However, they generally operate under the closed-set adaptation setting assuming an identical label set between the source and target domains, which is over-restrictive in clinical practice where new classes commonly exist across datasets due to taxonomic inconsistency. While several methods have been presented to tackle both domain shifts and incoherent label sets, none of them take into account the common characteristics of the two issues and consider the learning dynamics along network training. In this work, we propose optimization trajectory distillation, a unified approach to address the two technical challenges from a new perspective. It exploits the low-rank nature of gradient space and devises a dual-stream distillation algorithm to regularize the learning dynamics of insufficiently annotated domain and classes with the external guidance obtained from reliable sources. Our approach resolves the issue of inadequate navigation along network optimization, which is the major obstacle in the taxonomy adaptive cross-domain adaptation scenario. We evaluate the proposed method extensively on several tasks towards various endpoints with clinical and open-world significance. The results demonstrate its effectiveness and improvements over previous methods. Code is available at https://github.com/camwew/TADA-MI.
Jianan Fan, Dongnan Liu, Hang Chang, Heng Huang 0001, Tom Weidong Cai
ICCV2
2023 Unsupervised Domain Adaptation for Neuron Membrane Segmentation based on Structural Features
abstract
AI-enhanced segmentation of neuronal boundaries in electron microscopy (EM) images is crucial for automatic and accurate neuroinformatics studies. To enhance the limited generalization ability of typical deep learning frameworks for medical image analysis, unsupervised domain adaptation (UDA) methods have been applied. In this work, we propose to improve the performance of UDA methods on cross-domain neuron membrane segmentation in EM images. First, we designed a feature weight module considering the structural features during adaptation. Second, we introduced a structural feature-based super-resolution approach to alleviating the domain gap by adjusting the cross-domain image resolutions. Third, we proposed an orthogonal decomposition module to facilitate the extraction of domain-invariant features. Extensive experiments on two domain adaptive membrane segmentation applications have indicated the effectiveness of our method.
Yuxiang An, Dongnan Liu, Tom Weidong Cai
ICME2
2023 ASRCD: Adaptive Serial Relation-Based Model for Cognitive Diagnosis
Zhuonan Liang, Dongnan Liu, Caiyun Sun, Tom Weidong Cai, Peng Fu 0003
ICONIP (14)2
2023 Topology Repairing of Disconnected Pulmonary Airways and Vessels: Baselines and a Dataset
Ziqiao Weng, Jiancheng Yang, Dongnan Liu, Tom Weidong Cai
MICCAI (7)3
2023 Decompose to Adapt: Cross-Domain Object Detection Via Feature Disentanglement
abstract
Recent advances in unsupervised domain adaptation (UDA) techniques have witnessed great success in cross-domain computer vision tasks, enhancing the generalization ability of data-driven deep learning architectures by bridging the domain distribution gaps. For the UDA-based cross-domain object detection methods, the majority of them alleviate the domain bias by inducing the domain-invariant feature generation via adversarial learning strategy. However, their domain discriminators have limited classification ability due to the unstable adversarial training process. Therefore, the extracted features induced by them cannot be perfectly domain-invariant and still contain domain-private factors, bringing obstacles to further alleviate the cross-domain discrepancy. To tackle this issue, we design a Domain Disentanglement Faster-RCNN (DDF) to eliminate the source-specific information in the features for detection task learning. Our DDF method facilitates the feature disentanglement at the global and local stages, with a Global Triplet Disentanglement (GTD) module and an Instance Similarity Disentanglement (ISD) module, respectively. By outperforming state-of-the-art methods on four benchmark UDA object detection tasks, our DDF method is demonstrated to be effective with wide applicability.
Dongnan Liu, Chaoyi Zhang, Yang Song 0001, Heng Huang 0001, Chenyu Wang 0001, Michael Barnett 0006, Tom Weidong Cai
IEEE Trans. Multim.1
2022 Unsupervised Domain Adaptive Fundus Image Segmentation with Few Labeled Source Data
Qianbi Yu, Dongnan Liu, Chaoyi Zhang, Xinwen Zhang, Tom Weidong Cai
BMVC2
2022 Domain Adaptive Nuclei Instance Segmentation and Classification via Category-Aware Feature Alignment and Pseudo-Labelling
Canran Li, Dongnan Liu, Haoran Li 0024, Zheng Zhang 0006, Guangming Lu 0002, Xiaojun Chang, Tom Weidong Cai
MICCAI (8)2
2022 Towards bi-directional skip connections in encoder-decoder architectures and beyond
Tiange Xiang, Chaoyi Zhang, Xinyi Wang 0015, Yang Song 0001, Dongnan Liu, Heng Huang 0001, Tom Weidong Cai
Medical Image Anal.5
2022 Multiple Sclerosis Lesion Analysis in Brain Magnetic Resonance Images: Techniques and Clinical Applications
abstract
Multiple sclerosis (MS) is a chronic inflammatory and degenerative disease of the central nervous system, characterized by the appearance of focal lesions in the white and gray matter that topographically correlate with an individual patient's neurological symptoms and signs. Magnetic resonance imaging (MRI) provides detailed in-vivo structural information, permitting the quantification and categorization of MS lesions that critically inform disease management. Traditionally, MS lesions have been manually annotated on 2D MRI slices, a process that is inefficient and prone to inter-/intra-observer errors. Recently, automated statistical imaging analysis techniques have been proposed to detect and segment MS lesions based on MRI voxel intensity. However, their effectiveness is limited by the heterogeneity of both MRI data acquisition techniques and the appearance of MS lesions. By learning complex lesion representations directly from images, deep learning techniques have achieved remarkable breakthroughs in the MS lesion segmentation task. Here, we provide a comprehensive review of state-of-the-art automatic statistical and deep-learning MS segmentation methods and discuss current and future clinical applications. Further, we review technical strategies, such as domain adaptation, to enhance MS lesion segmentation in real-world clinical settings.
Chaoyi Zhang, Mariano Cabezas, Yang Song 0001, Zihao Tang 0002, Dongnan Liu, Tom Weidong Cai, Michael Barnett 0006, Chenyu Wang 0001
IEEE J. Biomed. Health Informatics6
2022 DSNet: A Dual-Stream Framework for Weakly-Supervised Gigapixel Pathology Image Analysis
abstract
We present a novel weakly-supervised framework for classifying whole slide images (WSIs). WSIs, due to their gigapixel resolution, are commonly processed by patch-wise classification with patch-level labels. However, patch-level labels require precise annotations, which is expensive and usually unavailable on clinical data. With image-level labels only, patch-wise classification would be sub-optimal due to inconsistency between the patch appearance and image-level label. To address this issue, we posit that WSI analysis can be effectively conducted by integrating information at both high magnification (local) and low magnification (regional) levels. We auto-encode the visual signals in each patch into a latent embedding vector representing local information, and down-sample the raw WSI to hardware-acceptable thumbnails representing regional information. The WSI label is then predicted with a Dual-Stream Network (DSNet), which takes the transformed local patch embeddings and multi-scale thumbnail images as inputs and can be trained by the image-level label only. Experiments conducted on three large-scale public datasets demonstrate that our method outperforms all recent state-of-the-art weakly-supervised WSI classification methods.
Tiange Xiang, Yang Song 0001, Chaoyi Zhang, Dongnan Liu, Fan Zhang 0013, Heng Huang 0001, Lauren O'Donnell, Tom Weidong Cai
IEEE Trans. Medical Imaging4
2021 LG-Net: Lesion Gate Network for Multiple Sclerosis Lesion Inpainting
Zihao Tang 0002, Mariano Cabezas, Dongnan Liu, Michael Barnett 0006, Tom Weidong Cai, Chenyu Wang 0001
MICCAI (7)3
2021 BiX-NAS: Searching Efficient Bi-directional Architecture for Medical Image Segmentation
Xinyi Wang 0015, Tiange Xiang, Chaoyi Zhang, Yang Song 0001, Dongnan Liu, Heng Huang 0001, Tom Weidong Cai
MICCAI (1)5
2021 Panoptic Feature Fusion Net: A Novel Instance Segmentation Paradigm for Biomedical and Biological Images
abstract
Instance segmentation is an important task for biomedical and biological image analysis. Due to the complicated background components, the high variability of object appearances, numerous overlapping objects, and ambiguous object boundaries, this task still remains challenging. Recently, deep learning based methods have been widely employed to solve these problems and can be categorized into proposal-free and proposal-based methods. However, both proposal-free and proposal-based methods suffer from information loss, as they focus on either global-level semantic or local-level instance features. To tackle this issue, we present a Panoptic Feature Fusion Net (PFFNet) that unifies the semantic and instance features in this work. Specifically, our proposed PFFNet contains a residual attention feature fusion mechanism to incorporate the instance prediction with the semantic features, in order to facilitate the semantic contextual information learning in the instance branch. Then, a mask quality sub-branch is designed to align the confidence score of each object with the quality of the mask prediction. Furthermore, a consistency regularization mechanism is designed between the semantic segmentation tasks in the semantic and instance branches, for the robust learning of both tasks. Extensive experiments demonstrate the effectiveness of our proposed PFFNet, which outperforms several state-of-the-art methods on various biomedical and biological datasets.
Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Heng Huang 0001, Tom Weidong Cai
IEEE Trans. Image Process.1
2021 PDAM: A Panoptic-Level Feature Alignment Framework for Unsupervised Domain Adaptive Instance Segmentation in Microscopy Images
abstract
In this work, we present an unsupervised domain adaptation (UDA) method, named Panoptic Domain Adaptive Mask R-CNN (PDAM), for unsupervised instance segmentation in microscopy images. Since there currently lack methods particularly for UDA instance segmentation, we first design a Domain Adaptive Mask R-CNN (DAM) as the baseline, with cross-domain feature alignment at the image and instance levels. In addition to the image- and instance-level domain discrepancy, there also exists domain bias at the semantic level in the contextual information. Next, we, therefore, design a semantic segmentation branch with a domain discriminator to bridge the domain gap at the contextual level. By integrating the semantic- and instance-level feature adaptation, our method aligns the cross-domain features at the panoptic level. Third, we propose a task re-weighting mechanism to assign trade-off weights for the detection and segmentation loss functions. The task re-weighting mechanism solves the domain bias issue by alleviating the task learning for some iterations when the features contain source-specific factors. Furthermore, we design a feature similarity maximization mechanism to facilitate instance-level feature adaptation from the perspective of representational learning. Different from the typical feature alignment methods, our feature similarity maximization mechanism separates the domain-invariant and domain-specific features by enlarging their feature distribution dependency. Experimental results on three UDA instance segmentation scenarios with five datasets demonstrate the effectiveness of our proposed PDAM method, which outperforms state-of-the-art UDA methods by a large margin.
Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Fan Zhang 0013, Lauren O'Donnell, Heng Huang 0001, Tom Weidong Cai
IEEE Trans. Medical Imaging1
2020 Unsupervised Instance Segmentation in Microscopy Images via Panoptic Domain Adaptation and Task Re-Weighting
abstract
Unsupervised domain adaptation (UDA) for nuclei instance segmentation is important for digital pathology, as it alleviates the burden of labor-intensive annotation and domain shift across datasets. In this work, we propose a Cycle Consistency Panoptic Domain Adaptive Mask R-CNN (CyC-PDAM) architecture for unsupervised nuclei segmentation in histopathology images, by learning from fluorescence microscopy images. More specifically, we first propose a nuclei inpainting mechanism to remove the auxiliary generated objects in the synthesized images. Secondly, a semantic branch with a domain discriminator is designed to achieve panoptic-level domain adaptation. Thirdly, in order to avoid the influence of the source-biased features, we propose a task re-weighting mechanism to dynamically add trade-off weights for the task-specific loss functions. Experimental results on three datasets indicate that our proposed method outperforms state-of-the-art UDA methods significantly, and demonstrates a similar performance as fully supervised methods.
Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Fan Zhang 0013, Lauren O'Donnell, Heng Huang 0001, Tom Weidong Cai
CVPR1
2020 BiO-Net: Learning Recurrent Bi-directional Connections for Encoder-Decoder Architecture
Tiange Xiang, Chaoyi Zhang, Dongnan Liu, Yang Song 0001, Heng Huang 0001, Tom Weidong Cai
MICCAI (1)3
2019 Nuclei Segmentation via a Deep Panoptic Model with Semantic Feature Fusion
abstract
Automated detection and segmentation of individual nuclei in histopathology images is important for cancer diagnosis and prognosis. Due to the high variability of nuclei appearances and numerous overlapping objects, this task still remains challenging. Deep learning based semantic and instance segmentation models have been proposed to address the challenges, but these methods tend to concentrate on either the global or local features and hence still suffer from information loss. In this work, we propose a panoptic segmentation model which incorporates an auxiliary semantic segmentation branch with the instance branch to integrate global and local features. Furthermore, we design a feature map fusion mechanism in the instance branch and a new mask generator to prevent information loss. Experimental results on three different histopathology datasets demonstrate that our method outperforms the state-of-the-art nuclei segmentation methods and popular semantic and instance segmentation models by a large margin.
Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Chaoyi Zhang, Fan Zhang 0013, Lauren O'Donnell, Tom Weidong Cai
IJCAI1
2019 Vessel-Net: Retinal Vessel Segmentation Under Multi-path Supervision
Yicheng Wu 0001, Yong Xia 0001, Yang Song 0001, Donghao Zhang 0004, Dongnan Liu, Chaoyi Zhang, Tom Weidong Cai
MICCAI (1)5
2018 Densely Connected Large Kernel Convolutional Network for Semantic Membrane Segmentation in Microscopy Images
abstract
Structural analysis of neurons can provide valuable insights of brain function. Semantic segmentation of neurons thus becomes an important technique in bioinformatics. Deep learning approaches have shown promising performance in various semantic segmentation problems. However, segmentation of neurons in Electron Microscopy (EM) images has some differences compared with typical segmentation tasks due to the image noise and the disturbance of the intracellular structures. In our work, we propose a network with a ResNet encoder and densely connected decoder with large kernels, and then refinement with simple morphological post-possessing. Two main advantages of our method are: 1) the network can prevent the loss of high-resolution information and enlarge the reception field; 2) the post-processing method is simple and can be directly applied to the probability map from the network to enhance the unconfident area. Evaluated on the ISBI2012 EM membrane segmentation challenge, the proposed method achieves competitive performance.
Dongnan Liu, Donghao Zhang 0004, Siqi Liu 0001, Yang Song 0001, Haozhe Jia, David Dagan Feng, Yong Xia 0001, Tom Weidong Cai
ICIP1
2018 3D Large Kernel Anisotropic Network for Brain Tumor Segmentation
Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Fan Zhang 0013, Lauren O'Donnell, Tom Weidong Cai
ICONIP (7)1
2018 Panoptic Segmentation with an End-to-End Cell R-CNN for Pathology Image Analysis
Donghao Zhang 0004, Yang Song 0001, Dongnan Liu, Haozhe Jia, Siqi Liu 0001, Yong Xia 0001, Heng Huang 0001, Tom Weidong Cai
MICCAI (2)3