VLDB 2026 Research / reviewers in the wild / expert
Jiahua Dong 0001
dblp:247/5746
· DBLP profile ↗
58ranked-venue papers
18as first author
55since 2021 · last 2026
0000-0001-8545-4447ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 16 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 11 first-author · 26 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Video SimpleQA: Towards Factuality Evaluation in Large Video Language ModelsabstractRecent advancements in Large Video Language Models (LVLMs) have highlighted their potential for multi-modal understanding, yet evaluating their factual grounding in videos remains a critical unsolved challenge. To address this gap, we introduce Video SimpleQA, the first comprehensive benchmark tailored for factuality evaluation in video contexts. Our work differs from existing video benchmarks through the following key features: 1) Knowledge required: demanding integration of external knowledge beyond the video’s explicit narrative; 2) Multi-hop fact-seeking question: Each question involves multiple explicit facts and requires strict factual grounding without hypothetical or subjective inferences. We include per-hop single-fact-based sub-QAs alongside final QAs to enable fine-grained, step-by-step evaluation; 3) Short-form definitive answer: Answers are crafted as unambiguous and definitively correct in a short format with minimal scoring variance; 4) Temporal grounded required: Requiring answers to rely on one or more temporal segments in videos, rather than single frames. We extensively evaluate 33 state-of-the-art LVLMs and summarize key findings as follows: 1) Current LVLMs exhibit notable deficiencies in factual adherence, with the best-performing model o3 merely achieving an F-score of 66.3%; 2) Most LVLMs are overconfident in what they generate, with self-stated confidence exceeding actual accuracy; 3) Retrieval-Augmented Generation demonstrates consistent improvements at the cost of additional inference time overhead; 4) Multi-hop QA demonstrates substantially degraded performance compared to single-hop sub-QAs, with first-hop object/event recognition emerging as the primary bottleneck. We position Video SimpleQA as the cornerstone benchmark for video factuality assessment, aiming to steer LVLM development toward verifiable grounding in real-world contexts. Meng Cao 0002, Yingyao Wang, Jihao Gu, Haoze Zhao, Jiahua Dong 0001, Wangbo Yu, Ge Zhang 0009, Xiang Li 0117, Ian Reid 0003, Xiaodan Liang |
AAAI | 8 |
| 2026 | Towards Efficient and Effective Interactive 3D SegmentationabstractInteractive 3D segmentation embodies an advanced human-in-the-loop paradigm, where a model iteratively refines the segmentation of interested objects within a 3D point cloud through user feedback. Existing methods have achieved notable advancements at the expense of substantial resource consumption. To address this challenge, we introduce E2I3D, an efficient and effective model for interactive 3D segmentation. Specifically, we propose a two-stage efficiency-to-effectiveness framework to decouple efficiency and effectiveness, avoiding the high training cost of joint optimization. For efficiency in the first stage, we present heterogeneous pruning, which reliably compresses the model by ranking and pruning the constructed heterogeneous groups separately based on gradient compensation. For effectiveness in the second stage, we design hierarchical click-aware attention that integrates geometric details from high-resolution features with global context from low-resolution features to enhance click-guided interaction. Extensive experiments across public datasets demonstrate that E2I3D exceeds state-of-the-art methods in both efficiency and effectiveness. For instance, on the KITTI-360 dataset, E2I3D boosts the IoU for interactive single-object segmentation from 44.4% to 49.0% with 5 user clicks, while simultaneously reducing parameters from 39.3M to 5.7M. Wei Cong, Yang Cong, Jiahua Dong 0001, Gan Sun |
AAAI | 3 |
| 2026 | Bring Your Dreams to Life: Continual Text-to-Video CustomizationabstractCustomized text-to-video generation (CTVG) has recently witnessed great progress in generating tailored videos from user-specific text. However, most CTVG methods assume that personalized concepts remain static and do not expand incrementally over time. Additionally, they struggle with forgetting and concept neglect when continuously learning new concepts, including subjects and motions. To resolve the above challenges, we develop a novel Continual Customized Video Diffusion (CCVD) model, which can continuously learn new concepts to generate videos across various text-to-video generation tasks by tackling forgetting and concept neglect. To address catastrophic forgetting, we introduce a concept-specific attribute retention module and a task-aware concept aggregation strategy. They can capture the unique characteristics and identities of old concepts during training, while combining all subject and motion adapters of old concepts based on their relevance during testing. Besides, to tackle concept neglect, we develop a controllable conditional synthesis to enhance regional features and align video contexts with user conditions, by incorporating layer-specific region attention-guided noise estimation. Extensive experimental comparisons demonstrate that our CCVD outperforms existing CTVG models. Jiahua Dong 0001, Wenqi Liang, Zongyan Han, Meng Cao 0002, Duzhen Zhang, Hanbin Zhao, Zhi Han, Salman Khan 0001, Fahad Shahbaz Khan |
AAAI | 1 |
| 2026 | SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical PlanningabstractSequential-Horizon Vision-and-Language Navigation (SH-VLN) presents a challenging scenario where agents should sequentially execute multi-task trajectory navigation guided by complex, long-horizon natural language instructions. Current vision-and-language navigation models exhibit significant performance degradation with such instructions, as information overload impairs the agent's ability to attend to observationally relevant details. To address this problem, we propose SeqWalker, a novel navigation model built on a hierarchical planning framework. Our SeqWalker features: (1) A High-Level Planner that dynamically selects global instructions into contextually relevant sub-instructions based on the agent's current visual observations, thus reducing cognitive load; (2) A Low-Level Planner incorporating an Exploration-Verification strategy that leverages the inherent logical structure of instructions for trajectory error correction. To evaluate SH-VLN performance, we also extend the IVLN dataset and establish a new benchmark. Extensive experiments are performed to demonstrate the effectiveness and superiority of SeqWalker. Zebin Han, Baichen Liu, Qi Lyu, Zhenduo Shang, Jiahua Dong 0001, Lianqing Liu, Zhi Han |
AAAI | 6 |
| 2026 | Lifelong Language-Conditioned Robotic Manipulation Learning
Zebin Han, Gan Li, Jiahua Dong 0001, Baichen Liu, Lianqing Liu, Zhi Han |
AAAI | 5 |
| 2026 | CE-SDWV: Effective and Efficient Concept Erasure for Text-to-Image Diffusion Models via a Semantic-Driven Word Vocabulary
Jiahang Tu, Jiahua Dong 0001, Hanbin Zhao, Chao Zhang 0001, Nicu Sebe, Hui Qian 0001 |
Int. J. Comput. Vis. | 3 |
| 2026 | Learning From Each Other: Generalized Federated Incremental Semantic SegmentationabstractFederated learning (FL) has advanced semantic segmentation through decentralized training to reduce annotation costs. However, most FL-based semantic segmentation methods assume fixed foreground classes, resulting in catastrophic forgetting of old categories when local clients continually collect streaming data of new classes without storing old categories. Moreover, the irregular participation of new local clients with novel classes unseen by others may exacerbate heterogeneous forgetting across clients during global FL training. To resolve the above challenges, we propose a Hierarchical Forgetting Alleviation (HFA) model. By tackling forgetting within and across local clients, our model ensures that all local clients learn from each other as they continuously learn new categories. Specifically, to alleviate class-imbalanced forgetting within local clients induced by background shift, we develop a confidence-regularized pseudo labeling strategy to produce class-balanced soft pseudo labels for old categories that are labeled as background. Guided by soft pseudo labels, we design a graph-induced relation matching loss and a forgetting-balanced gradient propagation module to tackle ambiguous inter-class relations and class-imbalanced gradient propagation among old classes. Besides, a novel task detection module and an adaptive DBSCAN clustering are devised to address inter-client heterogeneous forgetting. They detect the arrival of new tasks to store the old global model for local pseudo labeling and distillation, while supplying global class prototypes for modeling inter-class relations and warm-starting global classifier. Experiments on multiple datasets verify our model's superiority over other methods. Jiahua Dong 0001, Wenqi Liang, Yang Cong, Gan Sun, Lixu Wang, Henghui Ding, Yulun Zhang 0001, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | From System 1 to System 2: A Survey of Reasoning Large Language ModelsabstractAchieving human-level intelligence requires refining the transition from the fast, intuitive System 1 to the slower, more deliberate System 2 reasoning. While System 1 excels in quick, heuristic decisions, System 2 relies on logical reasoning for more accurate judgments and reduced biases. Foundational Large Language Models (LLMs) excel at fast decision-making but lack the depth for complex reasoning, as they have not yet fully embraced the step-by-step analysis characteristic of true System 2 thinking. Recently, reasoning LLMs like OpenAI's o1/o3 and DeepSeek's R1 have demonstrated expert-level performance in fields such as mathematics and coding, closely mimicking the deliberate reasoning of System 2 and showcasing human-like cognitive abilities. This survey begins with a brief overview of the progress in foundational LLMs and the early development of System 2 technologies, exploring how their combination has paved the way for reasoning LLMs. Next, we discuss how to construct reasoning LLMs, trace the evolution of various reasoning models, and examine the core methods that enable advanced reasoning behind them. Additionally, we provide an overview of reasoning benchmarks, offering an in-depth comparison of the performance of representative reasoning LLMs. Finally, we explore promising directions for advancing reasoning LLMs and maintain a real-time GitHub Repository to track the latest developments. We hope this survey will serve as a valuable resource to inspire innovation and drive progress in this rapidly evolving field. Duzhen Zhang, Zhongzhi Li, Jiaxin Zhang 0024, Zengyan Liu, Junhao Zheng, Xiuyi Chen, Jiahua Dong 0001, Zhijiang Guo, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 12 |
| 2026 | IAP: Improving Continual Learning of Vision-Language Models via Instance-Aware PromptingabstractRecent pre-trained vision-language models (PT-VLMs) often face a Multi-Domain Task Incremental Learning (MTIL) scenario in practice, where several classes and domains of multi-modal tasks are arrive incrementally. Without access to previously seen tasks and unseen tasks, memory-constrained MTIL suffers from forward and backward forgetting. To alleviate the above challenges, parameter-efficient fine-tuning techniques (PEFT), such as prompt tuning, are employed to adapt the PT-VLM to the diverse incrementally learned tasks. To achieve effective new task adaptation, existing methods only consider the effect of PEFT strategy selection, but neglect the influence of PEFT parameter setting (e.g., prompting). In this paper, we tackle the challenge of optimizing prompt designs for diverse tasks in MTIL and propose an Instance-Aware Prompting (IAP) framework. Specifically, our Instance-Aware Gated Prompting (IA-GP) strategy enhances adaptation to new tasks while mitigating forgetting by adaptively assigning prompts across transformer layers at the instance level. Our Instance-Aware Class-Distribution-Driven Prompting (IA-CDDP) improves the task adaptation process by determining an accurate task-label-related confidence score for each instance. Experimental evaluations across 11 datasets, using three performance metrics, demonstrate the effectiveness of our proposed method. The source codes are available at https://github.com/FerdinandZJU/IAP. Hao Fu 0023, Hanbin Zhao, Jiahua Dong 0001, Henghui Ding, Chao Zhang 0001, Hui Qian 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance SegmentationabstractContinual video instance segmentation (CVIS) requires the plasticity to absorb new categories while maintaining the stability to retain previously learned knowledge. Crucially, the model must also preserve temporal consistency of instances across video frames. In this work, we introduce Contrastive Residual Injection and Semantic Prompting (CRISP), a framework tailored to address instance-wise, category-wise, and task-wise confusion in CVIS. For instance-wise learning, we model instance tracking and construct instance correlation loss, which emphasizes the correlation with the prior query space while strengthening the specificity of the current task query. For category-wise learning, we build an adaptive residual semantic prompt (ARSP) learning framework, which constructs a learnable semantic residual prompt pool generated by category text and uses an adjustive query-prompt matching mechanism to build a mapping relationship between the query of the current task and the semantic residual prompt. Meanwhile, a semantic consistency loss based on the contrastive learning is introduced to maintain semantic coherence between object queries and residual prompts during incremental training. For task-wise learning, to ensure the correlation at the inter-task level within the query space, we introduce a concise yet powerful initialization strategy for incremental prompts. Extensive experiments on YouTube-VIS-2019 and YouTube-VIS-2021 datasets demonstrate that CRISP significantly outperforms existing continual segmentation methods in the long-term continual video instance segmentation task, avoiding catastrophic forgetting and effectively improving segmentation and classification performance. The code is available at https://github.com/LyuQi127/CRISP. Baichen Liu, Qi Lyu, Jiahua Dong 0001, Lianqing Liu, Zhi Han |
IEEE Trans. Image Process. | 4 |
| 2025 | Enhancing Multimodal Continual Instruction Tuning with BranchLoRAabstractMultimodal Continual Instruction Tuning (MCIT) aims to finetune Multimodal Large Language Models (MLLMs) to continually align with human intent across sequential tasks. Existing approaches often rely on the Mixture-of-Experts (MoE) LoRA framework to preserve previous instruction alignments. However, these methods are prone to Catastrophic Forgetting (CF), as they aggregate all LoRA blocks via simple summation, which compromises performance over time. In this paper, we identify a critical parameter inefficiency in the MoELoRA framework within the MCIT context. Based on this insight, we propose BranchLoRA, an asymmetric framework to enhance both efficiency and performance. To mitigate CF, we introduce a flexible tuning-freezing mechanism within BranchLoRA, enabling branches to specialize in intra-task knowledge while fostering inter-task collaboration. Moreover, we incrementally incorporate task-specific routers to ensure an optimal branch distribution over time, rather than favoring the most recent task. To streamline inference, we introduce a task selector that automatically routes test inputs to the appropriate router without requiring task identity. Extensive experiments on the latest MCIT benchmark demonstrate that BranchLoRA significantly outperforms MoELoRA and maintains its superiority across various MLLM sizes. Duzhen Zhang, Yong Ren 0006, Zhongzhi Li, Yahan Yu, Jiahua Dong 0001, Chenxing Li, Zhilong Ji, Jinfeng Bai |
ACL (1) | 5 |
| 2025 | Hierarchical Visual Prompt Learning for Continual Video Instance SegmentationabstractVideo instance segmentation (VIS) has gained significant attention for its capability in tracking and segmenting object instances across video frames. However, most of the existing VIS approaches unrealistically assume that the categories of object instances remain fixed over time. Moreover, they experience catastrophic forgetting of old classes when required to continuously learn object instances belonging to new categories. To resolve these challenges, we develop a novel Hierarchical Visual Prompt Learning (HVPL) model that overcomes catastrophic forgetting of previous categories from both frame-level and video-level perspectives. Specifically, to mitigate forgetting at the frame level, we devise a task-specific frame prompt and an orthogonal gradient correction (OGC) module. The OGC module helps the frame prompt encode task-specific global instance information for new classes in each individual frame by projecting its gradients onto the orthogonal feature space of old classes. Furthermore, to address forgetting at the video level, we design a task-specific video prompt and a video context decoder. This decoder first embeds structural inter-class relationships across frames into the frame prompt features, and then propagates task-specific global video contexts from the frame prompt features to the video prompt. Through rigorous comparisons, our HVPL model proves to be more effective than baseline approaches. The code is available at https://github.com/JiahuaDong/HVPL. Jiahua Dong 0001, Wenqi Liang, Hanbin Zhao, Henghui Ding, Nicu Sebe, Salman Khan 0001, Fahad Shahbaz Khan |
ICCV | 1 |
| 2025 | All in One: Visual-Description-Guided Unified Point Cloud SegmentationabstractUnified segmentation of 3D point clouds is crucial for scene understanding, but is hindered by its sparse structure, limited annotations, and the challenge of distinguishing fine-grained object classes in complex environments. Existing methods often struggle to capture rich semantic and contextual information due to limited supervision and a lack of diverse multimodal cues, leading to suboptimal differentiation of classes and instances. To address these challenges, we propose VDG-Uni3DSeg, a novel framework that integrates pre-trained vision-language models (e.g., CLIP) and large language models (LLMs) to enhance 3D segmentation. By leveraging LLM-generated textual descriptions and reference images from the internet, our method incorporates rich multimodal cues, facilitating fine-grained class and instance separation. We further design a Semantic-Visual Contrastive Loss to align point features with multimodal queries and a Spatial Enhanced Module to model scene-wide relationships efficiently. Operating within a closed-set paradigm that utilizes multimodal knowledge generated offline, VDG-Uni3DSeg achieves state-of-the-art results in semantic, instance, and panoptic segmentation, offering a scalable and practical solution for 3D understanding. Our code is available at https://github.com/Hanzy1996/VDG-Uni3DSeg. Zongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong 0001, Rao Muhammad Anwer |
ICCV | 3 |
| 2025 | Rehearsal-free Federated Domain-incremental LearningabstractWe introduce a rehearsal-free federated domain incremental learning framework, RefFiL, based on a global prompt-sharing paradigm to alleviate catastrophic forgetting challenges in federated domain-incremental learning, where unseen domains are continually learned. Typical methods for mitigating forgetting, such as the use of additional datasets and the retention of private data from earlier tasks, are not viable in federated learning (FL) due to devices’ limited resources. Our method, RefFiL, addresses this by learning domain-invariant knowledge and incorporating various domain-specific prompts from the domains represented by different FL participants. A key feature of RefFiL is the generation of local finegrained prompts by our domain adaptive prompt generator, which effectively learns from local domain knowledge while maintaining distinctive boundaries on a global scale. We also introduce a domain-specific prompt contrastive learning loss that differentiates between locally generated prompts and those from other domains, enhancing RefFiL’s precision and effectiveness. Compared to existing methods, RefFiL significantly alleviates catastrophic forgetting without requiring extra memory space, making it ideal for privacy-sensitive and resource-constrained devices. Rui Sun 0010, Haoran Duan 0001, Jiahua Dong 0001, Varun Ojha 0001, Tejal Shah, Rajiv Ranjan 0001 |
ICDCS | 3 |
| 2025 | Complementary Information Guided Occupancy Prediction via Multi-Level Representation FusionabstractCamera-based occupancy prediction is a main-stream approach for 3D perception in autonomous driving, aiming to infer complete 3D scene geometry and semantics from 2D images. Almost existing methods focus on improving performance through structural modifications, such as lightweight backbones and complex cascaded frameworks, with good yet limited performance. Few studies explore from the perspective of representation fusion, leaving the rich diversity of features in 2D images underutilized. Motivated by this, we propose CIGOcc, a two-stage occupancy prediction framework based on multi-level representation fusion. CIGOcc extracts segmentation, graphics, and depth features from an input image and introduces a deformable multi-level fusion mechanism to fuse these three multi-level features. Additionally, CIGOcc incorporates knowledge distilled from SAM to further enhance prediction accuracy. Without increasing training costs, CIGOcc achieves state-of-the-art performance on the SemanticKITTI benchmark. The code is provided in the supplementary material and will be released project page. Rongtao Xu, Jinzhou Lin 0001, Jialei Zhou, Jiahua Dong 0001, Changwei Wang 0001, Ruisheng Wang 0001, Li Guo 0004, Shibiao Xu, Xiaodan Liang |
ICRA | 4 |
| 2025 | Resource-Constrained Federated Continual Learning: What Does Matter?abstractFederated Continual Learning (FCL) aims to enable sequential privacy-preserving model training on streams of incoming data that vary in edge devices by preserving previous knowledge while adapting to new data. Current FCL literature focuses on restricted data privacy and access to previously seen data while imposing no constraints on the training overhead. This is unreasonable for FCL applications in real-world scenarios, where edge devices are primarily constrained by resources such as storage, computational budget, and label rate. We revisit this problem with a large-scale benchmark and analyze the performance of state-of-the-art FCL approaches under different resource-constrained settings. Various typical FCL techniques and six datasets in two incremental learning scenarios (Class-IL and Domain-IL) are involved in our experiments. Through extensive experiments amounting to a total of over 1,000+ GPU hours, we find that, under limited resource-constrained settings, existing FCL approaches, with no exception, fail to achieve the expected performance. Our conclusions are consistent in the sensitivity analysis. This suggests that most existing FCL methods are particularly too resource-dependent for real-world deployment. Moreover, we study the performance of typical FCL techniques with resource constraints and shed light on future research directions in FCL. Yichen Li 0006, Jiahua Dong 0001, Haozhao Wang, Yining Qi, Rui Zhang 0003, Ruixuan Li 0001 |
NeurIPS | 3 |
| 2025 | Feature Distillation is the Better Choice for Model-Heterogeneous Federated LearningabstractModel-Heterogeneous Federated Learning (Hetero-FL) has attracted growing attention for its ability to aggregate knowledge from heterogeneous models while keeping private data locally. To better aggregate knowledge from clients, ensemble distillation, as a widely used and effective technique, is often employed after global aggregation to enhance the performance of the global model. However, simply combining Hetero-FL and ensemble distillation does not always yield promising results and can make the training process unstable. The reason is that existing methods primarily focus on logit distillation, which, while being model-agnostic with softmax predictions, fails to compensate for the knowledge bias arising from heterogeneous models.
To tackle this challenge, we propose a stable and efficient Feature Distillation for model-heterogeneous Federated learning, dubbed FedFD, that can incorporate aligned feature information via orthogonal projection to integrate knowledge from heterogeneous models better. Specifically, a new feature-based ensemble federated knowledge distillation paradigm is proposed. The global model on the server needs to maintain a projection layer for each client-side model architecture to align the features separately. Orthogonal techniques are employed to re-parameterize the projection layer to mitigate knowledge bias from heterogeneous models and thus maximize the distilled knowledge. Extensive experiments show that FedFD achieves superior performance compared to state-of-the-art methods. Yichen Li 0006, Xiuying Wang 0015, Wenchao Xu 0001, Haozhao Wang, Yining Qi, Jiahua Dong 0001, Ruixuan Li 0001 |
NeurIPS | 6 |
| 2025 | DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time seriesabstractMedical time-series data play a vital role in disease diagnosis but suffer from limited labeled samples and single-center bias, which hinder model generalization and lead to overfitting. To address these challenges, we propose DAAC (Discrepancy-Aware Adaptive Contrastive learning), a learnable multi-view contrastive framework that integrates external normal samples and enhances feature learning through adaptive contrastive strategies. DAAC consists of two key modules: (1) a Discrepancy Estimator, built upon a GAN-enhanced encoder-decoder architecture, captures the distribution of normal data and computes reconstruction errors as indicators of abnormality. These discrepancy features augment the target dataset to mitigate overfitting. (2) an Adaptive Contrastive Learner uses multi-head attention to extract discriminative representations by contrasting embeddings across multiple views and data granularities (subject, trial, epoch, and temporal levels), eliminating the need for handcrafted positive-negative sample pairs. Extensive experiments on three clinical datasets—covering Alzheimer’s disease, Parkinson’s disease, and myocardial infarction—demonstrate that DAAC significantly outperforms existing methods, even when only 10\% of labeled data is available, showing strong generalization and diagnostic performance. Our code is available
at https://github.com/CUHKSZ-MED-BioE/DAAC. Hongfeng Ai, Ruiqi Li 0004, Maowei Jiang, Quangao Liu, Jiahua Dong 0001, Ruiyuan Kang, Alan Liang, Ruikai Liu, Chenzhong Li |
NeurIPS | 6 |
| 2025 | Accurate Tracking of Arabidopsis Root Cortex Cell Nuclei in 3D Time-Lapse Microscopy Images Based on Genetic AlgorithmabstractArabidopsis is a widely used model plant to study physiology and development. Live imaging is an important technique to visualize and quantify processes in plant growth and cell division, where accurate cell tracking is essential. The commonly used software TrackMate adopts a tracking-by-detection approach, applying Laplacian of Gaussian (LoG) for blob detection and a Linear Assignment Problem (LAP) tracker for tracking. However, its performance declines when cells are densely arranged. To overcome this limitation, we propose an accurate tracking method based on a Genetic Algorithm (GA) that incorporates knowledge of Arabidopsis root cellular patterns and spatial relationships among volumes. Our method follows a coarse-to-fine strategy: first performing relatively simple line-level tracking of nuclei, then refining associations based on the linear arrangement of cell files and their spatial relationships. We evaluated the method on long-term live imaging datasets of Arabidopsis root tips, and with minor manual correction, it achieved accurate nuclear tracking. To the best of our knowledge, this represents the first successful attempt to address a long-standing problem in time-lapse microscopy of the root meristem by providing an accurate tracking method for Arabidopsis root nuclei. Yu Song 0008, Tatsuaki Goh, Yinhao Li 0002, Jiahua Dong 0001, Shunsuke Miyashima, Yutaro Iwamoto, Yohei Kondo, Keiji Nakajima, Yenwei Chen |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | PECTP: Parameter-Efficient Cross-Task Prompts for Incremental Vision TransformerabstractIncremental Learning (IL) aims to learn deep models on sequential tasks continually, where each new task includes a batch of new classes and deep models have no access to task ID information at the inference time. Recent vast pre-trained models (PTMs) have achieved outstanding performance by prompt technique in practical IL without the old samples (rehearsal-free) and with a memory constraint (memory-constrained): Prompt-extending and Prompt-fixed methods. However, prompt-extending methods need a large memory buffer to maintain an ever-expanding prompt pool and meet an extra challenging prompt selection problem. Prompt-fixed methods only learn a single set of prompts on one of the incremental tasks and can not handle all the incremental tasks effectively. To achieve a good balance between the memory cost and the performance on all the tasks, we propose a Parameter-Efficient Cross-Task Prompt (PECTP) framework with Prompt Retention Module (PRM) and classifier Head Retention Module (HRM). To make the final learned prompts effective on all incremental tasks, PRM constrains the evolution of cross-task prompts’ parameters from Outer Prompt Granularity and Inner Prompt Granularity. Besides, we employ HRM to inherit old knowledge in the previously learned classifier heads to facilitate the cross-task prompts’ generalization ability. Extensive experiments show the effectiveness of our method. The source codes are available at https://github.com/RAIAN08/PECTP. Hanbin Zhao, Chao Zhang 0001, Jiahua Dong 0001, Henghui Ding, Yu-Gang Jiang 0001, Hui Qian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | MuseumMaker: Continual Style Customization Without Catastrophic ForgettingabstractPre-trainedlarge text-to-image (T2I) models with an appropriate text prompt has attracted growing interests in customized image generation fields. However, catastrophic forgetting issue makes it hard to continually synthesize new user-provided styles while retaining the satisfying results amongst learned styles. In this paper, we propose MuseumMaker, a method that enables the synthesis of images by following a set of customized styles in a never-end manner, and gradually accumulates these creative artistic works as a Museum. When facing with a new customization style, we develop a style distillation loss module to extract and learn the styles of the training data for new image generation task. It can minimize the learning biases caused by content of new training images, and address the catastrophic overfitting issue induced by few-shot images. To deal with catastrophic forgetting issue amongst past learned styles, we devise a dual regularization for shared-LoRA module to optimize the direction of model update, which could regularize the diffusion model from both weight and feature aspects, respectively. Meanwhile, to further preserve historical knowledge from past styles and address the limited representability of LoRA, we design a task-wise token learning module where a unique token embedding is learned to denote a new style. As any new user-provided style come, our MuseumMaker can capture the nuances of the new styles while maintaining the details of learned styles. Experimental results on diverse style datasets validate the effectiveness of our proposed MuseumMaker method, showcasing its robustness and versatility across various scenarios. Gan Sun, Wenqi Liang, Jiahua Dong 0001, Can Qin, Yang Cong |
IEEE Trans. Image Process. | 4 |
| 2024 | Xformer: Hybrid X-Shaped Transformer for Image DenoisingabstractIn this paper, we present a hybrid X-shaped vision Transformer, named Xformer, which performs notably on image denoising tasks. We explore strengthening the global representation of tokens from different scopes. In detail, we adopt two types of Transformer blocks. The spatial-wise Transformer block performs fine-grained local patches interactions across tokens defined by spatial dimension. The channel-wise Transformer block performs direct global context interactions across tokens defined by channel dimension. Based on the concurrent network structure, we design two branches to conduct these two interaction fashions. Within each branch, we employ an encoder-decoder architecture to capture multi-scale features. Besides, we propose the Bidirectional Connection Unit (BCU) to couple the learned representations from these two branches while providing enhanced information fusion. The joint designs make our Xformer powerful to conduct global information modeling in both spatial and channel dimensions. Extensive experiments show that Xformer, under the comparable model complexity, achieves state-of-the-art performance on the synthetic and real-world image denoising tasks. We also provide code and models at https://github.com/gladzhang/Xformer. Yulun Zhang 0001, Jinjin Gu, Jiahua Dong 0001, Linghe Kong, Xiaokang Yang 0001 |
ICLR | 4 |
| 2024 | FedDGP: Disentangling Global and Personal Models for Federated LearningabstractFederated learning (FL) aims to construct a global model by collaboratively training local models on data with diverse distributions, emphasizing the exchange of model parameters rather than raw data sharing. In the medical domain, achieving high performance in local models, strong generalization capabilities in the global model, and minimizing communication costs are all crucial. However, current federated learning methods struggle to concurrently optimize these three aspects. This paper introduces FedDGP, an innovative FL framework addressing these challenges. FedDGP comprises three key modules: Domain-guided Model Disentanglement (MD), Heterogeneous Aggregation (HA), and Reciprocal Iterative Training (RIT). MD divides a client model into two parts with different task objectives, enhancing both client and server model performance while avoiding issues like catastrophic forgetting. HA assigns reduced weight to underperforming models, limiting their impact. RIT minimizes communication costs by uploading only one client model’s parameters per round, facilitating local optimal model transfers. Extensive experiments on six domain classification datasets demonstrate FedDGP’s effectiveness, showcasing improved performance and reduced communication costs compared to existing approaches. Zhenhu Zhang, Dan Song 0006, Jiahua Dong 0001, Ruofeng Tong 0001 |
ICME | 4 |
| 2024 | How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?abstractCustom diffusion models (CDMs) have attracted widespread attention due to their astonishing generative ability for personalized concepts. However, most existing CDMs unreasonably assume that personalized concepts are fixed and cannot change over time. Moreover, they heavily suffer from catastrophic forgetting and concept neglect on old personalized concepts when continually learning a series of new concepts. To address these challenges, we propose a novel Concept-Incremental text-to-image Diffusion Model (CIDM), which can resolve catastrophic forgetting and concept neglect to learn new customization tasks in a concept-incremental manner. Specifically, to surmount the catastrophic forgetting of old concepts, we develop a concept consolidation loss and an elastic weight aggregation module. They can explore task-specific and task-shared knowledge during training, and aggregate all low-rank weights of old concepts based on their contributions during inference. Moreover, in order to address concept neglect, we devise a context-controllable synthesis strategy that leverages expressive region features and noise estimation to control the contexts of generated images according to user conditions. Experiments validate that our CIDM surpasses existing custom diffusion models. The source codes are available at https://github.com/JiahuaDong/CIFC. Jiahua Dong 0001, Wenqi Liang, Hongliu Li, Duzhen Zhang, Meng Cao 0002, Henghui Ding, Salman Khan 0001, Fahad Shahbaz Khan |
NeurIPS | 1 |
| 2024 | Where and How to Transfer: Knowledge Aggregation-Induced Transferability Perception for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation without accessing expensive annotation processes of target data has achieved remarkable successes in semantic segmentation. However, most existing state-of-the-art methods cannot explore whether semantic representations across domains are transferable or not, which may result in the negative transfer brought by irrelevant knowledge. To tackle this challenge, in this paper, we develop a novel Knowledge Aggregation-induced Transferability Perception (KATP) for unsupervised domain adaptation, which is a pioneering attempt to distinguish transferable or untransferable knowledge across domains. Specifically, the KATP module is designed to quantify which semantic knowledge across domains is transferable, by incorporating transferability information propagation from global category-wise prototypes. Based on KATP, we design a novel KATP Adaptation Network (KATPAN) to determine where and how to transfer. The KATPAN contains a transferable appearance translation module T_A() and a transferable representation augmentation module T_R(), where both modules construct a virtuous circle of performance promotion. T_A() develops a transferability-aware information bottleneck to highlight where to adapt transferable visual characterizations and modality information; T_R() explores how to augment transferable representations while abandoning untransferable information, and promotes the translation performance of T_A() in return. Experiments on several representative datasets and a medical dataset support the state-of-the-art performance of our model. Jiahua Dong 0001, Yang Cong, Gan Sun, Zhen Fang 0001, Zhengming Ding |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | No One Left Behind: Real-World Federated Class-Incremental LearningabstractFederated learning (FL) is a hot collaborative training framework via aggregating model parameters of decentralized local clients. However, most FL methods unreasonably assume data categories of FL framework are known and fixed in advance. Moreover, some new local clients that collect novel categories unseen by other clients may be introduced to FL training irregularly. These issues render global model to undergo catastrophic forgetting on old categories, when local clients receive new categories consecutively under limited memory of storing old categories. To tackle the above issues, we propose a novelLocal-GlobalAnti-forgetting (LGA) model. It ensures no local clients are left behind as they learn new classes continually, by addressing local and global catastrophic forgetting. Specifically, considering tackling class imbalance of local client to surmount local forgetting, we develop a category-balanced gradient-adaptive compensation loss and a category gradient-induced semantic distillation loss. They can balance heterogeneous forgetting speeds of hard-to-forget and easy-to-forget old categories, while ensure consistent class-relations within different tasks. Moreover, a proxy server is designed to tackle global forgetting caused by Non-IID class imbalance between different clients. It augments perturbed prototype images of new categories collected from local clients via self-supervised prototype augmentation, thus improving robustness to choose the best old global model for local-side semantic distillation loss. Experiments on representative datasets verify superior performance of our model against comparison methods. The code is available athttps://github.com/JiahuaDong/LGA. Jiahua Dong 0001, Hongliu Li, Yang Cong, Gan Sun, Yulun Zhang 0001, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Create Your World: Lifelong Text-to-Image DiffusionabstractText-to-image generative models can produce diverse high-quality images of concepts with a text prompt, which have demonstrated excellent ability in image generation, image translation, etc. We in this work study the problem of synthesizing instantiations of a user's own concepts in a never-ending manner,i.e.,create your world, where the new concepts from user are quickly learned with a few examples. To achieve this goal, we propose aLifelong text-to-imageDiffusionModel (L$^{2}$DM), which intends to overcome knowledge “catastrophic forgetting” for the past encountered concepts, and semantic “catastrophic neglecting” for one or more concepts in the text prompt. In respect of knowledge “catastrophic forgetting”, our L$^{2}$DM framework devises a task-aware memory enhancement module and an elastic-concept distillation module, which could respectively safeguard the knowledge of both prior concepts and each past personalized concept. When generating images with a user text prompt, the solution to semantic “catastrophic neglecting” is that a concept attention artist module can alleviate the semantic neglecting from concept aspect, and an orthogonal attention module can reduce the semantic binding from attribute aspect. To the end, our model can generate more faithful image across a range of continual text prompts in terms of both qualitative and quantitative metrics, when comparing with the related state-of-the-art models. The code will be released athttps://wenqiliang.github.io/. Gan Sun, Wenqi Liang, Jiahua Dong 0001, Jun Li 0027, Zhengming Ding, Yang Cong |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Self-Paced Weight Consolidation for Continual LearningabstractContinual learning algorithms which keep the parameters of new tasks close to that of previous tasks, are popular in preventing catastrophic forgetting in sequential task learning settings. However, 1) the performance for the new continual learner will be degraded without distinguishing the contributions of previously learned tasks; 2) the computational cost will be greatly increased with the number of tasks, since most existing algorithms need to regularize all previous tasks when learning new tasks. To address the above challenges, we propose aself-pacedWeightConsolidation (spWC) framework to attain robust continual learning via evaluating the discriminative contributions of previous tasks. To be specific, we develop a self-paced regularization to reflect the priorities of past tasks via measuring difficulty based on key performance indicator (i.e., accuracy). When encountering a new task, all previous tasks are sorted from “difficult” to “easy” based on the priorities. Then the parameters of the new continual learner will be learned via selectively maintaining the knowledge amongst more difficult past tasks, which could well overcome catastrophic forgetting with less computational cost. We adopt an alternative convex search to iteratively update the model parameters and priority weights in the bi-convex formulation. The proposed spWC framework is plug-and-play, which is applicable to most continual learning algorithms (e.g., EWC, MAS and RCIL) in different directions (e.g., classification and segmentation). Experimental results on several public benchmark datasets demonstrate that our proposed framework can effectively improve performance when compared with other popular continual learning algorithms. Wei Cong, Yang Cong, Gan Sun, Jiahua Dong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Gradient-Semantic Compensation for Incremental Semantic SegmentationabstractIncremental semantic segmentation focuses on continually learning the segmentation of new coming classes without obtaining the training data from previously seen classes. However, most current methods fail to tackle catastrophic forgetting and background shift since they 1) treat all previous classes equally without considering different forgetting paces caused by imbalanced gradient back-propagation; 2) lack strong semantic guidance between classes. In this paper, to solve the aforementioned challenges, we propose aGradient-SemanticCompensation (GSC) model, which surmounts incremental semantic segmentation from both gradient and semantic perspectives. Specifically, to handle catastrophic forgetting from the gradient aspect, we develop a step-aware gradient compensation that can balance forgetting paces of previously seen classes by re-weighting gradient back-propagation. Meanwhile, we propose a soft-sharp semantic relation distillation to distill consistent inter-class semantic relations via soft labels for alleviating catastrophic forgetting from the semantic aspect. In addition, we design a prototypical pseudo re-labeling which provides strong semantic guidance to mitigate background shift. It produces high-quality pseudo labels for background pixels belonging to previous classes by assessing distances of pixels relative to class-wise prototypes. Experiments on three public segmentation datasets provide strong evidence for the effectiveness of our proposed GSC model. Wei Cong, Yang Cong, Jiahua Dong 0001, Gan Sun, Henghui Ding |
IEEE Trans. Multim. | 3 |
| 2024 | Open-Ended Online Learning for Autonomous Visual PerceptionabstractThe visual perception systems aim to autonomously collect consecutive visual data and perceive the relevant information online like human beings. In comparison with the classical static visual systems focusing on fixed tasks (e.g., face recognition for visual surveillance), the real-world visual systems (e.g., the robot visual system) often need to handle unpredicted tasks and dynamically changed environments, which need to imitate human-like intelligence with open-ended online learning ability. Therefore, we provide a comprehensive analysis of open-ended online learning problems for autonomous visual perception in this survey. Based on "what to online learn" among visual perception scenarios, we classify the open-ended online learning methods into five categories: instance incremental learning to handle data attributes changing, feature evolution learning for incremental and decremental features with the feature dimension changed dynamically, class incremental learning and task incremental learning aiming at online adding new coming classes/tasks, and parallel and distributed learning for large-scale data to reveal the computational and storage advantages. We discuss the characteristic of each method and introduce several representative works as well. Finally, we introduce some representative visual perception applications to show the enhanced performance when using various open-ended online learning models, followed by a discussion of several future directions. Yang Cong, Gan Sun, Dongdong Hou, Jiahua Dong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Topology-Aware Graph Convolution Network for Few-Shot Incremental 3-D Object LearningabstractThree-dimensional (3-D) object recognition has achieved satisfied achievement in both academia and industry. However, most traditional 3-D object classification methods implicitly assume that there are abundant training data from a static distribution. To relax the assumption, we target on a more challenging and realistic setting: few-shot incremental 3-D object learning (FSI3DL), which intends to incrementally classify the new coming 3-D objects with few training data. In order to achieve this, two key challenges need to be concerned: 1) the catastrophic forgetting issue caused by incremental 3-D data with irregular and redundant topological structures and 2) the overfitting issue caused by few-shot training data. To address the first challenge, we use Laplacian spectral analysis based on 3-D meshes to design an embedding network that consists of super-vertex graph convolution (SVGC) module and topology-aware graph attention (TAGA) module. The SVGC is designed to construct the discriminative local topological characteristics for representing the irregular 3-D meshes better. The TAGA is designed to identify redundant topological characteristics. To address the second challenge, a fine-tuning strategy with model alignment regularization is investigated. Furthermore, an embedding space selection and fusion (ESSF) strategy is proposed in the inference phase to mitigate catastrophic forgetting and overfitting further. Combining SVGC, TAGA, and alignment regularization with ESSF strategy, a novel topology-aware graph convolution network (TopGCN) is proposed to address the FSI3DL. Experiments on representative 3-D classification datasets validate the superiority of TopGCN. Bingtao Ma, Yang Cong, Jiahua Dong 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | Task Relation Distillation and Prototypical Pseudo Label for Incremental Named Entity RecognitionabstractIncremental Named Entity Recognition (INER) involves the sequential learning of new entity types without accessing the training data of previously learned types. However, INER faces the challenge of catastrophic forgetting specific for incremental learning, further aggravated by background shift (i.e., old and future entity types are labeled as the non-entity type in the current task). To address these challenges, we propose a method called task Relation Distillation and Prototypical pseudo label (RDP) for INER. Specifically, to tackle catastrophic forgetting, we introduce a task relation distillation scheme that serves two purposes: 1) ensuring inter-task semantic consistency across different incremental learning tasks by minimizing inter-task relation distillation loss, and 2) enhancing the model's prediction confidence by minimizing intra-task self-entropy loss. Simultaneously, to mitigate background shift, we develop a prototypical pseudo label strategy that distinguishes old entity types from the current non-entity type using the old model. This strategy generates high-quality pseudo labels by measuring the distances between token embeddings and type-wise prototypes. We conducted extensive experiments on ten INER settings of three benchmark datasets (i.e., CoNLL2003, I2B2, and OntoNotes5). The results demonstrate that our method achieves significant improvements over the previous state-of-the-art methods, with an average increase of 6.08% in Micro F1 score and 7.71% in Macro F1 score. Duzhen Zhang, Hongliu Li, Wei Cong, Rongtao Xu, Jiahua Dong 0001, Xiuyi Chen |
CIKM | 5 |
| 2023 | Federated Incremental Semantic SegmentationabstractFederated learning-based semantic segmentation (FSS) has drawn widespread attention via decentralized training on local clients. However, most FSS models assume categories are fixed in advance, thus heavily undergoing forgetting on old categories in practical applications where local clients receive new categories incrementally while have no memory storage to access old classes. Moreover, new clients collecting novel classes may join in the global training of FSS, which further exacerbates catastrophic forgetting. To surmount the above challenges, we propose a Forgetting-Balanced Learning (FBL) model to address heterogeneous forgetting on old classes from both intra-client and interclient aspects. Specifically, under the guidance of pseudo labels generated via adaptive class-balanced pseudo labeling, we develop a forgetting-balanced semantic compensation loss and a forgetting-balanced relation consistency loss to rectify intra-client heterogeneous forgetting of old categories with background shift. It performs balanced gradient propagation and relation consistency distillation within local clients. Moreover, to tackle heterogeneous forgetting from inter-client aspect, we propose a task transition monitor. It can identify new classes under privacy protection and store the latest old global model for relation distillation. Qualitative experiments reveal large improvement of our model against comparison methods. The code is available at https://github.com/JiahuaDong/FISS. Jiahua Dong 0001, Duzhen Zhang, Yang Cong, Wei Cong, Henghui Ding, Dengxin Dai |
CVPR | 1 |
| 2023 | Continual Named Entity Recognition without Catastrophic ForgettingabstractContinual Named Entity Recognition (CNER) is a burgeoning area, which involves updating an existing model by incorporating new entity types sequentially.Nevertheless, continual learning approaches are often severely afflicted by catastrophic forgetting.This issue is intensified in CNER due to the consolidation of old entity types from previous steps into the non-entity type at each step, leading to what is known as the semantic shift problem of the non-entity type.In this paper, we introduce a pooled feature distillation loss that skillfully navigates the trade-off between retaining knowledge of old entity types and acquiring new ones, thereby more effectively mitigating the problem of catastrophic forgetting.Additionally, we develop a confidence-based pseudo-labeling for the non-entity type, i.e., predicting entity types using the old model to handle the semantic shift of the non-entity type.Following the pseudo-labeling process, we suggest an adaptive re-weighting type-balanced learning strategy to handle the issue of biased type distribution.We carried out comprehensive experiments on ten CNER settings using three different datasets.The results illustrate that our method significantly outperforms prior state-of-the-art approaches, registering an average improvement of 6.3% and 8.0% in Micro and Macro F1 scores, respectively.1 * Equal contributions.† The corresponding author is Dr. Duzhen Zhang, Wei Cong, Jiahua Dong 0001, Yahan Yu, Xiuyi Chen, Yonggang Zhang 0003, Zhen Fang 0001 |
EMNLP | 3 |
| 2023 | Heterogeneous Forgetting Compensation for Class-Incremental LearningabstractClass-incremental learning (CIL) has achieved remarkable successes in learning new classes consecutively while overcoming catastrophic forgetting on old categories. However, most existing CIL methods unreasonably assume that all old categories have the same forgetting pace, and neglect negative influence of forgetting heterogeneity among different old classes on forgetting compensation. To surmount the above challenges, we develop a novel Heterogeneous Forgetting Compensation (HFC) model, which can resolve heterogeneous forgetting of easy-to-forget and hard-to-forget old categories from both representation and gradient aspects. Specifically, we design a task-semantic aggregation block to alleviate heterogeneous forgetting from representation aspect. It aggregates local category information within each task to learn task-shared global representations. Moreover, we develop two novel plug-and-play losses: a gradient-balanced forgetting compensation loss and a gradient-balanced relation distillation loss to alleviate forgetting from gradient aspect. They consider gradient-balanced compensation to rectify forgetting heterogeneity of old categories and heterogeneous relation consistency. Experiments on several representative datasets illustrate effectiveness of our HFC model. The code is available at https://github.com/JiahuaDong/HFC. Jiahua Dong 0001, Wenqi Liang, Yang Cong, Gan Sun |
ICCV | 1 |
| 2023 | I3DOD: Towards Incremental 3D Object Detection via Promptingabstract3D object detection have achieved significant performance in many fields, e.g., robotics system, autonomous driving, and augmented reality. However, most existing methods could cause catastrophic forgetting of old classes when performing on the class-incremental scenarios. Meanwhile, the current class-incremental 3D object detection methods neglect the relationships between the object localization information and category semantic information, and assume all the knowledge of old model is reliable. To address the above challenge, we present a novel Incremental 3D Object Detection framework with the guidance of prompting, i.e., I3DOD. Specifically, we propose a task-shared prompts mechanism to learn the matching relationships between the object localization information and category semantic information. After training on the current task, these prompts will be stored in our prompt pool, and perform the relationship of old classes in the next task. Moreover, we design a reliable distillation strategy to transfer knowledge from two aspects: a reliable dynamic distillation is developed to filter out the negative knowledge and transfer the reliable 3D knowledge to new detection model; the relation feature is proposed to capture the responses relation in feature space and protect plasticity of the model when learning novel 3D classes. To the end, we conduct comprehensive experiments on two benchmark datasets and our method outperforms the state-of-the-art object detection methods by 0.6% ∼ 2.7% in terms of [email protected]. Wenqi Liang, Gan Sun, Jiahua Dong 0001, Kangru Wang |
IROS | 4 |
| 2023 | Uni3DA: Universal 3D Domain Adaptation for Object RecognitionabstractTraditional 3D point cloud classification tasks focus on training a classifier in the closed-set scenario, where training and test data have the same label set and the same data distribution. In this work, we focus on a more challenging and realistic scenario in 3D point cloud classification task: universal domain adaptation (UniDA), where 1) data distributions for training and test data are different; and 2) for given label sets of training data and test data, they may have the shared classes and keep the private classes respectively, introducing an extra label set discrepancy. To solve UniDA problem, researchers have designed many methods based on 2D image datasets. However, due to the difficulty in capturing discriminative local geometric structures brought by the unordered and irregular 3D point cloud data, we cannot directly deploy the existing methods based on 2D image datasets to the 3D scenarios. To address UniDA in 3D scenarios, we develop a 3D universal domain adaptation framework, which consists of three modules: Self-Constructed Geometric (SCG) module, Local-to-Global Hypersphere Reasoning (LGHR) module and Self-Supervised Boundary Adaptation (SBA) module. SCG and LGHR generate the discriminative representation, which is used to acquire domain-invariant knowledge for training and test data. SBA is designed to automatically recognize whether a given label is from the shared label set or private label set, and adapts training and test data from the shared label set. To our best knowledge, this work is the first exploration of UniDA for 3D scenarios. Extensive experiments on public 3D point cloud datasets verify that the proposed method outperforms the existing UniDA methods. Yang Cong, Jiahua Dong 0001, Gan Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | InOR-Net: Incremental 3-D Object Recognition Network for Point Cloud Representationabstract3-D object recognition has successfully become an appealing research topic in the real world. However, most existing recognition models unreasonably assume that the categories of 3-D objects cannot change over time in the real world. This unrealistic assumption may result in significant performance degradation for them to learn new classes of 3-D objects consecutively due to the catastrophic forgetting on old learned classes. Moreover, they cannot explore which 3-D geometric characteristics are essential to alleviate the catastrophic forgetting on old classes of 3-D objects. To tackle the above challenges, we develop a novel Incremental 3-D Object Recognition Network (i.e., InOR-Net), which could recognize new classes of 3-D objects continuously by overcoming the catastrophic forgetting on old classes. Specifically, category-guided geometric reasoning is proposed to reason local geometric structures with distinctive 3-D characteristics of each class by leveraging intrinsic category information. We then propose a novel critic-induced geometric attention mechanism to distinguish which 3-D geometric characteristics within each class are beneficial to overcome the catastrophic forgetting on old classes of 3-D objects while preventing the negative influence of useless 3-D characteristics. In addition, a dual adaptive fairness compensations' strategy is designed to overcome the forgetting brought by class imbalance by compensating biased weights and predictions of the classifier. Comparison experiments verify the state-of-the-art performance of the proposed InOR-Net model on several public point cloud datasets. Jiahua Dong 0001, Yang Cong, Gan Sun, Lixu Wang, Lingjuan Lyu, Jun Li 0027, Ender Konukoglu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Federated Class-Incremental LearningabstractFederated learning (FL) has attracted growing attentions via data-private collaborative training on decentralized clients. However, most existing methods unrealistically assume object classes of the overall framework are fixed over time. It makes the global model suffer from significant catastrophic forgetting on old classes in real-world scenarios, where local clients often collect new classes continuously and have very limited storage memory to store old classes. Moreover, new clients with unseen new classes may participate in the FL training, further aggravating the catastrophic forgetting of global model. To address these challenges, we develop a novel Global-Local Forgetting Compensation (GLFC) model, to learn a global class-incremental model for alleviating the catastrophic forgetting from both local and global perspectives. Specifically, to address local forgetting caused by class imbalance at the local clients, we design a class-aware gradient compensation loss and a class-semantic relation distillation loss to balance the forgetting of old classes and distill consistent inter-class relations across tasks. To tackle the global forgetting brought by the non-i.i.d class imbalance across clients, we propose a proxy server that selects the best old global model to assist the local relation distillation. Moreover, a prototype gradient-based communication mechanism is developed to protect the privacy. Our model outperforms state-of-the-art methods by 4.4%~15.1% in terms of average accuracy on representative benchmark datasets. The code is available at https://github.com/conditionWang/FCIL. Jiahua Dong 0001, Lixu Wang, Zhen Fang 0001, Gan Sun, Shichao Xu, Xiao Wang 0012, Qi Zhu 0002 |
CVPR | 1 |
| 2022 | Is Out-of-Distribution Detection Learnable?abstractSupervised learning aims to train a classifier under the assumption that training and test data are from the same distribution. To ease the above assumption, researchers have studied a more realistic setting: out-of-distribution (OOD) detection, where test data may come from classes that are unknown during training (i.e., OOD data). Due to the unavailability and diversity of OOD data, good generalization ability is crucial for effective OOD detection algorithms. To study the generalization of OOD detection, in this paper, we investigate the probably approximately correct (PAC) learning theory of OOD detection, which is proposed by researchers as an open problem. First, we find a necessary condition for the learnability of OOD detection. Then, using this condition, we prove several impossibility theorems for the learnability of OOD detection under some scenarios. Although the impossibility theorems are frustrating, we find that some conditions of these impossibility theorems may not hold in some practical scenarios. Based on this observation, we next give several necessary and sufficient conditions to characterize the learnability of OOD detection in some practical scenarios. Lastly, we also offer theoretical supports for several representative OOD detection works based on our OOD theory. Zhen Fang 0001, Yixuan Li 0001, Jie Lu 0001, Jiahua Dong 0001, Bo Han 0003, Feng Liu 0003 |
NeurIPS | 4 |
| 2022 | Data Poisoning Attacks on Federated Machine LearningabstractFederated machine learning which enables resource-constrained node devices (e.g., Internet of Things (IoT) devices and smartphones) to establish a knowledge-shared model while keeping the raw data local, could provide privacy preservation, and economic benefit by designing an effective communication protocol. However, this communication protocol can be adopted by attackers to launch data poisoning attacks for different nodes, which has been shown as a big threat to most machine learning models. Therefore, we in this article intend to study the model vulnerability of federated machine learning, and even on IoT systems. To be specific, we here attempt to attacking a popular federated multitask learning framework, which uses a general multitask learning framework to handle statistical challenges in the federated learning setting. The problem of calculating optimal poisoning attacks on federated multitask learning is formulated as a bilevel program, which is adaptive to the arbitrary selection oftargetnodes andsource attackingnodes. We then propose a novel systems-aware optimization method, called as attack on federated learning (AT2FL), to efficiently derive the implicit gradients for poisoned data, and further attain optimal attack strategies in the federated machine learning. This is an earlier work, to our knowledge, that explores attacking federated machine learning via data poisoning. Finally, experiments on several real-world data sets demonstrate that when the attackers directly poison thetargetnodes or indirectly poison the related nodes via using the communication protocol, the federated multitask learning model is sensitive to both poisoning attacks. Gan Sun, Yang Cong, Jiahua Dong 0001, Qiang Wang 0015, Lingjuan Lyu, Ji Liu 0002 |
IEEE Internet Things J. | 3 |
| 2022 | What and How: Generalized Lifelong Spectral Clustering via Dual MemoryabstractSpectral clustering (SC) has become one of the most widely-adopted clustering algorithms, and been successfully applied into various applications. We in this work explore the problem of spectral clustering in a lifelong learning framework termed asGeneralizedLifelongSpectralClustering (GL$^2$SC). Different from most current studies, which concentrate on a fixed spectral clustering task set and cannot efficiently incorporate a new clustering task, the goal of our work is to establish a generalized model for new spectral clustering tasks by “What” and “How” to lifelong learn from past tasks. In respect of “what to lifelong learn”, our GL$^2$SC framework contains a dual memory mechanism with a deep orthogonal factorization manner: an orthogonal basis memory stores hidden and hierarchical clustering centers among learned tasks, and a feature embedding memory captures deep manifold representation common across multiple related tasks. When learning a new clustering task, the intuition here for “how to lifelong learn” is that GL$^2$SC can transfer intrinsic knowledge from dual memory mechanism to obtain task-specific encoding matrix. Then the encoding matrix can redefine the dual memory over time to provide maximal benefits when learning future tasks, and reversely maximize performance for past tasks. To achieve this, we propose an alternative optimization formulation with convergence guarantee for solving our GL$^2$SC model. To the end, empirical comparisons on several benchmark datasets show the effectiveness of our GL$^2$SC, in comparison with several state-of-the-art clustering models. Gan Sun, Yang Cong, Jiahua Dong 0001, Zhengming Ding |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Lifelong robotic visual-tactile perception learning
Jiahua Dong 0001, Yang Cong, Gan Sun, Tao Zhang 0084 |
Pattern Recognit. | 1 |
| 2022 | Fast Multi-View Outlier Detection via Deep EncoderabstractMulti-view outlier detection has a wide range of applications and has been well investigated in recent years. However, 1) most existing state-of-the-art methods cannot efficiently handle outlier detection problem for large-scale multi-view data, since exploring pairwise constraints among different views causes highly-computational cost; 2) the data collected from original heterogeneous feature spaces further increases the consistent difficulty of multi-view outlier detection. To address these issues, we present a fast multi-view outlier detection model via learning a low-rank latent subspace representation with deep encoder architecture, which can not only efficiently identify the outliers for large-scale data even with numerous data views, but also exploit a discriminative common latent subspace shared by all the views. First, we learn a set of orthogonal bases as view-specific dictionaries from a small dataset, which is randomly sampled from the original dataset. Benefitting from view-specific dictionaries, the sampled data is projected and decomposed as a shared and discriminative latent subspace representations, which correspond to the view-consistent and view-specific components across multiple views, respectively. Then, the obtained discriminative latent representations are applied to train the view-specific deep encoders, which can efficiently compute the abnormal score for the remaining instances. Our proposed model can cost-effectively identify the outliers in large-scale datasets from numerous data views with less computational complexity. Experiments conducted on eight real datasets and a synthesis dataset show that our proposed model outperforms the existing ones on effectiveness and efficiency. Dongdong Hou, Yang Cong, Gan Sun, Jiahua Dong 0001, Jun Li 0027, Kai Li 0012 |
IEEE Trans. Big Data | 4 |
| 2022 | Evolving Metric Learning for Incremental and Decremental FeaturesabstractOnline metric learning has been widely exploited for large-scale data classification due to the low computational cost. However, amongst online practical scenarios where the features are evolving (e.g., some features are vanished and some new features are augmented), most metric learning models cannot be successfully applied to these scenarios, although they can tackle the evolving instances efficiently. To address the challenge, we develop a new online Evolving Metric Learning (EML) model for incremental and decremental features, which can handle the instance and feature evolutions simultaneously by incorporating with a smoothed Wasserstein metric distance. Specifically, our model contains two essential stages: a Transforming stage (T-stage) and a Inheriting stage (I-stage). For the T-stage, we propose to extract important information from vanished features while neglecting non-informative knowledge, and forward it into survived features by transforming them into a low-rank discriminative metric space. It further explores the intrinsic low-rank structure of heterogeneous samples to reduce the computation and memory burden especially for highly-dimensional large-scale data. For the I-stage, we inherit the metric performance of survived features from the T-stage and then expand to include the new augmented features. Moreover, a smoothed Wasserstein distance is utilized to characterize the similarity relationships among the heterogeneous and complex samples, since the evolving features are not strictly aligned in the different stages. In addition to tackling the challenges in one-shot case, we also extend our model into multi-shot scenario. After deriving an efficient optimization strategy for both T-stage and I-stage, extensive experiments on several datasets verify the superior performance of our EML model. Jiahua Dong 0001, Yang Cong, Gan Sun, Tao Zhang 0084, Xiaowei Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Continuous Multi-View Human Action RecognitionabstractHuman action recognition which recognizes human actions in a video is a fundamental task in computer vision field. Although multiple existing methods with single-view or multi-view have been presented for human action recognition, these recognition approaches cannot be extended into new action recognition or action classification tasks, as well as discover underlying correlations among different views. To tackle the above problem, this paper proposes a new lifelong multi-view subspace learning framework for continuous human action recognition, which could exploit the complementary information amongst different views from a lifelong learning perspective. More specifically, a set of view-specific libraries is established to gradually store the useful information within multiple views. As a new action recognition task comes, we decompose the model parameters into a set of embedded parameters over view-specific libraries. A latent representation subspace is constructed via encouraging it to be close to different view-specific libraries, which can leverage the high-order correlations among different views and further avoid partial information for action recognition task. Meanwhile, we propose to employ an alternating direction strategy to optimize our proposed method. Empirical studies on real-world multi-view action recognition datasets have shown that our proposed framework attains the superior recognition performance and saves the computational time when continually learning new action recognition tasks. Qiang Wang 0015, Gan Sun, Jiahua Dong 0001, Qianqian Wang 0001, Zhengming Ding |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Visual-Tactile Fused Graph Learning for Object Clusteringabstracthighlights how to mitigate the differences between vision and touch, and further maximize the mutual information, which adopts a minimizing disagreement scheme to guide the modality-specific representations toward a unified affinity graph. To achieve ideal clustering performance, a Laplacian rank constraint is imposed to regularize the learned graph with ideal connected components, where noises that caused wrong connections are removed and clustering labels can be obtained directly. Finally, we propose an efficient alternating iterative minimization updating strategy, followed by a theoretical proof to prove framework convergence. Comprehensive experiments on five public datasets demonstrate the superiority of the proposed framework. Tao Zhang 0084, Yang Cong, Gan Sun, Jiahua Dong 0001 |
IEEE Trans. Cybern. | 4 |
| 2022 | PFDN: Pyramid Feature Decoupling Network for Single Image DerainingabstractRestoring images degraded by rain has attracted more academic attention since rain streaks could reduce the visibility of outdoor scenes. However, most existing deraining methods attempt to remove rain while recovering details in a unified framework, which is an ideal and contradictory target in the image deraining task. Moreover, the relative independence of rain streak features and background features is usually ignored in the feature domain. To tackle these challenges above, we propose an effective Pyramid Feature Decoupling Network (i.e., PFDN) for single image deraining, which could accomplish image deraining and details recovery with the corresponding features. Specifically, the input rainy image features are extracted via a recurrent pyramid module, where the features for the rainy image are divided into two parts, i.e., rain-relevant and rain-irrelevant features. Afterwards, we introduce a novel rain streak removal network for rain-relevant features and remove the rain streak from the rainy image by estimating the rain streak information. Benefiting from lateral outputs, we propose an attention module to enhance the rain-irrelevant features, which could generate spatially accurate and contextually reliable details for image recovery. For better disentanglement, we also enforce multiple causality losses at the pyramid features to encourage the decoupling of rain-relevant and rain-irrelevant features from the high to shallow layers. Extensive experiments demonstrate that our module can well model the rain-relevant information over the domain of the feature. Our framework empowered by PFDN modules significantly outperforms the state-of-the-art methods on single image deraining with multiple widely-used benchmarks, and also shows superiority in the fully-supervised domain. Qiang Wang 0015, Gan Sun, Jiahua Dong 0001, Yulun Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Partial Visual-Tactile Fused Learning for Robotic Object RecognitionabstractCurrently, visual-tactile fusion learning for robotic object recognition has achieved appealing performance, due to the fact that visual and tactile data can offer complementary information. However: 1) the distinct gap between vision and touch makes it difficult to fully explore the complementary information, which would further lead to performance degradation and 2) most of the existing visual-tactile fused learning methods assume that visual and tactile data are complete, which is often difficult to be satisfied in many real-world applications. In this article, we propose a partial visual-tactile fused (PVTF) framework for robotic object recognition to address these challenges. Specifically, we first employ two modality-specific (MS) encoders to encode partial visual-tactile data into two incomplete subspaces (i.e., visual subspace and tactile subspace). Then, a modality gap mitigated (MGM) network is adopted to discover modality-invariant high-level label information, which is utilized to generate gap loss and further help updating the MS encoders for relatively consistent visual and tactile subspaces generation. In this way, the huge gap between vision and touch is mitigated, which would further contribute to mine the complementary visual-tactile information. Finally, to achieve data completeness and complementary visual-tactile information exploration simultaneously, a cycle subspace leaning technique is proposed to project the incomplete subspaces into a complete subspace by fully exploiting all the obtainable samples, where complete latent representations with maximum complementary information can be learned. A lot of comparative experiments conducted on three visual-tactile datasets validate the advantage of the proposed PVTF framework, by comparing with state-of-the-art baselines. Tao Zhang 0084, Yang Cong, Jiahua Dong 0001, Dongdong Hou |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | I3DOL: Incremental 3D Object Learning without Catastrophic Forgettingabstract3D object classification has attracted appealing attentions in academic researches and industrial applications. However, most existing methods need to access the training data of past 3D object classes when facing the common real-world scenario: new classes of 3D objects arrive in a sequence. Moreover, the performance of advanced approaches degrades dramatically for past learned classes (i.e., catastrophic forgetting), due to the irregular and redundant geometric structures of 3D point cloud data. To address these challenges, we propose a new Incremental 3D Object Learning (i.e., I3DOL) model, which is the first exploration to learn new classes of 3D object continually. Specifically, an adaptive-geometric centroid module is designed to construct discriminative local geometric structures, which can better characterize the irregular point cloud representation for 3D object. Afterwards, to prevent the catastrophic forgetting brought by redundant geometric information, a geometric-aware attention mechanism is developed to quantify the contributions of local geometric structures, and explore unique 3D geometric characteristics with high contributions for classes incremental learning. Meanwhile, a score fairness compensation strategy is proposed to further alleviate the catastrophic forgetting caused by unbalanced data between past and new classes of 3D object, by compensating biased prediction for new classes in the validation phase. Experiments on 3D representative datasets validate the superiority of our I3DOL framework. Jiahua Dong 0001, Yang Cong, Gan Sun, Bingtao Ma, Lichen Wang |
AAAI | 1 |
| 2021 | Generative Partial Visual-Tactile Fused Object ClusteringabstractVisual-tactile fused sensing for object clustering has achieved significant progresses recently, since the involvement of tactile modality can effectively improve clustering performance. However, the missing data (i.e., partial data) issues always happen due to occlusion and noises during the data collecting process. This issue is not well solved by most existing partial multi-view clustering methods for the heterogeneous modality challenge. Naively employing these methods would inevitably induce a negative effect and further hurt the performance. To solve the mentioned challenges, we propose a Generative Partial Visual-Tactile Fused (i.e., GPVTF) framework for object clustering. More specifically, we first do partial visual and tactile features extraction from the partial visual and tactile data, respectively, and encode the extracted features in modality-specific feature subspaces. A conditional cross-modal clustering generative adversarial network is then developed to synthesize one modality conditioning on the other modality, which can compensate missing samples and align the visual and tactile modalities naturally by adversarial learning. To the end, two pseudo-label based KL-divergence losses are employed to update the corresponding modality-specific encoders. Extensive comparative experiments on three public visual-tactile datasets prove the effectiveness of our method. Tao Zhang 0084, Yang Cong, Gan Sun, Jiahua Dong 0001, Zhengming Ding |
AAAI | 4 |
| 2021 | Unsupervised Dense Deformation Embedding Network for Template-Free Shape CorrespondenceabstractShape correspondence from 3D deformation learning has attracted appealing academy interests recently. Nevertheless, current deep learning based methods require the supervision of dense annotations to learn per-point translations, which severely over-parameterize the deformation process. Moreover, they fail to capture local geometric details of original shape via global feature embedding. To address these challenges, we develop a new Unsupervised Dense Deformation Embedding Network (i.e., UD2E-Net), which learns to predict deformations between non-rigid shapes from dense local features. Since it is non-trivial to match deformation-variant local features for deformation prediction, we develop an Extrinsic-Intrinsic Autoencoder to first encode extrinsic geometric features from source into intrinsic coordinates in a shared canonical shape, with which the decoder then synthesizes corresponding target features. Moreover, a bounded maximum mean discrepancy loss is developed to mitigate the distribution divergence between the synthesized and original features. To learn natural deformation without dense supervision, we introduce a coarse parameterized deformation graph, for which a novel trace and propagation algorithm is proposed to improve both the quality and efficiency of the deformation. Our UD2E-Net outperforms state-of-the-art unsupervised methods by 24% on Faust Inter challenge and even supervised methods by 13% on Faust Intra challenge. Ronghan Chen, Yang Cong, Jiahua Dong 0001 |
ICCV | 3 |
| 2021 | Confident Anchor-Induced Multi-Source Free Domain AdaptationabstractUnsupervised domain adaptation has attracted appealing academic attentions by transferring knowledge from labeled source domain to unlabeled target domain. However, most existing methods assume the source data are drawn from a single domain, which cannot be successfully applied to explore complementarily transferable knowledge from multiple source domains with large distribution discrepancies. Moreover, they require access to source data during training, which are inefficient and unpractical due to privacy preservation and memory storage. To address these challenges, we develop a novel Confident-Anchor-induced multi-source-free Domain Adaptation (CAiDA) model, which is a pioneer exploration of knowledge adaptation from multiple source domains to the unlabeled target domain without any source data, but with only pre-trained source models. Specifically, a source-specific transferable perception module is proposed to automatically quantify the contributions of the complementary knowledge transferred from multi-source domains to the target domain. To generate pseudo labels for the target domain without access to the source data, we develop a confident-anchor-induced pseudo label generator by constructing a confident anchor group and assigning each unconfident target sample with a semantic-nearest confident anchor. Furthermore, a class-relationship-aware consistency loss is proposed to preserve consistent inter-class relationships by aligning soft confusion matrices across domains. Theoretical analysis answers why multi-source domains are better than a single source domain, and establishes a novel learning bound to show the effectiveness of exploiting multi-source domains. Experiments on several representative datasets illustrate the superiority of our proposed CAiDA model. The code is available at https://github.com/Learning-group123/CAiDA. Jiahua Dong 0001, Zhen Fang 0001, Anjin Liu, Gan Sun, Tongliang Liu |
NeurIPS | 1 |
| 2021 | Weakly-Supervised Cross-Domain Adaptation for Endoscopic Lesions SegmentationabstractWeakly-supervised learning has attracted growing research attention on medical lesions segmentation due to significant saving in pixel-level annotation cost. However, 1) most existing methods require effective prior and constraints to explore the intrinsic lesions characterization, which only generates incorrect and rough prediction; 2) they neglect the underlying semantic dependencies among weakly-labeled target enteroscopy diseases and fully-annotated source gastroscope lesions, while forcefully utilizing untransferable dependencies leads to the negative performance. To tackle above issues, we propose a new weakly-supervised lesions transfer framework, which can not only explore transferable domain-invariant knowledge across different datasets, but also prevent the negative transfer of untransferable representations. Specifically, a Wasserstein quantified transferability framework is developed to highlight wide-range transferable contextual dependencies, while neglecting the irrelevant semantic characterizations. Moreover, a novel self-supervised pseudo label generator is designed to equally provide confident pseudo pixel labels for both hard-to-transfer and easy-to-transfer target samples. It inhibits the enormous deviation of false pseudo pixel labels under the self-supervision manner. Afterwards, dynamically-searched feature centroids are aligned to narrow category-wise distribution shift. Comprehensive theoretical analysis and experiments show the superiority of our model on the endoscopic dataset and several public datasets. Jiahua Dong 0001, Yang Cong, Gan Sun, Yunsheng Yang, Xiaowei Xu 0001, Zhengming Ding |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | L3DOC: Lifelong 3D Object Classificationabstract3D object classification has been widely applied in both academic and industrial scenarios. However, most state-of-the-art algorithms rely on a fixed object classification task set, which cannot tackle the scenario when a new 3D object classification task is coming. Meanwhile, the existing lifelong learning models can easily destroy the learned tasks performance, due to the unordered, large-scale, and irregular 3D geometry data. To address these challenges, we propose a Lifelong 3D Object Classification (i.e., L3DOC) model, which can consecutively learn new 3D object classification tasks via imitating "human learning". More specifically, the core idea of our model is to capture and store the cross-task common knowledge of 3D geometry data in a 3D neural network, named as point-knowledge, through employing layer-wise point-knowledge factorization architecture. Afterwards, a task-relevant knowledge distillation mechanism is employed to connect the current task to previous relevant tasks and effectively prevent catastrophic forgetting. It consists of a point-knowledge distillation module and a transforming-space distillation module, which transfers the accumulated point-knowledge from previous tasks and soft-transfers the compact factorized representations of the transforming-space, respectively. To our best knowledge, the proposed L3DOC algorithm is the first attempt to perform deep learning on 3D object classification tasks in a lifelong learning way. Extensive experiments on several point cloud benchmarks illustrate the superiority of our L3DOC model over the state-of-the-art lifelong learning methods. Yang Cong, Gan Sun, Tao Zhang 0084, Jiahua Dong 0001, Hongsen Liu |
IEEE Trans. Image Process. | 5 |
| 2020 | What Can Be Transferred: Unsupervised Domain Adaptation for Endoscopic Lesions SegmentationabstractUnsupervised domain adaptation has attracted growing research attention on semantic segmentation. However, 1) most existing models cannot be directly applied into lesions transfer of medical images, due to the diverse appearances of same lesion among different datasets; 2) equal attention has been paid into all semantic representations instead of neglecting irrelevant knowledge, which leads to negative transfer of untransferable knowledge. To address these challenges, we develop a new unsupervised semantic transfer model including two complementary modules (i.e., T_D and T_F ) for endoscopic lesions segmentation, which can alternatively determine where and how to explore transferable domain-invariant knowledge between labeled source lesions dataset (e.g., gastroscope) and unlabeled target diseases dataset (e.g., enteroscopy). Specifically, T_D focuses on where to translate transferable visual information of medical lesions via residual transferability-aware bottleneck, while neglecting untransferable visual characterizations. Furthermore, T_F highlights how to augment transferable semantic features of various lesions and automatically ignore untransferable representations, which explores domain-invariant knowledge and in return improves the performance of T_D. To the end, theoretical analysis and extensive experiments on medical endoscopic dataset and several non-medical public datasets well demonstrate the superiority of our proposed model. Jiahua Dong 0001, Yang Cong, Gan Sun, Bineng Zhong 0001, Xiaowei Xu 0001 |
CVPR | 1 |
| 2020 | CSCL: Critical Semantic-Consistent Learning for Unsupervised Domain Adaptation
Jiahua Dong 0001, Yang Cong, Gan Sun, Xiaowei Xu 0001 |
ECCV (8) | 1 |
| 2019 | Semantic-Transferable Weakly-Supervised Endoscopic Lesions SegmentationabstractWeakly-supervised learning under image-level labels supervision has been widely applied to semantic segmentation of medical lesions regions. However, 1) most existing models rely on effective constraints to explore the internal representation of lesions, which only produces inaccurate and coarse lesions regions; 2) they ignore the strong probabilistic dependencies between target lesions dataset (e.g., enteroscopy images) and well-to-annotated source diseases dataset (e.g., gastroscope images). To better utilize these dependencies, we present a new semantic lesions representation transfer model for weakly-supervised endoscopic lesions segmentation, which can exploit useful knowledge from relevant fully-labeled diseases segmentation task to enhance the performance of target weakly-labeled lesions segmentation task. More specifically, a pseudo label generator is proposed to leverage seed information to generate highly-confident pseudo pixel labels by incorporating class balance and super-pixel spatial prior. It can iteratively include more hard-to-transfer samples from weakly-labeled target dataset into training set. Afterwards, dynamically-searched feature centroids for same class among different datasets are aligned by accumulating previously-learned features. Meanwhile, adversarial learning is also employed in this paper, to narrow the gap between the lesions among different datasets in output space. Finally, we build a new medical endoscopic dataset with 3659 images collected from more than 1100 volunteers. Extensive experiments on our collected dataset and several benchmark datasets validate the effectiveness of our model. Jiahua Dong 0001, Yang Cong, Gan Sun, Dongdong Hou |
ICCV | 1 |