VLDB 2026 Research / reviewers in the wild / expert
Gan Sun
dblp:191/2425
· DBLP profile ↗
69ranked-venue papers
12as first author
48since 2021 · last 2026
0000-0003-1111-6909ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 9 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 3 first-author · 24 since 2021Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Efficient and Effective Interactive 3D SegmentationabstractInteractive 3D segmentation embodies an advanced human-in-the-loop paradigm, where a model iteratively refines the segmentation of interested objects within a 3D point cloud through user feedback. Existing methods have achieved notable advancements at the expense of substantial resource consumption. To address this challenge, we introduce E2I3D, an efficient and effective model for interactive 3D segmentation. Specifically, we propose a two-stage efficiency-to-effectiveness framework to decouple efficiency and effectiveness, avoiding the high training cost of joint optimization. For efficiency in the first stage, we present heterogeneous pruning, which reliably compresses the model by ranking and pruning the constructed heterogeneous groups separately based on gradient compensation. For effectiveness in the second stage, we design hierarchical click-aware attention that integrates geometric details from high-resolution features with global context from low-resolution features to enhance click-guided interaction. Extensive experiments across public datasets demonstrate that E2I3D exceeds state-of-the-art methods in both efficiency and effectiveness. For instance, on the KITTI-360 dataset, E2I3D boosts the IoU for interactive single-object segmentation from 44.4% to 49.0% with 5 user clicks, while simultaneously reducing parameters from 39.3M to 5.7M. Wei Cong, Yang Cong, Jiahua Dong 0001, Gan Sun |
AAAI | 4 |
| 2026 | Learning From Each Other: Generalized Federated Incremental Semantic SegmentationabstractFederated learning (FL) has advanced semantic segmentation through decentralized training to reduce annotation costs. However, most FL-based semantic segmentation methods assume fixed foreground classes, resulting in catastrophic forgetting of old categories when local clients continually collect streaming data of new classes without storing old categories. Moreover, the irregular participation of new local clients with novel classes unseen by others may exacerbate heterogeneous forgetting across clients during global FL training. To resolve the above challenges, we propose a Hierarchical Forgetting Alleviation (HFA) model. By tackling forgetting within and across local clients, our model ensures that all local clients learn from each other as they continuously learn new categories. Specifically, to alleviate class-imbalanced forgetting within local clients induced by background shift, we develop a confidence-regularized pseudo labeling strategy to produce class-balanced soft pseudo labels for old categories that are labeled as background. Guided by soft pseudo labels, we design a graph-induced relation matching loss and a forgetting-balanced gradient propagation module to tackle ambiguous inter-class relations and class-imbalanced gradient propagation among old classes. Besides, a novel task detection module and an adaptive DBSCAN clustering are devised to address inter-client heterogeneous forgetting. They detect the arrival of new tasks to store the old global model for local pseudo labeling and distillation, while supplying global class prototypes for modeling inter-class relations and warm-starting global classifier. Experiments on multiple datasets verify our model's superiority over other methods. Jiahua Dong 0001, Wenqi Liang, Yang Cong, Gan Sun, Lixu Wang, Henghui Ding, Yulun Zhang 0001, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Self-Guided Discriminative Locality Preserving ProjectionsabstractLocality Preserving Projections (LPP) aims to find a projection matrix to map the high-dimensional data into a low-dimensional subspace while preserving the local manifold structure, which is a classical unsupervised subspace learning method. However, the lack of label guidance makes LPP not able to fully exploit the discriminative information of the data. To solve the problem, we propose a Self-Guided Discriminative LPP algorithm employing pseudo labels learned by K-Means to guide the subspace learning. In this way, it facilitates the discovery of discriminative cluster information while preserving inherent manifold structure. Besides, considering K-Means' sensitivity to selection of cluster centroids, we introduce a centerless K-Means method to improve robustness by eliminating the need of centroid initialization. We also discuss the internal relationship between K-Means and LPP, and prove that K-Means can be written in the form of LPP under certain conditions. Experiments on seven benchmark datasets demonstrate that our method greatly improves the clustering performance. Qianqian Wang 0001, Mengping Jiang, Gan Sun, Wei Feng 0010, Licheng Jiao |
IEEE Trans. Multim. | 3 |
| 2026 | Generalizable Multistage Assembly via One-Shot Category-Level DemonstrationabstractImitation learning offers a flexible approach for robot skill acquisition, enabling robots to learn complex tasks directly from demonstrations. However, most existing methods require a large number of demonstrations, whereas humans typically only need one or a few demonstrations. This discrepancy results in significant time consumption for data collection. Furthermore, these methods often assume that test scenarios will always be identical to the demonstration, which can lead to substantial performance degradation when facing novel scenarios, such as manipulating objects from the same category but with different shapes and sizes, or encountering object collisions during manipulation. To address these challenges, we propose a generalized multistage manipulation network for category-level robot assembly tasks. This network allows a robot to learn a multistage screw-nut assembly task from a single demonstration and generalize to new object instances with varying shapes and sizes. Specifically, the network uses category-level pose estimation to extract manipulation trajectories from the demonstration and applies manipulation-pose generalization to transfer these trajectories to novel instances. In addition, real-time action correction adjusts the trajectory based on real-time force feedback, enabling the robot to adapt to unexpected collisions during execution. We validate our method through experiments in both simulation and real-world environments, verifying its effectiveness and flexibility. Yang Cong, Ronghan Chen, Wei Cong, Gan Sun |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | GLAM: Global-Local Variation Awareness in Mamba-based World ModelabstractMimicking the real interaction trajectory in the inference of the world model has been shown to improve the sample efficiency of model-based reinforcement learning (MBRL) algorithms. Many methods directly use known state sequences for reasoning. However, this approach fails to enhance the quality of reasoning by capturing the subtle variation between states. Much like how humans infer trends in event development from this variation, in this work, we introduce Global-Local variation Awareness Mamba-based world model (GLAM) that improves reasoning quality by perceiving and predicting variation between states. GLAM comprises two Mamba-based parallel reasoning modules, GMamba and LMamba, which focus on perceiving variation from global and local perspectives, respectively, during the reasoning process. GMamba focuses on identifying patterns of variation between states in the input sequence and leverages these patterns to enhance the prediction of future state variation. LMamba emphasizes reasoning about unknown information, such as rewards, termination signals, and visual representations, by perceiving variation in adjacent states. By integrating the strengths of the two modules, GLAM accounts for higher-value variation in environmental changes, providing the agent with more efficient imagination-based training. We demonstrate that our method outperforms existing methods in normalized human scores on the Atari 100k benchmark. Wenqi Liang, Chunhui Hao, Gan Sun, Jiandong Tian |
AAAI | 4 |
| 2025 | Creative style transfer for image stylization via learning neural permutation
Zedong Zhang, Gan Sun, Li-Wei H. Lehman, Jian Yang 0003, Jun Li 0027 |
Knowl. Based Syst. | 3 |
| 2025 | MuseumMaker: Continual Style Customization Without Catastrophic ForgettingabstractPre-trainedlarge text-to-image (T2I) models with an appropriate text prompt has attracted growing interests in customized image generation fields. However, catastrophic forgetting issue makes it hard to continually synthesize new user-provided styles while retaining the satisfying results amongst learned styles. In this paper, we propose MuseumMaker, a method that enables the synthesis of images by following a set of customized styles in a never-end manner, and gradually accumulates these creative artistic works as a Museum. When facing with a new customization style, we develop a style distillation loss module to extract and learn the styles of the training data for new image generation task. It can minimize the learning biases caused by content of new training images, and address the catastrophic overfitting issue induced by few-shot images. To deal with catastrophic forgetting issue amongst past learned styles, we devise a dual regularization for shared-LoRA module to optimize the direction of model update, which could regularize the diffusion model from both weight and feature aspects, respectively. Meanwhile, to further preserve historical knowledge from past styles and address the limited representability of LoRA, we design a task-wise token learning module where a unique token embedding is learned to denote a new style. As any new user-provided style come, our MuseumMaker can capture the nuances of the new styles while maintaining the details of learned styles. Experimental results on diverse style datasets validate the effectiveness of our proposed MuseumMaker method, showcasing its robustness and versatility across various scenarios. Gan Sun, Wenqi Liang, Jiahua Dong 0001, Can Qin, Yang Cong |
IEEE Trans. Image Process. | 2 |
| 2024 | FTGraph: A Flexible Tree-Based Graph Store on Persistent Memory for Large-Scale Dynamic GraphsabstractTraditional in-memory graph systems often suffer from scalability due to the limited capacity and volatility of DRAM. Emerging non-volatile memory (NVM) provides an opportunity to achieve highly scalable and high-performance graph stores for its large capacity and persistence characteristics. However, directly deploying current in-memory graph storage systems on NVM would cause significant inefficiencies in NVM access, as their graph organization designed for DRAM may incur higher write amplification, crash inconsistency and costly concurrency control overhead in NVM for write-intensive work-loads. In this paper, we propose FTGraph, a Flexible Tree-based Graph storage system, for both efficient dynamical graph updates and analysis. To achieve this goal, we introduce a novel degree-aware suffix bit tree to effectively manage vertices and edges of the graph, enabling adaptability to real-world power-law degree distributions while significantly reducing NVM writes. Based on it, we adopt two optimization methods, logical vertex ID translation and sequential storage, for vertices with very high degrees within the tree to enhance graph analysis operations. We further integrate 8B NVM atomic writes with optimistic version-based concurrency control through a dual bitmap design to ensure low-overhead, log-free crash consistency and reduce read-writer contention. Experimental results show that FTGraph achieves up to$\mathbf{21.2}\times$higher update performance and up to$\mathbf{85.4}\times$higher analysis performance, compared with state-of-the-art dynamic graph systems implemented on NVMs. Gan Sun, Bo Li 0063, Xiaoyan Gu 0001, Weiping Wang 0005, Shuibing He |
CLUSTER | 1 |
| 2024 | Cs2K: Class-Specific and Class-Shared Knowledge Guidance for Incremental Semantic Segmentation
Wei Cong, Yang Cong, Gan Sun |
ECCV (5) | 4 |
| 2024 | GroupTrack: Multi-Object Tracking by Using Group Motion PatternsabstractThe main challenge of Multi-Object Tracking (MOT) lies in maintaining a distinctive identity for each target in dense crowds or occluded scenarios. Although the existing methods have achieved significantly progress by using robust object detectors or complex association strategies, they cannot effectively solve long-term tracking due to individually motion or appearance modeling for each single target. In this paper, we propose a novel 2D MOT tracker GroupTrack, to learn reliable motion state for each target using group motion patterns. Specifically, for each tracklet, we first choose its neighboring ones to form a group of motion patterns, which can provide informative clues for the motion estimation of the current tracklet. Then, we apply the group motion patterns to perform tracklet prediction and data association. By integrating prior from neighboring motion patterns into the data association process, GroupTrack provides a new paradigm for target motion modeling in extremely crowded and occluded scenarios. Through extensive experiments on the public MOT17 and MOT20 datasets, we demonstrate the effectiveness of our approach in challenging scenarios and show state-of-the-art performance at various MOT metrics. Xinglong Xu, Weihong Ren, Gan Sun, Haoyu Ji 0001, Yu Gao 0010, Honghai Liu 0001 |
IROS | 3 |
| 2024 | Novel Object Synthesis via Adaptive Text-Image HarmonyabstractIn this paper, we study an object synthesis task that combines an object text with an object image to create a new object image. However, most diffusion models struggle with this task, \textit{i.e.}, often generating an object that predominantly reflects either the text or the image due to an imbalance between their inputs. To address this issue, we propose a simple yet effective method called Adaptive Text-Image Harmony (ATIH) to generate novel and surprising objects.
First, we introduce a scale factor and an injection step to balance text and image features in cross-attention and to preserve image information in self-attention during the text-image inversion diffusion process, respectively. Second, to better integrate object text and image, we design a balanced loss function with a noise parameter, ensuring both optimal editability and fidelity of the object image. Third, to adaptively adjust these parameters, we present a novel similarity score function that not only maximizes the similarities between the generated object image and the input text/image but also balances these similarities to harmonize text and image integration.
Extensive experiments demonstrate the effectiveness of our approach, showcasing remarkable object creations such as colobus-glass jar. https://xzr52.github.io/ATIH/ Zeren Xiong, Zedong Zhang, Shuo Chen 0003, Xiang Li 0041, Gan Sun, Jian Yang 0003, Jun Li 0027 |
NeurIPS | 6 |
| 2024 | Where and How to Transfer: Knowledge Aggregation-Induced Transferability Perception for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation without accessing expensive annotation processes of target data has achieved remarkable successes in semantic segmentation. However, most existing state-of-the-art methods cannot explore whether semantic representations across domains are transferable or not, which may result in the negative transfer brought by irrelevant knowledge. To tackle this challenge, in this paper, we develop a novel Knowledge Aggregation-induced Transferability Perception (KATP) for unsupervised domain adaptation, which is a pioneering attempt to distinguish transferable or untransferable knowledge across domains. Specifically, the KATP module is designed to quantify which semantic knowledge across domains is transferable, by incorporating transferability information propagation from global category-wise prototypes. Based on KATP, we design a novel KATP Adaptation Network (KATPAN) to determine where and how to transfer. The KATPAN contains a transferable appearance translation module T_A() and a transferable representation augmentation module T_R(), where both modules construct a virtuous circle of performance promotion. T_A() develops a transferability-aware information bottleneck to highlight where to adapt transferable visual characterizations and modality information; T_R() explores how to augment transferable representations while abandoning untransferable information, and promotes the translation performance of T_A() in return. Experiments on several representative datasets and a medical dataset support the state-of-the-art performance of our model. Jiahua Dong 0001, Yang Cong, Gan Sun, Zhen Fang 0001, Zhengming Ding |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | No One Left Behind: Real-World Federated Class-Incremental LearningabstractFederated learning (FL) is a hot collaborative training framework via aggregating model parameters of decentralized local clients. However, most FL methods unreasonably assume data categories of FL framework are known and fixed in advance. Moreover, some new local clients that collect novel categories unseen by other clients may be introduced to FL training irregularly. These issues render global model to undergo catastrophic forgetting on old categories, when local clients receive new categories consecutively under limited memory of storing old categories. To tackle the above issues, we propose a novelLocal-GlobalAnti-forgetting (LGA) model. It ensures no local clients are left behind as they learn new classes continually, by addressing local and global catastrophic forgetting. Specifically, considering tackling class imbalance of local client to surmount local forgetting, we develop a category-balanced gradient-adaptive compensation loss and a category gradient-induced semantic distillation loss. They can balance heterogeneous forgetting speeds of hard-to-forget and easy-to-forget old categories, while ensure consistent class-relations within different tasks. Moreover, a proxy server is designed to tackle global forgetting caused by Non-IID class imbalance between different clients. It augments perturbed prototype images of new categories collected from local clients via self-supervised prototype augmentation, thus improving robustness to choose the best old global model for local-side semantic distillation loss. Experiments on representative datasets verify superior performance of our model against comparison methods. The code is available athttps://github.com/JiahuaDong/LGA. Jiahua Dong 0001, Hongliu Li, Yang Cong, Gan Sun, Yulun Zhang 0001, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Create Your World: Lifelong Text-to-Image DiffusionabstractText-to-image generative models can produce diverse high-quality images of concepts with a text prompt, which have demonstrated excellent ability in image generation, image translation, etc. We in this work study the problem of synthesizing instantiations of a user's own concepts in a never-ending manner,i.e.,create your world, where the new concepts from user are quickly learned with a few examples. To achieve this goal, we propose aLifelong text-to-imageDiffusionModel (L$^{2}$DM), which intends to overcome knowledge “catastrophic forgetting” for the past encountered concepts, and semantic “catastrophic neglecting” for one or more concepts in the text prompt. In respect of knowledge “catastrophic forgetting”, our L$^{2}$DM framework devises a task-aware memory enhancement module and an elastic-concept distillation module, which could respectively safeguard the knowledge of both prior concepts and each past personalized concept. When generating images with a user text prompt, the solution to semantic “catastrophic neglecting” is that a concept attention artist module can alleviate the semantic neglecting from concept aspect, and an orthogonal attention module can reduce the semantic binding from attribute aspect. To the end, our model can generate more faithful image across a range of continual text prompts in terms of both qualitative and quantitative metrics, when comparing with the related state-of-the-art models. The code will be released athttps://wenqiliang.github.io/. Gan Sun, Wenqi Liang, Jiahua Dong 0001, Jun Li 0027, Zhengming Ding, Yang Cong |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Self-Paced Weight Consolidation for Continual LearningabstractContinual learning algorithms which keep the parameters of new tasks close to that of previous tasks, are popular in preventing catastrophic forgetting in sequential task learning settings. However, 1) the performance for the new continual learner will be degraded without distinguishing the contributions of previously learned tasks; 2) the computational cost will be greatly increased with the number of tasks, since most existing algorithms need to regularize all previous tasks when learning new tasks. To address the above challenges, we propose aself-pacedWeightConsolidation (spWC) framework to attain robust continual learning via evaluating the discriminative contributions of previous tasks. To be specific, we develop a self-paced regularization to reflect the priorities of past tasks via measuring difficulty based on key performance indicator (i.e., accuracy). When encountering a new task, all previous tasks are sorted from “difficult” to “easy” based on the priorities. Then the parameters of the new continual learner will be learned via selectively maintaining the knowledge amongst more difficult past tasks, which could well overcome catastrophic forgetting with less computational cost. We adopt an alternative convex search to iteratively update the model parameters and priority weights in the bi-convex formulation. The proposed spWC framework is plug-and-play, which is applicable to most continual learning algorithms (e.g., EWC, MAS and RCIL) in different directions (e.g., classification and segmentation). Experimental results on several public benchmark datasets demonstrate that our proposed framework can effectively improve performance when compared with other popular continual learning algorithms. Wei Cong, Yang Cong, Gan Sun, Jiahua Dong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Gradient-Semantic Compensation for Incremental Semantic SegmentationabstractIncremental semantic segmentation focuses on continually learning the segmentation of new coming classes without obtaining the training data from previously seen classes. However, most current methods fail to tackle catastrophic forgetting and background shift since they 1) treat all previous classes equally without considering different forgetting paces caused by imbalanced gradient back-propagation; 2) lack strong semantic guidance between classes. In this paper, to solve the aforementioned challenges, we propose aGradient-SemanticCompensation (GSC) model, which surmounts incremental semantic segmentation from both gradient and semantic perspectives. Specifically, to handle catastrophic forgetting from the gradient aspect, we develop a step-aware gradient compensation that can balance forgetting paces of previously seen classes by re-weighting gradient back-propagation. Meanwhile, we propose a soft-sharp semantic relation distillation to distill consistent inter-class semantic relations via soft labels for alleviating catastrophic forgetting from the semantic aspect. In addition, we design a prototypical pseudo re-labeling which provides strong semantic guidance to mitigate background shift. It produces high-quality pseudo labels for background pixels belonging to previous classes by assessing distances of pixels relative to class-wise prototypes. Experiments on three public segmentation datasets provide strong evidence for the effectiveness of our proposed GSC model. Wei Cong, Yang Cong, Jiahua Dong 0001, Gan Sun, Henghui Ding |
IEEE Trans. Multim. | 4 |
| 2024 | Open-Ended Online Learning for Autonomous Visual PerceptionabstractThe visual perception systems aim to autonomously collect consecutive visual data and perceive the relevant information online like human beings. In comparison with the classical static visual systems focusing on fixed tasks (e.g., face recognition for visual surveillance), the real-world visual systems (e.g., the robot visual system) often need to handle unpredicted tasks and dynamically changed environments, which need to imitate human-like intelligence with open-ended online learning ability. Therefore, we provide a comprehensive analysis of open-ended online learning problems for autonomous visual perception in this survey. Based on "what to online learn" among visual perception scenarios, we classify the open-ended online learning methods into five categories: instance incremental learning to handle data attributes changing, feature evolution learning for incremental and decremental features with the feature dimension changed dynamically, class incremental learning and task incremental learning aiming at online adding new coming classes/tasks, and parallel and distributed learning for large-scale data to reveal the computational and storage advantages. We discuss the characteristic of each method and introduce several representative works as well. Finally, we introduce some representative visual perception applications to show the enhanced performance when using various open-ended online learning models, followed by a discussion of several future directions. Yang Cong, Gan Sun, Dongdong Hou, Jiahua Dong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Heterogeneous Forgetting Compensation for Class-Incremental LearningabstractClass-incremental learning (CIL) has achieved remarkable successes in learning new classes consecutively while overcoming catastrophic forgetting on old categories. However, most existing CIL methods unreasonably assume that all old categories have the same forgetting pace, and neglect negative influence of forgetting heterogeneity among different old classes on forgetting compensation. To surmount the above challenges, we develop a novel Heterogeneous Forgetting Compensation (HFC) model, which can resolve heterogeneous forgetting of easy-to-forget and hard-to-forget old categories from both representation and gradient aspects. Specifically, we design a task-semantic aggregation block to alleviate heterogeneous forgetting from representation aspect. It aggregates local category information within each task to learn task-shared global representations. Moreover, we develop two novel plug-and-play losses: a gradient-balanced forgetting compensation loss and a gradient-balanced relation distillation loss to alleviate forgetting from gradient aspect. They consider gradient-balanced compensation to rectify forgetting heterogeneity of old categories and heterogeneous relation consistency. Experiments on several representative datasets illustrate effectiveness of our HFC model. The code is available at https://github.com/JiahuaDong/HFC. Jiahua Dong 0001, Wenqi Liang, Yang Cong, Gan Sun |
ICCV | 4 |
| 2023 | I3DOD: Towards Incremental 3D Object Detection via Promptingabstract3D object detection have achieved significant performance in many fields, e.g., robotics system, autonomous driving, and augmented reality. However, most existing methods could cause catastrophic forgetting of old classes when performing on the class-incremental scenarios. Meanwhile, the current class-incremental 3D object detection methods neglect the relationships between the object localization information and category semantic information, and assume all the knowledge of old model is reliable. To address the above challenge, we present a novel Incremental 3D Object Detection framework with the guidance of prompting, i.e., I3DOD. Specifically, we propose a task-shared prompts mechanism to learn the matching relationships between the object localization information and category semantic information. After training on the current task, these prompts will be stored in our prompt pool, and perform the relationship of old classes in the next task. Moreover, we design a reliable distillation strategy to transfer knowledge from two aspects: a reliable dynamic distillation is developed to filter out the negative knowledge and transfer the reliable 3D knowledge to new detection model; the relation feature is proposed to capture the responses relation in feature space and protect plasticity of the model when learning novel 3D classes. To the end, we conduct comprehensive experiments on two benchmark datasets and our method outperforms the state-of-the-art object detection methods by 0.6% ∼ 2.7% in terms of [email protected]. Wenqi Liang, Gan Sun, Jiahua Dong 0001, Kangru Wang |
IROS | 2 |
| 2023 | Hierarchical Lifelong Machine Learning With "Watchdog"abstractMost existing lifelong machine learning works focus on how to exploit previously accumulated experiences (e.g., knowledge library) from earlier tasks, and transfer it to learn a new task. However, when a lifelong learning system encounters a large pool of candidate tasks, the knowledge among various coming tasks are imbalance, and the system should intelligently choose the next one to learn. In this paper, an effective “human cognition” strategy is taken into consideration via actively sorting the importance of new tasks in the process of unknown-to-known, and preferentially selecting the most valuable task with more information to learn. To be specific, we assess the importance of each new coming task (e.g., unknown or not) as an outlier detection issue, and propose to employ a “watchdog” knowledge library to reconstruct each task under$\ell _0$-norm constraint. The coming candidate tasks are then sorted depending on the sparse reconstruction scores in a descending order, which is referred to as a “watchdog” mechanism. Following this, we design a hierarchical knowledge library for the lifelong learning framework to encode new task with higher reconstruction score, where the library consists of two-level task descriptors, i.e., a high-dimensional one with low-rank constraint and a low-dimensional one. Both “watchdog” knowledge library and hierarchy knowledge library can be optimized with knowledge from both previously learned tasks and current task automatically. For model optimization, we explore an alternating method to iteratively update our proposed framework with a guaranteed convergence. Experimental results on several existing benchmarks demonstrate that our proposed model outperforms various state-of-the-art task selection methods. Gan Sun, Yang Cong, Changjun Gu, Zhengming Ding |
IEEE Trans. Big Data | 1 |
| 2023 | Lifelong Visual-Tactile Spectral Clustering for Robotic Object PerceptionabstractThis work presents a novel visual-tactile fused clustering framework, calledLifelongVisual-TactileSpectralClustering (i.e., LVTSC), to effectively learn consecutive object clustering tasks for robotic perception. Lifelong learning has become an important and hot topic in recent studies on machine learning, aiming to imitate “human learning” and reduce the computational cost when consecutively learning new tasks. Our proposed LVTSC model explores the knowledge transfer and representation correlation from a local modality-invariant perspective under modality-consistent constraint guidance. For the modality-invariant part, we design a set of modality-invariant basis libraries to capture the latent clustering centers of each modality and a set of modality-invariant feature libraries to forcibly embed the manifold information of each modality. A modal-consistent constraint reinforces the correlation between visual and tactile modalities by maximizing the feature manifold correspondences. When the object clustering task comes continuously, the overall objective is optimized by an effective alternating direction method with guaranteed convergence. Our proposed LVTSC framework has been extensively validated for its effectiveness and efficiency on the three challenging real-world robotic object perception datasets. Yang Cong, Gan Sun, Zhengming Ding |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Uni3DA: Universal 3D Domain Adaptation for Object RecognitionabstractTraditional 3D point cloud classification tasks focus on training a classifier in the closed-set scenario, where training and test data have the same label set and the same data distribution. In this work, we focus on a more challenging and realistic scenario in 3D point cloud classification task: universal domain adaptation (UniDA), where 1) data distributions for training and test data are different; and 2) for given label sets of training data and test data, they may have the shared classes and keep the private classes respectively, introducing an extra label set discrepancy. To solve UniDA problem, researchers have designed many methods based on 2D image datasets. However, due to the difficulty in capturing discriminative local geometric structures brought by the unordered and irregular 3D point cloud data, we cannot directly deploy the existing methods based on 2D image datasets to the 3D scenarios. To address UniDA in 3D scenarios, we develop a 3D universal domain adaptation framework, which consists of three modules: Self-Constructed Geometric (SCG) module, Local-to-Global Hypersphere Reasoning (LGHR) module and Self-Supervised Boundary Adaptation (SBA) module. SCG and LGHR generate the discriminative representation, which is used to acquire domain-invariant knowledge for training and test data. SBA is designed to automatically recognize whether a given label is from the shared label set or private label set, and adapts training and test data from the shared label set. To our best knowledge, this work is the first exploration of UniDA for 3D scenarios. Extensive experiments on public 3D point cloud datasets verify that the proposed method outperforms the existing UniDA methods. Yang Cong, Jiahua Dong 0001, Gan Sun |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | InOR-Net: Incremental 3-D Object Recognition Network for Point Cloud Representationabstract3-D object recognition has successfully become an appealing research topic in the real world. However, most existing recognition models unreasonably assume that the categories of 3-D objects cannot change over time in the real world. This unrealistic assumption may result in significant performance degradation for them to learn new classes of 3-D objects consecutively due to the catastrophic forgetting on old learned classes. Moreover, they cannot explore which 3-D geometric characteristics are essential to alleviate the catastrophic forgetting on old classes of 3-D objects. To tackle the above challenges, we develop a novel Incremental 3-D Object Recognition Network (i.e., InOR-Net), which could recognize new classes of 3-D objects continuously by overcoming the catastrophic forgetting on old classes. Specifically, category-guided geometric reasoning is proposed to reason local geometric structures with distinctive 3-D characteristics of each class by leveraging intrinsic category information. We then propose a novel critic-induced geometric attention mechanism to distinguish which 3-D geometric characteristics within each class are beneficial to overcome the catastrophic forgetting on old classes of 3-D objects while preventing the negative influence of useless 3-D characteristics. In addition, a dual adaptive fairness compensations' strategy is designed to overcome the forgetting brought by class imbalance by compensating biased weights and predictions of the classifier. Comparison experiments verify the state-of-the-art performance of the proposed InOR-Net model on several public point cloud datasets. Jiahua Dong 0001, Yang Cong, Gan Sun, Lixu Wang, Lingjuan Lyu, Jun Li 0027, Ender Konukoglu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Federated Class-Incremental LearningabstractFederated learning (FL) has attracted growing attentions via data-private collaborative training on decentralized clients. However, most existing methods unrealistically assume object classes of the overall framework are fixed over time. It makes the global model suffer from significant catastrophic forgetting on old classes in real-world scenarios, where local clients often collect new classes continuously and have very limited storage memory to store old classes. Moreover, new clients with unseen new classes may participate in the FL training, further aggravating the catastrophic forgetting of global model. To address these challenges, we develop a novel Global-Local Forgetting Compensation (GLFC) model, to learn a global class-incremental model for alleviating the catastrophic forgetting from both local and global perspectives. Specifically, to address local forgetting caused by class imbalance at the local clients, we design a class-aware gradient compensation loss and a class-semantic relation distillation loss to balance the forgetting of old classes and distill consistent inter-class relations across tasks. To tackle the global forgetting brought by the non-i.i.d class imbalance across clients, we propose a proxy server that selects the best old global model to assist the local relation distillation. Moreover, a prototype gradient-based communication mechanism is developed to protect the privacy. Our model outperforms state-of-the-art methods by 4.4%~15.1% in terms of average accuracy on representative benchmark datasets. The code is available at https://github.com/conditionWang/FCIL. Jiahua Dong 0001, Lixu Wang, Zhen Fang 0001, Gan Sun, Shichao Xu, Xiao Wang 0012, Qi Zhu 0002 |
CVPR | 4 |
| 2022 | Class-Incremental Gesture Recognition Learning with Out-of-Distribution DetectionabstractGesture recognition is a popular human-computer interaction technology, which has been widely applied in many fields (e.g., autonomous driving, medical care, VR and AR). However, 1) most existing gesture recognition methods focus on the fixed recognition scenarios with several gestures, which could lead to memory consumption and computational effort when continuously learning new gestures; 2) Meanwhile, the performance of popular class-incremental methods degrades significantly for previously learned classes (i.e., catastrophic forgetting) due to the ambiguity and variability of gestures. To tackle these challenges, we propose a novel class-incremental gesture recognition method with out-of-distribution (OOD) detection, which can continuously adapt to new gesture classes and achieve high performance for both learned and new gestures. Specifically, we construct an episodic memory with a subset of learned training samples to preserve the previous knowledge from forgetting. Moreover, the OOD detection-based memory management is developed for exploring the most representative and informative core set from the learned datasets. When a new gesture recognition task with strange classes comes, rehearsal enhancement is adopted to increase the diversity of memory exemplars for better fitting the real characteristics of gesture recognition. After deriving an effective class-incremental gesture recognition strategy, we perform experiments on two representative datasets to validate the superiority of our method. Evaluation experiments demonstrate that our proposed method substantially outperforms the state-of-the-art methods with about 2.17%-3.81% improvement under different class-incremental learning scenarios. Mingxue Li, Yang Cong, Gan Sun |
IROS | 4 |
| 2022 | Data Poisoning Attacks on Federated Machine LearningabstractFederated machine learning which enables resource-constrained node devices (e.g., Internet of Things (IoT) devices and smartphones) to establish a knowledge-shared model while keeping the raw data local, could provide privacy preservation, and economic benefit by designing an effective communication protocol. However, this communication protocol can be adopted by attackers to launch data poisoning attacks for different nodes, which has been shown as a big threat to most machine learning models. Therefore, we in this article intend to study the model vulnerability of federated machine learning, and even on IoT systems. To be specific, we here attempt to attacking a popular federated multitask learning framework, which uses a general multitask learning framework to handle statistical challenges in the federated learning setting. The problem of calculating optimal poisoning attacks on federated multitask learning is formulated as a bilevel program, which is adaptive to the arbitrary selection oftargetnodes andsource attackingnodes. We then propose a novel systems-aware optimization method, called as attack on federated learning (AT2FL), to efficiently derive the implicit gradients for poisoned data, and further attain optimal attack strategies in the federated machine learning. This is an earlier work, to our knowledge, that explores attacking federated machine learning via data poisoning. Finally, experiments on several real-world data sets demonstrate that when the attackers directly poison thetargetnodes or indirectly poison the related nodes via using the communication protocol, the federated multitask learning model is sensitive to both poisoning attacks. Gan Sun, Yang Cong, Jiahua Dong 0001, Qiang Wang 0015, Lingjuan Lyu, Ji Liu 0002 |
IEEE Internet Things J. | 1 |
| 2022 | What and How: Generalized Lifelong Spectral Clustering via Dual MemoryabstractSpectral clustering (SC) has become one of the most widely-adopted clustering algorithms, and been successfully applied into various applications. We in this work explore the problem of spectral clustering in a lifelong learning framework termed asGeneralizedLifelongSpectralClustering (GL$^2$SC). Different from most current studies, which concentrate on a fixed spectral clustering task set and cannot efficiently incorporate a new clustering task, the goal of our work is to establish a generalized model for new spectral clustering tasks by “What” and “How” to lifelong learn from past tasks. In respect of “what to lifelong learn”, our GL$^2$SC framework contains a dual memory mechanism with a deep orthogonal factorization manner: an orthogonal basis memory stores hidden and hierarchical clustering centers among learned tasks, and a feature embedding memory captures deep manifold representation common across multiple related tasks. When learning a new clustering task, the intuition here for “how to lifelong learn” is that GL$^2$SC can transfer intrinsic knowledge from dual memory mechanism to obtain task-specific encoding matrix. Then the encoding matrix can redefine the dual memory over time to provide maximal benefits when learning future tasks, and reversely maximize performance for past tasks. To achieve this, we propose an alternative optimization formulation with convergence guarantee for solving our GL$^2$SC model. To the end, empirical comparisons on several benchmark datasets show the effectiveness of our GL$^2$SC, in comparison with several state-of-the-art clustering models. Gan Sun, Yang Cong, Jiahua Dong 0001, Zhengming Ding |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Lifelong robotic visual-tactile perception learning
Jiahua Dong 0001, Yang Cong, Gan Sun, Tao Zhang 0084 |
Pattern Recognit. | 3 |
| 2022 | APAN: Across-Scale Progressive Attention Network for Single Image DerainingabstractRecent single image deraining works have achieved significant improvement using convolutional neural networks. However, the rain streaks in the rain image share similar patterns with its multi-scale versions, which are not fully exploited in recent works. In this paper, we propose anAcross-scaleProgressiveAttentionNetwork (i.e.,APAN) to explore the multi-scale collaborative representation for single image deraining. Specifically, we represent each rainy image via a multi-scale module. An across-scale attention module is then used to capture long-range feature correspondences from multi-scale features, which can model the rain streaks at an enlarging feature dimension. Afterwards, we construct a pyramid structure and further predict the rain streak progressively, which also guides the across-scale attention module to refine the feature representation from coarse to fine. The proposed model exploits self-similarity of features via an across-scale attention between different scales, which can well model the rain streak with long-range information. Experiments on several datasets show that our model achieves significant improvement compared with most state-of-the-art deraining models. Qiang Wang 0015, Gan Sun, Huijie Fan, Yandong Tang |
IEEE Signal Process. Lett. | 2 |
| 2022 | Fast Multi-View Outlier Detection via Deep EncoderabstractMulti-view outlier detection has a wide range of applications and has been well investigated in recent years. However, 1) most existing state-of-the-art methods cannot efficiently handle outlier detection problem for large-scale multi-view data, since exploring pairwise constraints among different views causes highly-computational cost; 2) the data collected from original heterogeneous feature spaces further increases the consistent difficulty of multi-view outlier detection. To address these issues, we present a fast multi-view outlier detection model via learning a low-rank latent subspace representation with deep encoder architecture, which can not only efficiently identify the outliers for large-scale data even with numerous data views, but also exploit a discriminative common latent subspace shared by all the views. First, we learn a set of orthogonal bases as view-specific dictionaries from a small dataset, which is randomly sampled from the original dataset. Benefitting from view-specific dictionaries, the sampled data is projected and decomposed as a shared and discriminative latent subspace representations, which correspond to the view-consistent and view-specific components across multiple views, respectively. Then, the obtained discriminative latent representations are applied to train the view-specific deep encoders, which can efficiently compute the abnormal score for the remaining instances. Our proposed model can cost-effectively identify the outliers in large-scale datasets from numerous data views with less computational complexity. Experiments conducted on eight real datasets and a synthesis dataset show that our proposed model outperforms the existing ones on effectiveness and efficiency. Dongdong Hou, Yang Cong, Gan Sun, Jiahua Dong 0001, Jun Li 0027, Kai Li 0012 |
IEEE Trans. Big Data | 3 |
| 2022 | Evolving Metric Learning for Incremental and Decremental FeaturesabstractOnline metric learning has been widely exploited for large-scale data classification due to the low computational cost. However, amongst online practical scenarios where the features are evolving (e.g., some features are vanished and some new features are augmented), most metric learning models cannot be successfully applied to these scenarios, although they can tackle the evolving instances efficiently. To address the challenge, we develop a new online Evolving Metric Learning (EML) model for incremental and decremental features, which can handle the instance and feature evolutions simultaneously by incorporating with a smoothed Wasserstein metric distance. Specifically, our model contains two essential stages: a Transforming stage (T-stage) and a Inheriting stage (I-stage). For the T-stage, we propose to extract important information from vanished features while neglecting non-informative knowledge, and forward it into survived features by transforming them into a low-rank discriminative metric space. It further explores the intrinsic low-rank structure of heterogeneous samples to reduce the computation and memory burden especially for highly-dimensional large-scale data. For the I-stage, we inherit the metric performance of survived features from the T-stage and then expand to include the new augmented features. Moreover, a smoothed Wasserstein distance is utilized to characterize the similarity relationships among the heterogeneous and complex samples, since the evolving features are not strictly aligned in the different stages. In addition to tackling the challenges in one-shot case, we also extend our model into multi-shot scenario. After deriving an efficient optimization strategy for both T-stage and I-stage, extensive experiments on several datasets verify the superior performance of our EML model. Jiahua Dong 0001, Yang Cong, Gan Sun, Tao Zhang 0084, Xiaowei Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Continuous Multi-View Human Action RecognitionabstractHuman action recognition which recognizes human actions in a video is a fundamental task in computer vision field. Although multiple existing methods with single-view or multi-view have been presented for human action recognition, these recognition approaches cannot be extended into new action recognition or action classification tasks, as well as discover underlying correlations among different views. To tackle the above problem, this paper proposes a new lifelong multi-view subspace learning framework for continuous human action recognition, which could exploit the complementary information amongst different views from a lifelong learning perspective. More specifically, a set of view-specific libraries is established to gradually store the useful information within multiple views. As a new action recognition task comes, we decompose the model parameters into a set of embedded parameters over view-specific libraries. A latent representation subspace is constructed via encouraging it to be close to different view-specific libraries, which can leverage the high-order correlations among different views and further avoid partial information for action recognition task. Meanwhile, we propose to employ an alternating direction strategy to optimize our proposed method. Empirical studies on real-world multi-view action recognition datasets have shown that our proposed framework attains the superior recognition performance and saves the computational time when continually learning new action recognition tasks. Qiang Wang 0015, Gan Sun, Jiahua Dong 0001, Qianqian Wang 0001, Zhengming Ding |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Visual-Tactile Fused Graph Learning for Object Clusteringabstracthighlights how to mitigate the differences between vision and touch, and further maximize the mutual information, which adopts a minimizing disagreement scheme to guide the modality-specific representations toward a unified affinity graph. To achieve ideal clustering performance, a Laplacian rank constraint is imposed to regularize the learned graph with ideal connected components, where noises that caused wrong connections are removed and clustering labels can be obtained directly. Finally, we propose an efficient alternating iterative minimization updating strategy, followed by a theoretical proof to prove framework convergence. Comprehensive experiments on five public datasets demonstrate the superiority of the proposed framework. Tao Zhang 0084, Yang Cong, Gan Sun, Jiahua Dong 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | PFDN: Pyramid Feature Decoupling Network for Single Image DerainingabstractRestoring images degraded by rain has attracted more academic attention since rain streaks could reduce the visibility of outdoor scenes. However, most existing deraining methods attempt to remove rain while recovering details in a unified framework, which is an ideal and contradictory target in the image deraining task. Moreover, the relative independence of rain streak features and background features is usually ignored in the feature domain. To tackle these challenges above, we propose an effective Pyramid Feature Decoupling Network (i.e., PFDN) for single image deraining, which could accomplish image deraining and details recovery with the corresponding features. Specifically, the input rainy image features are extracted via a recurrent pyramid module, where the features for the rainy image are divided into two parts, i.e., rain-relevant and rain-irrelevant features. Afterwards, we introduce a novel rain streak removal network for rain-relevant features and remove the rain streak from the rainy image by estimating the rain streak information. Benefiting from lateral outputs, we propose an attention module to enhance the rain-irrelevant features, which could generate spatially accurate and contextually reliable details for image recovery. For better disentanglement, we also enforce multiple causality losses at the pyramid features to encourage the decoupling of rain-relevant and rain-irrelevant features from the high to shallow layers. Extensive experiments demonstrate that our module can well model the rain-relevant information over the domain of the feature. Our framework empowered by PFDN modules significantly outperforms the state-of-the-art methods on single image deraining with multiple widely-used benchmarks, and also shows superiority in the fully-supervised domain. Qiang Wang 0015, Gan Sun, Jiahua Dong 0001, Yulun Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Representative Task Self-Selection for Flexible Clustered Lifelong LearningabstractConsider the lifelong machine learning paradigm whose objective is to learn a sequence of tasks depending on previous experiences, e.g., knowledge library or deep network weights. However, the knowledge libraries or deep networks for most recent lifelong learning models are of prescribed size and can degenerate the performance for both learned tasks and coming ones when facing with a new task environment (cluster). To address this challenge, we propose a novel incremental clustered lifelong learning framework with two knowledge libraries: feature learning library and model knowledge library, called Flexible Clustered Lifelong Learning (FCL3). Specifically, the feature learning library modeled by an autoencoder architecture maintains a set of representation common across all the observed tasks, and the model knowledge library can be self-selected by identifying and adding new representative models (clusters). When a new task arrives, our FCL3 model firstly transfers knowledge from these libraries to encode the new task, i.e., effectively and selectively soft-assigning this new task to multiple representative models over feature learning library. Then: 1) the new task with a higher outlier probability will be judged as a new representative, and used to redefine both feature learning library and representative models over time; or 2) the new task with lower outlier probability will only refine the feature learning library. For model optimization, we cast this lifelong learning problem as an alternating direction minimization problem as a new task comes. Finally, we evaluate the proposed framework by analyzing several multitask data sets, and the experimental results demonstrate that our FCL3 model can achieve better performance than most lifelong learning frameworks, even batch clustered multitask learning models. Gan Sun, Yang Cong, Qianqian Wang 0001, Bineng Zhong 0001, Yun Fu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | MedUCC: Medium-Driven Underwater Camera Calibration for Refractive 3-D ReconstructionabstractUnderwater camera calibration has attracted much attentions due to its significance in high-precision three-dimensional (3-D) pose estimation and scene reconstruction. However, most existing calibration methods focus on calibrating the underwater camera in a single scenario [e.g., air-glass-water], which can not well formulate the geometry constraint and further result in the complex calibration process. Moreover, the calibration precision of these methods is low, since multilayer transparent refractions with unknown layer orientation and distance make the task more difficult than that in air. To address these challenges, we develop a novel and efficient medium-driven method for underwater camera calibration (MedUCC), which can calibrate the underwater camera parameters, including the orientation and position of the transparent glass accurately. Our key idea of this article is to leverage the light-path changes formed by medium refractions between different media to acquire calibration data, which can better formulate the geometry constraint, and estimate the initial value of the underwater camera parameters. To improve the calibration accuracy of the underwater camera system, a quaternion-based solution is developed to refine the underwater camera parameters. To the end, we evaluate the calibration performance on an underwater camera system. Extensive experiment results demonstrate that our proposed method can obtain a better performance in comparison to the existing works. We also validate our proposed MedUCC method on our designed 3-D scanner prototype, which illustrates the superiority of our proposed calibration method. Changjun Gu, Yang Cong, Gan Sun, Yajun Gao, Tao Zhang 0084, Baojie Fan |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | I3DOL: Incremental 3D Object Learning without Catastrophic Forgettingabstract3D object classification has attracted appealing attentions in academic researches and industrial applications. However, most existing methods need to access the training data of past 3D object classes when facing the common real-world scenario: new classes of 3D objects arrive in a sequence. Moreover, the performance of advanced approaches degrades dramatically for past learned classes (i.e., catastrophic forgetting), due to the irregular and redundant geometric structures of 3D point cloud data. To address these challenges, we propose a new Incremental 3D Object Learning (i.e., I3DOL) model, which is the first exploration to learn new classes of 3D object continually. Specifically, an adaptive-geometric centroid module is designed to construct discriminative local geometric structures, which can better characterize the irregular point cloud representation for 3D object. Afterwards, to prevent the catastrophic forgetting brought by redundant geometric information, a geometric-aware attention mechanism is developed to quantify the contributions of local geometric structures, and explore unique 3D geometric characteristics with high contributions for classes incremental learning. Meanwhile, a score fairness compensation strategy is proposed to further alleviate the catastrophic forgetting caused by unbalanced data between past and new classes of 3D object, by compensating biased prediction for new classes in the validation phase. Experiments on 3D representative datasets validate the superiority of our I3DOL framework. Jiahua Dong 0001, Yang Cong, Gan Sun, Bingtao Ma, Lichen Wang |
AAAI | 3 |
| 2021 | Generative Partial Visual-Tactile Fused Object ClusteringabstractVisual-tactile fused sensing for object clustering has achieved significant progresses recently, since the involvement of tactile modality can effectively improve clustering performance. However, the missing data (i.e., partial data) issues always happen due to occlusion and noises during the data collecting process. This issue is not well solved by most existing partial multi-view clustering methods for the heterogeneous modality challenge. Naively employing these methods would inevitably induce a negative effect and further hurt the performance. To solve the mentioned challenges, we propose a Generative Partial Visual-Tactile Fused (i.e., GPVTF) framework for object clustering. More specifically, we first do partial visual and tactile features extraction from the partial visual and tactile data, respectively, and encode the extracted features in modality-specific feature subspaces. A conditional cross-modal clustering generative adversarial network is then developed to synthesize one modality conditioning on the other modality, which can compensate missing samples and align the visual and tactile modalities naturally by adversarial learning. To the end, two pseudo-label based KL-divergence losses are employed to update the corresponding modality-specific encoders. Extensive comparative experiments on three public visual-tactile datasets prove the effectiveness of our method. Tao Zhang 0084, Yang Cong, Gan Sun, Jiahua Dong 0001, Zhengming Ding |
AAAI | 3 |
| 2021 | Confident Anchor-Induced Multi-Source Free Domain AdaptationabstractUnsupervised domain adaptation has attracted appealing academic attentions by transferring knowledge from labeled source domain to unlabeled target domain. However, most existing methods assume the source data are drawn from a single domain, which cannot be successfully applied to explore complementarily transferable knowledge from multiple source domains with large distribution discrepancies. Moreover, they require access to source data during training, which are inefficient and unpractical due to privacy preservation and memory storage. To address these challenges, we develop a novel Confident-Anchor-induced multi-source-free Domain Adaptation (CAiDA) model, which is a pioneer exploration of knowledge adaptation from multiple source domains to the unlabeled target domain without any source data, but with only pre-trained source models. Specifically, a source-specific transferable perception module is proposed to automatically quantify the contributions of the complementary knowledge transferred from multi-source domains to the target domain. To generate pseudo labels for the target domain without access to the source data, we develop a confident-anchor-induced pseudo label generator by constructing a confident anchor group and assigning each unconfident target sample with a semantic-nearest confident anchor. Furthermore, a class-relationship-aware consistency loss is proposed to preserve consistent inter-class relationships by aligning soft confusion matrices across domains. Theoretical analysis answers why multi-source domains are better than a single source domain, and establishes a novel learning bound to show the effectiveness of exploiting multi-source domains. Experiments on several representative datasets illustrate the superiority of our proposed CAiDA model. The code is available at https://github.com/Learning-group123/CAiDA. Jiahua Dong 0001, Zhen Fang 0001, Anjin Liu, Gan Sun, Tongliang Liu |
NeurIPS | 4 |
| 2021 | Weakly-Supervised Cross-Domain Adaptation for Endoscopic Lesions SegmentationabstractWeakly-supervised learning has attracted growing research attention on medical lesions segmentation due to significant saving in pixel-level annotation cost. However, 1) most existing methods require effective prior and constraints to explore the intrinsic lesions characterization, which only generates incorrect and rough prediction; 2) they neglect the underlying semantic dependencies among weakly-labeled target enteroscopy diseases and fully-annotated source gastroscope lesions, while forcefully utilizing untransferable dependencies leads to the negative performance. To tackle above issues, we propose a new weakly-supervised lesions transfer framework, which can not only explore transferable domain-invariant knowledge across different datasets, but also prevent the negative transfer of untransferable representations. Specifically, a Wasserstein quantified transferability framework is developed to highlight wide-range transferable contextual dependencies, while neglecting the irrelevant semantic characterizations. Moreover, a novel self-supervised pseudo label generator is designed to equally provide confident pseudo pixel labels for both hard-to-transfer and easy-to-transfer target samples. It inhibits the enormous deviation of false pseudo pixel labels under the self-supervision manner. Afterwards, dynamically-searched feature centroids are aligned to narrow category-wise distribution shift. Comprehensive theoretical analysis and experiments show the superiority of our model on the endoscopic dataset and several public datasets. Jiahua Dong 0001, Yang Cong, Gan Sun, Yunsheng Yang, Xiaowei Xu 0001, Zhengming Ding |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | L3DOC: Lifelong 3D Object Classificationabstract3D object classification has been widely applied in both academic and industrial scenarios. However, most state-of-the-art algorithms rely on a fixed object classification task set, which cannot tackle the scenario when a new 3D object classification task is coming. Meanwhile, the existing lifelong learning models can easily destroy the learned tasks performance, due to the unordered, large-scale, and irregular 3D geometry data. To address these challenges, we propose a Lifelong 3D Object Classification (i.e., L3DOC) model, which can consecutively learn new 3D object classification tasks via imitating "human learning". More specifically, the core idea of our model is to capture and store the cross-task common knowledge of 3D geometry data in a 3D neural network, named as point-knowledge, through employing layer-wise point-knowledge factorization architecture. Afterwards, a task-relevant knowledge distillation mechanism is employed to connect the current task to previous relevant tasks and effectively prevent catastrophic forgetting. It consists of a point-knowledge distillation module and a transforming-space distillation module, which transfers the accumulated point-knowledge from previous tasks and soft-transfers the compact factorized representations of the transforming-space, respectively. To our best knowledge, the proposed L3DOC algorithm is the first attempt to perform deep learning on 3D object classification tasks in a lifelong learning way. Extensive experiments on several point cloud benchmarks illustrate the superiority of our L3DOC model over the state-of-the-art lifelong learning methods. Yang Cong, Gan Sun, Tao Zhang 0084, Jiahua Dong 0001, Hongsen Liu |
IEEE Trans. Image Process. | 3 |
| 2021 | Adversarial Multi-Path Residual Network for Image Super-ResolutionabstractRecently, deep convolutional neural networks have demonstrated remarkable progresses on single image super-resolution (SR) problem. However, most of them use more deeper and wider networks to improve SR performance, which is not practical in real-world applications due to large complexity, high computation cost, and low efficiency. In addition, they cannot provide high perception quality and guarantee objective quality simultaneously. To address these limitations, we in this paper propose a novel Adversarial Multipath Residual Network (AMPRN), which can largely suppress the number of network parameters and achieve a higher SR performance compared with the state-of-the-art methods. More specifically, we propose a multi-path residual block (MPRB) for multi-path residual network (MPRN) with fewer network parameters, which can extract abundant local features by fully using features from different paths generated by channel slices. These hierarchical features from all the MPRBs are then jointly aggregated by global gradual feature fusion. Following MPRN, we construct an adversarial gradient network with a gradient loss to make the gradient distribution of the generated SR images and ground truth image closer. In this way, the generated SR images of our model can provide high perception quality and objective quality. Finally, several experimental results demonstrate that our AMPRN achieves better performance in comparison with fewer parameters than the state-of-the-art methods. Qianqian Wang 0001, Quanxue Gao, Linlu Wu, Gan Sun, Licheng Jiao |
IEEE Trans. Image Process. | 4 |
| 2021 | Semi-Supervised Dual Relation Learning for Multi-Label ClassificationabstractIn a real-world scenario, an object could contain multiple tags instead of a single categorical label. To this end, multi-label learning (MLL) emerged. In MLL, the feature distributions are long-tailed and the complex semantic label relation and the long-tailed training samples are the main challenges. Semi-supervised learning is a potential solution. While, existing methods are mainly designed for single class scenario while ignoring the latent label relations. In addition, they cannot well handle the distribution shift commonly existing across source and target domains. To this end, a Semi-supervised Dual Relation Learning (SDRL) framework for multi-label classification is proposed. SDRL utilizes a few labeled samples as well as large scale unlabeled samples in the training stage. It jointly explores the inter-instance feature-level relation and the intra-instance label-level relation even from the unlabeled samples. In our model, a dual-classifier structure is deployed to obtain domain invariant representations. The prediction results from the classifiers are further compared and the most confident predictions are extracted as pseudo labels. A trainable label relation tensor is designed to explicitly explore the pairwise latent label relations and refine the predicted labels. SDRL is able to effectively and efficiently explore the feature-label relation as well as the label-label relation knowledge without any extra semantic knowledge. We evaluated SDRL in general and zero-shot multi-label classification tasks and we concluded that SDRL is superior to other SOTA baselines. Furthermore, extensive ablation studies have been done which reveal the effectiveness of each component in our framework. Lichen Wang, Yunyu Liu, Hang Di, Can Qin, Gan Sun, Yun Fu 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | iCmSC: Incomplete Cross-Modal Subspace ClusteringabstractCross-modal clustering aims to cluster the high-similar cross-modal data into one group while separating the dissimilar data. Despite the promising cross-modal methods have developed in recent years, existing state-of-the-arts cannot effectively capture the correlations between cross-modal data when encountering with incomplete cross-modal data, which can gravely degrade the clustering performance. To well tackle the above scenario, we propose a novel incomplete cross-modal clustering method that integrates canonical correlation analysis and exclusive representation, named incomplete Cross-modal Subspace Clustering (i.e., iCmSC). To learn a consistent subspace representation among incomplete cross-modal data, we maximize the intrinsic correlations among different modalities by deep canonical correlation analysis (DCCA), while an exclusive self-expression layer is proposed after the output layers of DCCA. We exploit a ℓ1,2-norm regularization in the learned subspace to make the learned representation more discriminative, which makes samples between different clusters mutually exclusive and samples among the same cluster attractive to each other. Meanwhile, the decoding networks are employed to reconstruct the feature representation, and further preserve the structural information among the original cross-modal data. To the end, we demonstrate the effectiveness of the proposed iCmSC via extensive experiments, which can justify that iCmSC achieves consistently large improvement compared with the state-of-the-arts. Qianqian Wang 0001, Huanhuan Lian, Gan Sun, Quanxue Gao, Licheng Jiao |
IEEE Trans. Image Process. | 3 |
| 2021 | Accurate and Fast Image Denoising via Attention Guided ScalingabstractImage denoising is a classical topic yet still a challenging problem, especially for reducing noise from the texture information. Feature scaling (e.g., downscale and upscale) is a widely practice in image denoising to enlarge receptive field size and save resources. However, such a common operation would lose some visual informative details. To address those problems, we propose fast and accurate image denoising via attention guided scaling (AGS). We find that the main informative feature channel and visual primitives during the scaling should keep similar. We then propose to extract the global channel-wise attention to maintain main channel information. Moreover, we propose to collect global descriptors by considering the entire spatial feature. And we then distribute the global descriptors to local positions of the scaled feature, based on their specific needs. We further introduce AGS for adversarial training, resulting in a more powerful discriminator. Extensive experiments show the effectiveness of our proposed method, where we clearly surpass all the state-of-the-art methods on most popular synthetic and real-world denoising benchmarks quantitatively and visually. We further show that our network contributes to other high-level vision applications and improves their performances significantly. Yulun Zhang 0001, Kai Li 0012, Gan Sun, Yu Kong 0001, Yun Fu 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Recurrent Generative Adversarial Network for Face CompletionabstractMost recently-proposed face completion algorithms use high-level features extracted from convolutional neural networks (CNNs) to recover semantic texture content. Although the completed face is natural-looking, the synthesized content still lacks lots of high-frequency details, since the high-level features cannot supply sufficient spatial information for details recovery. To tackle this limitation, in this paper, we propose aRecurrentGenerativeAdversarialNetwork (RGAN) for face completion. Unlike previous algorithms, RGAN can take full advantage of multi-level features, and further provide advanced representations from multiple perspectives, which can well restore spatial information and details in face completion. Specifically, our RGAN model is composed of a CompletionNet and a DisctiminationNet, where the CompletionNet consists of two deep CNNs and a recurrent neural network (RNN). The first deep CNN is presented to learn the internal regulations of a masked image and represent it with multi-level features. The RNN model then exploits the relationships among the multi-level features and transfers these features in another domain, which can be used to complete the face image. Benefiting from bidirectional short links, another CNN is used to fuse multi-level features transferred from RNN and reconstruct the face image in different scales. Meanwhile, two context discrimination networks in the DisctiminationNet are adopted to ensure the completed image consistency globally and locally. Experimental results on benchmark datasets demonstrate qualitatively and quantitatively that our model performs better than the state-of-the-art face completion models, and simultaneously generates realistic image content and high-frequency details. The code will be released available soon. Qiang Wang 0015, Huijie Fan, Gan Sun, Weihong Ren, Yandong Tang |
IEEE Trans. Multim. | 3 |
| 2021 | Continual Multiview Task Learning via Deep Matrix FactorizationabstractThe state-of-the-art multitask multiview (MTMV) learning tackles a scenario where multiple tasks are related to each other via multiple shared feature views. However, in many real-world scenarios where a sequence of the multiview task comes, the higher storage requirement and computational cost of retraining previous tasks with MTMV models have presented a formidable challenge for this lifelong learning scenario. To address this challenge, in this article, we propose a new continual multiview task learning model that integrates deep matrix factorization and sparse subspace learning in a unified framework, which is termed deep continual multiview task learning (DCMvTL). More specifically, as a new multiview task arrives, DCMvTL first adopts a deep matrix factorization technique to capture hidden and hierarchical representations for this new coming multiview task while accumulating the fresh multiview knowledge in a layerwise manner. Then, a sparse subspace learning model is employed for the extracted factors at each layer and further reveals cross-view correlations via a self-expressive constraint. For model optimization, we derive a general multiview learning formulation when a new multiview task comes and apply an alternating minimization strategy to achieve lifelong learning. Extensive experiments on benchmark data sets demonstrate the effectiveness of our proposed DCMvTL model compared with the existing state-of-the-art MTMV and lifelong multiview task learning models. Gan Sun, Yang Cong, Yulun Zhang 0001, Guoshuai Zhao 0001, Yun Fu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Robust 3-D Object Recognition via View-Specific ConstraintabstractThree-dimensional (3-D) object recognition task focuses on detecting the objects of a scene and estimating their 6-DOF pose via effective feature extraction methods. Most recent feature extraction methods are based on the deep neural networks and show good performances. However, these methods require rendering engine to assist in generating a large amount of training data, which need much time to converge and further lead to the block in a rapid industrial production line. Besides, for the common hand-crafted features, the lack of discriminant feature-points amongst various texture-less and surface-smooth objects can cause ambiguity in the process of feature-points matching. To address these challenges above, a hand-crafted 3-D feature descriptor with center offset and pose annotations is proposed in this article, which is called view-specific local projection statistics (VSLPSs). By relying on these annotations as seeds, a voting strategy is then used to transform the feature-points matching problem into the problem of voting an optimal model-view in the 6-DOF space. In this way, the ambiguity of feature-points matching caused by poor feature discrimination is eliminated. To the end, various experiments on three public datasets and our built 3-D bin-picking dataset demonstrate that our proposed VSLPS method performs well in comparison with the state-of-the-art. Hongsen Liu, Yang Cong, Gan Sun, Yandong Tang |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | Cross-Modal Subspace Clustering via Deep Canonical Correlation AnalysisabstractFor cross-modal subspace clustering, the key point is how to exploit the correlation information between cross-modal data. However, most hierarchical and structural correlation information among cross-modal data cannot be well exploited due to its high-dimensional non-linear property. To tackle this problem, in this paper, we propose an unsupervised framework named Cross-Modal Subspace Clustering via Deep Canonical Correlation Analysis (CMSC-DCCA), which incorporates the correlation constraint with a self-expressive layer to make full use of information among the inter-modal data and the intra-modal data. More specifically, the proposed model consists of three components: 1) deep canonical correlation analysis (Deep CCA) model; 2) self-expressive layer; 3) Deep CCA decoders. The Deep CCA model consists of convolutional encoders and correlation constraint. Convolutional encoders are used to obtain the latent representations of cross-modal data, while adding the correlation constraint for the latent representations can make full use of the information of the inter-modal data. Furthermore, self-expressive layer works on latent representations and constrain it perform self-expression properties, which makes the shared coefficient matrix could capture the hierarchical intra-modal correlations of each modality. Then Deep CCA decoders reconstruct data to ensure that the encoded features can preserve the structure of the original data. Experimental results on several real-world datasets demonstrate the proposed method outperforms the state-of-the-art methods. Quanxue Gao, Huanhuan Lian, Qianqian Wang 0001, Gan Sun |
AAAI | 4 |
| 2020 | Robust Low-Rank Discovery of Data-Driven Partial Differential EquationsabstractPartial differential equations (PDEs) are essential foundations to model dynamic processes in natural sciences. Discovering the underlying PDEs of complex data collected from real world is key to understanding the dynamic processes of natural laws or behaviors. However, both the collected data and their partial derivatives are often corrupted by noise, especially from sparse outlying entries, due to measurement/process noise in the real-world applications. Our work is motivated by the observation that the underlying data modeled by PDEs are in fact often low rank. We thus develop a robust low-rank discovery framework to recover both the low-rank data and the sparse outlying entries by integrating double low-rank and sparse recoveries with a (group) sparse regression method, which is implemented as a minimization problem using mixed nuclear norms with ℓ1 and ℓ0 norms. We propose a low-rank sequential (grouped) threshold ridge regression algorithm to solve the minimization problem. Results from several experiments on seven canonical models (i.e., four PDEs and three parametric PDEs) verify that our framework outperforms the state-of-art sparse and group sparse regression methods. Code is available at https://github.com/junli2019/Robust-Discovery-of-PDEs Jun Li 0027, Gan Sun, Guoshuai Zhao 0001, Li-Wei H. Lehman |
AAAI | 2 |
| 2020 | Lifelong Spectral ClusteringabstractIn the past decades, spectral clustering (SC) has become one of the most effective clustering algorithms. However, most previous studies focus on spectral clustering tasks with a fixed task set, which cannot incorporate with a new spectral clustering task without accessing to previously learned tasks. In this paper, we aim to explore the problem of spectral clustering in a lifelong machine learning framework, i.e., Lifelong Spectral Clustering (L2SC). Its goal is to efficiently learn a model for a new spectral clustering task by selectively transferring previously accumulated experience from knowledge library. Specifically, the knowledge library of L2SC contains two components: 1) orthogonal basis library: capturing latent cluster centers among the clusters in each pair of tasks; 2) feature embedding library: embedding the feature manifold information shared among multiple related tasks. As a new spectral clustering task arrives, L2SC firstly transfers knowledge from both basis library and feature library to obtain encoding matrix, and further redefines the library base over time to maximize performance across all the clustering tasks. Meanwhile, a general online update formulation is derived to alternatively update the basis library and feature library. Finally, the empirical experiments on several real-world benchmark datasets demonstrate that our L2SC model can effectively improve the clustering performance when comparing with other state-of-the-art spectral clustering algorithms. Gan Sun, Yang Cong, Qianqian Wang 0001, Jun Li 0027, Yun Fu 0001 |
AAAI | 1 |
| 2020 | Dual Relation Semi-Supervised Multi-Label LearningabstractMulti-label learning (MLL) solves the problem that one single sample corresponds to multiple labels. It is a challenging task due to the long-tail label distribution and the sophisticated label relations. Semi-supervised MLL methods utilize a small-scale labeled samples and large-scale unlabeled samples to enhance the performance. However, these approaches mainly focus on exploring the data distribution in feature space while ignoring mining the label relation inside of each instance. To this end, we proposed a Dual Relation Semi-supervised Multi-label Learning (DRML) approach which jointly explores the feature distribution and the label relation simultaneously. A dual-classifier domain adaptation strategy is proposed to align features while generating pseudo labels to improve learning performance. A relation network is proposed to explore the relation knowledge. As a result, DRML effectively explores the feature-label and label-label relations in both labeled and unlabeled samples. It is an end-to-end model without any extra knowledge. Extensive experiments illustrate the effectiveness and efficiency of our method1. Lichen Wang, Yunyu Liu, Can Qin, Gan Sun, Yun Fu 0001 |
AAAI | 4 |
| 2020 | Visual Tactile Fusion Object ClusteringabstractObject clustering, aiming at grouping similar objects into one cluster with an unsupervised strategy, has been extensively-studied among various data-driven applications. However, most existing state-of-the-art object clustering methods (e.g., single-view or multi-view clustering methods) only explore visual information, while ignoring one of most important sensing modalities, i.e., tactile information which can help capture different object properties and further boost the performance of object clustering task. To effectively benefit both visual and tactile modalities for object clustering, in this paper, we propose a deep Auto-Encoder-like Non-negative Matrix Factorization framework for visual-tactile fusion clustering. Specifically, deep matrix factorization constrained by an under-complete Auto-Encoder-like architecture is employed to jointly learn hierarchical expression of visual-tactile fusion data, and preserve the local structure of data generating distribution of visual and tactile modalities. Meanwhile, a graph regularizer is introduced to capture the intrinsic relations of data samples within each modality. Furthermore, we propose a modality-level consensus regularizer to effectively align the visual and tactile data in a common subspace in which the gap between visual and tactile data is mitigated. For the model optimization, we present an efficient alternating minimization strategy to solve our proposed model. Finally, we conduct extensive experiments on public datasets to verify the effectiveness of our framework. Tao Zhang 0084, Yang Cong, Gan Sun, Qianqian Wang 0001, Zhengming Ding |
AAAI | 3 |
| 2020 | What Can Be Transferred: Unsupervised Domain Adaptation for Endoscopic Lesions SegmentationabstractUnsupervised domain adaptation has attracted growing research attention on semantic segmentation. However, 1) most existing models cannot be directly applied into lesions transfer of medical images, due to the diverse appearances of same lesion among different datasets; 2) equal attention has been paid into all semantic representations instead of neglecting irrelevant knowledge, which leads to negative transfer of untransferable knowledge. To address these challenges, we develop a new unsupervised semantic transfer model including two complementary modules (i.e., T_D and T_F ) for endoscopic lesions segmentation, which can alternatively determine where and how to explore transferable domain-invariant knowledge between labeled source lesions dataset (e.g., gastroscope) and unlabeled target diseases dataset (e.g., enteroscopy). Specifically, T_D focuses on where to translate transferable visual information of medical lesions via residual transferability-aware bottleneck, while neglecting untransferable visual characterizations. Furthermore, T_F highlights how to augment transferable semantic features of various lesions and automatically ignore untransferable representations, which explores domain-invariant knowledge and in return improves the performance of T_D. To the end, theoretical analysis and extensive experiments on medical endoscopic dataset and several non-medical public datasets well demonstrate the superiority of our proposed model. Jiahua Dong 0001, Yang Cong, Gan Sun, Bineng Zhong 0001, Xiaowei Xu 0001 |
CVPR | 3 |
| 2020 | CSCL: Critical Semantic-Consistent Learning for Unsupervised Domain Adaptation
Jiahua Dong 0001, Yang Cong, Gan Sun, Xiaowei Xu 0001 |
ECCV (8) | 3 |
| 2020 | Double robust principal component analysis
Qianqian Wang 0001, Quanxue Gao, Gan Sun, Chris Ding |
Neurocomputing | 3 |
| 2020 | Fine-Grained Spatial Alignment Model for Person Re-Identification With Focal Triplet LossabstractRecent advances of person re-identification have well advocated the usage of human body cues to boost performance. However, most existing methods still retain on exploiting a relatively coarse-grained local information. Such information may include redundant backgrounds that are sensitive to the apparently similar persons when facing challenging scenarios like complex poses, inaccurate detection, occlusion and misalignment. In this paper we propose a novel Fine-Grained Spatial Alignment Model (FGSAM) to mine fine-grained local information to handle the aforementioned challenge effectively. In particular, we first design a pose resolve net with channel parse blocks (CPB) to extract pose information in pixel-level. This network allows the proposed model to be robust to complex pose variations while suppressing the redundant backgrounds caused by inaccurate detection and occlusion. Given the extracted pose information, a locally reinforced alignment mode is further proposed to address the misalignment problem between different local parts by considering different local parts along with attribute information in a fine-grained way. Finally, a focal triplet loss is designed to effectively train the entire model, which imposes a constraint on the intra-class and an adaptively weight adjustment mechanism to handle the hard sample problem. Extensive evaluations and analysis on Market1501, DukeMTMC-reid and PETA datasets demonstrate the effectiveness of FGSAM in coping with the problems of misalignment, occlusion and complex poses. Qinqin Zhou 0001, Bineng Zhong 0001, Xiangyuan Lan, Gan Sun, Yulun Zhang 0001, Baochang Zhang 0001, Rongrong Ji |
IEEE Trans. Image Process. | 4 |
| 2019 | Semantic-Transferable Weakly-Supervised Endoscopic Lesions SegmentationabstractWeakly-supervised learning under image-level labels supervision has been widely applied to semantic segmentation of medical lesions regions. However, 1) most existing models rely on effective constraints to explore the internal representation of lesions, which only produces inaccurate and coarse lesions regions; 2) they ignore the strong probabilistic dependencies between target lesions dataset (e.g., enteroscopy images) and well-to-annotated source diseases dataset (e.g., gastroscope images). To better utilize these dependencies, we present a new semantic lesions representation transfer model for weakly-supervised endoscopic lesions segmentation, which can exploit useful knowledge from relevant fully-labeled diseases segmentation task to enhance the performance of target weakly-labeled lesions segmentation task. More specifically, a pseudo label generator is proposed to leverage seed information to generate highly-confident pseudo pixel labels by incorporating class balance and super-pixel spatial prior. It can iteratively include more hard-to-transfer samples from weakly-labeled target dataset into training set. Afterwards, dynamically-searched feature centroids for same class among different datasets are aligned by accumulating previously-learned features. Meanwhile, adversarial learning is also employed in this paper, to narrow the gap between the lesions among different datasets in output space. Finally, we build a new medical endoscopic dataset with 3659 images collected from more than 1100 volunteers. Extensive experiments on our collected dataset and several benchmark datasets validate the effectiveness of our model. Jiahua Dong 0001, Yang Cong, Gan Sun, Dongdong Hou |
ICCV | 3 |
| 2019 | Memory-Based Parameterized Skills Learning for Mapless Visual NavigationabstractThe recently-proposed reinforcement learning for mapless visual navigation can generate an optimal policy for searching different targets. However, most state-of-the-art deep reinforcement learning (DRL) models depend on hard rewards to learn the optimal policy, which can lead to the lack of previous diverse experiences. Moreover, these pre-trained DRL models cannot generalize well to un-trained tasks. To overcome these problems above, in this paper, we propose a Memory-based Parameterized Skills Learning (MPSL) model for mapless visual navigation. The parameterized skills in our MPSL are learned to predict critic parameters for un-trained tasks in actor-critic reinforcement learning, which can be achieved by transferring memory sequence knowledge from long short term memory network. In order to generalize into un-trained tasks, MPSL aims to capture more discriminative features by using a scene-specific layer. Finally, experiment results on an indoor photographic simulation framework AI2THOR demonstrate the effectiveness of our proposed MPSL model, and the generalization ability to un-trained tasks. Yang Cong, Gan Sun |
ICIP | 3 |
| 2019 | Environment Driven Underwater Camera-IMU Calibration for Monocular Visual-Inertial SLAMabstractMost state-of-the-art underwater vision systems are calibrated manually in shallow water and used in open seas without changing. However, the refractivity of the water is adaptively changed depending on the salinity, temperature, depth or other underwater environmental indexes, which inevitably generate the calibration errors and induces incorrectness e.g., for underwater Simultaneously Localization and Mapping (SLAM). To address this issue, in this paper, we propose a new underwater Camera-Inertial Measurement Unit (IMU) calibration model, which just needs to be calibrated once in the air, and then both the intrinsic parameters and extrinsic parameters between the camera and IMU could be automatically calculated depending on the environment indexes. To our best knowledge, this is the first work to consider the underwater Camera-IMU calibration via environmental indexes. We also build a verification platform to validate the effectiveness of our proposed method on real experiments, and use it for underwater monocular Visual-Inertial SLAM. Changjun Gu, Yang Cong, Gan Sun |
ICRA | 3 |
| 2019 | LRDNN: Local-refining based Deep Neural Network for Person Re-Identification with Attribute DiscerningabstractRecently, pose or attribute information has been widely used to solve person re-identification (re-ID) problem. However, the inaccurate output from pose or attribute modules will impair the final person re-ID performance. Since re-ID, pose estimation and attribute recognition are all based on the person appearance information, we propose a Local-refining based Deep Neural Network (LRDNN) to aggregate pose estimation and attribute recognition to improve the re-ID performance. To this end, we add a pose branch to extract the local spatial information and optimize the whole network on both person identity and attribute objectives. To diminish the negative affect from unstable pose estimation, a novel structure called channel parse block (CPB) is introduced to learn weights on different feature channels in pose branch. Then two branches are combined with compact bilinear pooling. Experimental results on Market1501 and DukeMTMC-reid datasets illustrate the effectiveness of the proposed method. Qinqin Zhou 0001, Bineng Zhong 0001, Xiangyuan Lan, Gan Sun, Yulun Zhang 0001, Mengran Gou |
IJCAI | 4 |
| 2019 | Anomaly detection via adaptive greedy model
Dongdong Hou, Yang Cong, Gan Sun, Ji Liu 0002, Xiaowei Xu 0001 |
Neurocomputing | 3 |
| 2019 | Laplacian pyramid adversarial network for face completion
Qiang Wang 0015, Huijie Fan, Gan Sun, Yang Cong, Yandong Tang |
Pattern Recognit. | 3 |
| 2019 | Lifelong Metric LearningabstractThe state-of-the-art online learning approaches are only capable of learning the metric for predefined tasks. In this paper, we consider a lifelong learning problem to mimic "human learning," i.e., endowing a new capability to the learned metric for a new task from new online samples and incorporating the previous experiences. Therefore, we propose a new metric learning framework: lifelong metric learning (LML), which only utilizes the data of the new task to train the metric model while preserving the original capabilities. More specifically, the proposed LML maintains a common subspace for all learned metrics, named lifelong dictionary, transfers knowledge from the common subspace to learn each new metric learning task with task-specific idiosyncrasy, and redefines the common subspace over time to maximize performance across all metric tasks. For model optimization, we apply online passive aggressive optimization algorithm to achieve lifelong metric task learning, where the lifelong dictionary and task-specific partition are optimized alternatively and consecutively. Finally, we evaluate our approach by analyzing several multitask metric learning datasets. Extensive experimental results demonstrate effectiveness and efficiency of the proposed framework. Gan Sun, Yang Cong, Ji Liu 0002, Lianqing Liu, Xiaowei Xu 0001 |
IEEE Trans. Cybern. | 1 |
| 2018 | Active Lifelong Learning With "Watchdog"abstractLifelong learning intends to learn new consecutive tasks depending on previously accumulated experiences, i.e., knowledge library. However, the knowledge among different new coming tasks are imbalance. Therefore, in this paper, we try to mimic an effective "human cognition" strategy by actively sorting the importance of new tasks in the process of unknown-to-known and selecting to learn the important tasks with more information preferentially. To achieve this, we consider to assess the importance of the new coming task, i.e., unknown or not, as an outlier detection issue, and design a hierarchical dictionary learning model consisting of two-level task descriptors to sparse reconstruct each task with the l0 norm constraint. The new coming tasks are sorted depending on the sparse reconstruction score in descending order, and the task with high reconstruction score will be permitted to pass, where this mechanism is called as "watchdog." Next, the knowledge library of the lifelong learning framework encode the selected task by transferring previous knowledge, and then can also update itself with knowledge from both previously learned task and current task automatically. For model optimization, the alternating direction method is employed to solve our model and converges to a fixed point. Extensive experiments on both benchmark datasets and our own dataset demonstrate the effectiveness of our proposed model especially in task selection and dictionary learning. Gan Sun, Yang Cong, Xiaowei Xu 0001 |
AAAI | 1 |
| 2018 | Clustered Lifelong Learning Via Representative Task SelectionabstractConsider the lifelong machine learning problem where the objective is to learn new consecutive tasks depending on previously accumulated experiences, i.e., knowledge library. In comparison with most state-of-the-arts which adopt knowledge library with prescribed size, in this paper, we propose a new incremental clustered lifelong learning model with two libraries: feature library and model library, called Clustered Lifelong Learning (CL3), in which the feature library maintains a set of learned features common across all the encountered tasks, and the model library is learned by identifying and adding representative models (clusters). When a new task arrives, the original task model can be firstly reconstructed by representative models measured by capped l2-norm distance, i.e., effectively assigning the new task model to multiple representative models under feature library. Based on this assignment knowledge of new task, the objective of our CL3 model is to transfer the knowledge from both feature library and model library to learn the new task. The new task 1) with a higher outlier probability will then be judged as a new representative, and used to refine both feature library and representative models over time; 2) with lower outlier probability will only update the feature library. For the model optimisation, we cast this problem as an alternating direction minimization problem. To this end, the performance of CL3 is evaluated through comparing with most lifelong learning models, even some batch clustered multi-task learning models. Gan Sun, Yang Cong, Yu Kong 0001, Xiaowei Xu 0001 |
ICDM | 1 |
| 2018 | Online Low-Rank Metric Learning via Parallel Coordinate Descent Methodabstract11The corresponding author is Prof. Yang Cong. This work is supported by Nature Science Foundation of China under Grant (61722311, U1613214, 61533015) and CAS-Youth Innovation Promotion Association Scholarship (2012163)Recently, many machine learning problems rely on a valuable tool: metric learning. However, in many applications, large-scale applications embedded in high-dimensional feature space may induce both computation and storage requirements to grow quadratically. In order to tackle these challenges, in this paper, we intend to establish a robust metric learning formulation with the expectation that online metric learning and parallel optimization can solve large-scale and high-dimensional data efficiently, respectively. Specifically, based on the matrix factorization strategy, the first step aims to learn a similarity function in the objective formulation for similarity measurement; in the second step, we derive a variational trace norm to promote low-rankness on the transformation matrix. After converting this variational regularization into its separable form, for the model optimization, we present an parallel block coordinate descent method to learn the optimal metric parameters, which can handle the high-dimensional data in an efficient way. Crucially, our method shares the efficiency and flexibility of block coordinate descent method, and it is also guaranteed to converge to the optimal solution. Finally, we evaluate our approach by analyzing scene categorization dataset with tens of thousands of dimensions, and the experimental results show the effectiveness of our proposed model. Gan Sun, Yang Cong, Qiang Wang 0015, Xiaowei Xu 0001 |
ICPR | 1 |
| 2018 | User attribute discovery with missing labels
Yang Cong, Gan Sun, Ji Liu 0002, Jiebo Luo 0001 |
Pattern Recognit. | 2 |
| 2017 | Adaptive Greedy Dictionary Selection for Web Media SummarizationabstractInitializing an effective dictionary is an indispensable step for sparse representation. In this paper, we focus on the dictionary selection problem with the objective to select a compact subset of basis from original training data instead of learning a new dictionary matrix as dictionary learning models do. We first design a new dictionary selection model via l2,0norm. For model optimization, we propose two methods: one is the standard forward-backward greedy algorithm, which is not suitable for large-scale problems; the other is based on the gradient cues at each forward iteration and speeds up the process dramatically. In comparison with the state-of-the-art dictionary selection models, our model is not only more effective and efficient, but also can control the sparsity. To evaluate the performance of our new model, we select two practical web media summarization problems: 1) we build a new data set consisting of around 500 users, 3000 albums, and 1 million images, and achieve effective assisted albuming based on our model and 2) by formulating the video summarization problem as a dictionary selection issue, we employ our model to extract keyframes from a video sequence in a more flexible way. Generally, our model outperforms the state-of-the-art methods in both these two tasks. Yang Cong, Ji Liu 0002, Gan Sun, Quanzeng You, Yuncheng Li, Jiebo Luo 0001 |
IEEE Trans. Image Process. | 3 |