Yang Cong

dblp:76/3700 · DBLP profile ↗
← Back
111ranked-venue papers
18as first author
56since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 62 · 10 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 55 · 7 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 6 since 2021Systems, architecture and hardware · 7 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Computer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 Towards Efficient and Effective Interactive 3D Segmentation
abstract
Interactive 3D segmentation embodies an advanced human-in-the-loop paradigm, where a model iteratively refines the segmentation of interested objects within a 3D point cloud through user feedback. Existing methods have achieved notable advancements at the expense of substantial resource consumption. To address this challenge, we introduce E2I3D, an efficient and effective model for interactive 3D segmentation. Specifically, we propose a two-stage efficiency-to-effectiveness framework to decouple efficiency and effectiveness, avoiding the high training cost of joint optimization. For efficiency in the first stage, we present heterogeneous pruning, which reliably compresses the model by ranking and pruning the constructed heterogeneous groups separately based on gradient compensation. For effectiveness in the second stage, we design hierarchical click-aware attention that integrates geometric details from high-resolution features with global context from low-resolution features to enhance click-guided interaction. Extensive experiments across public datasets demonstrate that E2I3D exceeds state-of-the-art methods in both efficiency and effectiveness. For instance, on the KITTI-360 dataset, E2I3D boosts the IoU for interactive single-object segmentation from 44.4% to 49.0% with 5 user clicks, while simultaneously reducing parameters from 39.3M to 5.7M.
Wei Cong, Yang Cong, Jiahua Dong 0001, Gan Sun
AAAI2
2026 Learning From Each Other: Generalized Federated Incremental Semantic Segmentation
abstract
Federated learning (FL) has advanced semantic segmentation through decentralized training to reduce annotation costs. However, most FL-based semantic segmentation methods assume fixed foreground classes, resulting in catastrophic forgetting of old categories when local clients continually collect streaming data of new classes without storing old categories. Moreover, the irregular participation of new local clients with novel classes unseen by others may exacerbate heterogeneous forgetting across clients during global FL training. To resolve the above challenges, we propose a Hierarchical Forgetting Alleviation (HFA) model. By tackling forgetting within and across local clients, our model ensures that all local clients learn from each other as they continuously learn new categories. Specifically, to alleviate class-imbalanced forgetting within local clients induced by background shift, we develop a confidence-regularized pseudo labeling strategy to produce class-balanced soft pseudo labels for old categories that are labeled as background. Guided by soft pseudo labels, we design a graph-induced relation matching loss and a forgetting-balanced gradient propagation module to tackle ambiguous inter-class relations and class-imbalanced gradient propagation among old classes. Besides, a novel task detection module and an adaptive DBSCAN clustering are devised to address inter-client heterogeneous forgetting. They detect the arrival of new tasks to store the old global model for local pseudo labeling and distillation, while supplying global class prototypes for modeling inter-class relations and warm-starting global classifier. Experiments on multiple datasets verify our model's superiority over other methods.
Jiahua Dong 0001, Wenqi Liang, Yang Cong, Gan Sun, Lixu Wang, Henghui Ding, Yulun Zhang 0001, Luc Van Gool
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Domain Consistency Representation Learning for Lifelong Person Re-Identification
abstract
Lifelong person re-identification (LReID) exhibits a contradictory relationship between intra-domain discrimination and inter-domain gaps when learning from continuous data. Intra-domain discrimination focuses on individual nuances (i.e., clothing type, accessories,etc.), while inter-domain gaps emphasize domain consistency. Achieving a trade-off between maximizing intra-domain discrimination and minimizing inter-domain gaps is a crucial challenge for improving LReID performance. Most existing methods strive to reduce inter-domain gaps through knowledge distillation to maintain domain consistency. However, they often ignore intra-domain discrimination. To address this challenge, we propose a novel domain consistency representation learning (DCR) model that explores global and attribute-wise representations as a bridge to balance intra-domain discrimination and inter-domain gaps. At the intra-domain level, we explore the complementary relationship between global and attribute-wise representations to improve discrimination among similar identities. Excessive learning intra-domain discrimination can lead to catastrophic forgetting. We further develop an attribute-oriented anti-forgetting (AF) strategy that explores attribute-wise representations to enhance inter-domain consistency, and propose a knowledge consolidation (KC) strategy to facilitate knowledge transfer. Extensive experiments show that our DCR achieves superior performance compared to state-of-the-art LReID methods. Our code is available at https://github.com/LiuShiBen/DCR.
Shiben Liu, Huijie Fan, Qiang Wang 0015, Weihong Ren, Yandong Tang, Yang Cong
IEEE Trans. Circuits Syst. Video Technol.6
2026 Generalizable Multistage Assembly via One-Shot Category-Level Demonstration
abstract
Imitation learning offers a flexible approach for robot skill acquisition, enabling robots to learn complex tasks directly from demonstrations. However, most existing methods require a large number of demonstrations, whereas humans typically only need one or a few demonstrations. This discrepancy results in significant time consumption for data collection. Furthermore, these methods often assume that test scenarios will always be identical to the demonstration, which can lead to substantial performance degradation when facing novel scenarios, such as manipulating objects from the same category but with different shapes and sizes, or encountering object collisions during manipulation. To address these challenges, we propose a generalized multistage manipulation network for category-level robot assembly tasks. This network allows a robot to learn a multistage screw-nut assembly task from a single demonstration and generalize to new object instances with varying shapes and sizes. Specifically, the network uses category-level pose estimation to extract manipulation trajectories from the demonstration and applies manipulation-pose generalization to transfer these trajectories to novel instances. In addition, real-time action correction adjusts the trajectory based on real-time force feedback, enabling the robot to adapt to unexpected collisions during execution. We validate our method through experiments in both simulation and real-world environments, verifying its effectiveness and flexibility.
Yang Cong, Ronghan Chen, Wei Cong, Gan Sun
IEEE Trans. Neural Networks Learn. Syst.2
2025 DetailRefine: Towards Fine-Grained and Efficient Online Monocular 3D Reconstruction
abstract
Online monocular 3D reconstruction has attracted widespread attention as it promotes the application of robots in interactive scenarios. Most existing methods focus on 1) real-time reconstruction, 2) accurate voxel featuring learning, and 3) effective voxel sparsification algorithm. To this end, 1) they adopt a coarse-to-fine pipeline, where all non-empty voxels are sent to the next level for refinement. However, this results in over-refinement of flat regions, leading to unnecessary computational overhead. Furthermore, 2) advanced methods focus on exploring view visibility but overlook the discriminability among visible views, which limits the representation of learned voxel features. Moreover, 3) existing sparsification algorithms struggle to distinguish detailed and empty voxels, resulting in either the loss of detailed voxels or the retention of empty voxels. To tackle these challenges, 1) we present Dynamic Detail Refinement (DDR) to allocate more voxels to detailed regions for refinement, which could alleviate the computational burden. Furthermore, 2) we propose Discriminability-Aware Fusion (DAF) to focus on discriminative views, which helps to capture accurate voxel features. In addition, 3) we propose Hierarchical Hybrid Sparsification (HHS) to balance global completeness and local refinement, which helps to preserve detailed voxels at hierarchical levels effectively. Extensive experiments conducted on the representative ScanNet (V2) and 7-Scenes datasets demonstrate the superiority of the proposed method.
Fupeng Chu, Yang Cong, Ronghan Chen
ICRA2
2025 Learning Generalizable 3D Manipulation With 10 Demonstrations
abstract
Learning robust and generalizable manipulation skills from few demonstrations remains a key challenge in robotics, with broad applications in industrial automation and service robotics. Although recent imitation learning methods have achieved impressive results, they often require a large amount of demonstration data and struggle to generalize across different spatial variants. In this work, we propose a framework that learns 3D manipulation policies from only 10 demonstrations while achieving robust generalization to unseen spatial configurations through semantic-guided perception and spatial-equivariant policy learning. Our framework consists of two key modules: a Semantic Guided Perception module that extracts task-aware 3D representations from RGB-D inputs using semantic priors and a Spatial Generalized Decision module implementing a diffusion-based policy that preserves spatial equivariance through denoising. Central to our framework is a spatially equivariant training strategy, which adapts 2D data augmentation principles to 3D manipulation by maintaining gripper-object spatial relationships during trajectory augmentation. We validate our framework through extensive experiments on both simulation benchmarks and real-world robotic systems. Our method demonstrates a significant improvement in success rates over state-of-the-art approaches on a series of challenging tasks, particularly under significant object pose variations. This work shows significant potential to advance efficient and generalizable manipulation skill learning in real-world applications.
Yang Cong, Bohao Huang, Jiahao Long, Ronghan Chen, Huijie Fan
IROS2
2025 Multifeature Fusion-Based Closed-Box Attack Detection Method for Automatic Modulation Classification
abstract
The application of deep learning (DL) technologies in automatic modulation classification (AMC) faces significant challenges from adversarial attacks. Existing attack methods often rely on idealized assumptions and are inadequate for complex real-world scenarios. To address this issue, we propose an innovative multifeature fusion-based closed-box attack detection method (MFCA). This method enhances the effectiveness and quality of adversarial examples by utilizing a multimodel fusion strategy to extract diverse features. MFCA consists of three key steps: 1) we obtain an initial dataset from the target model and use a multifeature fusion model to approximate its decision boundary; 2) we enhance the model’s stability by applying data augmentation and incremental learning techniques through multiple rounds of training; and 3) we generate adversarial examples using the optimized fusion model. By constructing this comprehensive attack framework, we demonstrate how integrating diverse model characteristics enhances adversarial attack performance and evaluate the robustness of the MFCA method against defense strategies. Experimental results show that adversarial examples generated by MFCA exhibit significantly higher transferability across different perturbation levels compared to those from single substitute models and other attack methods, with superior attack efficacy. This research provides an effective solution for closed-box adversarial attacks and validates the advantages of multifeature fusion models in adversarial example generation.
Rui Gao 0005, Yang Cong, Xuhao Zhang
IEEE Internet Things J.2
2025 Lightweight Class Incremental Semantic Segmentation Without Catastrophic Forgetting
abstract
Class incremental semantic segmentation (CISS) aims to progressively segment newly introduced classes while preserving the memory of previously learned ones. Traditional CISS methods directly employ advanced semantic segmentation models (e.g., Deeplab-v3) as continual learners. However, these methods require substantial computational and memory resources, limiting their deployment on edge devices. In this paper, we propose a Lightweight Class Incremental Semantic Segmentation (LISS) model tailored for resource-constrained scenarios. Specifically, we design an automatic knowledge-preservation pruning strategy based on the Hilbert-Schmidt Independence Criterion (HSIC) Lasso, which automatically compresses the CISS model by searching for global penalty coefficients. Nonetheless, reducing model parameters exacerbates catastrophic forgetting during incremental learning. To mitigate this challenge, we develop a clustering-based pseudo labels generator to obtain high-quality pseudo labels by considering the feature space structure of old classes. It adjusts predicted probabilities from the old model according to the feature proximity to nearest sub-cluster centers for each class. Additionally, we introduce a customized soft labels module that distills the semantic relationships between classes separately. It decomposes soft labels into target probabilities, background probabilities, and other probabilities, thereby maintaining knowledge of previously learned classes in a fine-grained manner. Extensive experiments on two benchmark datasets demonstrate that our LISS model outperforms state-of-the-art approaches in both effectiveness and efficiency.
Wei Cong, Yang Cong
IEEE Trans. Image Process.2
2025 MuseumMaker: Continual Style Customization Without Catastrophic Forgetting
abstract
Pre-trainedlarge text-to-image (T2I) models with an appropriate text prompt has attracted growing interests in customized image generation fields. However, catastrophic forgetting issue makes it hard to continually synthesize new user-provided styles while retaining the satisfying results amongst learned styles. In this paper, we propose MuseumMaker, a method that enables the synthesis of images by following a set of customized styles in a never-end manner, and gradually accumulates these creative artistic works as a Museum. When facing with a new customization style, we develop a style distillation loss module to extract and learn the styles of the training data for new image generation task. It can minimize the learning biases caused by content of new training images, and address the catastrophic overfitting issue induced by few-shot images. To deal with catastrophic forgetting issue amongst past learned styles, we devise a dual regularization for shared-LoRA module to optimize the direction of model update, which could regularize the diffusion model from both weight and feature aspects, respectively. Meanwhile, to further preserve historical knowledge from past styles and address the limited representability of LoRA, we design a task-wise token learning module where a unique token embedding is learned to denote a new style. As any new user-provided style come, our MuseumMaker can capture the nuances of the new styles while maintaining the details of learned styles. Experimental results on diverse style datasets validate the effectiveness of our proposed MuseumMaker method, showcasing its robustness and versatility across various scenarios.
Gan Sun, Wenqi Liang, Jiahua Dong 0001, Can Qin, Yang Cong
IEEE Trans. Image Process.6
2025 DetailRecon: Focusing on Detailed Regions for Online Monocular 3D Reconstruction
abstract
Learning-based online monocular 3D reconstruction has emerged with great potential recently. Most state-of-the-art methods focus on two key questions, namely 1) how to exploit accurate voxel features and 2) how to preserve detailed voxels in the sparsification process. However, 1) most methods adopt the same receptive field to extract features for both informative and uninformative regions, which struggle to capture geometric details. Furthermore, 2) they mainly utilize a fixed threshold or a straightforward ray-based algorithm to discard voxels in the sparsification process. However, some detailed regions (especially thin regions) may be discarded incorrectly. To tackle these challenges, we present a novel method named DetailRecon to focus on detailed regions that contain more geometric information. Specifically, we first propose an Adaptive Hybrid Fusion (AHF) module and a Connectivity-Aware Sparsification (CAS) module for voxel feature learning and voxel sparsification, respectively. 1) The AHF receives multiple feature maps with different receptive fields as input, and adaptively adopts a smaller receptive field for regions with fine structures to exploit accurate geometric details. 2) The CAS updates the occupancy value of voxels based on the connected voxels within its neighbor space, which could expand the radiation range of reliable voxels in detailed regions and eventually reduce their probability of being discarded. Moreover, 3) we introduce a lightweight yet effective pipeline named Focus On Fine (FOF) to accelerate our DetailRecon. In addition, 4) we propose a Hierarchical Consistency Loss (HCL) to align multi-level volume features, which assists in exploring accurate volume features for recovering more details. Extensive experiments conducted on the ScanNet (V2) and 7-Scenes datasets demonstrate the superiority of our DetailRecon.
Fupeng Chu, Yang Cong, Ronghan Chen
IEEE Trans. Multim.2
2024 Cs2K: Class-Specific and Class-Shared Knowledge Guidance for Incremental Semantic Segmentation
Wei Cong, Yang Cong, Gan Sun
ECCV (5)2
2024 Marrying NeRF with Feature Matching for One-step Pose Estimation
abstract
Given the image collection of an object, we aim at building a real-time image-based pose estimation method, which requires neither its CAD model nor hours of object-specific training. Recent NeRF-based methods provide a promising solution by directly optimizing the pose from pixel loss between rendered and target images. However, during inference, they require long converging time, and suffer from local minima, making them impractical for real-time robot applications. We aim at solving this problem by marrying image matching with NeRF. With 2D matches and depth rendered by NeRF, we directly solve the pose in one step by building 2D-3D correspondences between target and initial view, thus allowing for real-time prediction. Moreover, to improve the accuracy of 2D-3D correspondences, we propose a 3D consistent point mining strategy, which effectively discards unfaithful points reconstruted by NeRF. Moreover, current NeRF-based methods naively optimizing pixel loss fail at occluded images. Thus, we further propose a 2D matches based sampling strategy to preclude the occluded area. Experimental results on representative datasets prove that our method outperforms state-of-the-art methods, and improves inference efficiency by 90×, achieving real-time prediction at 6 FPS.
Ronghan Chen, Yang Cong
ICRA2
2024 Underwater RGB-D imaging system with millimetric precision
Yajun Gao, Yang Cong, Mingxue Li
Sci. China Inf. Sci.2
2024 Where and How to Transfer: Knowledge Aggregation-Induced Transferability Perception for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation without accessing expensive annotation processes of target data has achieved remarkable successes in semantic segmentation. However, most existing state-of-the-art methods cannot explore whether semantic representations across domains are transferable or not, which may result in the negative transfer brought by irrelevant knowledge. To tackle this challenge, in this paper, we develop a novel Knowledge Aggregation-induced Transferability Perception (KATP) for unsupervised domain adaptation, which is a pioneering attempt to distinguish transferable or untransferable knowledge across domains. Specifically, the KATP module is designed to quantify which semantic knowledge across domains is transferable, by incorporating transferability information propagation from global category-wise prototypes. Based on KATP, we design a novel KATP Adaptation Network (KATPAN) to determine where and how to transfer. The KATPAN contains a transferable appearance translation module T_A() and a transferable representation augmentation module T_R(), where both modules construct a virtuous circle of performance promotion. T_A() develops a transferability-aware information bottleneck to highlight where to adapt transferable visual characterizations and modality information; T_R() explores how to augment transferable representations while abandoning untransferable information, and promotes the translation performance of T_A() in return. Experiments on several representative datasets and a medical dataset support the state-of-the-art performance of our model.
Jiahua Dong 0001, Yang Cong, Gan Sun, Zhen Fang 0001, Zhengming Ding
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 No One Left Behind: Real-World Federated Class-Incremental Learning
abstract
Federated learning (FL) is a hot collaborative training framework via aggregating model parameters of decentralized local clients. However, most FL methods unreasonably assume data categories of FL framework are known and fixed in advance. Moreover, some new local clients that collect novel categories unseen by other clients may be introduced to FL training irregularly. These issues render global model to undergo catastrophic forgetting on old categories, when local clients receive new categories consecutively under limited memory of storing old categories. To tackle the above issues, we propose a novelLocal-GlobalAnti-forgetting (LGA) model. It ensures no local clients are left behind as they learn new classes continually, by addressing local and global catastrophic forgetting. Specifically, considering tackling class imbalance of local client to surmount local forgetting, we develop a category-balanced gradient-adaptive compensation loss and a category gradient-induced semantic distillation loss. They can balance heterogeneous forgetting speeds of hard-to-forget and easy-to-forget old categories, while ensure consistent class-relations within different tasks. Moreover, a proxy server is designed to tackle global forgetting caused by Non-IID class imbalance between different clients. It augments perturbed prototype images of new categories collected from local clients via self-supervised prototype augmentation, thus improving robustness to choose the best old global model for local-side semantic distillation loss. Experiments on representative datasets verify superior performance of our model against comparison methods. The code is available athttps://github.com/JiahuaDong/LGA.
Jiahua Dong 0001, Hongliu Li, Yang Cong, Gan Sun, Yulun Zhang 0001, Luc Van Gool
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Create Your World: Lifelong Text-to-Image Diffusion
abstract
Text-to-image generative models can produce diverse high-quality images of concepts with a text prompt, which have demonstrated excellent ability in image generation, image translation, etc. We in this work study the problem of synthesizing instantiations of a user's own concepts in a never-ending manner,i.e.,create your world, where the new concepts from user are quickly learned with a few examples. To achieve this goal, we propose aLifelong text-to-imageDiffusionModel (L$^{2}$DM), which intends to overcome knowledge “catastrophic forgetting” for the past encountered concepts, and semantic “catastrophic neglecting” for one or more concepts in the text prompt. In respect of knowledge “catastrophic forgetting”, our L$^{2}$DM framework devises a task-aware memory enhancement module and an elastic-concept distillation module, which could respectively safeguard the knowledge of both prior concepts and each past personalized concept. When generating images with a user text prompt, the solution to semantic “catastrophic neglecting” is that a concept attention artist module can alleviate the semantic neglecting from concept aspect, and an orthogonal attention module can reduce the semantic binding from attribute aspect. To the end, our model can generate more faithful image across a range of continual text prompts in terms of both qualitative and quantitative metrics, when comparing with the related state-of-the-art models. The code will be released athttps://wenqiliang.github.io/.
Gan Sun, Wenqi Liang, Jiahua Dong 0001, Jun Li 0027, Zhengming Ding, Yang Cong
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 OPEN: Occlusion-Invariant Perception Network for Single Image-Based 3D Shape Retrieval
abstract
Single image-based 3D shape retrieval (IBSR) has attracted appealing academic interests recently, which aims to find the corresponding 3D shape from a shape repository for a given single 2D image. However, state-of-the-art methods neglect the discrepancy in the image domain due to unavoidable occlusion. The occluded image representations acting as noise, may perturb the alignment of the normal 2D representations with the 3D representations, resulting in occlusion-sensitive image-shape retrieval. To tackle this crucial challenge, in this paper, we propose a novel Occlusion-invariant PErception Network (OPEN) to learn occlusion-invariant image representations and image-shape correspondence. Specifically, we propose a hard occlusion example mining strategy to sample a hard image pair. Hereafter, to enforce the consistency between normal and occluded 2D images, we propose an Occlusion-invariant Image Consistency (OIC) based on hard image pairs, which gathers 2D image representations of the same instance while pushing away other 2D image representations. In addition, to prevent the 3D representations from perturbation by the occluded 2D representations, we design an Occlusion-invariant Correspondence Consistency (OCC) based on hard image pairs, which pulls the image-specific 3D shape embedding derived by attention mechanism close to the other 2D image representation of the same instance. The combination of OIC and OCC leads to accurate 2D-3D shape matching in challenging occluded scenarios. Our OPEN outperforms state-of-the-art methods by 6%~11% in terms of Top-1 retrieval accuracy on several representative benchmark datasets.
Fupeng Chu, Yang Cong, Ronghan Chen
IEEE Trans. Circuits Syst. Video Technol.2
2024 Self-Paced Weight Consolidation for Continual Learning
abstract
Continual learning algorithms which keep the parameters of new tasks close to that of previous tasks, are popular in preventing catastrophic forgetting in sequential task learning settings. However, 1) the performance for the new continual learner will be degraded without distinguishing the contributions of previously learned tasks; 2) the computational cost will be greatly increased with the number of tasks, since most existing algorithms need to regularize all previous tasks when learning new tasks. To address the above challenges, we propose aself-pacedWeightConsolidation (spWC) framework to attain robust continual learning via evaluating the discriminative contributions of previous tasks. To be specific, we develop a self-paced regularization to reflect the priorities of past tasks via measuring difficulty based on key performance indicator (i.e., accuracy). When encountering a new task, all previous tasks are sorted from “difficult” to “easy” based on the priorities. Then the parameters of the new continual learner will be learned via selectively maintaining the knowledge amongst more difficult past tasks, which could well overcome catastrophic forgetting with less computational cost. We adopt an alternative convex search to iteratively update the model parameters and priority weights in the bi-convex formulation. The proposed spWC framework is plug-and-play, which is applicable to most continual learning algorithms (e.g., EWC, MAS and RCIL) in different directions (e.g., classification and segmentation). Experimental results on several public benchmark datasets demonstrate that our proposed framework can effectively improve performance when compared with other popular continual learning algorithms.
Wei Cong, Yang Cong, Gan Sun, Jiahua Dong 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 IOSL: Incremental Open Set Learning
abstract
Class incremental learning (CIL) has drawn wide attention in academic researches. However, most existing methods cannot be applied to some practical scenarios in which unknown classes occur during the inference stage. To solve this problem, we target a more challenging and realistic setting:Incremental Open Set Learning(IOSL), which needs to reject unknown classes from test data while incrementally learning new classes. IOSL has two coupled key challenges: 1) overcoming the catastrophic forgetting of old classes when learning new classes incrementally due to the rarity of old training samples, and 2) minimizing the empirical classification risk on known classes and the open space risk on unknown classes. To address these challenges, we propose an incremental open-set learning method with a “future-look” ability. This ability reserves embedding space for incrementally arriving new classes and potential unknown classes simultaneously to alleviate the catastrophic forgetting indirectly and recognize unknown classes well. Specifically, a normalized prototype learning strategy is designed to minimize the empirical classification risk and implicitly reserve some space. Moreover, we design an extra classes synthesizing module to explicitly reserve more suitable space. This further minimizes the empirical classification risk while reducing the open space risk. Furthermore, we develop an adaptive metric learning loss to mitigate the class imbalance between old and new classes, which focuses on exploiting exemplars fully and selects an adaptive margin for pairs of old and new classes. Extensive experiments on representative classification datasets validate the superiority of our method.
Bingtao Ma, Yang Cong
IEEE Trans. Circuits Syst. Video Technol.2
2024 Gradient-Semantic Compensation for Incremental Semantic Segmentation
abstract
Incremental semantic segmentation focuses on continually learning the segmentation of new coming classes without obtaining the training data from previously seen classes. However, most current methods fail to tackle catastrophic forgetting and background shift since they 1) treat all previous classes equally without considering different forgetting paces caused by imbalanced gradient back-propagation; 2) lack strong semantic guidance between classes. In this paper, to solve the aforementioned challenges, we propose aGradient-SemanticCompensation (GSC) model, which surmounts incremental semantic segmentation from both gradient and semantic perspectives. Specifically, to handle catastrophic forgetting from the gradient aspect, we develop a step-aware gradient compensation that can balance forgetting paces of previously seen classes by re-weighting gradient back-propagation. Meanwhile, we propose a soft-sharp semantic relation distillation to distill consistent inter-class semantic relations via soft labels for alleviating catastrophic forgetting from the semantic aspect. In addition, we design a prototypical pseudo re-labeling which provides strong semantic guidance to mitigate background shift. It produces high-quality pseudo labels for background pixels belonging to previous classes by assessing distances of pixels relative to class-wise prototypes. Experiments on three public segmentation datasets provide strong evidence for the effectiveness of our proposed GSC model.
Wei Cong, Yang Cong, Jiahua Dong 0001, Gan Sun, Henghui Ding
IEEE Trans. Multim.2
2024 Open-Ended Online Learning for Autonomous Visual Perception
abstract
The visual perception systems aim to autonomously collect consecutive visual data and perceive the relevant information online like human beings. In comparison with the classical static visual systems focusing on fixed tasks (e.g., face recognition for visual surveillance), the real-world visual systems (e.g., the robot visual system) often need to handle unpredicted tasks and dynamically changed environments, which need to imitate human-like intelligence with open-ended online learning ability. Therefore, we provide a comprehensive analysis of open-ended online learning problems for autonomous visual perception in this survey. Based on "what to online learn" among visual perception scenarios, we classify the open-ended online learning methods into five categories: instance incremental learning to handle data attributes changing, feature evolution learning for incremental and decremental features with the feature dimension changed dynamically, class incremental learning and task incremental learning aiming at online adding new coming classes/tasks, and parallel and distributed learning for large-scale data to reveal the computational and storage advantages. We discuss the characteristic of each method and introduce several representative works as well. Finally, we introduce some representative visual perception applications to show the enhanced performance when using various open-ended online learning models, followed by a discussion of several future directions.
Yang Cong, Gan Sun, Dongdong Hou, Jiahua Dong 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Topology-Aware Graph Convolution Network for Few-Shot Incremental 3-D Object Learning
abstract
Three-dimensional (3-D) object recognition has achieved satisfied achievement in both academia and industry. However, most traditional 3-D object classification methods implicitly assume that there are abundant training data from a static distribution. To relax the assumption, we target on a more challenging and realistic setting: few-shot incremental 3-D object learning (FSI3DL), which intends to incrementally classify the new coming 3-D objects with few training data. In order to achieve this, two key challenges need to be concerned: 1) the catastrophic forgetting issue caused by incremental 3-D data with irregular and redundant topological structures and 2) the overfitting issue caused by few-shot training data. To address the first challenge, we use Laplacian spectral analysis based on 3-D meshes to design an embedding network that consists of super-vertex graph convolution (SVGC) module and topology-aware graph attention (TAGA) module. The SVGC is designed to construct the discriminative local topological characteristics for representing the irregular 3-D meshes better. The TAGA is designed to identify redundant topological characteristics. To address the second challenge, a fine-tuning strategy with model alignment regularization is investigated. Furthermore, an embedding space selection and fusion (ESSF) strategy is proposed in the inference phase to mitigate catastrophic forgetting and overfitting further. Combining SVGC, TAGA, and alignment regularization with ESSF strategy, a novel topology-aware graph convolution network (TopGCN) is proposed to address the FSI3DL. Experiments on representative 3-D classification datasets validate the superiority of TopGCN.
Bingtao Ma, Yang Cong, Jiahua Dong 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2023 Federated Incremental Semantic Segmentation
abstract
Federated learning-based semantic segmentation (FSS) has drawn widespread attention via decentralized training on local clients. However, most FSS models assume categories are fixed in advance, thus heavily undergoing forgetting on old categories in practical applications where local clients receive new categories incrementally while have no memory storage to access old classes. Moreover, new clients collecting novel classes may join in the global training of FSS, which further exacerbates catastrophic forgetting. To surmount the above challenges, we propose a Forgetting-Balanced Learning (FBL) model to address heterogeneous forgetting on old classes from both intra-client and interclient aspects. Specifically, under the guidance of pseudo labels generated via adaptive class-balanced pseudo labeling, we develop a forgetting-balanced semantic compensation loss and a forgetting-balanced relation consistency loss to rectify intra-client heterogeneous forgetting of old categories with background shift. It performs balanced gradient propagation and relation consistency distillation within local clients. Moreover, to tackle heterogeneous forgetting from inter-client aspect, we propose a task transition monitor. It can identify new classes under privacy protection and store the latest old global model for relation distillation. Qualitative experiments reveal large improvement of our model against comparison methods. The code is available at https://github.com/JiahuaDong/FISS.
Jiahua Dong 0001, Duzhen Zhang, Yang Cong, Wei Cong, Henghui Ding, Dengxin Dai
CVPR3
2023 Autonomous Manipulation Learning for Similar Deformable Objects via Only One Demonstration
abstract
In comparison with most methods focusing on$3D$rigid object recognition and manipulation, deformable objects are more common in our real life but attract less attention. Generally, most existing methods for deformable object manipulation suffer two issues, 1) Massive demonstration: repeating thousands of robot-object demonstrations for model training of one specific instance; 2) Poor generalization: inevitably re-training for transferring the learned skill to a similar/new instance from the same category. Therefore, we propose a category-level deformable$3D$object manipulation framework, which could manipulate deformable$3D$objects with only one demonstration and generalize the learned skills to new similar instances without re-training. Specifically, our proposed framework consists of two modules. The Nocs State Transform$(NST)$module transfers the observed point clouds of the target to a pre-defined unified pose state (i.e.,Nocs state), which is the foundation for the category-level manipulation learning; the Neural Spatial Encoding$(NSE)$module generalizes the learned skill to novel instances by encoding the category-level spatial information to pursue the expected grasping point without re-training. The relative motion path is then planned to achieve autonomous manipulation. Both the simulated results via our$\text{Cap}_{40}$dataset and real robotic experiments justify the effectiveness of our framework.
Ronghan Chen, Yang Cong
CVPR3
2023 Heterogeneous Forgetting Compensation for Class-Incremental Learning
abstract
Class-incremental learning (CIL) has achieved remarkable successes in learning new classes consecutively while overcoming catastrophic forgetting on old categories. However, most existing CIL methods unreasonably assume that all old categories have the same forgetting pace, and neglect negative influence of forgetting heterogeneity among different old classes on forgetting compensation. To surmount the above challenges, we develop a novel Heterogeneous Forgetting Compensation (HFC) model, which can resolve heterogeneous forgetting of easy-to-forget and hard-to-forget old categories from both representation and gradient aspects. Specifically, we design a task-semantic aggregation block to alleviate heterogeneous forgetting from representation aspect. It aggregates local category information within each task to learn task-shared global representations. Moreover, we develop two novel plug-and-play losses: a gradient-balanced forgetting compensation loss and a gradient-balanced relation distillation loss to alleviate forgetting from gradient aspect. They consider gradient-balanced compensation to rectify forgetting heterogeneity of old categories and heterogeneous relation consistency. Experiments on several representative datasets illustrate effectiveness of our HFC model. The code is available at https://github.com/JiahuaDong/HFC.
Jiahua Dong 0001, Wenqi Liang, Yang Cong, Gan Sun
ICCV3
2023 Augmented Box Replay: Overcoming Foreground Shift for Incremental Object Detection
abstract
In incremental learning, replaying stored samples from previous tasks together with current task samples is one of the most efficient approaches to address catastrophic forgetting. However, unlike incremental classification, image replay has not been successfully applied to incremental object detection (IOD). In this paper, we identify the overlooked problem of foreground shift as the main reason for this. Foreground shift only occurs when replaying images of previous tasks and refers to the fact that their background might contain foreground objects of the current task. To overcome this problem, a novel and efficient Augmented Box Replay (ABR) method is developed that only stores and replays foreground objects and thereby circumvents the foreground shift problem. In addition, we propose an innovative Attentive RoI Distillation loss that uses spatial attention from region-of-interest (RoI) features to constrain current model to focus on the most important information from old model. ABR significantly reduces forgetting of previous classes while maintaining high plasticity in current classes. Moreover, it considerably reduces the storage requirements when compared to standard image replay. Comprehensive experiments on Pascal-VOC and COCO datasets support the state-of-the-art performance of our model1.
Yang Cong, Dipam Goswami, Xialei Liu, Joost van de Weijer 0001
ICCV2
2023 Angular Penalty for Few-Shot Incremental 3D Object Learning
abstract
3D object recognition has garnered notable success in both academic and industrial contexts. However, the majority of existing 3D object recognition approaches are tailor-made for static scenarios that have ample availability of training instances. Therefore, we target a more arduous and pragmatic task: few-shot incremental 3D object learning (FSI3DL), which aims to learn the new 3D objects in an incremental manner, yet new classes only have a few training instances. However, this task presents two challenges: overfitting to few-shot, and catastrophic forgetting on previous classes due to the absence of previous training instances. To mitigate these challenges, we introduce an Angular Penalty method to consume irregular mesh directly and reserve embedding space for incoming new classes. Specifically, we design a graph convolution network that can take advantage of the better representation ability of mesh and overcome its irregularity. Moreover, we use an angular penalty loss to increase the similarity between inra-class instances and reduce the similarity between instances from different classes. This can leave space in embedding space for incoming new classes. Experiments on representative 3D object classification datasets demonstrate the better efficacy of our method.
Bingtao Ma, Yang Cong
IJCNN2
2023 Hierarchical Lifelong Machine Learning With "Watchdog"
abstract
Most existing lifelong machine learning works focus on how to exploit previously accumulated experiences (e.g., knowledge library) from earlier tasks, and transfer it to learn a new task. However, when a lifelong learning system encounters a large pool of candidate tasks, the knowledge among various coming tasks are imbalance, and the system should intelligently choose the next one to learn. In this paper, an effective “human cognition” strategy is taken into consideration via actively sorting the importance of new tasks in the process of unknown-to-known, and preferentially selecting the most valuable task with more information to learn. To be specific, we assess the importance of each new coming task (e.g., unknown or not) as an outlier detection issue, and propose to employ a “watchdog” knowledge library to reconstruct each task under$\ell _0$-norm constraint. The coming candidate tasks are then sorted depending on the sparse reconstruction scores in a descending order, which is referred to as a “watchdog” mechanism. Following this, we design a hierarchical knowledge library for the lifelong learning framework to encode new task with higher reconstruction score, where the library consists of two-level task descriptors, i.e., a high-dimensional one with low-rank constraint and a low-dimensional one. Both “watchdog” knowledge library and hierarchy knowledge library can be optimized with knowledge from both previously learned tasks and current task automatically. For model optimization, we explore an alternating method to iteratively update our proposed framework with a guaranteed convergence. Experimental results on several existing benchmarks demonstrate that our proposed model outperforms various state-of-the-art task selection methods.
Gan Sun, Yang Cong, Changjun Gu, Zhengming Ding
IEEE Trans. Big Data2
2023 Lifelong Visual-Tactile Spectral Clustering for Robotic Object Perception
abstract
This work presents a novel visual-tactile fused clustering framework, calledLifelongVisual-TactileSpectralClustering (i.e., LVTSC), to effectively learn consecutive object clustering tasks for robotic perception. Lifelong learning has become an important and hot topic in recent studies on machine learning, aiming to imitate “human learning” and reduce the computational cost when consecutively learning new tasks. Our proposed LVTSC model explores the knowledge transfer and representation correlation from a local modality-invariant perspective under modality-consistent constraint guidance. For the modality-invariant part, we design a set of modality-invariant basis libraries to capture the latent clustering centers of each modality and a set of modality-invariant feature libraries to forcibly embed the manifold information of each modality. A modal-consistent constraint reinforces the correlation between visual and tactile modalities by maximizing the feature manifold correspondences. When the object clustering task comes continuously, the overall objective is optimized by an effective alternating direction method with guaranteed convergence. Our proposed LVTSC framework has been extensively validated for its effectiveness and efficiency on the three challenging real-world robotic object perception datasets.
Yang Cong, Gan Sun, Zhengming Ding
IEEE Trans. Circuits Syst. Video Technol.2
2023 Uni3DA: Universal 3D Domain Adaptation for Object Recognition
abstract
Traditional 3D point cloud classification tasks focus on training a classifier in the closed-set scenario, where training and test data have the same label set and the same data distribution. In this work, we focus on a more challenging and realistic scenario in 3D point cloud classification task: universal domain adaptation (UniDA), where 1) data distributions for training and test data are different; and 2) for given label sets of training data and test data, they may have the shared classes and keep the private classes respectively, introducing an extra label set discrepancy. To solve UniDA problem, researchers have designed many methods based on 2D image datasets. However, due to the difficulty in capturing discriminative local geometric structures brought by the unordered and irregular 3D point cloud data, we cannot directly deploy the existing methods based on 2D image datasets to the 3D scenarios. To address UniDA in 3D scenarios, we develop a 3D universal domain adaptation framework, which consists of three modules: Self-Constructed Geometric (SCG) module, Local-to-Global Hypersphere Reasoning (LGHR) module and Self-Supervised Boundary Adaptation (SBA) module. SCG and LGHR generate the discriminative representation, which is used to acquire domain-invariant knowledge for training and test data. SBA is designed to automatically recognize whether a given label is from the shared label set or private label set, and adapts training and test data from the shared label set. To our best knowledge, this work is the first exploration of UniDA for 3D scenarios. Extensive experiments on public 3D point cloud datasets verify that the proposed method outperforms the existing UniDA methods.
Yang Cong, Jiahua Dong 0001, Gan Sun
IEEE Trans. Circuits Syst. Video Technol.2
2023 A Comprehensive Study of 3-D Vision-Based Robot Manipulation
abstract
Robot manipulation, for example, pick-and-place manipulation, is broadly used for intelligent manufacturing with industrial robots, ocean engineering with underwater robots, service robots, or even healthcare with medical robots. Most traditional robot manipulations adopt 2-D vision systems with plane hypotheses and can only generate 3-DOF (degrees of freedom) pose accordingly. To mimic human intelligence and endow the robot with more flexible working capabilities, 3-D vision-based robot manipulation has been studied. However, this task is still challenging in the open world especially for general object recognition and pose estimation with occlusion in cluttered backgrounds and human-like flexible manipulation. In this article, we propose a comprehensive analysis of recent progress about the 3-D vision for robot manipulation, including 3-D data acquisition and representation, robot-vision calibration, 3-D object detection/recognition, 6-DOF pose estimation, grasping estimation, and motion planning. We then present some public datasets, evaluation criteria, comparisons, and challenges. Finally, the related application domains of robot manipulation are given, and some future directions and open problems are studied as well.
Yang Cong, Ronghan Chen, Bingtao Ma, Hongsen Liu, Dongdong Hou, Chenguang Yang 0001
IEEE Trans. Cybern.1
2023 InOR-Net: Incremental 3-D Object Recognition Network for Point Cloud Representation
abstract
3-D object recognition has successfully become an appealing research topic in the real world. However, most existing recognition models unreasonably assume that the categories of 3-D objects cannot change over time in the real world. This unrealistic assumption may result in significant performance degradation for them to learn new classes of 3-D objects consecutively due to the catastrophic forgetting on old learned classes. Moreover, they cannot explore which 3-D geometric characteristics are essential to alleviate the catastrophic forgetting on old classes of 3-D objects. To tackle the above challenges, we develop a novel Incremental 3-D Object Recognition Network (i.e., InOR-Net), which could recognize new classes of 3-D objects continuously by overcoming the catastrophic forgetting on old classes. Specifically, category-guided geometric reasoning is proposed to reason local geometric structures with distinctive 3-D characteristics of each class by leveraging intrinsic category information. We then propose a novel critic-induced geometric attention mechanism to distinguish which 3-D geometric characteristics within each class are beneficial to overcome the catastrophic forgetting on old classes of 3-D objects while preventing the negative influence of useless 3-D characteristics. In addition, a dual adaptive fairness compensations' strategy is designed to overcome the forgetting brought by class imbalance by compensating biased weights and predictions of the classifier. Comparison experiments verify the state-of-the-art performance of the proposed InOR-Net model on several public point cloud datasets.
Jiahua Dong 0001, Yang Cong, Gan Sun, Lixu Wang, Lingjuan Lyu, Jun Li 0027, Ender Konukoglu
IEEE Trans. Neural Networks Learn. Syst.2
2022 The Devil is in the Pose: Ambiguity-free 3D Rotation-invariant Learning via Pose-aware Convolution
abstract
Recent progress in introducing rotation invariance (RI) to 3D deep learning methods is mainly made by designing RI features to replace 3D coordinates as input. The key to this strategy lies in how to restore the global information that is lost by the input RI features. Most state-of-the-arts achieve this by incurring additional blocks or complex global representations, which is time-consuming and ineffective. In this paper, we real that the global information loss stems from an unexplored pose information loss problem, i.e., common convolution layers cannot capture the relative poses between RI features, thus hindering the global information to be hierarchically aggregated in the deep networks. To address this problem, we develop a Poseaware Rotation Invariant Convolution (i.e., PaRI-Conv), which dynamically adapts its kernels based on the relative poses. Specifically, in each PaRI-Conv layer, a lightweight Augmented Point Pair Feature (APPF) is designed to fully encode the RI relative pose information. Then, we propose to synthesize a factorized dynamic kernel, which reduces the computational cost and memory burden by decomposing it into a shared basis matrix and a pose-aware diagonal matrix that can be learned from the APPF. Extensive experiments on shape classification and part segmentation tasks show that our PaRI-Conv surpasses the state-of-the-art RI methods while being more compact and efficient.
Ronghan Chen, Yang Cong
CVPR2
2022 The phase-only null beamforming synthesis via manifold optimization
abstract
The phase-only beamforming synthesis is widely applied in millimeter wave communication, radar and sonar. Due to the CMC, the problem is non-convex. The most current methods solve the problem by designing the phase, which either degrades the performance or needs huge complexity. To address this issue, a low-complexity Riemannian Manifold Optimization based Conjugate Gradient (RMOCG) method is proposed. First, the original problem is transformed into an unconstrained prob-lem on a complex circle manifold. Then, a RMOCG algorithm is derived, by deriving the gradient descent direction and the step size for ensuring the cost function non-increasing. Comparing with the existing methods, the proposed method has the following advantages: 1) the null depth is respectively 8 dB deeper than [6] and 3 dB deeper than [12]. 2) The computational cost is 2 magnitude lower than [6] and 1 magnitude lower than [12].
Yang Cong, Jinfeng Hu, Kai Zhong 0002, Jie Wu 0044
IGARSS1
2022 Constant Modulus Waveform Design for Integrated Sensing and Communication Systems
abstract
The constant modulus (CM) waveform design for integrated sensing and communication (ISAC) systems is a key technology. The joint design of maximizing the Signal-to-Interference-and-Noise-Ratio (SINR) for radar and minimizing the Multiple User Interference (MUI) for communication is studied. The problem is nonconvex and NP-hard, due to the fractional expression and the CM constraint. To address this issue, a low-complexity Accelerated Coordinate Descent (ACD) method is proposed. First, the problem is simplified to a quadratic function with CMC by dinkelbatchs method. Then, a CD method is derived, by transforming the problem into a decomposable problem with multiple one-dimensional subproblems. Finally, the ACD algorithm is derived to accelerate the convergence by using the square iterative technique. Simulation results show that the proposed method obtains favorable trade-off performance between SINR and MUI.
Kai Zhong 0002, Jinfeng Hu, Yang Cong, Jie Wu 0044, Yaya Pei
IGARSS3
2022 Class-Incremental Gesture Recognition Learning with Out-of-Distribution Detection
abstract
Gesture recognition is a popular human-computer interaction technology, which has been widely applied in many fields (e.g., autonomous driving, medical care, VR and AR). However, 1) most existing gesture recognition methods focus on the fixed recognition scenarios with several gestures, which could lead to memory consumption and computational effort when continuously learning new gestures; 2) Meanwhile, the performance of popular class-incremental methods degrades significantly for previously learned classes (i.e., catastrophic forgetting) due to the ambiguity and variability of gestures. To tackle these challenges, we propose a novel class-incremental gesture recognition method with out-of-distribution (OOD) detection, which can continuously adapt to new gesture classes and achieve high performance for both learned and new gestures. Specifically, we construct an episodic memory with a subset of learned training samples to preserve the previous knowledge from forgetting. Moreover, the OOD detection-based memory management is developed for exploring the most representative and informative core set from the learned datasets. When a new gesture recognition task with strange classes comes, rehearsal enhancement is adopted to increase the diversity of memory exemplars for better fitting the real characteristics of gesture recognition. After deriving an effective class-incremental gesture recognition strategy, we perform experiments on two representative datasets to validate the superiority of our method. Evaluation experiments demonstrate that our proposed method substantially outperforms the state-of-the-art methods with about 2.17%-3.81% improvement under different class-incremental learning scenarios.
Mingxue Li, Yang Cong, Gan Sun
IROS2
2022 Data Poisoning Attacks on Federated Machine Learning
abstract
Federated machine learning which enables resource-constrained node devices (e.g., Internet of Things (IoT) devices and smartphones) to establish a knowledge-shared model while keeping the raw data local, could provide privacy preservation, and economic benefit by designing an effective communication protocol. However, this communication protocol can be adopted by attackers to launch data poisoning attacks for different nodes, which has been shown as a big threat to most machine learning models. Therefore, we in this article intend to study the model vulnerability of federated machine learning, and even on IoT systems. To be specific, we here attempt to attacking a popular federated multitask learning framework, which uses a general multitask learning framework to handle statistical challenges in the federated learning setting. The problem of calculating optimal poisoning attacks on federated multitask learning is formulated as a bilevel program, which is adaptive to the arbitrary selection oftargetnodes andsource attackingnodes. We then propose a novel systems-aware optimization method, called as attack on federated learning (AT2FL), to efficiently derive the implicit gradients for poisoned data, and further attain optimal attack strategies in the federated machine learning. This is an earlier work, to our knowledge, that explores attacking federated machine learning via data poisoning. Finally, experiments on several real-world data sets demonstrate that when the attackers directly poison thetargetnodes or indirectly poison the related nodes via using the communication protocol, the federated multitask learning model is sensitive to both poisoning attacks.
Gan Sun, Yang Cong, Jiahua Dong 0001, Qiang Wang 0015, Lingjuan Lyu, Ji Liu 0002
IEEE Internet Things J.2
2022 What and How: Generalized Lifelong Spectral Clustering via Dual Memory
abstract
Spectral clustering (SC) has become one of the most widely-adopted clustering algorithms, and been successfully applied into various applications. We in this work explore the problem of spectral clustering in a lifelong learning framework termed asGeneralizedLifelongSpectralClustering (GL$^2$SC). Different from most current studies, which concentrate on a fixed spectral clustering task set and cannot efficiently incorporate a new clustering task, the goal of our work is to establish a generalized model for new spectral clustering tasks by “What” and “How” to lifelong learn from past tasks. In respect of “what to lifelong learn”, our GL$^2$SC framework contains a dual memory mechanism with a deep orthogonal factorization manner: an orthogonal basis memory stores hidden and hierarchical clustering centers among learned tasks, and a feature embedding memory captures deep manifold representation common across multiple related tasks. When learning a new clustering task, the intuition here for “how to lifelong learn” is that GL$^2$SC can transfer intrinsic knowledge from dual memory mechanism to obtain task-specific encoding matrix. Then the encoding matrix can redefine the dual memory over time to provide maximal benefits when learning future tasks, and reversely maximize performance for past tasks. To achieve this, we propose an alternative optimization formulation with convergence guarantee for solving our GL$^2$SC model. To the end, empirical comparisons on several benchmark datasets show the effectiveness of our GL$^2$SC, in comparison with several state-of-the-art clustering models.
Gan Sun, Yang Cong, Jiahua Dong 0001, Zhengming Ding
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Lifelong robotic visual-tactile perception learning
Jiahua Dong 0001, Yang Cong, Gan Sun, Tao Zhang 0084
Pattern Recognit.2
2022 Fast Multi-View Outlier Detection via Deep Encoder
abstract
Multi-view outlier detection has a wide range of applications and has been well investigated in recent years. However, 1) most existing state-of-the-art methods cannot efficiently handle outlier detection problem for large-scale multi-view data, since exploring pairwise constraints among different views causes highly-computational cost; 2) the data collected from original heterogeneous feature spaces further increases the consistent difficulty of multi-view outlier detection. To address these issues, we present a fast multi-view outlier detection model via learning a low-rank latent subspace representation with deep encoder architecture, which can not only efficiently identify the outliers for large-scale data even with numerous data views, but also exploit a discriminative common latent subspace shared by all the views. First, we learn a set of orthogonal bases as view-specific dictionaries from a small dataset, which is randomly sampled from the original dataset. Benefitting from view-specific dictionaries, the sampled data is projected and decomposed as a shared and discriminative latent subspace representations, which correspond to the view-consistent and view-specific components across multiple views, respectively. Then, the obtained discriminative latent representations are applied to train the view-specific deep encoders, which can efficiently compute the abnormal score for the remaining instances. Our proposed model can cost-effectively identify the outliers in large-scale datasets from numerous data views with less computational complexity. Experiments conducted on eight real datasets and a synthesis dataset show that our proposed model outperforms the existing ones on effectiveness and efficiency.
Dongdong Hou, Yang Cong, Gan Sun, Jiahua Dong 0001, Jun Li 0027, Kai Li 0012
IEEE Trans. Big Data2
2022 Evolving Metric Learning for Incremental and Decremental Features
abstract
Online metric learning has been widely exploited for large-scale data classification due to the low computational cost. However, amongst online practical scenarios where the features are evolving (e.g., some features are vanished and some new features are augmented), most metric learning models cannot be successfully applied to these scenarios, although they can tackle the evolving instances efficiently. To address the challenge, we develop a new online Evolving Metric Learning (EML) model for incremental and decremental features, which can handle the instance and feature evolutions simultaneously by incorporating with a smoothed Wasserstein metric distance. Specifically, our model contains two essential stages: a Transforming stage (T-stage) and a Inheriting stage (I-stage). For the T-stage, we propose to extract important information from vanished features while neglecting non-informative knowledge, and forward it into survived features by transforming them into a low-rank discriminative metric space. It further explores the intrinsic low-rank structure of heterogeneous samples to reduce the computation and memory burden especially for highly-dimensional large-scale data. For the I-stage, we inherit the metric performance of survived features from the T-stage and then expand to include the new augmented features. Moreover, a smoothed Wasserstein distance is utilized to characterize the similarity relationships among the heterogeneous and complex samples, since the evolving features are not strictly aligned in the different stages. In addition to tackling the challenges in one-shot case, we also extend our model into multi-shot scenario. After deriving an efficient optimization strategy for both T-stage and I-stage, extensive experiments on several datasets verify the superior performance of our EML model.
Jiahua Dong 0001, Yang Cong, Gan Sun, Tao Zhang 0084, Xiaowei Xu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Visual-Tactile Fused Graph Learning for Object Clustering
abstract
highlights how to mitigate the differences between vision and touch, and further maximize the mutual information, which adopts a minimizing disagreement scheme to guide the modality-specific representations toward a unified affinity graph. To achieve ideal clustering performance, a Laplacian rank constraint is imposed to regularize the learned graph with ideal connected components, where noises that caused wrong connections are removed and clustering labels can be obtained directly. Finally, we propose an efficient alternating iterative minimization updating strategy, followed by a theoretical proof to prove framework convergence. Comprehensive experiments on five public datasets demonstrate the superiority of the proposed framework.
Tao Zhang 0084, Yang Cong, Gan Sun, Jiahua Dong 0001
IEEE Trans. Cybern.2
2022 Dual Aligned Siamese Dense Regression Tracker
abstract
Anchor or anchor-free based Siamese trackers have achieved the astonishing advancement. However, their parallel regression and classification branches lack the tracked target information link and interaction, and the corresponding independent optimization maybe lead to task-misalignment, such as the reliable classification prediction with imprecisely localization and vice versa. To address this problem, we develop a general Siamese dense regression tracker (SDRT) with both task and feature alignments. It consists of two cooperative and mutual-guidance core branches: dense local regression with RepPoint representation, the global and local multi-classifier fusion with aligned features. They complement and boost each other to constrain the results with well-localized followed to also be well-classified. Specifically, a dense local regression with RepPoint representation, directly estimates and averages multiple dense local bounding box offsets for accurate localization. And then, the refined bounding boxes can be used to learn the global and local affine alignment features for reliable multi-classifier fusion. The classified scores in turn guide the assigned positive bounding boxes for the regression task. The mutual guidance operations can bridge the connection between classification and regression substantially, since the assigned labels of one task depend on the prediction quality of the other task. The proposed tracking module is general, and it can boost both the anchor or anchor-free based Siamese trackers to some extent. The extensive tracking comparisons on six tracking benchmarks verify its favorable and competitive performance over states-of-the-arts tracking modules.
Baojie Fan, Hui Zhang 0023, Yang Cong, Yandong Tang, Huijie Fan, Jiandong Tian
IEEE Trans. Image Process.3
2022 Representative Task Self-Selection for Flexible Clustered Lifelong Learning
abstract
Consider the lifelong machine learning paradigm whose objective is to learn a sequence of tasks depending on previous experiences, e.g., knowledge library or deep network weights. However, the knowledge libraries or deep networks for most recent lifelong learning models are of prescribed size and can degenerate the performance for both learned tasks and coming ones when facing with a new task environment (cluster). To address this challenge, we propose a novel incremental clustered lifelong learning framework with two knowledge libraries: feature learning library and model knowledge library, called Flexible Clustered Lifelong Learning (FCL3). Specifically, the feature learning library modeled by an autoencoder architecture maintains a set of representation common across all the observed tasks, and the model knowledge library can be self-selected by identifying and adding new representative models (clusters). When a new task arrives, our FCL3 model firstly transfers knowledge from these libraries to encode the new task, i.e., effectively and selectively soft-assigning this new task to multiple representative models over feature learning library. Then: 1) the new task with a higher outlier probability will be judged as a new representative, and used to redefine both feature learning library and representative models over time; or 2) the new task with lower outlier probability will only refine the feature learning library. For model optimization, we cast this lifelong learning problem as an alternating direction minimization problem as a new task comes. Finally, we evaluate the proposed framework by analyzing several multitask data sets, and the experimental results demonstrate that our FCL3 model can achieve better performance than most lifelong learning frameworks, even batch clustered multitask learning models.
Gan Sun, Yang Cong, Qianqian Wang 0001, Bineng Zhong 0001, Yun Fu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2022 MedUCC: Medium-Driven Underwater Camera Calibration for Refractive 3-D Reconstruction
abstract
Underwater camera calibration has attracted much attentions due to its significance in high-precision three-dimensional (3-D) pose estimation and scene reconstruction. However, most existing calibration methods focus on calibrating the underwater camera in a single scenario [e.g., air-glass-water], which can not well formulate the geometry constraint and further result in the complex calibration process. Moreover, the calibration precision of these methods is low, since multilayer transparent refractions with unknown layer orientation and distance make the task more difficult than that in air. To address these challenges, we develop a novel and efficient medium-driven method for underwater camera calibration (MedUCC), which can calibrate the underwater camera parameters, including the orientation and position of the transparent glass accurately. Our key idea of this article is to leverage the light-path changes formed by medium refractions between different media to acquire calibration data, which can better formulate the geometry constraint, and estimate the initial value of the underwater camera parameters. To improve the calibration accuracy of the underwater camera system, a quaternion-based solution is developed to refine the underwater camera parameters. To the end, we evaluate the calibration performance on an underwater camera system. Extensive experiment results demonstrate that our proposed method can obtain a better performance in comparison to the existing works. We also validate our proposed MedUCC method on our designed 3-D scanner prototype, which illustrates the superiority of our proposed calibration method.
Changjun Gu, Yang Cong, Gan Sun, Yajun Gao, Tao Zhang 0084, Baojie Fan
IEEE Trans. Syst. Man Cybern. Syst.2
2022 Partial Visual-Tactile Fused Learning for Robotic Object Recognition
abstract
Currently, visual-tactile fusion learning for robotic object recognition has achieved appealing performance, due to the fact that visual and tactile data can offer complementary information. However: 1) the distinct gap between vision and touch makes it difficult to fully explore the complementary information, which would further lead to performance degradation and 2) most of the existing visual-tactile fused learning methods assume that visual and tactile data are complete, which is often difficult to be satisfied in many real-world applications. In this article, we propose a partial visual-tactile fused (PVTF) framework for robotic object recognition to address these challenges. Specifically, we first employ two modality-specific (MS) encoders to encode partial visual-tactile data into two incomplete subspaces (i.e., visual subspace and tactile subspace). Then, a modality gap mitigated (MGM) network is adopted to discover modality-invariant high-level label information, which is utilized to generate gap loss and further help updating the MS encoders for relatively consistent visual and tactile subspaces generation. In this way, the huge gap between vision and touch is mitigated, which would further contribute to mine the complementary visual-tactile information. Finally, to achieve data completeness and complementary visual-tactile information exploration simultaneously, a cycle subspace leaning technique is proposed to project the incomplete subspaces into a complete subspace by fully exploiting all the obtainable samples, where complete latent representations with maximum complementary information can be learned. A lot of comparative experiments conducted on three visual-tactile datasets validate the advantage of the proposed PVTF framework, by comparing with state-of-the-art baselines.
Tao Zhang 0084, Yang Cong, Jiahua Dong 0001, Dongdong Hou
IEEE Trans. Syst. Man Cybern. Syst.2
2021 I3DOL: Incremental 3D Object Learning without Catastrophic Forgetting
abstract
3D object classification has attracted appealing attentions in academic researches and industrial applications. However, most existing methods need to access the training data of past 3D object classes when facing the common real-world scenario: new classes of 3D objects arrive in a sequence. Moreover, the performance of advanced approaches degrades dramatically for past learned classes (i.e., catastrophic forgetting), due to the irregular and redundant geometric structures of 3D point cloud data. To address these challenges, we propose a new Incremental 3D Object Learning (i.e., I3DOL) model, which is the first exploration to learn new classes of 3D object continually. Specifically, an adaptive-geometric centroid module is designed to construct discriminative local geometric structures, which can better characterize the irregular point cloud representation for 3D object. Afterwards, to prevent the catastrophic forgetting brought by redundant geometric information, a geometric-aware attention mechanism is developed to quantify the contributions of local geometric structures, and explore unique 3D geometric characteristics with high contributions for classes incremental learning. Meanwhile, a score fairness compensation strategy is proposed to further alleviate the catastrophic forgetting caused by unbalanced data between past and new classes of 3D object, by compensating biased prediction for new classes in the validation phase. Experiments on 3D representative datasets validate the superiority of our I3DOL framework.
Jiahua Dong 0001, Yang Cong, Gan Sun, Bingtao Ma, Lichen Wang
AAAI2
2021 Generative Partial Visual-Tactile Fused Object Clustering
abstract
Visual-tactile fused sensing for object clustering has achieved significant progresses recently, since the involvement of tactile modality can effectively improve clustering performance. However, the missing data (i.e., partial data) issues always happen due to occlusion and noises during the data collecting process. This issue is not well solved by most existing partial multi-view clustering methods for the heterogeneous modality challenge. Naively employing these methods would inevitably induce a negative effect and further hurt the performance. To solve the mentioned challenges, we propose a Generative Partial Visual-Tactile Fused (i.e., GPVTF) framework for object clustering. More specifically, we first do partial visual and tactile features extraction from the partial visual and tactile data, respectively, and encode the extracted features in modality-specific feature subspaces. A conditional cross-modal clustering generative adversarial network is then developed to synthesize one modality conditioning on the other modality, which can compensate missing samples and align the visual and tactile modalities naturally by adversarial learning. To the end, two pseudo-label based KL-divergence losses are employed to update the corresponding modality-specific encoders. Extensive comparative experiments on three public visual-tactile datasets prove the effectiveness of our method.
Tao Zhang 0084, Yang Cong, Gan Sun, Jiahua Dong 0001, Zhengming Ding
AAAI2
2021 Unsupervised Dense Deformation Embedding Network for Template-Free Shape Correspondence
abstract
Shape correspondence from 3D deformation learning has attracted appealing academy interests recently. Nevertheless, current deep learning based methods require the supervision of dense annotations to learn per-point translations, which severely over-parameterize the deformation process. Moreover, they fail to capture local geometric details of original shape via global feature embedding. To address these challenges, we develop a new Unsupervised Dense Deformation Embedding Network (i.e., UD2E-Net), which learns to predict deformations between non-rigid shapes from dense local features. Since it is non-trivial to match deformation-variant local features for deformation prediction, we develop an Extrinsic-Intrinsic Autoencoder to first encode extrinsic geometric features from source into intrinsic coordinates in a shared canonical shape, with which the decoder then synthesizes corresponding target features. Moreover, a bounded maximum mean discrepancy loss is developed to mitigate the distribution divergence between the synthesized and original features. To learn natural deformation without dense supervision, we introduce a coarse parameterized deformation graph, for which a novel trace and propagation algorithm is proposed to improve both the quality and efficiency of the deformation. Our UD2E-Net outperforms state-of-the-art unsupervised methods by 24% on Faust Inter challenge and even supervised methods by 13% on Faust Intra challenge.
Ronghan Chen, Yang Cong, Jiahua Dong 0001
ICCV2
2021 Dynamic and reliable subtask tracker with general schatten p-norm regularization
Baojie Fan, Yang Cong, Jiandong Tian, Yandong Tang
Pattern Recognit.2
2021 Weakly-Supervised Cross-Domain Adaptation for Endoscopic Lesions Segmentation
abstract
Weakly-supervised learning has attracted growing research attention on medical lesions segmentation due to significant saving in pixel-level annotation cost. However, 1) most existing methods require effective prior and constraints to explore the intrinsic lesions characterization, which only generates incorrect and rough prediction; 2) they neglect the underlying semantic dependencies among weakly-labeled target enteroscopy diseases and fully-annotated source gastroscope lesions, while forcefully utilizing untransferable dependencies leads to the negative performance. To tackle above issues, we propose a new weakly-supervised lesions transfer framework, which can not only explore transferable domain-invariant knowledge across different datasets, but also prevent the negative transfer of untransferable representations. Specifically, a Wasserstein quantified transferability framework is developed to highlight wide-range transferable contextual dependencies, while neglecting the irrelevant semantic characterizations. Moreover, a novel self-supervised pseudo label generator is designed to equally provide confident pseudo pixel labels for both hard-to-transfer and easy-to-transfer target samples. It inhibits the enormous deviation of false pseudo pixel labels under the self-supervision manner. Afterwards, dynamically-searched feature centroids are aligned to narrow category-wise distribution shift. Comprehensive theoretical analysis and experiments show the superiority of our model on the endoscopic dataset and several public datasets.
Jiahua Dong 0001, Yang Cong, Gan Sun, Yunsheng Yang, Xiaowei Xu 0001, Zhengming Ding
IEEE Trans. Circuits Syst. Video Technol.2
2021 Structured and Consistent Multi-Layer Multi-Kernel Subtask Correction Filter Tracker
abstract
Some multi-task correlation filter trackers achieve the top-ranked performance in terms of accuracy and robustness. However, they directly fuse multiple types of features into a single kernel space. This operation fails to fully explore the discriminative strength and diversity of different features, and also ignores the structured correspondence of different tasks. To solve these issues, we propose a structured multi-kernel subtask correlation filter tracker with temporal-spatial consistency, which enjoys the merits of both layered multi-kernel subtask learning and structured correlation filter. Specifically, we firstly assign one kernel space to each channel feature. Multi-channel features correspond to multi-kernel spaces to boost their powerful discriminability. And then, we divide the target into multi-layer patches with different sizes, and regard the correlation filter trace of each patch with one channel feature as a subtask. In the following, we incorporate globally and locally structured correlation filters into a unified multi-kernel subtask particle tracking framework. The global and local subtasks complement and enhance each other with similar motion model. The proposed tracker not only exploits the cooperation and complementarity of layered multi-kernel subtask correlation filters, but also mines the underlying geometric structure of global subtasks, and the inner spatial locality correspondences of local subtasks inside the target. This operation is achieved by dual group sparsity regularized terms with mixed-norm lp,q, which decomposes the multi-kernel subtask filter matrix into two collaborative components. They correspond to the adaptive filter feature selection and outlier subtask detection, respectively. Besides, the developed tracking model maintains the temporal coherence and spatial consistency of multi-layer subtask filters via the smooth regularizer. Finally, the tracking formulation is optimized by the accelerated proximal gradient approach (APG). Encouraging analyses on six benchmark datasets, verify the favorable effectiveness and robustness of our method against state-of-the-art trackers.
Baojie Fan, Yang Cong, Yandong Tang, Jiandong Tian, Chenliang Xu
IEEE Trans. Circuits Syst. Video Technol.2
2021 L3DOC: Lifelong 3D Object Classification
abstract
3D object classification has been widely applied in both academic and industrial scenarios. However, most state-of-the-art algorithms rely on a fixed object classification task set, which cannot tackle the scenario when a new 3D object classification task is coming. Meanwhile, the existing lifelong learning models can easily destroy the learned tasks performance, due to the unordered, large-scale, and irregular 3D geometry data. To address these challenges, we propose a Lifelong 3D Object Classification (i.e., L3DOC) model, which can consecutively learn new 3D object classification tasks via imitating "human learning". More specifically, the core idea of our model is to capture and store the cross-task common knowledge of 3D geometry data in a 3D neural network, named as point-knowledge, through employing layer-wise point-knowledge factorization architecture. Afterwards, a task-relevant knowledge distillation mechanism is employed to connect the current task to previous relevant tasks and effectively prevent catastrophic forgetting. It consists of a point-knowledge distillation module and a transforming-space distillation module, which transfers the accumulated point-knowledge from previous tasks and soft-transfers the compact factorized representations of the transforming-space, respectively. To our best knowledge, the proposed L3DOC algorithm is the first attempt to perform deep learning on 3D object classification tasks in a lifelong learning way. Extensive experiments on several point cloud benchmarks illustrate the superiority of our L3DOC model over the state-of-the-art lifelong learning methods.
Yang Cong, Gan Sun, Tao Zhang 0084, Jiahua Dong 0001, Hongsen Liu
IEEE Trans. Image Process.2
2021 Multi-Scale Context-Guided Deep Network for Automated Lesion Segmentation With Endoscopy Images of Gastrointestinal Tract
abstract
Accurate lesion segmentation based on endoscopy images is a fundamental task for the automated diagnosis of gastrointestinal tract (GI Tract) diseases. Previous studies usually use hand-crafted features for representing endoscopy images, while feature definition and lesion segmentation are treated as two standalone tasks. Due to the possible heterogeneity between features and segmentation models, these methods often result in sub-optimal performance. Several fully convolutional networks have been recently developed to jointly perform feature learning and model training for GI Tract disease diagnosis. However, they generally ignore local spatial details of endoscopy images, as down-sampling operations (e.g., pooling and convolutional striding) may result in irreversible loss of image spatial information. To this end, we propose a multi-scale context-guided deep network (MCNet) for end-to-end lesion segmentation of endoscopy images in GI Tract, where both global and local contexts are captured as guidance for model training. Specifically, one global subnetwork is designed to extract the global structure and high-level semantic context of each input image. Then we further design two cascaded local subnetworks based on output feature maps of the global subnetwork, aiming to capture both local appearance information and relatively high-level semantic information in a multi-scale manner. Those feature maps learned by three subnetworks are further fused for the subsequent task of lesion segmentation. We have evaluated the proposed MCNet on 1,310 endoscopy images from the public EndoVis-Ab and CVC-ClinicDB datasets for abnormal segmentation and polyp segmentation, respectively. Experimental results demonstrate that MCNet achieves [Formula: see text] and [Formula: see text] mean intersection over union (mIoU) on two datasets, respectively, outperforming several state-of-the-art approaches in automated lesion segmentation with endoscopy images of GI Tract.
Shuai Wang 0003, Yang Cong, Hancan Zhu, Xianyi Chen, Liangqiong Qu, Huijie Fan, Qiang Zhang 0008, Mingxia Liu 0001
IEEE J. Biomed. Health Informatics2
2021 Continual Multiview Task Learning via Deep Matrix Factorization
abstract
The state-of-the-art multitask multiview (MTMV) learning tackles a scenario where multiple tasks are related to each other via multiple shared feature views. However, in many real-world scenarios where a sequence of the multiview task comes, the higher storage requirement and computational cost of retraining previous tasks with MTMV models have presented a formidable challenge for this lifelong learning scenario. To address this challenge, in this article, we propose a new continual multiview task learning model that integrates deep matrix factorization and sparse subspace learning in a unified framework, which is termed deep continual multiview task learning (DCMvTL). More specifically, as a new multiview task arrives, DCMvTL first adopts a deep matrix factorization technique to capture hidden and hierarchical representations for this new coming multiview task while accumulating the fresh multiview knowledge in a layerwise manner. Then, a sparse subspace learning model is employed for the extracted factors at each layer and further reveals cross-view correlations via a self-expressive constraint. For model optimization, we derive a general multiview learning formulation when a new multiview task comes and apply an alternating minimization strategy to achieve lifelong learning. Extensive experiments on benchmark data sets demonstrate the effectiveness of our proposed DCMvTL model compared with the existing state-of-the-art MTMV and lifelong multiview task learning models.
Gan Sun, Yang Cong, Yulun Zhang 0001, Guoshuai Zhao 0001, Yun Fu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2021 Robust 3-D Object Recognition via View-Specific Constraint
abstract
Three-dimensional (3-D) object recognition task focuses on detecting the objects of a scene and estimating their 6-DOF pose via effective feature extraction methods. Most recent feature extraction methods are based on the deep neural networks and show good performances. However, these methods require rendering engine to assist in generating a large amount of training data, which need much time to converge and further lead to the block in a rapid industrial production line. Besides, for the common hand-crafted features, the lack of discriminant feature-points amongst various texture-less and surface-smooth objects can cause ambiguity in the process of feature-points matching. To address these challenges above, a hand-crafted 3-D feature descriptor with center offset and pose annotations is proposed in this article, which is called view-specific local projection statistics (VSLPSs). By relying on these annotations as seeds, a voting strategy is then used to transform the feature-points matching problem into the problem of voting an optimal model-view in the 6-DOF space. In this way, the ambiguity of feature-points matching caused by poor feature discrimination is eliminated. To the end, various experiments on three public datasets and our built 3-D bin-picking dataset demonstrate that our proposed VSLPS method performs well in comparison with the state-of-the-art.
Hongsen Liu, Yang Cong, Gan Sun, Yandong Tang
IEEE Trans. Syst. Man Cybern. Syst.2
2020 Lifelong Spectral Clustering
abstract
In the past decades, spectral clustering (SC) has become one of the most effective clustering algorithms. However, most previous studies focus on spectral clustering tasks with a fixed task set, which cannot incorporate with a new spectral clustering task without accessing to previously learned tasks. In this paper, we aim to explore the problem of spectral clustering in a lifelong machine learning framework, i.e., Lifelong Spectral Clustering (L2SC). Its goal is to efficiently learn a model for a new spectral clustering task by selectively transferring previously accumulated experience from knowledge library. Specifically, the knowledge library of L2SC contains two components: 1) orthogonal basis library: capturing latent cluster centers among the clusters in each pair of tasks; 2) feature embedding library: embedding the feature manifold information shared among multiple related tasks. As a new spectral clustering task arrives, L2SC firstly transfers knowledge from both basis library and feature library to obtain encoding matrix, and further redefines the library base over time to maximize performance across all the clustering tasks. Meanwhile, a general online update formulation is derived to alternatively update the basis library and feature library. Finally, the empirical experiments on several real-world benchmark datasets demonstrate that our L2SC model can effectively improve the clustering performance when comparing with other state-of-the-art spectral clustering algorithms.
Gan Sun, Yang Cong, Qianqian Wang 0001, Jun Li 0027, Yun Fu 0001
AAAI2
2020 Visual Tactile Fusion Object Clustering
abstract
Object clustering, aiming at grouping similar objects into one cluster with an unsupervised strategy, has been extensively-studied among various data-driven applications. However, most existing state-of-the-art object clustering methods (e.g., single-view or multi-view clustering methods) only explore visual information, while ignoring one of most important sensing modalities, i.e., tactile information which can help capture different object properties and further boost the performance of object clustering task. To effectively benefit both visual and tactile modalities for object clustering, in this paper, we propose a deep Auto-Encoder-like Non-negative Matrix Factorization framework for visual-tactile fusion clustering. Specifically, deep matrix factorization constrained by an under-complete Auto-Encoder-like architecture is employed to jointly learn hierarchical expression of visual-tactile fusion data, and preserve the local structure of data generating distribution of visual and tactile modalities. Meanwhile, a graph regularizer is introduced to capture the intrinsic relations of data samples within each modality. Furthermore, we propose a modality-level consensus regularizer to effectively align the visual and tactile data in a common subspace in which the gap between visual and tactile data is mitigated. For the model optimization, we present an efficient alternating minimization strategy to solve our proposed model. Finally, we conduct extensive experiments on public datasets to verify the effectiveness of our framework.
Tao Zhang 0084, Yang Cong, Gan Sun, Qianqian Wang 0001, Zhengming Ding
AAAI2
2020 What Can Be Transferred: Unsupervised Domain Adaptation for Endoscopic Lesions Segmentation
abstract
Unsupervised domain adaptation has attracted growing research attention on semantic segmentation. However, 1) most existing models cannot be directly applied into lesions transfer of medical images, due to the diverse appearances of same lesion among different datasets; 2) equal attention has been paid into all semantic representations instead of neglecting irrelevant knowledge, which leads to negative transfer of untransferable knowledge. To address these challenges, we develop a new unsupervised semantic transfer model including two complementary modules (i.e., T_D and T_F ) for endoscopic lesions segmentation, which can alternatively determine where and how to explore transferable domain-invariant knowledge between labeled source lesions dataset (e.g., gastroscope) and unlabeled target diseases dataset (e.g., enteroscopy). Specifically, T_D focuses on where to translate transferable visual information of medical lesions via residual transferability-aware bottleneck, while neglecting untransferable visual characterizations. Furthermore, T_F highlights how to augment transferable semantic features of various lesions and automatically ignore untransferable representations, which explores domain-invariant knowledge and in return improves the performance of T_D. To the end, theoretical analysis and extensive experiments on medical endoscopic dataset and several non-medical public datasets well demonstrate the superiority of our proposed model.
Jiahua Dong 0001, Yang Cong, Gan Sun, Bineng Zhong 0001, Xiaowei Xu 0001
CVPR2
2020 CSCL: Critical Semantic-Consistent Learning for Unsupervised Domain Adaptation
Jiahua Dong 0001, Yang Cong, Gan Sun, Xiaowei Xu 0001
ECCV (8)2
2020 Dual Refinement Underwater Object Detection Network
Baojie Fan, Yang Cong, Jiandong Tian
ECCV (20)3
2020 TPN: Topological Perception Network For 3d Mesh Representation
abstract
As an important type of geometric data for 3D shapes, the unique topological connection makes the mesh more powerful than other types of data, but it also introduces complexity and irregularity. In this paper, we propose a Topological Perception Network (TPN) that consumes meshes directly to learn 3D shape representation via informative topology property. More specifically, to tackle the complexity and irregularity problem, a Topological Perception Attention (TPA) is designed that could incorporate local topological information efficiently via focusing on more important edges of the local topological neighborhood. Meanwhile, it could be stacked to produce global shape representation. Compared with the state-of-the-art, the proposed TPN uses less than half of the vertex number to get better performance, while costing less memory and computational time. Experiments on ModelNet40 and ShapeNet Core55 datasets demonstrate the effectiveness of our method on classification and retrieval.
Bingtao Ma, Yang Cong, Hongsen Liu
ICIP2
2020 Reliable Multi-Kernel Subtask Graph Correlation Tracker
abstract
Many astonishing correlation filter trackers pay limited concentration on the tracking reliability and locating accuracy. To solve the issues, we propose a reliable and accurate cross correlation particle filter tracker via graph regularized multi-kernel multi-subtask learning. Specifically, multiple non-linear kernels are assigned to multi-channel features with reliable feature selection. Each kernel space corresponds to one type of reliable and discriminative features. Then, we define the trace of each target subregion with one feature as a single view, and their multi-view cooperations and interdependencies are exploited to jointly learn multi-kernel subtask cross correlation particle filters, and make them complement and boost each other. The learned filters consist of two complementary parts: weighted combination of base kernels and reliable integration of base filters. The former is associated to feature reliability with importance map, and the weighted information reflects different tracking contribution to accurate location. The second part is to find the reliable target subtasks via the response map, to exclude the distractive subtasks or backgrounds. Besides, the proposed tracker constructs the Laplacian graph regularization via cross similarity of different subtasks, which not only exploits the intrinsic structure among subtasks, and preserves their spatial layout structure, but also maintains the temporal-spatial consistency of subtasks. Comprehensive experiments on five datasets demonstrate its remarkable and competitive performance against state-of-the-art methods.
Baojie Fan, Yang Cong, Jiandong Tian, Yandong Tang
IEEE Trans. Image Process.2
2019 Semantic-Transferable Weakly-Supervised Endoscopic Lesions Segmentation
abstract
Weakly-supervised learning under image-level labels supervision has been widely applied to semantic segmentation of medical lesions regions. However, 1) most existing models rely on effective constraints to explore the internal representation of lesions, which only produces inaccurate and coarse lesions regions; 2) they ignore the strong probabilistic dependencies between target lesions dataset (e.g., enteroscopy images) and well-to-annotated source diseases dataset (e.g., gastroscope images). To better utilize these dependencies, we present a new semantic lesions representation transfer model for weakly-supervised endoscopic lesions segmentation, which can exploit useful knowledge from relevant fully-labeled diseases segmentation task to enhance the performance of target weakly-labeled lesions segmentation task. More specifically, a pseudo label generator is proposed to leverage seed information to generate highly-confident pseudo pixel labels by incorporating class balance and super-pixel spatial prior. It can iteratively include more hard-to-transfer samples from weakly-labeled target dataset into training set. Afterwards, dynamically-searched feature centroids for same class among different datasets are aligned by accumulating previously-learned features. Meanwhile, adversarial learning is also employed in this paper, to narrow the gap between the lesions among different datasets in output space. Finally, we build a new medical endoscopic dataset with 3659 images collected from more than 1100 volunteers. Extensive experiments on our collected dataset and several benchmark datasets validate the effectiveness of our model.
Jiahua Dong 0001, Yang Cong, Gan Sun, Dongdong Hou
ICCV2
2019 Memory-Based Parameterized Skills Learning for Mapless Visual Navigation
abstract
The recently-proposed reinforcement learning for mapless visual navigation can generate an optimal policy for searching different targets. However, most state-of-the-art deep reinforcement learning (DRL) models depend on hard rewards to learn the optimal policy, which can lead to the lack of previous diverse experiences. Moreover, these pre-trained DRL models cannot generalize well to un-trained tasks. To overcome these problems above, in this paper, we propose a Memory-based Parameterized Skills Learning (MPSL) model for mapless visual navigation. The parameterized skills in our MPSL are learned to predict critic parameters for un-trained tasks in actor-critic reinforcement learning, which can be achieved by transferring memory sequence knowledge from long short term memory network. In order to generalize into un-trained tasks, MPSL aims to capture more discriminative features by using a scene-specific layer. Finally, experiment results on an indoor photographic simulation framework AI2THOR demonstrate the effectiveness of our proposed MPSL model, and the generalization ability to un-trained tasks.
Yang Cong, Gan Sun
ICIP2
2019 Environment Driven Underwater Camera-IMU Calibration for Monocular Visual-Inertial SLAM
abstract
Most state-of-the-art underwater vision systems are calibrated manually in shallow water and used in open seas without changing. However, the refractivity of the water is adaptively changed depending on the salinity, temperature, depth or other underwater environmental indexes, which inevitably generate the calibration errors and induces incorrectness e.g., for underwater Simultaneously Localization and Mapping (SLAM). To address this issue, in this paper, we propose a new underwater Camera-Inertial Measurement Unit (IMU) calibration model, which just needs to be calibrated once in the air, and then both the intrinsic parameters and extrinsic parameters between the camera and IMU could be automatically calculated depending on the environment indexes. To our best knowledge, this is the first work to consider the underwater Camera-IMU calibration via environmental indexes. We also build a verification platform to validate the effectiveness of our proposed method on real experiments, and use it for underwater monocular Visual-Inertial SLAM.
Changjun Gu, Yang Cong, Gan Sun
ICRA2
2019 Anomaly detection via adaptive greedy model
Dongdong Hou, Yang Cong, Gan Sun, Ji Liu 0002, Xiaowei Xu 0001
Neurocomputing2
2019 Novel event analysis for human-machine collaborative underwater exploration
Yang Cong, Baojie Fan, Dongdong Hou, Huijie Fan, Kaizhou Liu, Jiebo Luo 0001
Pattern Recognit.1
2019 Efficient 3D object recognition via geometric information preservation
Hongsen Liu, Yang Cong, Chenguang Yang 0001, Yandong Tang
Pattern Recognit.2
2019 Laplacian pyramid adversarial network for face completion
Qiang Wang 0015, Huijie Fan, Gan Sun, Yang Cong, Yandong Tang
Pattern Recognit.4
2019 Speedup 3-D Texture-Less Object Recognition Against Self-Occlusion for Intelligent Manufacturing
abstract
Realtime 3-D object detection and 6-DOF pose estimation in clutter background is crucial for intelligent manufacturing, for example, robot feeding and assembly, where robustness and efficiency are the two most desirable goals. Especially for various metal parts with a textless surface, it is hard for most state of the arts to extract robust feature from the clutter background with various occlusions. To overcome this, in this paper, we propose an online 3-D object detection and pose estimation method to overcome self-occlusion for textureless objects. For feature representation, we only adopt the raw 3-D point clouds with normal cues to define our local reference frame and we automatically learn the compact 3-D feature from the simple local normal statistics via autoencoder. For a similarity search, a new basis buffer k-d tree method is designed without suffering branch divergence; therefore, ours can maximize the GPU parallel processing capabilities especially in practice. We then generate the hypothesis candidates via the hough voting, filter the false hypotheses, and refine the pose estimation via the iterative closest point strategy. For the experiments, we build a new 3-D dataset including industrial objects with heavy self-occlusions and conduct various comparisons with the state of the arts to justify the effectiveness and efficiency of our method.
Yang Cong, Dongying Tian, Baojie Fan
IEEE Trans. Cybern.1
2019 Lifelong Metric Learning
abstract
The state-of-the-art online learning approaches are only capable of learning the metric for predefined tasks. In this paper, we consider a lifelong learning problem to mimic "human learning," i.e., endowing a new capability to the learned metric for a new task from new online samples and incorporating the previous experiences. Therefore, we propose a new metric learning framework: lifelong metric learning (LML), which only utilizes the data of the new task to train the metric model while preserving the original capabilities. More specifically, the proposed LML maintains a common subspace for all learned metrics, named lifelong dictionary, transfers knowledge from the common subspace to learn each new metric learning task with task-specific idiosyncrasy, and redefines the common subspace over time to maximize performance across all metric tasks. For model optimization, we apply online passive aggressive optimization algorithm to achieve lifelong metric task learning, where the lifelong dictionary and task-specific partition are optimized alternatively and consecutively. Finally, we evaluate our approach by analyzing several multitask metric learning datasets. Extensive experimental results demonstrate effectiveness and efficiency of the proposed framework.
Gan Sun, Yang Cong, Ji Liu 0002, Lianqing Liu, Xiaowei Xu 0001
IEEE Trans. Cybern.2
2019 A Learning Framework of Adaptive Manipulative Skills From Human to Robot
abstract
Robots are often required to generalize the skills learned from human demonstrations to fulfil new task requirements. However, skill generalization will be difficult to realize when facing with the following situations: the skill for a complex multistep task includes a number of features; some special constraints are imposed on the robots during the process of task reproduction; and a completely new situation quite different with the one in which demonstrations are given to the robot. This work proposes a new framework to facilitate robot skill generalization. The basic idea lies in that the learned skills are first segmented into a sequence of subskills automatically, then each individual subskill is encoded and regulated accordingly. Specifically, we adapt each set of the segmented movement trajectories individually instead of the whole movement profiles, thus, making it more convenient for the realization of skill generalization. In addition, human limb stiffness estimated from surface electromyographic signals is considered in the framework for the realization of human-to-robot variable impedance control skill transfer, as well as the generalization of both movement trajectories and stiffness profiles. Experimental study has been performed to verify the effectiveness of the proposed framework.
Chenguang Yang 0001, Chao Zeng 0002, Yang Cong, Ning Wang 0009, Min Wang 0003
IEEE Trans. Ind. Informatics3
2018 Active Lifelong Learning With "Watchdog"
abstract
Lifelong learning intends to learn new consecutive tasks depending on previously accumulated experiences, i.e., knowledge library. However, the knowledge among different new coming tasks are imbalance. Therefore, in this paper, we try to mimic an effective "human cognition" strategy by actively sorting the importance of new tasks in the process of unknown-to-known and selecting to learn the important tasks with more information preferentially. To achieve this, we consider to assess the importance of the new coming task, i.e., unknown or not, as an outlier detection issue, and design a hierarchical dictionary learning model consisting of two-level task descriptors to sparse reconstruct each task with the l0 norm constraint. The new coming tasks are sorted depending on the sparse reconstruction score in descending order, and the task with high reconstruction score will be permitted to pass, where this mechanism is called as "watchdog." Next, the knowledge library of the lifelong learning framework encode the selected task by transferring previous knowledge, and then can also update itself with knowledge from both previously learned task and current task automatically. For model optimization, the alternating direction method is employed to solve our model and converges to a fixed point. Extensive experiments on both benchmark datasets and our own dataset demonstrate the effectiveness of our proposed model especially in task selection and dictionary learning.
Gan Sun, Yang Cong, Xiaowei Xu 0001
AAAI2
2018 Clustered Lifelong Learning Via Representative Task Selection
abstract
Consider the lifelong machine learning problem where the objective is to learn new consecutive tasks depending on previously accumulated experiences, i.e., knowledge library. In comparison with most state-of-the-arts which adopt knowledge library with prescribed size, in this paper, we propose a new incremental clustered lifelong learning model with two libraries: feature library and model library, called Clustered Lifelong Learning (CL3), in which the feature library maintains a set of learned features common across all the encountered tasks, and the model library is learned by identifying and adding representative models (clusters). When a new task arrives, the original task model can be firstly reconstructed by representative models measured by capped l2-norm distance, i.e., effectively assigning the new task model to multiple representative models under feature library. Based on this assignment knowledge of new task, the objective of our CL3 model is to transfer the knowledge from both feature library and model library to learn the new task. The new task 1) with a higher outlier probability will then be judged as a new representative, and used to refine both feature library and representative models over time; 2) with lower outlier probability will only update the feature library. For the model optimisation, we cast this problem as an alternating direction minimization problem. To this end, the performance of CL3 is evaluated through comparing with most lifelong learning models, even some batch clustered multi-task learning models.
Gan Sun, Yang Cong, Yu Kong 0001, Xiaowei Xu 0001
ICDM2
2018 Online Low-Rank Metric Learning via Parallel Coordinate Descent Method
abstract
11The corresponding author is Prof. Yang Cong. This work is supported by Nature Science Foundation of China under Grant (61722311, U1613214, 61533015) and CAS-Youth Innovation Promotion Association Scholarship (2012163)Recently, many machine learning problems rely on a valuable tool: metric learning. However, in many applications, large-scale applications embedded in high-dimensional feature space may induce both computation and storage requirements to grow quadratically. In order to tackle these challenges, in this paper, we intend to establish a robust metric learning formulation with the expectation that online metric learning and parallel optimization can solve large-scale and high-dimensional data efficiently, respectively. Specifically, based on the matrix factorization strategy, the first step aims to learn a similarity function in the objective formulation for similarity measurement; in the second step, we derive a variational trace norm to promote low-rankness on the transformation matrix. After converting this variational regularization into its separable form, for the model optimization, we present an parallel block coordinate descent method to learn the optimal metric parameters, which can handle the high-dimensional data in an efficient way. Crucially, our method shares the efficiency and flexibility of block coordinate descent method, and it is also guaranteed to converge to the optimal solution. Finally, we evaluate our approach by analyzing scene categorization dataset with tens of thousands of dimensions, and the experimental results show the effectiveness of our proposed model.
Gan Sun, Yang Cong, Qiang Wang 0015, Xiaowei Xu 0001
ICPR2
2018 User attribute discovery with missing labels
Yang Cong, Gan Sun, Ji Liu 0002, Jiebo Luo 0001
Pattern Recognit.1
2018 Structured and weighted multi-task low rank tracker
Baojie Fan, Xiaomao Li, Yang Cong, Yandong Tang
Pattern Recognit.3
2018 Online Similarity Learning for Big Data with Overfitting
abstract
In this paper, we propose a general model to address the overfitting problem in online similarity learning for big data, which is generally generated by two kinds of redundancies: 1) feature redundancy, that is there exists redundant (irrelevant) features in the training data; 2) rank redundancy, that is non-redundant (or relevant) features lie in a low rank space. To overcome these, our model is designed to obtain a simple and robust metric matrix through detecting the redundant rows and columns in the metric matrix and constraining the remaining matrix to a low rank space. To reduce feature redundancy, we employ the group sparsity regularization, i.e., the `2;1 norm, to encourage a sparse feature set. To address rank redundancy, we adopt the low rank regularization, the max norm, instead of calculating the SVD as in traditional models using the nuclear norm. Therefore, our model can not only generate a low rank metric matrix to avoid overfitting, but also achieves feature selection simultaneously. For model optimization, an online algorithm based on the stochastic proximal method is derived to solve this problem efficiently with the complexity of O(d2). To validate the effectiveness and efficiency of our algorithms, we apply our model to online scene categorization and synthesized data and conduct experiments on various benchmark datasets with comparisons to several state-of-the-art methods. Our model is as efficient as the fastest online similarity learning model OASIS, while performing generally as well as the accurate model OMLLR. Moreover, our model can exclude irrelevant / redundant feature dimension simultaneously.
Yang Cong, Ji Liu 0002, Baojie Fan, Peng Zeng 0001, Jiebo Luo 0001
IEEE Trans. Big Data1
2018 Dual-Graph Regularized Discriminative Multitask Tracker
abstract
Multitask and low-rank learning methods have attracted increasing attention for visual tracking. However, most trackers only focus on learning appearance subspace basis or the sparse low rankness of representation and, thus, do not make full use of the structure information among and inside target candidates (or samples). In this paper, we propose a dual-graph regularized discriminative low-rank learning for a multitask tracker, which integrates the discriminative subspace and intrinsic geometric structures among tasks. By constructing dual-graph regulations from two views of multitask observation, the developed model not only exploits the intrinsic relationship among tasks, and preserves the spatial layout structure among the local patches inside each candidate, but also learns the salient features of the target samples. This operation has the benefit of having good target representation and improving the performance of the tracker. Moreover, our developed tracker is a collaborate multitask tracking model and learns the discriminative subspace with adaptive dimension and optimal classifier simultaneously. Then, a collaborate metric is developed to find the best candidate, which integrates both classification reliability and representation accuracy. Encouraging experimental results on a large set of public video sequences justify that our tracker performs favorably against many other state-of-the-art trackers.
Baojie Fan, Yang Cong, Yandong Tang
IEEE Trans. Multim.2
2017 Large receptive field convolutional neural network for image super-resolution
abstract
This paper presents a new approach to Single Image Super Resolution (SISR), based upon Convolutional Neural Network (CNN). Although the SISR is ill-posed which can be seen as finding a non-linear mapping from a low to high-dimensional space. Deep learning techniques have been successfully applied in many areas of computer vision, including low-level image restoration and non-linear mapping problems. We consider the single image Super-Resolution (SR) problem as convolution operators and develop a CNN to capture the characteristics of Low-Resolution (LR) input image. We find that increasing the receptive field shows the improvement in accuracy. Our solution is to establish the connection between traditional optimization-based schemes and neural network architectures. In the paper a novel, separable structure is introduced as a reliable support for robust convolution against artifacts. Our proposed method performs better than existing methods in terms of accuracy and visual improvements in our results are easily noticeable.
Qiang Wang 0015, Huijie Fan, Yang Cong, Yandong Tang
ICIP3
2017 Deep learning of directional truncated signed distance function for robust 3D object recognition
abstract
In this paper, we develop a novel 3D object recognition algorithm to perform detection and pose estimation jointly. We focus on analyzing the advantages of the 3D point cloud relative to the RGB-D image and try to eliminate the unpredictability of output values that inevitably occurs in regression tasks. To achieve this, we first adopt the Truncated Signed Distance Function (TSDF) to encode the point cloud and extract low compact discriminative feature via unsupervised deep learning network. This approach can not only eliminate the dense scale sampling for offline model training but also reduce the distortion by mapping the 3D shape to the 2D plane and overcome the dependence on color cues. Then, we train a Hough forests to achieve multi-object detection and 6-DoF pose estimation simultaneously. In addition, we propose a robust multilevel verification strategy that effectively reduces the unpredictability of output values which occurs in the hough regression module. Experiments on public datasets demonstrate that our approach provides effective results comparable to the state-of-the-arts.
Hongsen Liu, Yang Cong, Shuai Wang 0003, Huijie Fan, Dongying Tian, Yandong Tang
IROS2
2017 Consistent multi-layer subtask tracker via hyper-graph regularization
Baojie Fan, Yang Cong
Pattern Recognit.2
2017 Layered Multitask Tracker via Spatial-Temporal Laplacian Graph
abstract
Most multitask trackers define the trace of each candidate as one task, and assume all tasks are equally related. Multitask learning is only evaluated on the current frame. In fact, these assumptions are limited, and ignore the multitask relationship in consecutive frames. In this letter, we propose a discriminative layered multitask tracker via spatial-temporal Laplacian graphs, which defines the layered tasks from a novel view, and naturally incorporates the global and local target information into reverse multitask tracking process. The spatial-temporal Laplacian graphs not only exploit the sequential consistent information of the target, but also make full use of the geometric structure corresponding to the tasks among the adjacent frames. Besides, l0norm constraint and labeling information are used to improve the tracking robustness. Encouraging experimental results on challenging sequences justify that the proposed method performs well both in accuracy and robustness against some related trackers.
Baojie Fan, Xiaomao Li, Yang Cong
IEEE Signal Process. Lett.3
2017 Adaptive Greedy Dictionary Selection for Web Media Summarization
abstract
Initializing an effective dictionary is an indispensable step for sparse representation. In this paper, we focus on the dictionary selection problem with the objective to select a compact subset of basis from original training data instead of learning a new dictionary matrix as dictionary learning models do. We first design a new dictionary selection model via l2,0norm. For model optimization, we propose two methods: one is the standard forward-backward greedy algorithm, which is not suitable for large-scale problems; the other is based on the gradient cues at each forward iteration and speeds up the process dramatically. In comparison with the state-of-the-art dictionary selection models, our model is not only more effective and efficient, but also can control the sparsity. To evaluate the performance of our new model, we select two practical web media summarization problems: 1) we build a new data set consisting of around 500 users, 3000 albums, and 1 million images, and achieve effective assisted albuming based on our model and 2) by formulating the video summarization problem as a dictionary selection issue, we employ our model to extract keyframes from a video sequence in a more flexible way. Generally, our model outperforms the state-of-the-art methods in both these two tasks.
Yang Cong, Ji Liu 0002, Gan Sun, Quanzeng You, Yuncheng Li, Jiebo Luo 0001
IEEE Trans. Image Process.1
2017 Multi-Class Latent Concept Pooling for Computer-Aided Endoscopy Diagnosis
abstract
Successful computer-aided diagnosis systems typically rely on training datasets containing sufficient and richly annotated images. However, detailed image annotation is often time consuming and subjective, especially for medical images, which becomes the bottleneck for the collection of large datasets and then building computer-aided diagnosis systems. In this article, we design a novel computer-aided endoscopy diagnosis system to deal with the multi-classification problem of electronic endoscopy medical records (EEMRs) containing sets of frames, while labels of EEMRs can be mined from the corresponding text records using an automatic text-matching strategy without human special labeling. With unambiguous EEMR labels and ambiguous frame labels, we propose a simple but effective pooling scheme called Multi-class Latent Concept Pooling, which learns a codebook from EEMRs with different classes step by step and encodes EEMRs based on a soft weighting strategy. In our method, a computer-aided diagnosis system can be extended to new unseen classes with ease and applied to the standard single-instance classification problem even though detailed annotated images are unavailable. In order to validate our system, we collect 1,889 EEMRs with more than 59K frames and successfully mine labels for 348 of them. The experimental results show that our proposed system significantly outperforms the state-of-the-art methods. Moreover, we apply the learned latent concept codebook to detect the abnormalities in endoscopy images and compare it with a supervised learning classifier, and the evaluation shows that our codebook learning method can effectively extract the true prototypes related to different classes from the ambiguous data.
Shuai Wang 0003, Yang Cong, Huijie Fan, Baojie Fan, Lianqing Liu, Yunsheng Yang, Yandong Tang, Huaici Zhao
ACM Trans. Multim. Comput. Commun. Appl.2
2016 A design of phase-closed-loop nanomachining control based ultrasonic vibration-assisted AFM
abstract
This paper proposed a phase-closed-loop nanomachining control method to realize the directly control of machining depth based on ultrasonic vibration-assisted AFM. By using applied force to control the machining depth, conventional AFM machining approaches unable to machining a nanostructure with specified machined depth. With the proposed method, the vibration phase of micro-cantilever has a specific relationship with machining depth. Therefore, the nano-grooves with desired depth can be machined by using phase value as feedback of PID control. In this paper, the theoretical analysis and simulation are carried out, and the experiments of phase-closed-loop control method are conducted. The experimental results verify the primary feasibility of the proposed method. The present method also demonstrates the potential on the fabrication of three-dimension nanostructures and nanoelectronic device.
Jialin Shi, Lianqing Liu, Yang Cong
IROS4
2016 Structured low rank tracker with smoothed regularization
abstract
In this paper, we propose a structured low rank learning algorithm with smoothed regularization for robust object tracking, under particle filter framework. Specifically, the relationships among the particles are exploited with structured low rank regularization term, and simultaneously handle the outlier using a group sparsity regularization. The label information from training data is incorporated into the tracking objective function as the classification error term and idea coding regularization term respectively. By the smoothed regularization, the developed structured low rank learning based tracker can be efficiently solved by iterative reweighed least squares algorithm(IRLS), and avoids svd operation. Moreover, the collaborate normalized metric is developed to find the best candidate. Compared with some state-of-the-art tracking methods on 50 challenging sequences, the proposed algorithms perform well in terms of accuracy, robustness.
Baojie Fan, Yang Cong, Xiaomao Li, Yandong Tang
VCIP2
2016 Scalable gastroscopic video summarization via similar-inhibition dictionary selection
Shuai Wang 0003, Yang Cong, Jun Cao 0002, Yunsheng Yang, Yandong Tang, Huaici Zhao
Artif. Intell. Medicine2
2016 UDSFS: Unsupervised deep sparse feature selection
Yang Cong, Shuai Wang 0003, Baojie Fan, Yunsheng Yang
Neurocomputing1
2015 User-curated image collections: Modeling and recommendation
abstract
Most state-of-the-art image retrieval and recommendation systems predominantly focus on individual images. In contrast, socially curated image collections, condensing distinctive yet coherent images into one set, are largely overlooked by the research communities. In this paper, we aim to design a novel recommendation system that can provide users with image collections relevant to individual personal preferences and interests. To this end, two key issues need to be addressed, i.e., image collection modeling and similarity measurement. For image collection modeling, we consider each image collection as a whole in a group sparse reconstruction framework and extract concise collection descriptors given the pretrained dictionaries. We then consider image collection recommendation as a dynamic similarity measurement problem in response to user's clicked image set, and employ a metric learner to measure the similarity between the image collection and the clicked image set. As there is no previous work directly comparable to this study, we implement several competitive baselines and related methods for comparison. The evaluations on a large scale Pinterest data set have validated the effectiveness of our proposed methods for modeling and recommending image collections.
Yuncheng Li, Tao Mei 0001, Yang Cong, Jiebo Luo 0001
IEEE BigData3
2015 Computer aided endoscope diagnosis via weakly labeled data mining
abstract
In comparison to most computer aided endoscope diagnosis methods using pixel-wise groundtruth by physicians manually, it is easy to get lots of endoscope images with corresponding diagnostic reports. In this paper, we intend to mine pixel-wise label information from these reports with weak frame-level labels automatically. To achieve this, we formulate our computer aided diagnosis problem as a Multiple Instance Learning (MIL) issue, where we represent each image as superpixels. Each image and each superpixel is cast as bag and instance, respectively. We then evaluate and select the most positive instances from positive bags automatically which helps us transform the frame-level classification problem into a standard supervised learning problem. In the experiment, we build a new gastroscopic image dataset with more than 3000 weakly labeled images, and ours outperforms the state-of-the-art methods, which verifies the effectiveness of our model.
Shuai Wang 0003, Yang Cong, Huijie Fan, Yunsheng Yang, Yandong Tang, Huaici Zhao
ICIP2
2015 Real-time one-dimensional motion estimation and its application in computer vision
Yang Cong, Haifeng Gong, Yandong Tang, Shuzhi Sam Ge, Jiebo Luo 0001
Mach. Vis. Appl.1
2015 Object detection based on scale-invariant partial shape matching
Huijie Fan, Yang Cong, Yandong Tang
Mach. Vis. Appl.2
2015 Deep sparse feature selection for computer aided endoscopy diagnosis
Yang Cong, Shuai Wang 0003, Ji Liu 0002, Jun Cao 0002, Yunsheng Yang, Jiebo Luo 0001
Pattern Recognit.1
2015 Speeded Up Low-Rank Online Metric Learning for Object Tracking
abstract
Visual object tracking can be considered as an online procedure to adaptively measure the foreground object similarity itself. However, many previous works usually adopt a fixed metric or offline metric learning to evaluate this dynamic process; even with some online metric learning (OML) trackers, their models often suffer from overfitting issues. To overcome these deficiencies, we propose a self-supervised tracking method that incorporates adaptive metric learning and semisupervised learning into a unified framework. For similarity measurement, we design a new OML model via low-rank constraint to handle overfitting. In particular, we employ the max norm instead of the trace norm used in our previous work. This not only maintains the low-rank property to overcome overfitting, but also reduces the computational complexity from O(n3) to O(n2), such that the new model is more suitable for object tracking. Moreover, by associating the information from stored training templates with unlabeled testing samples, a bilinear graph is defined accordingly to propagate the label of each sample. High-confidence samples are then collected for self-training the model and updating the templates concurrently to handle large scale. Experiments on various benchmark data sets and comparisons to several state-of-the-art methods demonstrate the effectiveness and efficiency of our algorithm.
Yang Cong, Baojie Fan, Ji Liu 0002, Jiebo Luo 0001
IEEE Trans. Circuits Syst. Video Technol.1
2015 A Multifaceted Approach to Social Multimedia-Based Prediction of Elections
abstract
Compared with real-world polling, election prediction based on social media can be far more timely and cost-effective due to the immediate availability of fast evolving Web contents. However, information from social media may suffer from noise and sampling bias that are caused by various factors and thus pose one of biggest challenges in social media-based data analytics. This paper presents a new model, named competitive vector auto regression (CVAR), to build a reliable forecasting system for the US presidential elections and US House race. Our CVAR model is designed to analyze the correlation between image-centric social multimedia and real-world phenomena. By introducing the competition mechanism, CVAR compares the popularity among multiple competing candidates. More importantly , CVAR is able to combine visual information with textual information from rich and multifaceted social multimedia, which helps extract reliable signals and mitigate sampling bias. As a result, our proposed system can 1) accurately predict the election outcome, 2) infer the sentiment of the candidate photos shared in the social media communities, and 3) account for the sentiment of viewer comments towards the candidates on the related images. The experiments on the 2012 US presidential election at both national and state levels, as well as the 2014 US House race, have demonstrated the power and promise of the proposed approach.
Quanzeng You, Liangliang Cao, Yang Cong, Xianchao Zhang 0001, Jiebo Luo 0001
IEEE Trans. Multim.3
2014 A Unified Online Dictionary Learning Framework with Label Information for Robust Object Tracking
abstract
In this paper, a supervised approach to online learn a structured sparse and discriminative representation for object tracking is presented. Label information from training data is incorporated into the dictionary learning process to construct a robust and discriminative dictionary. This is accomplished by adding an ideal-code regularization term and classification error term to the unified objective function. By minimizing the unified objective function we learn the high quality dictionary and optimal linear multi-classifier jointly. Combined with robust sparse coding, the learned classifier is employed directly to separate the object from background. As the tracking continues, the proposed algorithm alternates between robust sparse coding and dictionary updating. Experimental evaluations on the challenging sequences show that the proposed algorithm performs favorably against state-of-the-art methods in terms of effectiveness, accuracy and robustness.
Baojie Fan, Yang Cong, Yingkui Du
ICPR3
2014 Discriminative multi-task objects tracking with active feature selection and drift correction
Baojie Fan, Yang Cong, Yingkui Du
Pattern Recognit.2
2013 Robust and accurate online pose estimation algorithm via efficient three-dimensional collinearity model
abstract
In this study, the authors propose a robust and high accurate pose estimation algorithm to solve the perspective‐ N ‐point problem in real time. This algorithm does away with the distinction between coplanar and non‐coplanar point configurations, and provides a unified formulation for the configurations. Based on the inverse projection ray, an efficient collinearity model in object–space is proposed as the cost function. The principle depth and the relative depth of reference points are introduced to remove the residual error of the cost function and to improve the robustness and the accuracy of the authors pose estimation method. The authors solve the pose information and the depth of the points iteratively by minimising the cost function, and then reconstruct their coordinates in camera coordinate system. In the following, the optimal absolute orientation solution gives the relative pose information between the estimated three‐dimensional (3D) point set and the 3D mode point set. This procedure with the above two steps is repeated until the result converges. The experimental results on simulated and real data show that the superior performance of the proposed algorithm: its accuracy is higher than the state‐of‐the‐art algorithms, and has best anti‐noise property and least deviation by the influence of outlier among the tested algorithms.
Baojie Fan, Yingkui Du, Yang Cong
IET Comput. Vis.3
2013 Abnormal event detection in crowded scenes using sparse representation
Yang Cong, Junsong Yuan 0001, Ji Liu 0002
Pattern Recognit.1
2013 Video Anomaly Search in Crowded Scenes via Spatio-Temporal Motion Context
abstract
Video anomaly detection plays a critical role for intelligent video surveillance. We present an abnormal video event detection system that considers both spatial and temporal contexts. To characterize the video, we first perform the spatio-temporal video segmentation and then propose a new region-based descriptor called “Motion Context,” to describe both motion and appearance information of the spatio-temporal segment. For anomaly measurements, we formulate the abnormal event detection as a matching problem, which is more robust than statistic model-based methods, especially when the training dataset is of limited size. For each testing spatio-temporal segment, we search for its best match in the training dataset, and determine how normal it is using a dynamic threshold. To speed up the search process, compact random projections are also adopted. Experiments on the benchmark dataset and comparisons with the state-of-the-art methods validate the advantages of our algorithm.
Yang Cong, Junsong Yuan 0001, Yandong Tang
IEEE Trans. Inf. Forensics Secur.1
2013 Self-Supervised Online Metric Learning With Low Rank Constraint for Scene Categorization
abstract
Conventional visual recognition systems usually train an image classifier in a bath mode with all training data provided in advance. However, in many practical applications, only a small amount of training samples are available in the beginning and many more would come sequentially during online recognition. Because the image data characteristics could change over time, it is important for the classifier to adapt to the new data incrementally. In this paper, we present an online metric learning method to address the online scene recognition problem via adaptive similarity measurement. Given a number of labeled data followed by a sequential input of unseen testing samples, the similarity metric is learned to maximize the margin of the distance among different classes of samples. By considering the low rank constraint, our online metric learning model not only can provide competitive performance compared with the state-of-the-art methods, but also guarantees convergence. A bi-linear graph is also defined to model the pair-wise similarity, and an unseen sample is labeled depending on the graph-based label propagation, while the model can also self-update using the more confident new samples. With the ability of online learning, our methodology can well handle the large-scale streaming video data with the ability of incremental self-updating. We evaluate our model to online scene categorization and experiments on various benchmark datasets and comparisons with state-of-the-art methods demonstrate the effectiveness and efficiency of our algorithm.
Yang Cong, Ji Liu 0002, Junsong Yuan 0001, Jiebo Luo 0001
IEEE Trans. Image Process.1
2012 Object tracking via online metric learning
abstract
By considering visual tracking as a similarity matching problem, we propose a self-supervised tracking method that incorporates adaptive metric learning and semi-supervised learning into the framework of object tracking. For object representation, the spatial-pyramid structure is applied by fusing both the shape and texture cues as descriptors. A metric learner is adaptively trained online to best distinguish the foreground object and background, and a new bi-linear graph is defined accordingly to propagate the label of each sample. Then high-confident samples are collected to self-update the model to handle large-scale issue. Experiments on the benchmark dataset and comparisons with the state-of-the-art methods validate the advantages of our algorithm.
Yang Cong, Junsong Yuan 0001, Yandong Tang
ICIP1
2012 Self-closed partial shape descriptor for shape retrieval
abstract
We propose a discriminative partial-based algorithm for shape recognition and retrieval. A key distinction of our approach is that we use pairwise geometric relations between contour fragments containing important and salient shape information to establish self-closed partial descriptor (SCPD), it can capture similar local parts in matching shape contours and meanwhile overcome part occlusion and distortion. We establish local coordinate system for each fragment to make sure SCPD is invariant to RST (rotation, scaling, and translation) transformation. In the matching stage, a scale approximation scheme is used to get rid of invalid matches. We experiment on MPEG7 shape database, and experimental results illustrate that our algorithm performs well on shape retrieval.
Huijie Fan, Yang Cong, Yandong Tang
ICIP2
2012 Active drift correction template tracking algorithm
abstract
This paper presents a novel active drift correction template tracking algorithm. Compared to Matthews' algorithm in [8], the proposed algorithm achieves synchronously object tracking and drift correction, and save half running time. For the template drift problem during long sequential object tracking, we introduce the active drift correction term into inverse compositional affine image alignment algorithm. This operation can avoid the template drift before it occurs, or reduce the drift after it happens. The total energy function consists of two terms: the tracking term and the active drift correction term. By minimizing the total energy function with the steepest descent algorithm, the proposed algorithm can decrease the accumulative tracking error, and prevent the drift during the tracking process effectively. Various object tracking experiments show that our method has super performance than the passive drift correction algorithm in [8].
Baojie Fan, Yingkui Du, Yang Cong, Yandong Tang
ICIP3
2012 Local Line Derivative Pattern for face recognition
abstract
In this paper, we propose a novel face descriptor for face recognition, named Local Line Derivative Pattern (LLDP). High-order derivative images in two directions are obtained by convolving original images with Sobel Masks. A revised binary coding function is proposed and three standards on arranging the weights are also proposed. Based on the standards, the weights of a line neighborhood in two directions are arranged. The LLDP labels in two directions are calculated with the proposed binary coding function and weights. The labeled image is divided into blocks where spatial histograms are extracted separately and concatenated into an entire histogram as features for recognition. The experiments on the FERET and Extended Yale B show superior performances of the proposed LLDP compared to other existing methods based on the LBP. The results prove that the LLDP has good robustness against expression, illumination and aging variations.
Zhichao Lian, Meng Joo Er, Yang Cong
ICIP3
2012 Towards Scalable Summarization of Consumer Videos Via Sparse Dictionary Selection
abstract
The rapid growth of consumer videos requires an effective and efficient content summarization method to provide a user-friendly way to manage and browse the huge amount of video data. Compared with most previous methods that focus on sports and news videos, the summarization of personal videos is more challenging because of its unconstrained content and the lack of any pre-imposed video structures. We formulate video summarization as a novel dictionary selection problem using sparsity consistency, where a dictionary of key frames is selected such that the original video can be best reconstructed from this representative dictionary. An efficient global optimization algorithm is introduced to solve the dictionary selection model with the convergence rates asO(1/K2) (whereKis the iteration counter), in contrast to traditional sub-gradient descent methods ofO(1/√K). Our method provides a scalable solution for both key frame extraction and video skim generation, because one can select an arbitrary number of key frames to represent the original videos. Experiments on a human labeled benchmark dataset and comparisons to the state-of-the-art methods demonstrate the advantages of our algorithm.
Yang Cong, Junsong Yuan 0001, Jiebo Luo 0001
IEEE Trans. Multim.1
2011 Sparse reconstruction cost for abnormal event detection
abstract
We propose to detect abnormal events via a sparse reconstruction over the normal bases. Given an over-complete normal basis set (e.g., an image sequence or a collection of local spatio-temporal patches), we introduce the sparse reconstruction cost (SRC) over the normal dictionary to measure the normalness of the testing sample. To condense the size of the dictionary, a novel dictionary selection method is designed with sparsity consistency constraint. By introducing the prior weight of each basis during sparse reconstruction, the proposed SRC is more robust compared to other outlier detection criteria. Our method provides a unified solution to detect both local abnormal events (LAE) and global abnormal events (GAE). We further extend it to support online abnormal event detection by updating the dictionary incrementally. Experiments on three benchmark datasets and the comparison to the state-of-the-art methods validate the advantages of our algorithm.
Yang Cong, Junsong Yuan 0001, Ji Liu 0002
CVPR1
2009 Flow mosaicking: Real-time pedestrian counting without scene-specific learning
abstract
In this paper, we present a novel algorithm based on flow velocity field estimation to count the number of pedestrians across a detection line or inside a specified region. We regard pedestrians across the line as fluid flow, and design a novel model to estimate the flow velocity field. By integrating over time, the dynamic mosaics are constructed to count the number of pixels and edges passed through the line. Consequentially, the number of pedestrians can be estimated by quadratic regression, with the number of weighted pixels and edges as input. The regressors are learned off line from several camera tilt angles, and have taken the calibration information into account.We use tilt-angle-specific learning to ensure direct deployment and avoid overfitting while the commonly used scene-specific learning scheme needs on-site annotation and always trends to overfitting. Experiments on a variety of videos verified that the proposed method can give accurate estimation under different camera setup in real-time.
Yang Cong, Haifeng Gong, Song-Chun Zhu, Yandong Tang
CVPR1
2008 Lunar terrain reconstruction using PDEs
abstract
Based on the geometry features of lunar terrain, this paper treats lunar terrain reconstruction as a surface reconstruction problem. We define an energy functional model consisting of local energy term and smooth energy term for lunar terrain reconstruction. The solution to minimize the functional (by partial differential equations) is defined as the optimal surface. In the smooth energy term, we design a vector field of depth discontinuousness likelihood (VFDDL) to control the direction and degree of smoothing. Experiments indicate that accurate VFDDL can lead to an exact reconstructed surface. Thus, VFDDL transfers 3D terrain reconstruction into a 2D image processing problem. An innovative method is proposed to estimate VFDDL, using image local and statistical features. Experiments verify our method and show a good performance in terrain reconstruction.
Ji Liu 0002, Yang Cong, Xiaomao Li, Yuechao Wang, Yandong Tang, Chuan Zhou 0010
ICIP2