VLDB 2026 Research / reviewers in the wild / expert
Chunyan Xu
dblp:70/8453
· DBLP profile ↗
93ranked-venue papers
14as first author
42since 2021 · last 2026
0000-0002-0814-4362ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 60 · 6 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 50 · 6 first-author · 23 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Noisy Candidates to Reliable Grounding in Weakly-Supervised Referring Expression ComprehensionabstractWeakly-Supervised Referring Expression Comprehension (WREC) aims to ground natural language expressions in image regions, using only image-text pairs without bounding-box annotations. However, the absence of explicit localization supervision results in noisy training signals and unreliable candidate selection. We observe that modern query-based detectors exhibit a high-recall, low-precision behavior in WREC: although the top-ranked query often localizes inaccurately, the ground-truth region is frequently recalled among high-confidence candidates. Motivated by this observation, we propose a Noisy-to-Reliable Grounding (NRG) framework that progressively transforms noisy high-recall candidates into reliable grounding supervision. Given an paired image-text data, we leverage a large Vision-Language Model (VLM) to generate confidence-aware soft pseudo labels, which provide robust semantic and spatial guidance under weak supervision. To disentangle true positives from noisy candidates, we introduce a contrastive candidate mining module that jointly exploits the detector’s prediction and VLM-derived cues to identify positive and hard negative queries, progressively enhancing grounding discriminability. Furthermore, the confidence-aware pseudo labels are integrated to construct a regression objective, enabling effective localization learning through low-rank adaptation of a pretrained open-vocabulary detector. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg datasets demonstrate that the proposed NRG consistently outperforms existing WREC methods. Ziqi Gu, Tong Zhang 0021, Zhen Cui 0001, Chunyan Xu |
ICMR | 5 |
| 2026 | Branch-adaptive mean-teacher: Reliable pseudo-labeling for semi-supervised medical image segmentation
Lei Li 0065, Yuanbin Zhou, Chunyan Xu, Zhuoli Dong, Tianli Liao, Yun Wang 0009 |
Expert Syst. Appl. | 3 |
| 2026 | Text-vision fusion and semantic adaptive labeling for compositional zero-shot learning
Run Shi, Chenyi Jiang, Chunyan Xu, Haofeng Zhang 0001 |
Expert Syst. Appl. | 3 |
| 2026 | A cross-guided multimodal fusion network for Alzheimer's disease classification
Lei Li 0065, Ruimin Guo, Tianli Liao, Zhuoli Dong, Chunyan Xu |
Pattern Recognit. Lett. | 5 |
| 2026 | Boosting Few-Shot Continual Learning via Self-Adaptive EvolutionabstractFew-shot continual learning (FSCL) has attracted increasing attention for real-world applications, where models must continuously adapt to new classes with only a few labeled samples while retaining prior knowledge. These abilities are essential in dynamic environments where data availability is often sparse and nonstationary. However, traditional FSCL methods are largely confined to closed data spaces, which limits their generalizability when diverse and evolving distributions are involved. Inspired by the paradigm of human lifelong learning, we propose a new self-adaptive evolution framework for FSCL that enables continuous interaction with and adaptation to external environments. To exploit latent knowledge in large-scale models, we use an adaptive diffusion-based generator that not only implicitly captures the distribution of new few-shot samples but also produces more high-quality samples. To mitigate the inevitable variability in generation quality, we also use a reinforced sample selection module, comprising a generated sample explorer and a selection evaluator, which explicitly guides the retained distributions toward alignment with the large-scale models. Integrated with the continual model, these components are optimized in an iterative self-adaptive evolution framework, ensuring stable knowledge retention while improving adaptability to newly emerging classes. We validate our approach through experiments on three benchmarks, revealing its effectiveness in exploiting external distributions and achieving notable performance improvements. Ziqi Gu, Chunyan Xu, Yuanzhi Wang, Cao Han, Di Xia, Zhen Cui 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Multi-clue Consistency Learning to Bridge Gaps Between General and Oriented Object in Semi-supervised DetectionabstractWhile existing semi-supervised object detection (SSOD) methods perform well in general scenes, they encounter challenges in handling oriented objects in aerial images. We experimentally find three gaps between general and oriented object detection in semi-supervised learning: 1) Sampling inconsistency: the common center sampling is not suitable for oriented objects with larger aspect ratios when selecting positive labels from labeled data. 2) Assignment inconsistency: balancing the precision and localization quality of oriented pseudo-boxes poses greater challenges which introduces more noise when selecting positive labels from unlabeled data. 3) Confidence inconsistency: there exists more mismatch between the predicted classification and localization qualities when considering oriented objects, affecting the selection of pseudo-labels. Therefore, we propose a Multi-clue Consistency Learning (MCL) framework to bridge gaps between general and oriented objects in semi-supervised detection. Specifically, considering various shapes of rotated objects, the Gaussian Center Assignment is specially designed to select the pixel-level positive labels from labeled data. We then introduce the Scale-aware Label Assignment to select pixel-level pseudo-labels instead of unreliable pseudo-boxes, which is a divide-and-rule strategy suited for objects with various scales. The Consistent Confidence Soft Label is adopted to further boost the detector by maintaining the alignment of the predicted results. Comprehensive experiments on DOTA-v1.5 and DOTA-v1.0 benchmarks demonstrate that our proposed MCL can achieve state-of-the-art performance in the semi-supervised oriented object detection task. Chunyan Xu, Xiang Li 0041, YuXuan Li, Ziqi Gu, Zhen Cui 0001 |
AAAI | 2 |
| 2025 | LLM-Assisted Semantic Guidance for Sparsely Annotated Remote Sensing Object DetectionabstractSparse annotation in remote sensing object detection poses significant challenges due to dense object distributions and category imbalances. Although existing Dense Pseudo-Label methods have demonstrated substantial potential in pseudo-labeling tasks, they remain constrained by selection ambiguities and inconsistencies in confidence estimation.In this paper, we introduce an LLM-assisted semantic guidance framework tailored for sparsely annotated remote sensing object detection, exploiting the advanced semantic reasoning capabilities of large language models (LLMs) to distill high-confidence pseudo-labels.By integrating LLM-generated semantic priors, we propose a Class-Aware Dense Pseudo-Label Assignment mechanism that adaptively assigns pseudo-labels for both unlabeled and sparsely labeled data, ensuring robust supervision across varying data distributions. Additionally, we develop an Adaptive Hard-Negative Reweighting Module to stabilize the supervised learning branch by mitigating the influence of confounding background information. Extensive experiments on DOTA and HRSC2016 demonstrate that the proposed method outperforms existing single-stage detector-based frameworks, significantly improving detection performance under sparse annotations. Chunyan Xu, Zhen Cui 0001 |
ICCV | 2 |
| 2025 | Semantic Discrepancy-Aware Detector for Image Forgery IdentificationabstractWith the rapid advancement of image generation techniques, robust forgery detection has become increasingly imperative to ensure the trustworthiness of digital media. Recent research indicates that the learned semantic concepts of pre-trained models are critical for identifying fake images. However, the misalignment between the forgery and semantic concept spaces hinders the model's forgery detection performance. To address this problem, we propose a novel Semantic Discrepancy-aware Detector (SDD) that leverages reconstruction learning to align the two spaces at a fine-grained visual level. By exploiting the conceptual knowledge embedded in the pre-trained vision language model, we specifically design a semantic token sampling module to mitigate the space shifts caused by features irrelevant to both forgery traces and semantic concepts. A concept-level forgery discrepancy learning module, built upon a visual reconstruction paradigm, is proposed to strengthen the interaction between visual semantic concepts and forgery traces, effectively capturing discrepancies under the concepts' guidance. Finally, the low-level forgery feature enhancemer integrates the learned concept level forgery discrepancies to minimize redundant forgery information. Experiments conducted on two standard image forgery datasets demonstrate the efficacy of the proposed SDD, which achieves superior results compared to existing methods. The code is available at https://github.com/wzy1111111/SSD. Minghang Yu, Chunyan Xu, Zhen Cui 0001 |
ICCV | 3 |
| 2025 | CLIP-driven Few-Shot Continual LearningabstractFew-Shot continual learning (FSCL) has garnered significant attention and has made notable progress in recent years. Drawing inspiration from the continuous interaction of humans with their environment in the lifelonglearning [1], we propose a novel CLIP-driven Few-Shot Continual Learning (C-FSCL) framework that aims to progressively enhance the continual network by leveraging the capabilities of a foundational CLIP model. When the continual model encounters new-coming data, we introduce a hierarchical feature-aware alignment module to better generalize knowledge gained from the foundational CLIP model, allowing it to apply learned concepts to new tasks more effectively. Recognizing the instability of inter-class structural relationships, which leads to catastrophic forgetting, we further introduce a relation-aware structure alignment module to align the consistent relationships from the CLIP model, thereby enhancing the reliability of the continual model’s knowledge structure. Experiments demonstrate the feasibility of CLIP-driven FSCL task and its remarkable performance on three benchmarks. Ziqi Gu, Chunyan Xu, Zhen Cui 0001 |
ICME | 2 |
| 2025 | Dual-Perspective United Transformer for Object Segmentation in Optical Remote Sensing ImagesabstractAutomatically segmenting objects from optical remote sensing images (ORSIs) is an important task. Most existing models are primarily based on either convolutional or Transformer features, each offering distinct advantages. Exploiting both advantages is valuable research, but it presents several challenges, including the heterogeneity between the two types of features, high complexity, and large parameters of the model. However, these issues are often overlooked in existing the ORSIs methods, causing sub-optimal segmentation. For that, we propose a novel Dual-Perspective United Transformer (DPU-Former) with a unique structure designed to simultaneously integrate long-range dependencies and spatial details. In particular, we design the global-local mixed attention, which captures diverse information through two perspectives and introduces a Fourier-space merging strategy to obviate deviations for efficient fusion. Furthermore, we present a gated linear feed-forward network to increase the expressive ability. Additionally, we construct a DPU-Former decoder to aggregate and strength features at different layers. Consequently, the DPU-Former model outperforms the state-of-the-art methods on multiple datasets. Code: https://github.com/CSYSI/DPU-Former. Jiexi Yan, Jianjun Qian, Chunyan Xu, Jian Yang 0003, Lei Luo 0001 |
IJCAI | 4 |
| 2025 | Learn and Ensemble Bridge Adapters for Multi-domain Task Incremental LearningabstractMulti-domain task incremental learning (MTIL) demands models to master domain-specific expertise while preserving generalization capabilities.
Inspired by human lifelong learning, which relies on revisiting, aligning, and integrating past experiences, we propose a Learning and Ensembling Bridge Adapters (LEBA) framework.
To facilitate cohesive knowledge transfer across domains, specifically, we propose a continuous-domain bridge adaptation module, leveraging the distribution transfer capabilities of Schrödinger bridge for stable progressive learning.
To strengthen memory consolidation, we further propose a progressive knowledge ensemble strategy that revisits past task representations via a diffusion model and dynamically integrates historical adapters.
For efficiency, LEBA maintains a compact adapter pool through similarity-based selection and employs learnable weights to align replayed samples with current task semantics.
Together, these components effectively mitigate catastrophic forgetting and enhance generalization across tasks.
Extensive experiments across multiple benchmarks validate the effectiveness and superiority of LEBA over state-of-the-art methods. Ziqi Gu, Chunyan Xu, Xin Liu 0011, Yide Qiu, Zhen Cui 0001 |
NeurIPS | 2 |
| 2025 | A Generalized Contour Vibration Model for Building Extraction
Chunyan Xu, Shuaizhen Yao, Zhen Cui 0001, Jian Yang 0003 |
Int. J. Comput. Vis. | 1 |
| 2025 | Correction: A Generalized Contour Vibration Model for Building Extraction
Chunyan Xu, Shuaizhen Yao, Zhen Cui 0001, Jian Yang 0003 |
Int. J. Comput. Vis. | 1 |
| 2025 | Adaptive exploration for few-shot incremental learning
Cao Han, Ziqi Gu, Chunyan Xu, Zhen Cui 0001 |
Knowl. Based Syst. | 3 |
| 2025 | Collaborative contrastive learning for cross-domain gaze estimation
Lifan Xia, Yong Li 0032, Zhen Cui 0001, Chunyan Xu, Antoni B. Chan |
Pattern Recognit. | 5 |
| 2024 | Frequency-Spatial Entanglement Learning for Camouflaged Object Detection
Chunyan Xu, Jian Yang 0003, Hanyu Xuan, Lei Luo 0001 |
ECCV (6) | 2 |
| 2024 | Progressive Exploration-Conformal Learning for Sparsely Annotated Object Detection in Aerial ImagesabstractThe ability to detect aerial objects with limited annotation is pivotal to the development of real-world aerial intelligence systems. In this work, we focus on a demanding but practical sparsely annotated object detection (SAOD) in aerial images, which encompasses a wider variety of aerial scenes with the same number of annotated objects. Although most existing SAOD methods rely on fixed thresholding to filter pseudo-labels for enhancing detector performance, adapting to aerial objects proves challenging due to the imbalanced probabilities/confidences associated with predicted aerial objects. To address this problem, we propose a novel Progressive Exploration-Conformal Learning (PECL) framework to address the SAOD task, which can adaptively perform the selection of high-quality pseudo-labels in aerial images. Specifically, the pseudo-label exploration can be formulated as a decision-making paradigm by adopting a conformal pseudo-label explorer and a multi-clue selection evaluator. The conformal pseudo-label explorer learns an adaptive policy by maximizing the cumulative reward, which can decide how to select these high-quality candidates by leveraging their essential characteristics and inter-instance contextual information. The multi-clue selection evaluator is designed to evaluate the explorer-guided pseudo-label selections by providing an instructive feedback for policy optimization. Finally, the explored pseudo-labels can be adopted to guide the optimization of aerial object detector in a closed-looping progressive fashion. Comprehensive evaluations on two public datasets demonstrate the superiority of our PECL when compared with other state-of-the-art methods in the sparsely annotated aerial object detection task. Zihan Lu, Chunyan Xu, Xiangwei Zheng 0001, Zhen Cui 0001 |
NeurIPS | 3 |
| 2024 | MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image GenerationabstractRecently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote sensing (RS) images that are tremendously different from general images in terms of scale and perspective remains a formidable challenge due to the lack of a comprehensive remote sensing image generation dataset with various modalities, ground sample distances (GSD), and scenes. In this paper, we propose a Multi-modal, Multi-GSD, Multi-scene Remote Sensing (MMM-RS) dataset and benchmark for text-to-image generation in diverse remote sensing scenarios. Specifically, we first collect nine publicly available RS datasets and conduct standardization for all samples. To bridge RS images to textual semantic information, we utilize a large-scale pretrained vision-language model to automatically output text prompts and perform hand-crafted rectification, resulting in information-rich text-image pairs (including multi-modal images). In particular, we design some methods to obtain the images with different GSD and various environments (e.g., low-light, foggy) in a single sample. With extensive manual screening and refining annotations, we ultimately obtain a MMM-RS dataset that comprises approximately 2.1 million text-image pairs. Extensive experimental results verify that our proposed MMM-RS dataset allows off-the-shelf diffusion models to generate diverse RS images across various modalities, scenes, weather conditions, and GSD. The dataset is available at https://github.com/ljl5261/MMM-RS. Jialin Luo, Yuanzhi Wang, Ziqi Gu, Yide Qiu, Shuaizhen Yao, Fuyun Wang, Chunyan Xu, Zhen Cui 0001 |
NeurIPS | 7 |
| 2024 | An attribution graph-based interpretable method for CNNs
Xiangwei Zheng 0001, Chunyan Xu, Xuanchi Chen, Zhen Cui 0001 |
Neural Networks | 3 |
| 2024 | Linear Regression Problem Relaxations Solved by Nonconvex ADMM With Convergence AnalysisabstractIn this work, we focus on studying the differentiable relaxations of several linear regression problems, where the original formulations are usually both nonsmooth with one nonconvex term. Unfortunately, in most cases, the standard alternating direction method of multipliers (ADMM) cannot guarantee global convergence when addressing these kinds of problems. To address this issue, by smoothing the convex term and applying a linearization technique before designing the iteration procedures, we employ nonconvex ADMM to optimize challenging nonconvex-convex composite problems. In our theoretical analysis, we prove the boundedness of the generated variable sequence and then guarantee that it converges to a stationary point. Meanwhile, a potential function is derived from the augmented Lagrange function, and we further verify that the objective function is monotonically nonincreasing. Under the Kurdyka-Łojasiewicz (KŁ) property, the global convergence is analyzed step by step. Finally, experiments on face reconstruction, image classification, and subspace clustering tasks are conducted to show the superiority of our algorithms over several state-of-the-art ones. Hengmin Zhang, Junbin Gao, Jianjun Qian, Jian Yang 0003, Chunyan Xu, Bob Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Exploratory Inference Learning for Scribble Supervised Semantic SegmentationabstractScribble supervised semantic segmentation has achieved great advances in pseudo label exploitation, yet suffers insufficient label exploration for the mass of unannotated regions. In this work, we propose a novel exploratory inference learning (EIL) framework, which facilitates efficient probing on unlabeled pixels and promotes selecting confident candidates for boosting the evolved segmentation. The exploration of unannotated regions is formulated as an iterative decision-making process, where a policy searcher learns to infer in the unknown space and the reward to the exploratory policy is based on a contrastive measurement of candidates. In particular, we devise the contrastive reward with the intra-class attraction and the inter-class repulsion in the feature space w.r.t the pseudo labels. The unlabeled exploration and the labeled exploitation are jointly balanced to improve the segmentation, and framed in a close-looping end-to-end network. Comprehensive evaluations on the benchmark datasets (PASCAL VOC 2012 and PASCAL Context) demonstrate the superiority of our proposed EIL when compared with other state-of-the-art methods for the scribble-supervised semantic segmentation problem. Chuanwei Zhou, Zhen Cui 0001, Chunyan Xu, Cao Han, Jian Yang 0003 |
AAAI | 3 |
| 2023 | Progressive Bayesian Inference for Scribble-Supervised Semantic SegmentationabstractThe scribble-supervised semantic segmentation is an important yet challenging task in the field of computer vision. To deal with the pixel-wise sparse annotation problem, we propose a Progressive Bayesian Inference (PBI) framework to boost the performance of the scribble-supervised semantic segmentation, which can effectively infer the semantic distribution of these unlabeled pixels to guide the optimization of the segmentation network. The PBI dynamically improves the model learning from two aspects: the Bayesian inference module (i.e., semantic distribution learning) and the pixel-wise segmenter (i.e., model updating). Specifically, we effectively infer the semantic probability distribution of these unlabeled pixels with our designed Bayesian inference module, where its guidance is estimated through the Bayesian expectation maximization under the situation of partially observed data. The segmenter can be progressively improved under the joint guidance of the original scribble information and the learned semantic distribution. The segmenter optimization and semantic distribution promotion are encapsulated into a unified architecture where they could improve each other with mutual evolution in a progressive fashion. Comprehensive evaluations of several benchmark datasets demonstrate the effectiveness and superiority of our proposed PBI when compared with other state-of-the-art methods applied to the scribble-supervised semantic segmentation task. Chuanwei Zhou, Chunyan Xu, Zhen Cui 0001 |
AAAI | 2 |
| 2023 | Few-shot Continual Infomax LearningabstractFew-shot continual learning is the ability to continually train a neural network from a sequential stream of few-shot data. In this paper, we propose a Few-shot Continual Infomax Learning (FCIL) framework that makes a deep model to continually/incrementally learn new concepts from few labeled samples, relieving the catastrophic forgetting of past knowledge. Specifically, inspired by the theoretical definition of transfer entropy, we introduce a feature embedding infomax to effectively perform the few-shot learning, which can transfer the strong encoding capability of the base network to learn the feature embedding of these novel classes by maximizing the mutual information of different-level feature distributions. Further, considering that the learned knowledge in the human brain is a generalization of actual information and exists in a certain relational structure, we perform continual structure infomax learning to relieve the catastrophic forgetting problem in the continual learning process. The information structure of this learned knowledge can be preserved through maximizing the mutual information across these continual-changing relations of inter-classes. Comprehensive evaluations on CIFAR100, miniImageNet, and CUB200 datasets demonstrate the superiority of our FCIL when compared against state-of-the-art methods on the few-shot continual learning task. Ziqi Gu, Chunyan Xu, Jian Yang 0003, Zhen Cui 0001 |
ICCV | 2 |
| 2023 | Grassmann Graph Embedding for Few-Shot Class Incremental Learning
Ziqi Gu, Chunyan Xu, Zhen Cui 0001 |
PRCV (8) | 2 |
| 2023 | Learning cross-modal interaction for RGB-T tracking
Chunyan Xu, Zhen Cui 0001, Chaoqun Wang 0012, Chuanwei Zhou, Jian Yang 0003 |
Sci. China Inf. Sci. | 1 |
| 2023 | Prototype-guided Instance matching for multiple pedestrian tracking
Qiang Wang 0023, Wankou Yang, Chunyan Xu, Zhen Cui 0001 |
Neurocomputing | 4 |
| 2023 | A graph-based interpretability method for deep neural networks
Xiangwei Zheng 0001, Zhen Cui 0001, Chunyan Xu |
Neurocomputing | 5 |
| 2023 | Quality-aware pattern diffusion for video object segmentation
Chuanwei Zhou, Chunyan Xu, Jun Li 0027, Zhen Cui 0001, Jian Yang 0003 |
Neurocomputing | 2 |
| 2023 | Visual Micro-Pattern PropagationabstractStatistic observations demonstrate that visual feature patterns or structure patterns recur high-frequently within/across homo/heterogeneous images. Motivated by the interdependencies of visual patterns, we propose visual micro-pattern propagation (VMPP) to facilitate universal visual pattern learning. Especially, we present a graph framework to unify the conventional micro-pattern propagations in spatial, temporal, cross-modal and cross-task domains. A general formulation of pattern propagation named cross-graph model is presented under this framework, and accordingly a factorized version is derived for more efficient computation as well as better understanding. To correlate homo/heterogeneous patterns, in cross-graph we introduce two types of pattern relations from feature-level and structure-level. The structure pattern relation defines second-order visual connections for heterogeneous patterns by measuring first-order visual relations of homogeneous feature patterns. In virtue of the constructed first-/second-order connections, we design feature pattern diffusion and structure pattern diffusion to prop up various pattern propagation cases. To fulfill different pattern diffusions involved, further, we deeply study two fundamental visual problems, multi-task pixel-level prediction and online dual-modal object tracking, and accordingly propose two end-to-end pattern propagation networks by encapsulating and integrating some necessary diffusion modules therein. We conduct extensive experiments by dissecting every diffusion component as well as comparing numerous advanced methods. The experiments validate the effectiveness of our proposed various pattern diffusion ways and meantime report the state-of-the-art results on the two representative visual problems. Zhen Cui 0001, Chaoqun Wang 0012, Chunyan Xu, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Fast subspace clustering by learning projective block diagonal representation
Yesong Xu, Shuo Chen 0003, Jun Li 0027, Chunyan Xu, Jian Yang 0003 |
Pattern Recognit. | 4 |
| 2023 | Progressive Context-Dependent Inference for Object Detection in Remote Sensing ImageryabstractInspired by our observation that numerous objects of remote sensing imageries are extremely consistent in geometric characteristics (e.g., object sizes/angles/layouts), in this work, we propose a novel Progressive Context-dependent Inference (PCI) method to make full use of large-scope contextual cues for better localizing objects in remote sensing imagery. Especially, to represent candidate objects and their geometric distributions, we build all of them into candidate object graphs, and subsequently perform inference learning by diffusing contextual object information. To make the inference more credible, we progressively accumulate these historical learning experiences on both label prediction and location regression processes into the next stage of network evolution, where topology structures and attributes of candidate object graphs would be dynamically updated. The graph update and ground object detection are jointly encapsulated as a closed-looping learning process. Hereby the problem of multi-object localization is converted into a progressive construction of dynamic graphs. Extensive experiments on three public datasets demonstrate the superiority of our proposed method over other state-of-the-art methods for ground object detection in remote sensing imagery. Chunyan Xu, Zhen Cui 0001, Jian Yang 0003 |
IEEE Trans. Image Process. | 2 |
| 2023 | Spatial-Temporal Tensor Graph Convolutional Network for Traffic Speed PredictionabstractAccurate traffic speed prediction is crucial for the guidance and management of urban traffic, which at the same time requires a model with a satisfactory computational burden and memory space in applications. In this paper, we propose a factorized Spatial-Temporal Tensor Graph Convolutional Network for traffic speed prediction. Traffic networks are modeled and unified into a graph tensor that integrates spatial and temporal information simultaneously. We extend graph convolution into tensor space and propose a tensor graph convolution network to extract more discriminating features from spatial-temporal graph data. We further introduce Tucker decomposition and derive a factorized tensor convolution to reduce the computational burden, which performs separate filtering in small-scale space, time, and feature modes. Besides, we can benefit from noise suppression of traffic data when discarding those trivial components in the process of tensor decomposition. Extensive experiments on the three real-world datasets demonstrate that our method is more effective than traditional prediction methods, and achieves state-of-the-art performance. Xuran Xu, Tong Zhang 0021, Chunyan Xu, Zhen Cui 0001, Jian Yang 0003 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | CVNet: Contour Vibration Network for Building ExtractionabstractThe classic active contour model raises a great promising solution to polygon-based object extraction with the progress of deep learning recently. Inspired by the physical vibration theory, we propose a contour vibration network (CVNet) for automatic building boundary delineation. Different from the previous contour models, the CVNet originally roots in the force and motion principle of contour string. Through the infinitesimal analysis and Newton's second law, we derive the spatial-temporal contour vibration model of object shapes, which is mathematically reduced to second-order differential equation. To concretize the dynamic model, we transform the vibration model into the space of image features, and reparameterize the equation coefficients as the learnable state from feature domain. The contour changes are finally evolved in a progressive mode through the computation of contour vibration equation. Both the polygon contour evolution and the model optimization are modulated to form a close-looping end-to-end network. Comprehensive experiments on three datasets demonstrate the effectiveness and superiority of our CVNet over other baselines and state-of-the-art methods for the polygon-based building extraction. The code is available at https://github.com/xzq-njust/CVNet. Chunyan Xu, Zhen Cui 0001, Xiangwei Zheng 0001, Jian Yang 0003 |
CVPR | 2 |
| 2022 | Temporal Discriminative Micro-Expression Recognition via Graph Contrastive LearningabstractMicro-Expressions (MEs) are involuntary and im-perceptible facial movements that reflect the underlying emotions and inner activities. Recently, ME recognition technology has been widely used in several fields such as medical treatment. Due to the subtle variations among the video sequence and the limited training data, the ME recognition task still remains a challenging problem. Existing methods tend to address the ME recognition problem from two aspects: (1) Data augmentation and (2) Expression signal amplification. Few works realize the importance of temporal variation hidden in the ME sequence. Based on the above observation, we propose a Graph Contrastive Learning (GCL) framework to effectively perceive subtle temporal variation for robust ME recognition. Specifically, the strong spatial feature representation is captured through the transformer-based ME feature encoder. Then, the proposed GCL builds the graph structure for the ME sequence and introduces the graph convolution to model the temporal relationship. To capture and highlight the temporal variation hidden in the ME sequence, a contrastive learning framework is designed to discriminately learn the differences between the normal and the abnormal ME samples. Both quantitative and qualitative experimental results show the effectiveness and superiority of our method compared with the prior state-of-the-arts. Lingjie Lao, Menglin Liu, Chunyan Xu, Zhen Cui 0001 |
ICPR | 4 |
| 2022 | Instance-Wise Contrastive Learning for Multi-object Tracking
Qiyu Luo, Chunyan Xu |
PRCV (4) | 2 |
| 2022 | Direction-induced convolution for point cloud analysis
Chunyan Xu, Chuanwei Zhou, Zhen Cui 0001, Chunlong Hu |
Multim. Syst. | 2 |
| 2022 | Self-Teaching Video Object SegmentationabstractVideo object segmentation (VOS) is one of the most fundamental tasks for numerous sequent video applications. The crucial issue of online VOS is the drifting of segmenter when incrementally updated on continuous video frames under unconfident supervision constraints. In this work, we propose a self-teaching VOS (ST-VOS) method to make segmenter to learn online adaptation confidently as much as possible. In the segmenter learning at each time slice, the segment hypothesis and segmenter update are enclosed into a self-looping optimization circle such that they can be mutually improved for each other. To reduce error accumulation of the self-looping process, we specifically introduce a metalearning strategy to learn how to do this optimization within only a few iteration steps. To this end, the learning rates of segmenter are adaptively derived through metaoptimization in the channel space of convolutional kernels. Furthermore, to better launch the self-looping process, we calculate an initial mask map through part detectors and motion flow to well-establish a foundation for subsequent refinement, which could result in the robustness of the segmenter update. Extensive experiments demonstrate that this ST idea can boost the performance of baselines, and in the meantime, our ST-VOS achieves encouraging performance on the DAVIS16, Youtube-objects, DAVIS17, and SegTrackV2 data sets, where, in particular, the accuracy of 75.7% in J-mean metric is obtained on the multi-instance DAVIS17 data set. Chuanwei Zhou, Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Learning Normal Dynamics in Videos With Meta Prototype NetworkabstractFrame reconstruction (current or future frame) based on Auto-Encoder (AE) is a popular method for video anomaly detection. With models trained on the normal data, the reconstruction errors of anomalous scenes are usually much larger than those of normal ones. Previous methods introduced the memory bank into AE, for encoding diverse normal patterns across the training videos. However, they are memory-consuming and cannot cope with unseen new scenarios in the testing data. In this work, we propose a dynamic prototype unit (DPU) to encode the normal dynamics as prototypes in real time, free from extra memory cost. In addition, we introduce meta-learning to our DPU to form a novel few-shot normalcy learner, namely Meta-Prototype Unit (MPU). It enables the fast adaption capability on new scenes by only consuming a few iterations of update. Extensive experiments are conducted on various benchmarks. The superior performance over the state-of-the-art demonstrates the effectiveness of our method. Our code is available at https://github.com/ktr-hubrt/MPN/. Chen Chen 0001, Zhen Cui 0001, Chunyan Xu, Yong Li 0044, Jian Yang 0003 |
CVPR | 4 |
| 2021 | Scribble-Supervised Semantic Segmentation InferenceabstractIn this paper, we propose a progressive segmentation inference (PSI) framework to tackle with scribble-supervised semantic segmentation. In virtue of latent contextual dependency, we encapsulate two crucial cues, contextual pattern propagation and semantic label diffusion, to enhance and refine pixel-level segmentation results from partially known seeds. In contextual pattern propagation, different-granular contextual patterns are correlated and leveraged to properly diffuse pattern information based on graphical model, so as to increase the inference confidence of pixel label prediction. Further, depending on high-confidence scores of estimated pixels, the initial annotated seeds are progressively spread over the image through dynamically learning an adaptive decision strategy. The two cues are finally modularized to form a close-looping update process during pixel-wise label inference. Extensive experiments demonstrate that our proposed progressive segmentation inference can benefit from the combination of spatial and semantic context cues, and meantime achieve the state-of-the-art performance on two public scribble segmentation datasets. Jingshan Xu, Chuanwei Zhou, Zhen Cui 0001, Chunyan Xu, Yuge Huang, Pengcheng Shen, Shaoxin Li 0001, Jian Yang 0003 |
ICCV | 4 |
| 2021 | Localizing Anomalies From Weakly-Labeled VideosabstractVideo anomaly detection under video-level labels is currently a challenging task. Previous works have made progresses on discriminating whether a video sequence contains anomalies. However, most of them fail to accurately localize the anomalous events within videos in the temporal domain. In this paper, we propose a Weakly Supervised Anomaly Localization (WSAL) method focusing on temporally localizing anomalous segments within anomalous videos. Inspired by the appearance difference in anomalous videos, the evolution of adjacent temporal segments is evaluated for the localization of anomalous segments. To this end, a high-order context encoding model is proposed to not only extract semantic representations but also measure the dynamic variations so that the temporal context could be effectively utilized. In addition, in order to fully utilize the spatial context information, the immediate semantics are directly derived from the segment representations. The dynamic variations as well as the immediate semantics, are efficiently aggregated to obtain the final anomaly scores. An enhancement strategy is further proposed to deal with noise interference and the absence of localization guidance in anomaly detection. Moreover, to facilitate the diversity requirement for anomaly detection benchmarks, we also collect a new traffic anomaly (TAD) dataset which specifies in the traffic conditions, differing greatly from the current popular anomaly detection evaluation benchmarks. Thedataset and the benchmark test codes, as well as experimental results, are made public on http://vgg-ai.cn/pages/Resource/ and https://github.com/ktr-hubrt/WSAL. Extensive experiments are conducted to verify the effectiveness of different components, and our proposed method achieves new state-of-the-art performance on the UCF-Crime and TAD datasets. Chuanwei Zhou, Zhen Cui 0001, Chunyan Xu, Yong Li 0032, Jian Yang 0003 |
IEEE Trans. Image Process. | 4 |
| 2021 | Meta-VOS: Learning to Adapt Online Target-Specific SegmentationabstractThe task of video object segmentation is a fundamental but challenging problem in the field of computer vision. To deal with large variations in target objects and background clutter, we propose an online adaptive video object segmentation (VOS) framework, named Meta-VOS, that learns to adapt the target-specific segmentation. Meta-VOS builds an online adaptive learning process by exploiting cumulative expertise after searching for confidence patterns across different videos/frames, and then dynamically improves the model learning from two aspects: Meta-seg learner (i.e., module updating) and Meta-seg criterion (i.e., rule of expertise). As our goal is to rapidly determine which patterns best represent the essential characteristics of specific targets in a video, Meta-seg learner is introduced to adaptively learn to update the parameters and hyperparameters of segmentation network in very few gradient descent steps. Furthermore, a Meta-seg criterion of learned expertise, which is constructed to evaluate the Meta-seg learner for the online adaptation of the segmentation network, can confidently online update positive/negative patterns under the guidance of motion cues, object appearances and learned knowledge. Comprehensive evaluations on several benchmark datasets demonstrate the superiority of our proposed Meta-VOS when compared with other state-of-the-art methods applied to the VOS problem. Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003 |
IEEE Trans. Image Process. | 1 |
| 2021 | Dual-Stream Structured Graph Convolution Network for Skeleton-Based Action RecognitionabstractIn this work, we propose a dual-stream structured graph convolution network ( DS-SGCN ) to solve the skeleton-based action recognition problem. The spatio-temporal coordinates and appearance contexts of the skeletal joints are jointly integrated into the graph convolution learning process on both the video and skeleton modalities. To effectively represent the skeletal graph of discrete joints, we create a structured graph convolution module specifically designed to encode partitioned body parts along with their dynamic interactions in the spatio-temporal sequence. In more detail, we build a set of structured intra-part graphs, each of which can be adopted to represent a distinctive body part (e.g., left arm, right leg, head). The inter-part graph is then constructed to model the dynamic interactions across different body parts; here each node corresponds to an intra-part graph built above, while an edge between two nodes is used to express these internal relationships of human movement. We implement the graph convolution learning on both intra- and inter-part graphs in order to obtain the inherent characteristics and dynamic interactions, respectively, of human action. After integrating the intra- and inter-levels of spatial context/coordinate cues, a convolution filtering process is conducted on time slices to capture these temporal dynamics of human motion. Finally, we fuse two streams of graph convolution responses in order to predict the category information of human action in an end-to-end fashion. Comprehensive experiments on five single/multi-modal benchmark datasets (including NTU RGB+D 60, NTU RGB+D 120, MSR-Daily 3D, N-UCLA, and HDM05) demonstrate that the proposed DS-SGCN framework achieves encouraging performance on the skeleton-based action recognition task. Chunyan Xu, Tong Zhang 0021, Zhen Cui 0001, Jian Yang 0003, Chunlong Hu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Variational Pathway Reasoning for EEG Emotion RecognitionabstractResearch on human emotion cognition revealed that connections and pathways exist between spatially-adjacent and functional-related areas during emotion expression (Adolphs 2002a; Bullmore and Sporns 2009). Deeply inspired by this mechanism, we propose a heuristic Variational Pathway Reasoning (VPR) method to deal with EEG-based emotion recognition. We introduce random walk to generate a large number of candidate pathways along electrodes. To encode each pathway, the dynamic sequence model is further used to learn between-electrode dependencies. The encoded pathways around each electrode are aggregated to produce a pseudo maximum-energy pathway, which consists of the most important pair-wise connections. To find those most salient connections, we propose a sparse variational scaling (SVS) module to learn scaling factors of pseudo pathways by using the Bayesian probabilistic process and sparsity constraint, where the former endows good generalization ability while the latter favors adaptive pathway selection. Finally, the salient pathways from those candidates are jointly decided by the pseudo pathways and scaling factors. Extensive experiments on EEG emotion recognition demonstrate that the proposed VPR is superior to those state-of-the-art methods, and could find some interesting pathways w.r.t. different emotions. Tong Zhang 0021, Zhen Cui 0001, Chunyan Xu, Wenming Zheng, Jian Yang 0003 |
AAAI | 3 |
| 2020 | Cross-Modal Pattern-Propagation for RGB-T TrackingabstractMotivated by our observations on RGB-T data that pattern correlations are high-frequently recurred across modalities also along sequence frames, in this paper, we propose a cross-modal pattern-propagation (CMPP) tracking framework to diffuse instance patterns across RGB-T data on spatial domain as well as temporal domain. To bridge RGB-T modalities, the cross-modal correlations on intra-modal paired pattern-affinities are derived to reveal those latent cues between heterogenous modalities. Through the correlations, the useful patterns may be mutually propagated between RGB-T modalities so as to fulfill inter-modal pattern-propagation. Further, considering the temporal continuity of sequence frames, we adopt the spirit of pattern propagation to dynamic temporal domain, in which long-term historical contexts are adaptively correlated and propagated into the current frame for more effective information inheritance. Extensive experiments demonstrate that the effectiveness of our proposed CMPP, and the new state-of-the-art results are achieved with the significant improvements on two RGB-T object tracking benchmarks. Chaoqun Wang 0012, Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003 |
CVPR | 2 |
| 2020 | Pattern-Structure Diffusion for Multi-Task LearningabstractInspired by the observation that pattern structures high-frequently recur within intra-task also across tasks, we propose a pattern-structure diffusion (PSD) framework to mine and propagate task-specific and task-across pattern structures in the task-level space for joint depth estimation, segmentation and surface normal prediction. To represent local pattern structures, we model them as small-scale graphlets, and propagate them in two different ways, i.e., intra-task and inter-task PSD. For the former, to overcome the limit of the locality of pattern structures, we use the high-order recursive aggregation on neighbors to multiplicatively increase the spread scope, so that long-distance patterns are propagated in the intra-task space. In the inter-task PSD, we mutually transfer the counterpart structures corresponding to the same spatial position into the task itself based on the matching degree of paired pattern structures therein. Finally, the intra-task and inter-task pattern structures are jointly diffused among the task-level patterns, and encapsulated into an end-to-end PSD network to boost the performance of multi-task learning. Extensive experiments on two widely-used benchmarks demonstrate that our proposed PSD is more effective and also achieves the state-of-the-art or competitive results. Zhen Cui 0001, Chunyan Xu, Zhenyu Zhang 0005, Chaoqun Wang 0012, Tong Zhang 0021, Jian Yang 0003 |
CVPR | 3 |
| 2020 | Graph inference learning for semi-supervised classification
Chunyan Xu, Zhen Cui 0001, Xiaobin Hong 0002, Tong Zhang 0021, Jian Yang 0003, Wei Liu 0005 |
ICLR | 1 |
| 2020 | Global Information Guided Video Anomaly DetectionabstractVideo anomaly detection (VAD) is currently a challenging task due to the complexity of "anomaly" as well as the lack of labor-intensive temporal annotations. In this paper, we propose an end-to-end Global Information Guided (GIG) anomaly detection framework for anomaly detection using the video-level annotations (i.e., weak labels). We propose to first mine the global pattern cues by leveraging the weak labels in a GIG module. Then we build a spatial reasoning module to measure the relevance between vectors in spatial domain with the global cue vectors, and select the most related feature vectors for temporal anomaly detection. The experimental results on the CityScene challenge demonstrate the effectiveness of our model. Chunyan Xu, Zhen Cui 0001 |
ACM Multimedia | 2 |
| 2020 | Fast Hyper-walk Gridded Convolution on Graph
Xiaobin Hong 0002, Tong Zhang 0021, Zhen Cui 0001, Chunyan Xu, Liangfang Zhang, Jian Yang 0003 |
PRCV (3) | 4 |
| 2020 | Joint Task-Recursive Learning for RGB-D Scene UnderstandingabstractRGB-D scene understanding under monocular camera is an emerging and challenging topic with many potential applications. In this paper, we propose a novel Task-Recursive Learning (TRL) framework to jointly and recurrently conduct three representative tasks therein containing depth estimation, surface normal prediction and semantic segmentation. TRL recursively refines the prediction results through a series of task-level interactions, where one-time cross-task interaction is abstracted as one network block of one time stage. In each stage, we serialize multiple tasks into a sequence and then recursively perform their interactions. To adaptively enhance counterpart patterns, we encapsulate interactions into a specific Task-Attentional Module (TAM) to mutually-boost the tasks from each other. Across stages, the historical experiences of previous states of tasks are selectively propagated into the next stages by using Feature-Selection unit (FS-Unit), which takes advantage of complementary information across tasks. The sequence of task-level interactions is also evolved along a coarse-to-fine scale space such that the required details may be refined progressively. Finally the task-abstracted sequence problem of multi-task prediction is framed into a recursive network. Extensive experiments on NYU-Depth v2 and SUN RGB-D datasets demonstrate that our method can recursively refines the results of the triple tasks and achieves state-of-the-art performance. Zhenyu Zhang 0005, Zhen Cui 0001, Chunyan Xu, Zequn Jie, Xiang Li 0041, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Hierarchical Semantic Propagation for Object Detection in Remote Sensing ImageryabstractObject detection in remote sensing imagery is a critical yet challenging task in the field of computer vision due to the bird's-eye-view perspective. Although existing object detection approaches in remote sensing imagery have achieved great advances through the utilization of deep features or rotation proposals, but they give insufficient consideration to multilevel semantic information and its propagation for guiding the learning process. Accordingly, in this article, we propose a hierarchical semantic propagation (HSP) framework to boost object detection performance in remote sensing imagery, which is better able to propagate hierarchical semantic information among different components in a unified network. Given a remote sensing image as input, the HSP framework can detect instances of semantic objects belonging to certain categories in an end-to-end way. First, the multiscale representation is captured by a basic feature pyramid network, which can hierarchically combine spatial attention details and the global semantic structure in order to learn more discriminative visual features. Second, the soft-segmentation prediction is used as an auxiliary objective in the intermediate layer of our HSP; its output instance-aware semantic information can be propagated to suppress noisy background information and thereby guide the proposal generation in the region proposal network. By further propagating this hierarchical semantic information into the region of interest module, we can then predict the object category information and the corresponding horizontal and oriented bounding boxes. Comprehensive evaluations on three benchmark data sets demonstrate the superiority of our HSP to the existing state-of-the-art methods for object detection in remote sensing imagery. Chunyan Xu, Chengzheng Li, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Walk-Steered Convolution for Graph ClassificationabstractGraph classification is a fundamental but challenging issue for numerous real-world applications. Despite recent great progress in image/video classification, convolutional neural networks (CNNs) cannot yet cater to graphs well because of graphical non-Euclidean topology. In this article, we propose a walk-steered convolutional (WSC) network to assemble the essential success of standard CNNs, as well as the powerful representation ability of random walk. Instead of deterministic neighbor searching used in previous graphical CNNs, we construct multiscale walk fields (a.k.a. local receptive fields) with random walk paths to depict subgraph structures and advocate graph scalability. To express the internal variations of a walk field, Gaussian mixture models are introduced to encode the principal components of walk paths therein. As an analogy to a standard convolution kernel on image, Gaussian models implicitly coordinate those unordered vertices/nodes and edges in a local receptive field after projecting to the gradient space of Gaussian parameters. We further stack graph coarsening upon Gaussian encoding by using dynamic clustering, such that high-level semantics of graph can be well learned like the conventional pooling on image. The experimental results on several public data sets demonstrate the superiority of our proposed WSC method over many state of the arts for graph classification. Jiatao Jiang, Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Wenming Zheng, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Gaussian-Induced Convolution for GraphsabstractLearning representation on graph plays a crucial role in numerous tasks of pattern recognition. Different from gridshaped images/videos, on which local convolution kernels can be lattices, however, graphs are fully coordinate-free on vertices and edges. In this work, we propose a Gaussianinduced convolution (GIC) framework to conduct local convolution filtering on irregular graphs. Specifically, an edgeinduced Gaussian mixture model is designed to encode variations of subgraph region by integrating edge information into weighted Gaussian models, each of which implicitly characterizes one component of subgraph variations. In order to coarsen a graph, we derive a vertex-induced Gaussian mixture model to cluster vertices dynamically according to the connection of edges, which is approximately equivalent to the weighted graph cut. We conduct our multi-layer graph convolution network on several public datasets of graph classification. The extensive experiments demonstrate that our GIC is effective and can achieve the state-of-the-art results. Jiatao Jiang, Zhen Cui 0001, Chunyan Xu, Jian Yang 0003 |
AAAI | 3 |
| 2019 | Hashing Graph Convolution for Node ClassificationabstractConvolution on graphs has aroused great interest in AI due to its potential applications to non-gridded data. To bypass the influence of ordering and different node degrees, the summation/average diffusion/aggregation is often imposed on local receptive field in most prior works. However, the collapsing into one node in this way tends to cause signal entanglements of nodes, which would result in a sub-optimal feature and decrease the discriminability of nodes. To address this problem, in this paper, we propose a simple but effective Hashing Graph Convolution (HGC) method by using global-hashing and local-projection on node aggregation for the task of node classification. In contrast to the conventional aggregation with a full collision, the hash-projection can greatly reduce the collision probability during gathering neighbor nodes. Another incidental effect of hash-projection is that the receptive field of each node is normalized into a common-size bucket space, which not only staves off the trouble of different-size neighbors and their order but also makes a graph convolution run like the standard shape-gridded convolution. Considering the few training samples, also, we introduce a prediction-consistent regularization term into HGC to constrain the score consistency of unlabeled nodes in the graph. HGC is evaluated on both transductive and inductive experimental settings and achieves new state-of-the-art results on all datasets for node classification task. The extensive experiments demonstrate the effectiveness of hash-projection. Wenting Zhao 0001, Zhen Cui 0001, Chunyan Xu, Chengzheng Li, Tong Zhang 0021, Jian Yang 0003 |
CIKM | 3 |
| 2019 | Pattern-Affinitive Propagation Across Depth, Surface Normal and Semantic SegmentationabstractIn this paper, we propose a novel Pattern-Affinitive Propagation (PAP) framework to jointly predict depth, surface normal and semantic segmentation. The motivation behind it comes from the statistic observation that pattern-affinitive pairs recur much frequently across different tasks as well as within a task. Thus, we can conduct two types of propagations, cross-task propagation and task-specific propagation, to adaptively diffuse those similar patterns. The former integrates cross-task affinity patterns to adapt to each task therein through the calculation on non-local relationships. Next the latter performs an iterative diffusion in the feature space so that the cross-task affinity patterns can be widely-spread within the task. Accordingly, the learning of each task can be regularized and boosted by the complementary task-level affinities. Extensive experiments demonstrate the effectiveness and the superiority of our method on the joint three tasks. Meanwhile, we achieve the state-of-the-art or competitive results on the three related datasets, NYUD-v2, SUN-RGBD and KITTI. Zhenyu Zhang 0005, Zhen Cui 0001, Chunyan Xu, Yan Yan 0002, Nicu Sebe, Jian Yang 0003 |
CVPR | 3 |
| 2019 | Feature-Attentioned Object Detection in Remote Sensing ImageryabstractIn this work, we introduce a novel feature-attentioned object detection framework to boost its performance in remote sensing imagery, which can focus on learning these intrinsic representations from different aspects in an end-to-end framework. Firstly, when fusing multi-scale visual features of backbone network, we adopt the channel-wise and pixel-wise attentions to enhance these object-related representations and weaken the background/noise information. Secondly, an adaptive multiple receptive fields attention mechanism is employed to generate horizontal region proposals under the special situation where objects in the remote sensing imagery are always with different aspect ratios. Finally, the proposal-level feature attention is proposed to better consider both multi-layer convolutional and apparent representations so that the region of interest network can better predict the object-wise category and its corresponding location information. Comprehensive evaluations on DOTA and UCAS-AOD datasets well demonstrate the effectiveness of our feature-attentioned network for object detection in remote sensing imagery. Chengzheng Li, Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003 |
ICIP | 2 |
| 2019 | Si-GCN: Structure-induced Graph Convolution Network for Skeleton-based Action RecognitionabstractIn recent years, the graph-convolution networks have been used to solve the problem of skeleton-based action recognition. Previous works often adopted a structure-fixed graph to model the physical joints of human skeleton, but cannot well consider these interactions of different human parts (e.g., the right arm and the left leg) to some extent. To deal with this problem, we propose a novel structure-induced graph convolution network (Si-GCN) framework to boost the performance of the skeleton-based action recognition task. Given a video sequence of human skeletons, the Si-GCN can produce the sample-wise category in an end-to-end way. Specifically, according to the natural divisions of human body, we define a collection of intra-part graphs for each input human skeleton (i.e., each graph denotes a specific part/global of human skeleton), and then formulate an inter-graph to model the relationships of different intra-part graphs. The Si-GCN framework, which will then perform the spectral graph convolutions on these constructed intra/inter-part graphs, can not only capture the internal modalities of each human part/subgraph, but also consider the interactions/relationships between different human parts. A temporal convolution follows to model the temporal and spatial dynamics of the skeleton in combination with the characteristics of time and space. Comprehensive evaluations on two public datasets (including NTU RGB+D and HDM05) well demonstrate the superiority of our proposed Si-GCN when compared with existing skeleton-based action recognition approaches. Chunyan Xu, Tong Zhang 0021, Wenting Zhao 0001, Zhen Cui 0001, Jian Yang 0003 |
IJCNN | 2 |
| 2019 | A jointly learned deep embedding for person re-identification
Caihong Yuan, Jingjuan Guo, Ping Feng, Chunyan Xu, Tianjiang Wang, Gwang-Min Choe, Kui Duan |
Neurocomputing | 5 |
| 2019 | Coupled-learning convolutional neural networks for object recognition
Chunyan Xu, Jian Yang 0003, Junbin Gao |
Multim. Tools Appl. | 1 |
| 2019 | Learning deep embedding with mini-cluster loss for person re-identification
Caihong Yuan, Jingjuan Guo, Ping Feng, Yihao Luo, Chunyan Xu, Tianjiang Wang, Kui Duan |
Multim. Tools Appl. | 6 |
| 2019 | UP-CNN: Un-pooling augmented convolutional neural network
Chunyan Xu, Jian Yang 0003, Hanjiang Lai, Junbin Gao, LinLin Shen, Shuicheng Yan |
Pattern Recognit. Lett. | 1 |
| 2019 | Spectral Filter TrackingabstractVisual object tracking is a challenging computer vision task with numerous real-world applications. In this paper, we propose a simple but efficient Spectral Filter Tracking (SFT) method from the view of graph, where each candidate image region is modeled as a pixelwise grid graph. Instead of the conventional graph matching, we formulate the tracking as a plain least square regression problem of learning spectral filters on graphs to predict an optimal vertex, which indicates the center of the target. To bypass computationally expensive eigenvalue decomposition on graph Laplacian L, we parameterize spectral graph filters as a polynomial of L to aggregate local graph features according to spectral graph theory, in which Lk exactly encodes a k-hop local neighborhood of each vertex. Thus, different from the holistic regression in those correlation filter based methods, SFT can operate on localized regions around a pixel (i.e., a vertex), which can effectively reduce the influence of local variations and cluttered backgrounds. Furthermore, we observe that the correlation filter tracking may be viewed as a specific case of our proposed spectral filtering method. The implementation of SFT can simply boil down to only a few line codes, but surprisingly it beats the correlation filter based model with the same feature input, and achieves the state-of-the-art performance on OTB-2015 and VOT2016 under the same feature extraction strategy. Zhen Cui 0001, Youyi Cai, Wenming Zheng, Chunyan Xu, Jian Yang 0003 |
IEEE Trans. Image Process. | 4 |
| 2019 | Efficient Recovery of Low-Rank Matrix via Double Nonconvex Nonsmooth Rank MinimizationabstractRecently, there is a rapidly increasing attraction for the efficient recovery of low-rank matrix in computer vision and machine learning. The popular convex solution of rank minimization is nuclear norm-based minimization (NNM), which usually leads to a biased solution since NNM tends to overshrink the rank components and treats each rank component equally. To address this issue, some nonconvex nonsmooth rank (NNR) relaxations have been exploited widely. Different from these convex and nonconvex rank substitutes, this paper first introduces a general and flexible rank relaxation function named weighted NNR relaxation function, which is actually derived from the initial double NNR (DNNR) relaxations, i.e., DNNR relaxation function acts on the nonconvex singular values function (SVF). An iteratively reweighted SVF optimization algorithm with continuation technology through computing the supergradient values to define the weighting vector is devised to solve the DNNR minimization problem, and the closed-form solution of the subproblem can be efficiently obtained by a general proximal operator, in which each element of the desired weighting vector usually satisfies the nondecreasing order. We next prove that the objective function values decrease monotonically, and any limit point of the generated subsequence is a critical point. Combining the Kurdyka-Łojasiewicz property with some milder assumptions, we further give its global convergence guarantee. As an application in the matrix completion problem, experimental results on both synthetic data and real-world data can show that our methods are competitive with several state-of-the-art convex and nonconvex matrix completion methods. Hengmin Zhang, Chen Gong 0002, Jianjun Qian, Bob Zhang 0001, Chunyan Xu, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | Scalable Proximal Jacobian Iteration Method With Global Convergence Analysis for Nonconvex Unconstrained Composite Optimizationsabstract-norm and rank function minimization problems. However, due to the absence of convexity in these nonconvex problems, developing efficient algorithms with convergence guarantee becomes very challenging. Inspired by the basic ideas of both the Jacobian alternating direction method of multipliers (JADMMs) for solving linearly constrained problems with separable objectives and the proximal gradient methods (PGMs) for optimizing the unconstrained problems with one variable, this paper focuses on extending the PGMs to the proximal Jacobian iteration methods (PJIMs) for handling with a family of nonconvex composite optimization problems with two splitting variables. To reduce the total computational complexity by decreasing the number of iterations, we devise the accelerated version of PJIMs through the well-known Nesterov's acceleration strategy and further extend both to solve the multivariable cases. Most importantly, we provide a rigorous convergence analysis, in theory, to show that the generated variable sequence globally converges to a critical point by exploiting the Kurdyka-Łojasiewica (KŁ) property for a broad class of functions. Furthermore, we also establish the linear and sublinear convergence rates of the obtained variable sequence in the objective function. As the specific application to the nonconvex sparse and low-rank recovery problems, several numerical experiments can verify that the newly proposed algorithms not only keep fast convergence speed but also have high precision. Hengmin Zhang, Jianjun Qian, Junbin Gao, Jian Yang 0003, Chunyan Xu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | Spatio-Temporal Graph Convolution for Skeleton Based Action RecognitionabstractVariations of human body skeletons may be considered as dynamic graphs, which are generic data representation for numerous real-world applications. In this paper, we propose a spatio-temporal graph convolution (STGC) approach for assembling the successes of local convolutional filtering and sequence learning ability of autoregressive moving average. To encode dynamic graphs, the constructed multi-scale local graph convolution filters, consisting of matrices of local receptive fields and signal mappings, are recursively performed on structured graph data of temporal and spatial domain. The proposed model is generic and principled as it can be generalized into other dynamic models. We theoretically prove the stability of STGC and provide an upper-bound of the signal transformation to be learnt. Further, the proposed recursive model can be stacked into a multi-layer architecture. To evaluate our model, we conduct extensive experiments on four benchmark skeleton-based action datasets, including the large-scale challenging NTU RGB+D. The experimental results demonstrate the effectiveness of our proposed model and the improvement over the state-of-the-art. Chaolong Li, Zhen Cui 0001, Wenming Zheng, Chunyan Xu, Jian Yang 0003 |
AAAI | 4 |
| 2018 | Joint Task-Recursive Learning for Semantic Segmentation and Depth Estimation
Zhenyu Zhang 0005, Zhen Cui 0001, Chunyan Xu, Zequn Jie, Xiang Li 0041, Jian Yang 0003 |
ECCV (10) | 3 |
| 2018 | Action Recognition with Spatial-Temporal Representation Analysis Across Grassmannian Manifold and Euclidean SpaceabstractAction recognition plays an important character for numerous tasks of video area. Although previous works often learn the appearance and motion information with Convolutional Neural Networks (CNNs), they ignore the corresponding space structures of video representation. In this work, we address action recognition task with a Spatial-Temporal representation analysis algorithm Across Grassmannian manifold and Euclidean space (ST-AGE), which considers the appearance and motion information of video samples in an unified framework. For each video sample, we extract temporal features with classical CNNs (e.g., ConvNet, VGG, ResNet) and motion representation with the trajectory tracking method. Both spatial and temporal information can be then analyzed by embedding them on the Grassmannian manifold and Euclidean space, and an appropriate multi-kernel SVM is further conducted. Comprehensive evaluations on HMDB-51 and UCF-101 datasets demonstrate the significant superiority of STAGE over other state-of-the-art for human action recognition. Xinshu Qiao, Chuanwei Zhou, Chunyan Xu, Zhen Cui 0001, Jian Yang 0003 |
ICIP | 3 |
| 2018 | Context-Dependent Diffusion Network for Visual Relationship DetectionabstractVisual relationship detection can bridge the gap between computer vision and natural language for scene understanding of images. Different from pure object recognition tasks, the relation triplets of subject-predicate-object lie on an extreme diversity space, such asperson-behind-person andcar-behind-building, while suffering from the problem of combinatorial explosion. In this paper, we propose a context-dependent diffusion network (CDDN) framework to deal with visual relationship detection. To capture the interactions of different object instances, two types of graphs, word semantic graph and visual scene graph, are constructed to encode global context interdependency. The semantic graph is built through language priors to model semantic correlations across objects, whilst the visual scene graph defines the connections of scene objects so as to utilize the surrounding scene information. For the graph-structured data, we design a diffusion network to adaptively aggregate information from contexts, which can effectively learn latent representations of visual relationships and well cater to visual relationship detection in view of its isomorphic invariance to graphs. Experiments on two widely-used datasets demonstrate that our proposed method is more effective and achieves the state-of-the-art performance. Zhen Cui 0001, Chunyan Xu, Wenming Zheng, Jian Yang 0003 |
ACM Multimedia | 2 |
| 2018 | Joint Bayesian guided metric learning for end-to-end face verification
Chunyan Xu, Jian Yang 0003, Jianjun Qian, Yuhui Zheng, LinLin Shen |
Neurocomputing | 2 |
| 2018 | A deep features based generative model for visual tracking
Ping Feng, Chunyan Xu, Fang Liu 0011, Jingjuan Guo, Caihong Yuan, Tianjiang Wang, Kui Duan |
Neurocomputing | 2 |
| 2018 | SRNN: Self-regularized neural network
Chunyan Xu, Jian Yang 0003, Junbin Gao, Hanjiang Lai, Shuicheng Yan |
Neurocomputing | 1 |
| 2018 | Low-rank structure preserving for unsupervised feature selection
Chunyan Xu, Jian Yang 0003, Junbin Gao, Fa Zhu |
Neurocomputing | 2 |
| 2018 | Large-scale semantic web image retrieval using bimodal deep learning techniques
Changqin Huang, Haijiao Xu, Liang Xie 0001, Jia Zhu 0003, Chunyan Xu, Yong Tang 0001 |
Inf. Sci. | 5 |
| 2018 | Deep multi-instance learning for end-to-end person re-identification
Caihong Yuan, Chunyan Xu, Tianjiang Wang, Fang Liu 0011, Ping Feng, Jingjuan Guo |
Multim. Tools Appl. | 2 |
| 2018 | Deep hierarchical guidance and regularization learning for end-to-end depth estimation
Zhenyu Zhang 0005, Chunyan Xu, Jian Yang 0003, Ying Tai, Liang Chen 0003 |
Pattern Recognit. | 2 |
| 2018 | Deep Recurrent Regression for Facial Landmark DetectionabstractWe propose a novel end-to-end deep architecture for face landmark detection, based on a deep convolutional and deconvolutional network followed by carefully designed recurrent network structures. The pipeline of this architecture consists of three parts. Through the first part, we encode an input face image to resolution-preserved deconvolutional feature maps via a deep network with stacked convolutional and deconvolutional layers. Then, in the second part, we estimate the initial coordinates of the facial key points by an additional convolutional layer on top of these deconvolutional feature maps. In the last part, by using the deconvolutional feature maps and the initial facial key points as input, we refine the coordinates of the facial key points by a recurrent network that consists of multiple long short-term memory components. Extensive evaluations on several benchmark data sets show that the proposed deep architecture has superior performance against the state-of-the-art methods. Hanjiang Lai, Shengtao Xiao, Yan Pan 0002, Zhen Cui 0001, Jiashi Feng, Chunyan Xu, Jian Yin 0001, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2018 | Action-Attending Graphic Neural NetworkabstractThe motion analysis of human skeletons is crucial for human action recognition, which is one of the most active topics in computer vision. In this paper, we propose a fully end-to-end action-attending graphic neural network (A2GNN) for skeleton-based action recognition, in which each irregular skeleton is structured as an undirected attribute graph. To extract high-level semantic representation from skeletons, we perform the local spectral graph filtering on the constructed attribute graphs like the standard image convolution operation. Considering not all joints are informative for action analysis, we design an actionattending layer to detect those salient action units (AUs) by adaptively weighting skeletal joints. Herein the filtering responses are parameterized into a weighting function irrelevant to the order of input nodes. To further encode continuous motion variations, the deep features learnt from skeletal graphs are gathered along consecutive temporal slices and then fed into a recurrent gated network. Finally, the spectral graph filtering, action-attending and recurrent temporal encoding are integrated together to jointly train for the sake of robust action recognition as well as the intelligibility of human actions. To evaluate our A2GNN, we conduct extensive experiments on four benchmark skeletonbased action datasets, including the large-scale challenging NTU RGB+D dataset. The experimental results demonstrate that our network achieves the state-of-the-art performances. Chaolong Li, Zhen Cui 0001, Wenming Zheng, Chunyan Xu, Rongrong Ji, Jian Yang 0003 |
IEEE Trans. Image Process. | 4 |
| 2018 | Progressive Hard-Mining Network for Monocular Depth EstimationabstractDepth estimation from the monocular RGB image is a challenging task for computer vision due to no reliable cues as the prior knowledge. Most existing monocular depth estimation works including various geometric or network learning methods lack of an effective mechanism to preserve the cross-border details of depth maps, which yet is very important for the performance promotion. In this paper, we propose a novel end-to-end progressive hard-mining network (PHN) framework to address this problem. Specifically, we construct the hard-mining objective function, the intra-scale and inter-scale refinement subnetworks to accurately localize and refine those hard-mining regions. The intra-scale refining block recursively recovers details of depth maps from different semantic features in the same receptive field while the inter-scale block favors a complementary interaction among multi-scale depth cues of different receptive fields. For further reducing the uncertainty of the network, we design a difficulty-ware refinement loss function to guide the depth learning process, which can adaptively focus on mining these hard-regions where accumulated errors easily occur. All three modules collaborate together to progressively reduce the error propagation in the depth learning process, and then, boost the performance of monocular depth estimation to some extent. We conduct comprehensive evaluations on several public benchmark data sets (including NYU Depth V2, KITTI, and Make3D). The experiment results well demonstrate the superiority of our proposed PHN framework over other state of the arts for monocular depth estimation task. Zhenyu Zhang 0005, Chunyan Xu, Jian Yang 0003, Junbin Gao, Zhen Cui 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | On Selecting Effective Patterns for Fast Support Vector Regression TrainingabstractIt is time consuming to train support vector regression (SVR) for large-scale problems even with efficient quadratic programming solvers. This issue is particularly serious when tuning the model's parameters. One way to address the issue is to reduce the problem's scale by selecting a subset of the training set. This paper presents a fast pattern selection method by scanning the training data set to reduce a problem's scale. In particular, we find the k-nearest neighbors (kNNs) in a local region around each pattern's target value, and then determine to retain the pattern according to the distribution of its nearest neighbors. There is a high probability that the pattern locates outside the -tube. Since the kNNs of a pattern are found in a very small region, it is fast to scan the whole training data set. The proposed method deals with the year prediction Million Song Data set, which contains 463 715 patterns, within 10 s on a personal computer with an Intel Core i5-4690 CPU at 3.50 GHz and 8GB DRAM. An additional advantage of the proposed method is that it can predefine the size of the selected subset according to the training set. Comprehensive empirical evaluations demonstrate that the proposed method can significantly eliminate redundant patterns for SVR training with only a slight decrease in performance. Fa Zhu, Junbin Gao, Chunyan Xu, Jian Yang 0003, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | MemNet: A Persistent Memory Network for Image RestorationabstractRecently, very deep convolutional neural networks (CNNs) have been attracting considerable attention in image restoration. However, as the depth grows, the longterm dependency problem is rarely realized for these very deep models, which results in the prior states/layers having little influence on the subsequent ones. Motivated by the fact that human thoughts have persistency, we propose a very deep persistent memory network (MemNet) that introduces a memory block, consisting of a recursive unit and a gate unit, to explicitly mine persistent memory through an adaptive learning process. The recursive unit learns multi-level representations of the current state under different receptive fields. The representations and the outputs from the previous memory blocks are concatenated and sent to the gate unit, which adaptively controls how much of the previous states should be reserved, and decides how much of the current state should be stored. We apply MemNet to three image restoration tasks, i.e., image denosing, super-resolution and JPEG deblocking. Comprehensive experiments demonstrate the necessity of the MemNet and its unanimous superiority on all three tasks over the state of the arts. Code is available at https://github.com/tyshiwo/MemNet. Ying Tai, Jian Yang 0003, Xiaoming Liu 0002, Chunyan Xu |
ICCV | 4 |
| 2017 | Sparse representation combined with context information for visual tracking
Ping Feng, Chunyan Xu, Fang Liu 0011, Caihong Yuan, Tianjiang Wang, Kui Duan |
Neurocomputing | 2 |
| 2017 | Convolutional neural networks for hyperspectral image classification
Shiqi Yu 0001, Sen Jia 0001, Chunyan Xu |
Neurocomputing | 3 |
| 2017 | Finding the samples near the decision plane for support vector learning
Fa Zhu, Jian Yang 0003, Junbin Gao, Chunyan Xu, Sheng Xu 0003, Cong Gao 0001 |
Inf. Sci. | 4 |
| 2017 | Human Parsing with Contextualized Convolutional Neural NetworkabstractIn this work, we address the human parsing task with a novel Contextualized Convolutional Neural Network (Co-CNN) architecture, which well integrates the cross-layer context, global image-level context, semantic edge context, within-super-pixel context and cross-super-pixel neighborhood context into a unified network. Given an input human image, Co-CNN produces the pixelwise categorization in an end-to-end way. First, the cross-layer context is captured by our basic local-to-global-to-local structure, which hierarchically combines the global semantic information and the local fine details across different convolutional layers. Second, the global image-level label prediction is used as an auxiliary objective in the intermediate layer of the Co-CNN, and its outputs are further used for guiding the feature learning in subsequent convolutional layers to leverage the global image-level context. Third, semantic edge context is further incorporated into Co-CNN, where the high-level semantic boundaries are leveraged to guide pixel-wise labeling. Finally, to further utilize the local super-pixel contexts, the within-super-pixel smoothing and cross-super-pixel neighbourhood voting are formulated as natural sub-components of the Co-CNN to achieve the local label consistency in both training and testing process. Comprehensive evaluations on two public datasets well demonstrate the significant superiority of our Co-CNN over other state-of-the-arts for human parsing. In particular, the F-1 score on the large dataset [1] reaches 81.72 percent by Co-CNN, significantly higher than 62.81 percent and 64.38 percent by the state-of-the-art algorithms, M-CNN [2] and ATR [1], respectively. By utilizing our newly collected large dataset for training, our Co-CNN can achieve 85.36 percent in F-1 score. Xiaodan Liang, Chunyan Xu, Xiaohui Shen, Jianchao Yang, Jinhui Tang 0001, Liang Lin 0004, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Deep Learning with S-Shaped Rectified Linear Activation UnitsabstractRectified linear activation units are important components for state-of-the-art deep convolutional networks. In this paper, we propose a novel S-shaped rectifiedlinear activation unit (SReLU) to learn both convexand non-convex functions, imitating the multiple function forms given by the two fundamental laws, namely the Webner-Fechner law and the Stevens law, in psychophysics and neural sciences. Specifically, SReLU consists of three piecewise linear functions, which are formulated by four learnable parameters. The SReLU is learned jointly with the training of the whole deep network through back propagation. During the training phase, to initialize SReLU in different layers, we propose a “freezing” method to degenerate SReLU into a predefined leaky rectified linear unit in the initial several training epochs and then adaptively learn the good initial values. SReLU can be universally used in the existing deep networks with negligible additional parameters and computation cost. Experiments with two popular CNN architectures, Network in Network and GoogLeNet on scale-various benchmarks including CIFAR10, CIFAR100, MNIST and ImageNet demonstrate that SReLU achieves remarkable improvement compared to other activation functions. Xiaojie Jin 0004, Chunyan Xu, Jiashi Feng, Yunchao Wei, Junjun Xiong, Shuicheng Yan |
AAAI | 2 |
| 2016 | Image auto-annotation via concept interdependency network
Haijiao Xu, Peng Pan 0001, Chunyan Xu, Yansheng Lu, Deng Chen |
Multim. Tools Appl. | 3 |
| 2016 | Extended nearest neighbor chain induced instance-weights for SVMs
Fa Zhu, Jian Yang 0003, Junbin Gao, Chunyan Xu |
Pattern Recognit. | 4 |
| 2016 | Multi-loss Regularized Deep Neural NetworkabstractA proper strategy to alleviate overfitting is critical to a deep neural network (DNN). In this paper, we introduce the cross-loss-function regularization for boosting the generalization capability of the DNN, which results in the multi-loss regularized DNN (ML-DNN) framework. For a particular learning task, e.g., image classification, only a single-loss function is used for all previous DNNs, and the intuition behind the multiloss framework is that the extra loss functions with different theoretical motivations (e.g., pairwise loss and LambdaRank loss) may drag the algorithm away from overfitting to one particular single-loss function (e.g., softmax loss). In the training stage, we pretrain the model with the single-core-loss function and then warm start the whole ML-DNN with the convolutional parameters transferred from the pretrained model. In the testing stage, the outputs by the ML-DNN from different loss functions are fused with average pooling to produce the ultimate prediction. The experiments conducted on several benchmark datasets (CIFAR-10, CIFAR-100, MNIST, and SVHN) demonstrate that the proposed ML-DNN framework, instantiated by the recently proposed network in network, considerably outperforms all other state-of-the-art methods. Chunyan Xu, Canyi Lu, Xiaodan Liang, Junbin Gao, Tianjiang Wang, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Generalized Singular Value ThresholdingabstractThis work studies the Generalized Singular Value Thresholding (GSVT) operator associated with a nonconvex function g defined on the singular values of X. We prove that GSVT can be obtained by performing the proximal operator of g on the singular values since Proxg(.) is monotone when g is lower bounded. If the nonconvex g satisfies some conditions (many popular nonconvex surrogate functions, e.g., lp-norm, 0 < p < 1, of l0-norm are special cases), a general solver to find Proxg(b) is proposed for any b ≥ 0. GSVT greatly generalizes the known Singular Value Thresholding (SVT) which is a basic subroutine in many convex low rank minimization methods. We are able to solve the nonconvex low rank minimization problem by using GSVT in place of SVT. Canyi Lu, Changbo Zhu, Chunyan Xu, Shuicheng Yan, Zhouchen Lin |
AAAI | 3 |
| 2015 | Human Parsing with Contextualized Convolutional Neural NetworkabstractIn this work, we address the human parsing task with a novel Contextualized Convolutional Neural Network (Co-CNN) architecture, which well integrates the cross-layer context, global image-level context, within-super-pixel context and cross-super-pixel neighborhood context into a unified network. Given an input human image, Co-CNN produces the pixel-wise categorization in an end-to-end way. First, the cross-layer context is captured by our basic local-to-global-to-local structure, which hierarchically combines the global semantic structure and the local fine details within the cross-layers. Second, the global image-level label prediction is used as an auxiliary objective in the intermediate layer of the Co-CNN, and its outputs are further used for guiding the feature learning in subsequent convolutional layers to leverage the global image-level context. Finally, to further utilize the local super-pixel contexts, the within-super-pixel smoothing and cross-super-pixel neighbourhood voting are formulated as natural sub-components of the Co-CNN to achieve the local label consistency in both training and testing process. Comprehensive evaluations on two public datasets well demonstrate the significant superiority of our Co-CNN architecture over other state-of-the-arts for human parsing. In particular, the F-1 score on the large dataset [15] reaches 76.95% by Co-CNN, significantly higher than 62.81% and 64.38% by the state-of-the-art algorithms, M-CNN [21] and ATR [15], respectively. Xiaodan Liang, Chunyan Xu, Xiaohui Shen, Jianchao Yang, Si Liu 0001, Jinhui Tang 0001, Liang Lin 0004, Shuicheng Yan |
ICCV | 2 |
| 2015 | Image retrieval based on multi-concept detector and semantic correlation
Haijiao Xu, Changqin Huang, Peng Pan 0001, Gansen Zhao, Chunyan Xu, Yansheng Lu, Deng Chen, Jiyi Wu |
Sci. China Inf. Sci. | 5 |
| 2015 | Facial Analysis With a Lie Group KernelabstractTo efficiently deal with the complex nonlinear variations of face images, a novel Lie group (LG) kernel is proposed in this paper to address the facial analysis problems. First, we present a linear dynamic model (LDM)-based face representation to capture both the appearance and spatial information of the face image. Second, the derived LDM can be parameterized as a specially structured upper triangular matrix, the space of which is proved to constitute an LG. An LG kernel is then designed to characterize the similarity between the LDMs for any two face images and the kernel can be fed into classical kernel-based classifiers for different types of facial analysis. Finally, experimental evaluations on face recognition and head pose estimation are conducted on several challenging data sets and the results show that the proposed algorithm outperforms other facial analysis methods. Chunyan Xu, Canyi Lu, Junbin Gao, Tianjiang Wang, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Discriminative Analysis for Symmetric Positive Definite Matrices on Lie GroupsabstractIn this paper, we study discriminative analysis of symmetric positive definite (SPD) matrices on Lie groups (LGs), namely, transforming an LG into a dimension-reduced one by optimizing data separability. In particular, we take the space of SPD matrices, e.g., covariance matrices, as a concrete example of LGs, which has proved to be a powerful tool for high-order image feature representation. The discriminative transformation of an LG is achieved by optimizing the within-class compactness as well as the between-class separability based on the popular graph embedding framework. A new kernel based on the geodesic distance between two samples in the dimension-reduced LG is then defined and fed into classical kernel-based classifiers, e.g., support vector machine, for various visual classification tasks. Extensive experiments on five public datasets, i.e., Scene-15, Caltech101, UIUC-Sport, MIT-Indoor, and VOC07, well demonstrate the effectiveness of discriminative analysis for SPD matrices on LGs, and the state-of-the-art performances are reported. Chunyan Xu, Canyi Lu, Junbin Gao, Tianjiang Wang, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | An Ordered-Patch-Based Image Classification Approach on the Image Grassmannian ManifoldabstractThis paper presents an ordered-patch-based image classification framework integrating the image Grassmannian manifold to address handwritten digit recognition, face recognition, and scene recognition problems. Typical image classification methods explore image appearances without considering the spatial causality among distinctive domains in an image. To address the issue, we introduce an ordered-patch-based image representation and use the autoregressive moving average (ARMA) model to characterize the representation. First, each image is encoded as a sequence of ordered patches, integrating both the local appearance information and spatial relationships of the image. Second, the sequence of these ordered patches is described by an ARMA model, which can be further identified as a point on the image Grassmannian manifold. Then, image classification can be conducted on such a manifold under this manifold representation. Furthermore, an appropriate Grassmannian kernel for support vector machine classification is developed based on a distance metric of the image Grassmannian manifold. Finally, the experiments are conducted on several image data sets to demonstrate that the proposed algorithm outperforms other existing image classification methods. Chunyan Xu, Tianjiang Wang, Junbin Gao, Shougang Cao, Wenbing Tao, Fang Liu 0011 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |