Congyan Lang

dblp:89/4275 · DBLP profile ↗
← Back
116ranked-venue papers
11as first author
60since 2021 · last 2026
0000-0001-6059-7943ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 60 · 5 first-author · 27 since 2021Artificial intelligence and machine learning · 50 · 6 first-author · 27 since 2021Databases, data management, data science and information retrieval · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Diversity-Aware Multi-prompt Learning for Compositional Zero-Shot Learning
Xinru Zhao, Congyan Lang
ICPR (2)3
2026 Adaptive spatial-temporal graph ODE networks for traffic flow forecasting
Shixiang Han, Xu Wang 0053, Yi Jin 0001, Songhe Feng, Congyan Lang, Yidong Li
Multim. Syst.5
2026 Reinforcement Learning-Based Sequential Parameter Tuning for Image Signal Processing
abstract
Hardware image signal processing (ISP) transforms RAW inputs into high-quality RGB images through a series of processing modules, each with numerous tunable parameters. Traditionally, these parameters are manually tuned by imaging experts, a time-consuming and subjective process. Recent deep learning approaches predict ISP parameters, but often treat the process as a black box and overlook the intrinsic relationships among ISP modules. To address these fundamental issues, we introduce a novel ISP parameter optimization model based on single-agent reinforcement learning (RL) (i.e., SARL-ISP), formulating the hardware ISP parameter tuning as a sequential optimization problem. During the optimization process, the agent updates ISP parameter tuning strategies for different tasks through interaction with the environment. In order to explore the influence of the sequential structure of hardware ISP modules and the coupling relationships among ISP parameters on the tuning process, we further propose a sequential ISP framework based on collaborative multi-agent RL (i.e., MARL-ISP). Specifically, the serialized parameter tuning module (SPTM) realistically simulates the process of manual prediction and module pipeline. Additionally, the feature selection module (FSM) facilitates the transmission and fusion of agent features, thereby selecting more appropriate feature inputs for downstream tasks. Extensive experiments across various tasks (e.g., object detection, instance segmentation) validate the effectiveness and efficiency of our models. Even with minimal training data, our models also outperform current state-of-the-art methods in both quantitative metrics and qualitative evaluations.
Bing Li 0001, Congyan Lang, Zhikun Zhao, Juan Wang 0012, Weihua Xiong, Weiming Hu 0004, Long Cheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 SHC: Deeply Activating Human-Like Cognitive Ability for Visual Question Answering
abstract
Human cognitive mechanism depends on a sophisticated information processing framework, including perception, attention, memory, language, reasoning, problem solving and decision-making. However, current research only focuses on isolated process rather than systematically simulating human cognitive mechanism. Meanwhile, with the rapid development of large language models, related works have predominantly centered on language-level exploration, while in-depth mining of visual information remains insufficient. Here, to deeply activate the multi-modal understanding ability, a Systematic Human-like Cognitive (SHC) method is proposed for visual question answering, where the above mentioned sophisticated seven processes are systematically modeled as three core modules: hierarchical perception, semantic refinement and dynamic reasoning. The Hierarchical Perception Module (HPM) extracts hierarchical features from different levels to simulate the incremental integration mode of biological neural system. Based on the selective attention theory, one Semantic Refinement Module (SRM) is designed as a key-value accumulation optimization mechanism that enhances high-level semantics from low-level features via a multi-level cascaded attention structure. Finally, the Dynamic Reasoning Module (DRM), following the utility maximization decision theory, employs a dual weighting mechanism to dynamically fuse high-level semantic features and low-level fine-grained features, forming a unified high-quality visual representation that is then fed into the large language model for reasoning together with the text input. Experimental results demonstrate that SHC achieves competitive performance on multiple visual question answering benchmarks, including VQA-v2, Text-VQA, GQA, and ScienceQA, as well as multimodal evaluation benchmarks such as POPE, MMB, MME, and MM-Vet. Comparative experiments with multiple models of the same-scale validate the latent capacity of SHC to prompt the performance of multi-modal understanding tasks and its superiority in fine-grained visual information perception, and even surpasses multimodal models with larger-scale on certain tasks.
Zhenxue Wang, Gaoyun An, Congyan Lang, Dapeng Oliver Wu
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 GrassNet: State space model meets graph neural network
Gongpei Zhao, Tao Wang 0011, Yi Jin 0001, Congyan Lang, Yidong Li, Haibin Ling
Pattern Recognit.4
2026 SinColor: Uncertainty-Guided Single-Step Diffusion for Image Colorization
abstract
Image colorization is a fundamental yet challenging task in computer vision, aiming to recover plausible and spatially coherent colors from grayscale images. Recent advancements in diffusion models have enabled significant progress in this field, yet existing methods predominantly rely on multi-step diffusion processes. While effective for generating high-frequency details, these approaches are suboptimal for colorization, as color information is inherently low-frequency, spatially smooth, and globally consistent. This mismatch leads to two critical limitations: 1) color artifacts and inconsistency due to excessive noise in the color space, and 2) high computational cost that hinders practical application. In this work, we propose a novel single-step diffusion framework for efficient and high-quality image colorization. We introduce a color uncertainty estimation (CUE) module to identify reliable and uncertain regions in the image, allowing the model to prioritize local certainty while reasoning about confused regions. To focus the model on low-frequency color generation, we directly encode the grayscale image into a latent representation, remove structural components in the output, and reconstruct the final image via efficient decoding. Extensive experiments on ImageNet, COCO-Stuff, and Extended COCO-Stuff demonstrate that our approach achieves state-of-the-art performance while reducing inference time by 98% and trainable parameters by 97% compared to leading multi-step diffusion methods. Our contributions include a systematic analysis of diffusion-based colorization, a lightweight yet effective uncertainty-aware framework, and comprehensive validation of its efficiency and effectiveness.
Yutong Gao 0001, Congyan Lang, Yidian Liu, Fayao Liu, Guoshun Nan, Yunchao Wei
IEEE Trans. Image Process.3
2026 Investigate Interactive Semantic Segmentation via an Uncertainty Mining View
abstract
With the rapid development of intelligence media, traditional semantic segmentation has shown excellent potential in application scenarios like autonomous driving. However, due to limited performance, traditional segmentation models usually lead to poor user experiences in applications that require high segmentation precision. Therefore, interactive semantic segmentation (ISS) is gaining the attention is gaining attention due to its capability to generate high-precision semantic segmentation results through a few user-provided clicks for experience improvement, which thus has a promising development prospect in fine-grained application scenarios,e.g., virtual reality, smart medical, data annotation,etc.. For good interaction efficiency, most existing interactive methods make efforts to conduct suitable click simulation strategies and reasonable click encoding methods, aiming at the robust understanding of diverse user clicks and translating comprehensible user intent,i.e., assign the correct category to the clicked area, for the neural network. Though proved effective, their designs ignore the uncertainty hiding in the extracted interaction features, which reflects the interaction difficulty and the user clicking intents. This can lead to inappropriate click simulation and click encoding, limiting the interaction efficiency. Hence we focus on exploring a reasonable ISS scheme via an uncertainty mining view. Specifically, we propose an uncertainty-based class-balanced click sampling (UCCS) simulation strategy by considering both the uncertainty of the click simulation region and its semantic imbalance, to form a reasonable click distribution. Furthermore, we propose a semantic uncertainty residual encoding (SURE) method to better embed the user's intention into the localization maps, by mining semantic confusion between the click and misprediction classes. We prove the effectiveness of our design through extensive experiments and initially analyze the importance of uncertainty mining for the ISS. Our model can achieve state-of-the-art performance on three semantic segmentation benchmarks.
Yutong Gao 0001, Congyan Lang, Fayao Liu, Xun Xu 0002, Yuanzhouhan Cao, Yunchao Wei
IEEE Trans. Multim.2
2025 A Hubness Perspective on Representation Learning for Graph-Based Multi-View Clustering
abstract
Recent graph-based multi-view clustering (GMVC) methods typically encode view features into high-dimensional spaces and construct graphs based on distance similarity. However, the high dimensionality of the embeddings often leads to the hubness problem, where a few points repeatedly appear in the nearest neighbor lists of other points. We show that this negatively impacts the extracted graph structures and message passing, thus degrading clustering performance. To the best of our knowledge, we are the first to highlight the detrimental effect of hubness in GMVC methods and introduce the hubREP (hub-aware Representation Embedding and Pairing) framework. Specifically, we propose a simple yet effective encoder that reduces hubness while preserving neighborhood topology within each view. Additionally, we propose a hub-aware pairing module to maintain structure consistency across views, efficiently enhancing the view-specific representations. The proposed hubREP is lightweight compared to the conventional autoencoders used in state-of-the-art GMVC methods and can be integrated into existing GMVC methods that mostly focus on novel fusion mechanisms, further boosting their performance. Comprehensive experiments performed on eight benchmarks confirm the superiority of our method. The code is available at https://github.com/zmxu196/hubREP.
Zheming Xu, Congyan Lang, Tao Wang 0011, Yidong Li, Michael Kampffmeyer
CVPR3
2025 Multimodal Large Language Model-Guided ISP Hyperparameter Optimization with Dynamic Preference Learning
Zhikun Zhao, Congyan Lang, Bing Li 0001, Juan Wang 0012
ICCV3
2025 Dynamic Dictionary Learning for Remote Sensing Image Segmentation
Xuechao Zou, Kai Li 0023, Pin Tao, Junliang Xing, Congyan Lang
ICCV8
2025 Knowledge Transfer and Domain Adaptation for Fine-Grained Remote Sensing Image Segmentation
abstract
Fine-Grained remote sensing image segmentation is essential for accurately identifying detailed objects in remote sensing images. Recently, vision transformer models (VTMs) pretrained on large-scale datasets have demonstrated strong zero-shot generalization. However, directly applying them to specific tasks may lead to domain shift. We introduce a novel end-to-end learning paradigm combining knowledge guidance with domain refinement to enhance performance. We present two key components: the Feature Alignment Module (FAM) and the Feature Modulation Module (FMM). FAM aligns features from a CNN-based backbone with those from the pretrained VTM’s encoder using channel transformation and spatial interpolation, and transfers knowledge via KL divergence and L2 normalization constraint. FMM further adapts the knowledge to the specific domain to address domain shift. We also introduce a fine-grained grass segmentation dataset and demonstrate, through experiments on two datasets, that our method achieves a significant improvement of 2.57 mIoU on the grass dataset and 3.73 mIoU on the cloud dataset. The results highlight the potential of combining knowledge transfer and domain adaptation to overcome domain-related challenges and data limitations. The project page is available at https://xavierjiezou.github.io/KTDA/.
Xuechao Zou, Kai Li 0023, Congyan Lang, Pin Tao
ICME4
2025 Integrating cross-graph consensus constraints for label disambiguation in multi-view partial multi-label learning
Zheming Xu, Congyan Lang, Songhe Feng
Appl. Intell.3
2025 UNAGI: Unified neighbor-aware graph neural network for multi-view clustering
Zheming Xu, Congyan Lang, Liqian Liang, Tao Wang 0011, Yidong Li, Michael Kampffmeyer
Neural Networks2
2025 Accelerated Self-Supervised Multi-Illumination Color Constancy With Hybrid Knowledge Distillation
abstract
Color constancy, the human visual system's ability to perceive consistent colors under varying illumination conditions, is crucial for accurate color perception. Recently, deep learning algorithms have been introduced into this task and have achieved remarkable achievements. However, existing methods are limited by the scale of current multi-illumination datasets and model size, hindering their ability to learn discriminative features effectively and their practical value for deployment in cameras. To overcome these limitations, this paper proposes a multi-illumination color constancy approach based on self-supervised learning and knowledge distillation. This approach includes three phases: self-supervised pre-training, supervised fine-tuning, and knowledge distillation. During the pre-training phase, we train Transformer-based and U-Net based encoders by two pretext tasks: light normalization task to learn lighting color contextual representation and grayscale colorization task to acquire objects' inherent color information. For the downstream color constancy task, we fine-tune the encoders and design a lightweight decoder to obtain better illumination distributions with fewer parameters. During the knowledge distillation phase, we introduce a hybrid knowledge distillation technique to align CNN features with those of Transformer and U-Net respectively. Our proposed method outperforms state-of-the-art techniques on multi-illumination and single-illumination benchmarks. Extensive ablation studies and visualizations confirm the effectiveness of our model.
Ziyu Feng, Bing Li 0001, Congyan Lang, Zheming Xu, Haina Qin, Juan Wang 0012, Weihua Xiong
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 The Cascaded Forward algorithm for neural network training
Gongpei Zhao, Tao Wang 0011, Yi Jin 0001, Congyan Lang, Yidong Li, Haibin Ling
Pattern Recognit.4
2025 Learning Diversified Primitive Prompts for Compositional Zero-Shot Learning
Xinru Zhao, Congyan Lang, Tao Wang 0011, Yidong Li
IEEE Trans. Circuits Syst. Video Technol.3
2025 Adapting Vision Foundation Models for Robust Cloud Segmentation in Remote Sensing Images
abstract
Cloud segmentation is a critical challenge in remote sensing image interpretation, as its accuracy directly impacts the effectiveness of subsequent data processing and analysis. Recently, vision foundation models (VFM) have demonstrated powerful generalization capabilities across various visual tasks. In this paper, we present a parameter-efficient adaptive approach, termed Cloud-Adapter, designed to enhance the accuracy and robustness of cloud segmentation. Our method leverages a VFM pretrained on general domain data, which remains frozen, eliminating the need for additional training. Cloud-Adapter incorporates a lightweight spatial perception module that initially utilizes a convolutional neural network (ConvNet) to extract dense spatial representations. These multi-scale features are then aggregated and serve as contextual inputs to an adapting module, which modulates the frozen transformer layers within the VFM. Experimental results demonstrate that the Cloud-Adapter approach, utilizing only 0.6% of the trainable parameters of the frozen backbone, achieves substantial performance gains. Cloud-Adapter consistently achieves state-of-the-art performance across various cloud segmentation datasets from multiple satellite sources, sensor series, data processing levels, land cover scenarios, and annotation granularities. Code and model checkpoints are available at https://xavierjiezou.github.io/Cloud-Adapter/.
Xuechao Zou, Kai Li 0023, Junliang Xing, Lei Jin 0003, Congyan Lang, Pin Tao
IEEE Trans. Geosci. Remote. Sens.7
2025 Ada3DLane: Adaptive 3D Lane Detection From Monocular Images
abstract
3D lane detection is a fundamental task in autonomous driving, and detecting 3D lanes from monocular images has been widely adopted due to the low computational cost and the property of lanes. Recent progresses have been made based on surrogate representations such as bird’s eye view (BEV) features. However, monocular BEV construction strictly relies on flat groud assumption, and the misalignment between perspective view and BEV is inevitable. In this paper, we propose Ada3DLane, a BEV-free, query based 3D lane detector, which adaptively generates queries with rich semantic and geometric information as well as lane interactions; adaptively samples perspective view image features in spatial and temporal domain; adaptively decode the sampled features for fast and accurate 3D lane detection. We conduct extensive experiments on benchmark lane detection datasets and outperforms previous state-of-the-art methods.
Zhiming Hou, Yuanzhouhan Cao, Naiyue Chen, Chao Ren 0002, Chunyu Lin, Congyan Lang, Yidong Li
IEEE Trans. Intell. Transp. Syst.6
2025 Deep Probabilistic Graph Matching
abstract
Most previous learning-based graph matching algorithms solve the quadratic assignment problem (QAP) by dropping one or more of the matching constraints and adopting a relaxed assignment solver to obtain sub-optimal correspondences. Such relaxation may actually weaken the original graph matching problem, and in turn hurt the matching performance. In this paper, we propose a deep learning-based graph matching framework that works for the original QAP without compromising on the matching constraints. In particular, we design an affinityassignment prediction network to jointly learn the pairwise affinity and estimate the node assignments, and we then develop a differentiable solver inspired by the probabilistic perspective of the pairwise affinities. Aiming to obtain better matching results, the probabilistic solver refines the estimated assignments in an iterative manner to impose both discrete and one-to-one matching constraints. The proposed method is trained in a supervised manner, evaluated on several benchmarks related to semantic keypoint corresponding, matching of social networks and pure QAP instances. In all experiment, it exhibits state-of-the-art matching performance on all benchmarks.
Tao Wang 0011, Congyan Lang, Yidong Li, Haibin Ling
IEEE Trans. Knowl. Data Eng.3
2025 Few-Shot 3D Point Cloud Segmentation via Relation Consistency-Guided Heterogeneous Prototypes
abstract
Few-shot 3D point cloud semantic segmentation is a challenging task due to the lack of labeled point clouds (support set). To segment unlabeled query point clouds, existing prototype-based methods learn 3D prototypes from point features of the support set and then measure their distances to the query points. However, such homogeneous 3D prototypes are often of low quality because they overlook the valuable heterogeneous information buried in the support set, such as semantic labels and projected 2D depth maps. To address this issue, in this paper, we propose a novel Relation Consistency-guided Heterogeneous Prototype learning framework (RCHP), which improves prototype quality by integrating heterogeneous information using large multi-modal models (e.g.CLIP). RCHP achieves this through two core components: Heterogeneous Prototype Generation module which collaborates with 3D networks and CLIP to generate heterogeneous prototypes, and Heterogeneous Prototype Fusion module which effectively fuses heterogeneous prototypes to obtain high-quality prototypes. Furthermore, to bridge the gap between heterogeneous prototypes, we introduce a Heterogeneous Relation Consistency loss, which transfers more reliable inter-class relations (i.e., inter-prototype relations) from refined prototypes to heterogeneous ones. Extensive experiments conducted on five point cloud segmentation datasets, including four indoor datasets (S3DIS, ScanNet, SceneNN, NYU Depth V2) and one outdoor dataset (Semantic3D), demonstrate the superiority and generalization capability of our method, outperforming state-of-the-art approaches across all datasets. The code will be released as soon as the paper is accepted.
Congyan Lang, Zheming Xu, Liqian Liang, Jun Liu 0036
IEEE Trans. Multim.2
2025 Mining Semantic Correlations Between Mispredictions and Corrections for Interactive Semantic Segmentation
abstract
Interactive semantic segmentation pursues high-quality segmentation results at the cost of a small number of user clicks. It is attracting more and more research attention for its convenience in labeling semantic pixel-level data. Existing interactive segmentation methods often pursue higher interaction efficiency by mining the latent information of user clicks or exploring efficient interaction manners. However, these works neglect to explicitly exploit the semantic correlations between user corrections and model mispredictions, thus suffering from two flaws. First, similar prediction errors frequently occur in actual use, causing users to repeatedly correct them. Second, the interaction difficulty of different semantic classes varies across images, but existing models use monotonic parameters for all images which lack semantic pertinence. Therefore, in this article, we explore the semantic correlations existing in corrections and mispredictions by proposing a simple yet effective online learning solution to the above problems, named correction-misprediction correlation mining (CM2). Specifically, we leverage the correction-misprediction similarities to design a confusion memory module (CMM) for automatic correction when similar prediction errors reappear. Furthermore, we measure the semantic interaction difficulty by counting the correction-misprediction pairs and design a challenge adaptive convolutional layer (CACL), which can adaptively switch different parameters according to interaction difficulties to better segment the challenging classes. Our method requires no extra training besides the online learning process and can effectively improve interaction efficiency. Our proposed CM2 achieves state-of-the-art results on three public semantic segmentation benchmarks.
Yutong Gao 0001, Congyan Lang, Fayao Liu, Chuan-Sheng Foo, Yuanzhouhan Cao, Yunchao Wei
IEEE Trans. Neural Networks Learn. Syst.2
2024 RL-SeqISP: Reinforcement Learning-Based Sequential Optimization for Image Signal Processing
abstract
Hardware image signal processing (ISP), aiming at converting RAW inputs to RGB images, consists of a series of processing blocks, each with multiple parameters. Traditionally, ISP parameters are manually tuned in isolation by imaging experts according to application-specific quality and performance metrics, which is time-consuming and biased towards human perception due to complex interaction with the output image. Since the relationship between any single parameter’s variation and the output performance metric is a complex, non-linear function, optimizing such a large number of ISP parameters is challenging. To address this challenge, we propose a novel Sequential ISP parameter optimization model, called the RL-SeqISP model, which utilizes deep reinforcement learning to jointly optimize all ISP parameters for a variety of imaging applications. Concretely, inspired by the sequential tuning process of human experts, the proposed model can progressively enhance image quality by seamlessly integrating information from both the image feature space and the parameter space. Furthermore, a dynamic parameter optimization module is introduced to avoid ISP parameters getting stuck into local optima, which is able to more effectively guarantee the optimal parameters resulting from the sequential learning strategy. These merits of the RL-SeqISP model as well as its high efficiency are substantiated by comprehensive experiments on a wide range of downstream tasks, including two visual analysis tasks (instance segmentation and object detection), and image quality assessment (IQA), as compared with representative methods both quantitatively and qualitatively. In particular, even using only 10% of the training data, our model outperforms other SOTA methods by an average of 7% mAP on two visual analysis tasks.
Zhikun Zhao, Congyan Lang, Mingxuan Cai, Longfei Han, Juan Wang 0012, Bing Li 0001
AAAI4
2024 Enhancing Multimedia Applications by Removing Dynamic Objects in Neural Radiance Fields
XianBen Yang, Tao Wang 0011, Yi Jin 0001, Congyan Lang, Yidong Li
ACCV (10)5
2024 Joint Homophily and Heterophily Relational Knowledge Distillation for Efficient and Compact 3D Object Detection
Shidi Chen, Liqian Liang, Congyan Lang
ACM Multimedia4
2024 Generated and Pseudo Content guided Prototype Refinement for Few-shot Point Cloud Segmentation
abstract
Few-shot 3D point cloud semantic segmentation aims to segment query point clouds with only a few annotated support point clouds. Existing prototype-based methods learn prototypes from the 3D support set to guide the segmentation of query point clouds. However, they encounter the challenge of low prototype quality due to constrained semantic information in the 3D support set and class information bias between support and query sets. To address these issues, in this paper, we propose a novel framework called Generated and Pseudo Content guided Prototype Refinement (GPCPR), which explicitly leverages LLM-generated content and reliable query context to enhance prototype quality. GPCPR achieves prototype refinement through two core components: LLM-driven Generated Content-guided Prototype Refinement (GCPR) and Pseudo Query Context-guided Prototype Refinement (PCPR). Specifically, GCPR integrates diverse and differentiated class descriptions generated by large language models to enrich prototypes with comprehensive semantic knowledge. PCPR further aggregates reliable class-specific pseudo-query context to mitigate class information bias and generate more suitable query-specific prototypes. Furthermore, we introduce a dual-distillation regularization term, enabling knowledge transfer between early-stage entities (prototypes or pseudo predictions) and their deeper counterparts to enhance refinement. Extensive experiments demonstrate the superiority of our method, surpassing the state-of-the-art methods by up to 12.10% and 13.75% mIoU on S3DIS and ScanNet, respectively.
Congyan Lang, Tao Wang 0011, Yidong Li, Jun Liu 0036
NeurIPS2
2024 DFA-GNN: Forward Learning of Graph Neural Networks by Direct Feedback Alignment
abstract
Graph neural networks (GNNs) are recognized for their strong performance across various applications, with the backpropagation (BP) algorithm playing a central role in the development of most GNN models. However, despite its effectiveness, BP has limitations that challenge its biological plausibility and affect the efficiency, scalability and parallelism of training neural networks for graph-based tasks. While several non-backpropagation (non-BP) training algorithms, such as the direct feedback alignment (DFA), have been successfully applied to fully-connected and convolutional network components for handling Euclidean data, directly adapting these non-BP frameworks to manage non-Euclidean graph data in GNN models presents significant challenges. These challenges primarily arise from the violation of the independent and identically distributed (i.i.d.) assumption in graph data and the difficulty in accessing prediction errors for all samples (nodes) within the graph. To overcome these obstacles, in this paper we propose DFA-GNN, a novel forward learning framework tailored for GNNs with a case study of semi-supervised learning. The proposed method breaks the limitations of BP by using a dedicated forward training mechanism. Specifically, DFA-GNN extends the principles of DFA to adapt to graph data and unique architecture of GNNs, which incorporates the information of graph topology into the feedback links to accommodate the non-Euclidean characteristics of graph data. Additionally, for semi-supervised graph learning tasks, we developed a pseudo error generator that spreads residual errors from training data to create a pseudo error for each unlabeled node. These pseudo errors are then utilized to train GNNs using DFA. Extensive experiments on 10 public benchmarks reveal that our learning framework outperforms not only previous non-BP methods but also the standard BP methods, and it exhibits excellent robustness against various types of noise and attacks.
Gongpei Zhao, Tao Wang 0011, Congyan Lang, Yi Jin 0001, Yidong Li, Haibin Ling
NeurIPS3
2024 Enhancing Point Cloud Sampling Quality with Dual-Branch Fusion Networks
abstract
Task-oriented point cloud sampling methods have attracted considerable attention for their ability to adaptively select important point sets based on downstream tasks, achieving an excellent balance between data simplification and task performance. However, existing task-oriented sampling models, primarily based on single-branch designs, struggle to fully extract features from input point clouds that comprehensively reflect multi-dimensional key information, thus limiting their sampling performance. In this paper, we introduce a dual-branch sampling network, named DBS-NET, which conducts crucial point sampling from both the global and local importance perspectives separately before merging them, thereby preserving multi-dimensional key information of the input data during the sampling process. Qualitative and quantitative experimental results demonstrate the competitive performance of DBS-NET on the classification benchmark task.
Yi Jin 0001, Xu Wang 0053, Mengxia Hu, Hui Yu 0001, Yidong Li, Tao Wang 0011, Songhe Feng, Congyan Lang
SMC8
2024 GLAN: A graph-based linear assignment network
Tao Wang 0011, Congyan Lang, Songhe Feng, Yi Jin 0001, Yidong Li
Pattern Recognit.3
2024 Linkage-Based Object Re-Identification via Graph Learning
abstract
Object Re-identification (Re-ID), which includes person Re-ID and vehicle Re-ID, is one of the core technologies of the intelligent transportation system. Existing supervised Re-ID studies mainly focus on discriminative feature learning (e.g., attention-based methods) or metric learning (e.g., triplet-loss-based methods) to obtain more accurate matches between the probe object and the positive gallery. However, they both pay less attention to global structure information (GSI) buried in the overall datasets. In this paper, we go beyond the traditional methods that are either unaware of or locally perceiving to GSI, and consider exploring the structural relationships among all the object instances of a dataset via a graph. Specifically, we construct a graph across the entire dataset, where each object instance is treated as a node and edges are assigned with the help of a classic algorithm like KNN. Seeing that a binary edge label can be used to predict whether its associated nodes belong to the same identity, we naturally formulate the problem of Re-ID as a new link prediction problem. Inspired by the superior capacity of capturing structure information of graph convolutional networks (GCN), a GCN-based global structure embedded network (GSE-Net) is proposed to take the graph as input and output a set of linkage likelihoods. During testing, we perform the evaluation according to the node features or estimated linkage likelihood via a graph where nodes include query and gallery images. Extensive experiments demonstrate that our proposed method outperforms the state-of-the-arts on both person and vehicle Re-ID benchmarks.
Zhenxue Wang, Congyan Lang, Liqian Liang, Tao Wang 0011, Songhe Feng, Yidong Li
IEEE Trans. Intell. Transp. Syst.3
2024 Dynamic Interaction Dilation for Interactive Human Parsing
abstract
Interactive segmentation pursues generating high-quality pixel-level predictions with a few user-provided clicks, which is gaining attention for its convenience in segmentation data annotation. Users are allowed to iteratively refine the prediction by adding clicks until the result is satisfactory. Existing interactive methods usually transform the clicks into a set of localization maps by Euclidian distance computation or RGB texture extraction to guide the segmentation, which makes the click transformation a core module in interactive segmentation networks. However, when adopted in human images where large poses, occlusions, and bad illuminations are prevailing, prior transformation methods tend to cause uncorrectable overlapping across localization maps which are difficult to form a good match among human parts. Furthermore, the inappropriately transformed information is hard to be refined with the static transformation manner which is out of tune with the dynamically refined interaction process. Hence, we design a dynamic transformation scheme for interactive human parsing (IHP) named Dynamic Interaction Dilation Net (DID-Net), which serves as an initial attempt to break the limitations of static transformation while capturing long-range dependencies of clicks within each human part. Specifically, we construct a Dynamic Dilation Module (DD-Module) to dilate clicks radially in several directions assisted by human body edge detection to refine the dilation quality in each interaction iteration. Furthermore, we propose an Adaptive Interaction Excitation Block (AIE-Block) to exploit potential semantic clues buried in the dilated clicks. Our DID-Net achieves state-of-the-art performance on 3 public human parsing benchmarks.
Yutong Gao 0001, Congyan Lang, Fayao Liu, Yuanzhouhan Cao, Yunchao Wei
IEEE Trans. Multim.2
2024 Neighborhood Pattern Is Crucial for Graph Convolutional Networks Performing Node Classification
abstract
Graph convolutional networks (GCNs) are widely believed to perform well in the graph node classification task, and homophily assumption plays a core rule in the design of previous GCNs. However, some recent advances on this area have pointed out that homophily may not be a necessity for GCNs. For deeper analysis of the critical factor affecting the performance of GCNs, we first propose a metric, namely, neighborhood class consistency (NCC), to quantitatively characterize the neighborhood patterns of graph datasets. Experiments surprisingly illustrate that our NCC is a better indicator, in comparison to the widely used homophily metrics, to estimate GCN performance for node classification. Furthermore, we propose a topology augmentation graph convolutional network (TA-GCN) framework under the guidance of the NCC metric, which simultaneously learns an augmented graph topology with higher NCC score and a node classifier based on the augmented graph topology. Extensive experiments on six public benchmarks clearly show that the proposed TA-GCN derives ideal topology with higher NCC score given the original graph topology and raw features, and it achieves excellent performance for semi-supervised node classification in comparison to several state-of-the-art (SOTA) baseline algorithms.
Gongpei Zhao, Tao Wang 0011, Yidong Li, Yi Jin 0001, Congyan Lang, Songhe Feng
IEEE Trans. Neural Networks Learn. Syst.5
2023 SMM: Self-supervised Multi-Illumination Color Constancy Model with Multiple Pretext Tasks
abstract
Color constancy is an important ability of the human visual system to perceive constant colors across different illumination. In this paper, we study a more practical yet challenging task, removing color cast by multiple spatial-varying illumination. Previous methods are limited by the scale of the current multi-illumination datasets, which hinders them from learning more discriminative features. Instead, we first propose a self-supervised multi-illumination color constancy model that leverages multiple pretext tasks to fully explore lighting color contextual information and inherent color information without using any manual annotations. During the pre-training phase, we train multiple Transformer-based encoders by learning multiple pretext tasks: (i) the local color distortion recovery task, which is carefully designed to learn lighting color contextual representation, and (ii) the colorization task, which is utilized to acquire inherent knowledge. In the downstream color constancy task, we fine-tune the encoders and design a lightweight decoder to obtain better illumination distributions with fewer parameters. Our lightweight architecture outperforms the state-of-the-art methods on the multi-illuminant benchmark (LSMI) and got robust performance on the single illuminant benchmark (NUS-8). Additionally, extensive ablation studies and visualization results demonstrate the effectiveness of integrating lighting color contextual and inherent color information learning in a self-supervised manner.
Ziyu Feng, Zheming Xu, Haina Qin, Congyan Lang, Bing Li 0001, Weihua Xiong
ACM Multimedia4
2023 Depth guided feature selection for RGBD salient object detection
Zun Li 0001, Congyan Lang, Guanqin Li, Tao Wang 0011, Yidong Li
Neurocomputing2
2023 Integrating topology beyond descriptions for zero-shot learning
Yutong Gao 0001, Congyan Lang, Yidong Li, Hongzhe Liu 0001, Fayao Liu
Pattern Recognit.3
2023 Joint Graph Learning and Matching for Semantic Feature Correspondence
Tao Wang 0011, Yidong Li, Congyan Lang, Yi Jin 0001, Haibin Ling
Pattern Recognit.4
2023 Prior Knowledge Regularized Self-Representation Model for Partial Multilabel Learning
abstract
Partial multilabel learning (PML) aims to learn from training data, where each instance is associated with a set of candidate labels, among which only a part is correct. The common strategy to deal with such a problem is disambiguation, that is, identifying the ground-truth labels from the given candidate labels. However, the existing PML approaches always focus on leveraging the instance relationship to disambiguate the given noisy label space, while the potentially useful information in label space is not effectively explored. Meanwhile, the existence of noise and outliers in training data also makes the disambiguation operation less reliable, which inevitably decreases the robustness of the learned model. In this article, we propose a prior label knowledge regularized self-representation PML approach, called PAKS, where the self-representation scheme and prior label knowledge are jointly incorporated into a unified framework. Specifically, we introduce a self-representation model with a low-rank constraint, which aims to learn the subspace representations of distinct instances and explore the high-order underlying correlation among different instances. Meanwhile, we incorporate prior label knowledge into the above self-representation model, where the prior label knowledge is regarded as the complement of features to obtain an accurate self-representation matrix. The core of PAKS is to take advantage of the data membership preference, which is derived from the prior label knowledge, to purify the discovered membership of the data and accordingly obtain more representative feature subspace for model induction. Enormous experiments on both synthetic and real-world datasets show that our proposed approach can achieve superior or comparable performance to state-of-the-art approaches.
Gengyu Lyu, Songhe Feng, Yi Jin 0001, Tao Wang 0011, Congyan Lang, Yidong Li
IEEE Trans. Cybern.5
2023 Redundant Label Learning via Subspace Representation and Global Disambiguation
abstract
Redundant Label Learning (RLL) aims at inducing a robust model from training data, where each example is associated with a set of candidate labels, among which some of them are incorrect. Most existing approaches deal with such problem by disambiguating the candidate labels first and then inducing the predictive model from the disambiguated data. However, these approaches only focus on disambiguation for each instance’ candidate label set, while the global label context tends to be ignored. Meanwhile, these approaches usually induce the objective model by directly utilizing the original feature information, which may lead to the model overfitting due to high-dimensional redundant features. To tackle the above issues, we propose a novel feature S ubspac E R epresentation and label G lobal Disambiguat IO n ( SERGIO ) approach, which improves the generalization ability of the learning system from the perspective of both feature space and label space. Specifically, we project the original high-dimensional feature space into a low-dimensional subspace, where the projection matrix is regularized with an orthogonality constraint to make the subspace more compact. Meanwhile, we introduce a label confidence matrix and constrain it with ℓ 1 -norm and trace-norm regularization simultaneously, which are utilized to explore global label correlations and further well in accordance with the nature of single-label classification and multi-label classification problem, respectively. Extensive experiments on both single-label and multi-label RLL datasets demonstrate that our proposed method achieves competitive performance against state-of-the-art approaches.
Gengyu Lyu, Songhe Feng, Wei Liu 0207, Shuoyan Liu, Congyan Lang
ACM Trans. Intell. Syst. Technol.5
2023 Distance-Preserving Embedding Adaptive Bipartite Graph Multi-View Learning with Application to Multi-Label Classification
abstract
Graph-based multi-view learning has attracted much attention due to the efficacy of fusing the information from different views. However, most of them exhibit high computational complexity. We propose an anchor-based bipartite graph embedding approach to accelerate the learning process. Specifically, different from existing anchor-based methods where anchors are obtained from key samples by clustering or weighted averaging strategies, in this article, the anchors are learned in a principled fashion which aims at constructing a distance-preserving embedding for each view from samples to their representations, whose elements are the weights of the edges linking corresponding samples and anchors. In addition, the consistency among different views can be explored by imposing a low-rank constraint on the concatenated embedding representations. We further design a concise yet effective feature collinearity guided feature selection scheme to learn tight multi-label classifiers. The objective function is optimized in an alternating optimization fashion. Both theoretical analysis and experimental results on different multi-label image datasets verify the effectiveness and efficiency of the proposed method.
Songhe Feng, Gengyu Lyu, Yi Jin 0001, Congyan Lang
ACM Trans. Knowl. Discov. Data5
2023 Beyond Word Embeddings: Heterogeneous Prior Knowledge Driven Multi-Label Image Classification
abstract
Multi-Label Image Classification (MLIC) is a fundamental yet challenging task which aims to recognize multiple labels from given images. The key to solve MLIC lies in how to accurately model the correlation between labels. Recent studies often adopt Graph Convolutional Network (GCN) to model label dependencies with word embeddings as prior knowledge. However, classical word embeddings typically contain redundant information due to the imperfect distributional hypothesis it relies on, which may degrade model generalizability. To tackle this problem, we propose a novel deep learning framework termedVisual-Semantic basedGraphConvolutionalNetwork (VSGCN), which alleviates the negative impact of redundant information by utilizing heterogeneous sources of prior knowledge. Specifically, we construct both visual prototype and semantic prototype for each label as heterogeneous prior label representations, which are further mapped to multi-label classifiers via two Multi-Head GCNs separately. The Multi-Head GCN mechanism proposed in this paper aims to guide the information propagation between prototypes for each label, which constructs multiple correlation graphs to simultaneously model the label correlation in different subspaces. Notably, we alleviate the negative influence of needless information by decreasing the inconsistency of predictions that come from visual space and semantic space. Extensive experiments conducted on various multi-label image datasets demonstrate the superiority of our proposed method.
Songhe Feng, Gengyu Lyu, Tao Wang 0011, Congyan Lang
IEEE Trans. Multim.5
2023 Clicking Matters: Towards Interactive Human Parsing
abstract
In this work, we focus on Interactive Human Parsing (IHP), which aims to segment a human image into multiple human body parts with guidance from users’ interactions. This new task inherits the class-aware property of human parsing, which cannot be well solved by traditional interactive image segmentation approaches that are generally class-agnostic. To tackle this new task, we first exploit user clicks to identify different human parts in the given image. These clicks are subsequently transformed into semantic-aware localization maps, which are concatenated with the RGB image to form the input of the segmentation network and generate the initial parsing result. To enable the network to better perceive user's purpose during the correction process, we investigate several principal ways for the refinement, and reveal that random-sampling-based click augmentation is the best way for promoting the correction effectiveness. Furthermore, we also propose a semantic-perceiving loss (SP-loss) to augment the training, which can effectively exploit the semantic relationships of clicks for better optimization. To the best knowledge, this work is the first attempt to tackle the human parsing task under the interactive setting. Our IHP solution achieves 85% mIoU on the benchmark LIP, 80% mIoU on PASCAL-Person-Part and CIHP, 75% mIoU on Helen with only 1.95, 3.02, 2.84 and 1.09 clicks per class respectively. These results demonstrate that we can simply acquire high-quality human parsing masks with only a few human effort. We hope this work can motivate more researchers to develop data-efficient solutions to IHP in the future.
Yutong Gao 0001, Liqian Liang, Congyan Lang, Songhe Feng, Yidong Li, Yunchao Wei
IEEE Trans. Multim.3
2023 SSR-Net: A Spatial Structural Relation Network for Vehicle Re-identification
abstract
Vehicle re-identification (Re-ID) represents the task aiming to identify the same vehicle from images captured by different cameras. Recent years have seen various feature learning-based approaches merely focusing on feature representations including global features or local features to obtain more subtle details to identify highly similar vehicles. However, few such methods consider the spatial geometrical structure relationship among local regions or between the global and local regions. By contrast, in this study, we propose a Spatial Structural Relation Network (SSR-Net) that explores the above-mentioned two kinds of relations simultaneously to learn more discriminative features by modeling the spatial structure information and global context information. In this article, we propose to adopt a Graph Convolution Network (GCN), for modeling spatial structural relationships among characteristic features. The GCN model aggregating the local and global features is shown to be more discriminative and robust to several car image transformations. To improve the performance of our proposed network, we jointly combine the classification loss with metric learning loss. Extensive experiments conducted on the public VehicleID and VeRi-776 datasets validate the effectiveness of our approach in comparison with recent works.
Zheming Xu, Congyan Lang, Songhe Feng, Tao Wang 0011, Adrian G. Bors, Hongzhe Liu 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Pedestrian attribute recognition based on attribute correlation
Ruijie Zhao 0007, Congyan Lang, Zun Li 0001, Liqian Liang, Songhe Feng, Tao Wang 0011
Multim. Syst.2
2022 Object detection by crossing relational reasoning based on graph neural network
XiuTing You, Tao Wang 0011, Songhe Feng, Congyan Lang
Mach. Vis. Appl.5
2022 Fine-grained facial expression recognition via relational reasoning and hierarchical relation optimization
Congyan Lang, Songhe Feng, Yidong Li
Pattern Recognit. Lett.3
2022 Dense Attentive Feature Enhancement for Salient Object Detection
abstract
Attention mechanisms have been proven highly effective for salient object detection. Most previous works utilize attention as a self-gated module to reweigh the feature maps at different levels independently. However, they are limited to certain-level guidance and could not satisfy the need of both accurately detecting intact objects and maintaining their detailed boundaries. In this paper, we build dense attention upon features from multiple levels simultaneously and propose a novel Dense Attentive Feature Enhancement (DAFE) module for efficient feature enhancement in saliency detection. DAFE stacks several attentional units and densely connects attentive feature output from current unit to its all subsequent units. This allows feature maps at deep units to absorb attentive information from shallow units, thus more discriminative information can be efficiently selected at the final output. Note that DAFE is plug and play, which can be effortlessly inserted into any saliency or video saliency models for their performance improvements. We further instantiate a highly effective Dense Attentive Feature Enhancement Network (DAFE-Net) for accurate salient object detection. DAFE-Net constructs DAFE over the aggregation feature that contains both semantics and saliency details, the entire salient objects and their boundaries can be well retained through dense attentions. Extensive experiments demonstrate that the proposed DAFE module is highly effective, and the DAFE-Net performs favorably compared with state-of-the-art approaches.
Zun Li 0001, Congyan Lang, Liqian Liang, Jian Zhao 0006, Songhe Feng, Qibin Hou, Jiashi Feng
IEEE Trans. Circuits Syst. Video Technol.2
2022 A Self-Paced Regularization Framework for Partial-Label Learning
abstract
Partial-label learning (PLL) aims to solve the problem where each training instance is associated with a set of candidate labels, one of which is the correct label. Most PLL algorithms try to disambiguate the candidate label set, by either simply treating each candidate label equally or iteratively identifying the true label. Nonetheless, existing algorithms usually treat all labels and instances equally, and the complexities of both labels and instances are not taken into consideration during the learning stage. Inspired by the successful application of a self-paced learning strategy in the machine-learning field, we integrate the self-paced regime into the PLL framework and propose a novel self-paced PLL (SP-PLL) algorithm, which could control the learning process to alleviate the problem by ranking the priorities of the training examples together with their candidate labels during each learning iteration. Extensive experiments and comparisons with other baseline methods demonstrate the effectiveness and robustness of the proposed method.
Gengyu Lyu, Songhe Feng, Tao Wang 0011, Congyan Lang
IEEE Trans. Cybern.4
2022 Weakly Supervised Video Object Segmentation via Dual-attention Cross-branch Fusion
abstract
Recently, concerning the challenge of collecting large-scale explicitly annotated videos, weakly supervised video object segmentation (WSVOS) using video tags has attracted much attention. Existing WSVOS approaches follow a general pipeline including two phases, i.e., a pseudo masks generation phase and a refinement phase. To explore the intrinsic property and correlation buried in the video frames, most of them focus on the later phase by introducing optical flow as temporal information to provide more supervision. However, these optical flow-based studies are greatly affected by illumination and distortion and lack consideration of the discriminative capacity of multi-level deep features. In this article, with the goal of capturing more effective temporal information and investigating a temporal information fusion strategy accordingly, we propose a unified WSVOS model by adopting a two-branch architecture with a multi-level cross-branch fusion strategy, named as dual-attention cross-branch fusion network (DACF-Net). Concretely, the two branches of DACF-Net, i.e., a temporal prediction subnetwork (TPN) and a spatial segmentation subnetwork (SSN), are used for extracting temporal information and generating predicted segmentation masks, respectively. To perform the cross-branch fusion between TPN and SSN, we propose a dual-attention fusion module that can be plugged into the SSN flexibly. We also pose a cross-frame coherence loss (CFCL) to achieve smooth segmentation results by exploiting the coherence of masks produced by TPN and SSN. Extensive experiments demonstrate the effectiveness of proposed approach compared with the state-of-the-arts on two challenging datasets, i.e., Davis-2016 and YouTube-Objects.
Congyan Lang, Liqian Liang, Songhe Feng, Tao Wang 0011, Shidi Chen
ACM Trans. Intell. Syst. Technol.2
2022 CMAN: Leaning Global Structure Correlation for Monocular 3D Object Detection
abstract
The key to 3D object detection is proper utilization of depth data. Compared with LiDAR based approaches, 3D object detection from a single image remains a challenging task due to the lack of structure information. Recent methods leverage monocular depth estimation as a way to produce 2D depth maps, and adopt the depth maps as additional source of input to explore structure information. However, these methods either encode local structure correlations, or encode long range structure correlations by iteratively passing local messages. In this work, we propose a cross modal attention network (CMAN) for monocular 3D object detection. It is built upon the self-attention module which learns attention map from single modal data. Our CMAN is able to encode structure correlations from depth data, and embed the structure correlations with appearance information which is learned from RGB data. Thanks to the attention learning mechanism, our CMAN learns global structure correlations without iteration. In order to reduce the computational burden, our CMAN adopts a novel node sampler to eliminate redundant nodes during the attention map calculation. Experiment results on benchmark KITTI3D dataset show that our proposed CMAN outperforms the state-of-the-art methods.
Yuanzhouhan Cao, Hui Zhang 0091, Yidong Li, Chao Ren 0002, Congyan Lang
IEEE Trans. Intell. Transp. Syst.5
2022 Global-Local Label Correlation for Partial Multi-Label Learning
abstract
Partial Multi-label Learning (PML) addresses the scenario where each instance is assigned with multiple candidate labels, while only a subset of the labels are relevant. This task is very challenging because the training procedure can be misguided by the noisy (irrelevant) labels. Exploiting label correlations is useful for partial multi-label learning. However, the existing PML methods often ignore to explicitly and sufficiently leverage the label correlation information for handling the noisy labels. To this end, in this paper, we propose a novelGlobal-Local Label Correlation (GLC) approach for partial multi-label learning. On one hand, we introduce a label coefficient matrix to explicitly exploit the global structure information of labels from multiple subspaces. On the other hand, we present a new label manifold regularizer to capture the local label correlations to further improve the performance of our method. By jointly taking advantage of the global and local label correlations, our proposed approach achieves superior performance on both the synthetic and real-world data sets from diverse domains.
Songhe Feng, Jun Liu 0036, Gengyu Lyu, Congyan Lang
IEEE Trans. Multim.5
2022 Seeing Crucial Parts: Vehicle Model Verification via a Discriminative Representation Model
abstract
Widely used surveillance cameras have promoted large amounts of street scene data, which contains one important but long-neglected object: the vehicle. Here we focus on the challenging problem of vehicle model verification. Most previous works usually employ global features (e.g., fully connected features) to further perform vehicle-level deep metric learning (e.g., triplet-based network). However, we argue that it is noteworthy to investigate the distinctiveness of local features and consider vehicle-part-level metric learning by reducing the intra-class variance as much as possible. In this article, we introduce a simple yet powerful deep model—the enforced intra-class alignment network (EIA-Net)—which can learn a more discriminative image representation by localizing key vehicle parts and jointly incorporating two distance metrics: vehicle-level embedding and vehicle-part-sensitive embedding. For learning features, we propose an effective feature extraction module that is composed of two components: the regional proposal network (RPN)-based network and part-based CNN. The RPN is used to define key vehicle regions and aggregate local features on these regions, whereas part-based CNN offers supplementary global features for the RPN-based network. The fusion features learned by feature extraction module are cast into the deep metric learning module. Especially, we derived an enforced intra-class alignment loss by re-utilizing key vehicle part information to enhance reducing intra-class variance. Furthermore, we modify the coupled cluster loss to model the vehicle-level embedding by enlarging the inter-class variance while shortening intra-class variance. Extensive experiments over benchmark datasets VehicleID and CompCars have shown that the proposed EIA-Net significantly outperforms the state-of-the-art approaches for vehicle model verification. Furthermore, we also conduct comprehensive experiments on vehicle re-identification datasets (i.e., VehicleID and VeRi776) to validate the generalization ability effectiveness of our proposed method.
Liqian Liang, Congyan Lang, Zun Li 0001, Jian Zhao 0006, Tao Wang 0011, Songhe Feng
ACM Trans. Multim. Comput. Commun. Appl.2
2021 PST-NET: Point Cloud Sampling via Point-Based Transformer
Xu Wang 0053, Yi Jin 0001, Yi-Gang Cen, Congyan Lang, Yidong Li
ICIG (3)4
2021 MSO: Multi-Feature Space Joint Optimization Network for RGB-Infrared Person Re-Identification
abstract
The RGB-infrared cross-modality person re-identification (ReID) task aims to recognize the images of the same identity between the visible modality and the infrared modality. Existing methods mainly use a two-stream architecture to eliminate the discrepancy between the two modalities in the final common feature space, which ignore the single space of each modality in the shallow layers. To solve it, in this paper, we present a novel multi-feature space joint optimization (MSO) network, which can learn modality-sharable features in both the single-modality space and the common space. Firstly, based on the observation that edge information is modality-invariant, we propose an edge features enhancement module to enhance the modality-sharable features in each single-modality space. Specifically, we design a perceptual edge features (PEF) loss after the edge fusion strategy analysis. According to our knowledge, this is the first work that proposes explicit optimization in the single-modality feature space on cross-modality ReID task. Moreover, to increase the difference between cross-modality distance and class distance, we introduce a novel cross-modality contrastive-center (CMCC) loss into the modality-joint constraints in the common feature space. The PEF loss and CMCC loss jointly optimize the model in an end-to-end manner, which markedly improves the network's performance. Extensive experiments demonstrate that the proposed model significantly outperforms state-of-the-art methods on both the SYSU-MM01 and RegDB datasets.
Yajun Gao, Tengfei Liang, Yi Jin 0001, Xiaoyan Gu 0001, Wu Liu 0005, Yidong Li, Congyan Lang
ACM Multimedia7
2021 Deep spatio-frequency saliency detection
Zun Li 0001, Congyan Lang, Tao Wang 0011, Yidong Li, Jiashi Feng
Neurocomputing2
2021 Entropy-aware self-training for graph convolutional networks
Gongpei Zhao, Tao Wang 0011, Yidong Li, Yi Jin 0001, Congyan Lang
Neurocomputing5
2021 Text to photo-realistic image synthesis via chained deep recurrent generative adversarial network
Congyan Lang, Songhe Feng, Tao Wang 0011, Yi Jin 0001, Yidong Li
J. Vis. Commun. Image Represent.2
2021 Fine-Grained Facial Expression Recognition in the Wild
abstract
Over the past decades, researches on facial expression recognition have been restricted within six basic expressions (anger, fear, disgust, happiness, sadness and surprise). However, these six words can not fully describe the richness and diversity of human beings' emotions. To enhance the recognitive capabilities for computers, in this paper, we focus on fine-grained facial expression recognition in the wild and build a brand new benchmark FG-Emotions to push the research frontiers on this topic, which extends the original six classes to more elaborate thirty-three classes. Our FG-Emotions contains 10,371 images and 1,491 video clips annotated with corresponding fine-grained facial expression categories and landmarks. FG-Emotions also provides several features (e.g., LBP features and dense trajectories features) to facilitate related research. Moreover, on top of FG-Emotions, we propose a new end-to-end Multi-Scale Action Unit (AU)-based Network (MSAU-Net) for facial expression recognition with image which learns a more powerful facial representation by directly focusing on locating facial action units and utilizing “zoom in” operation to aggregate distinctive local features. As for recognition with video, we further extend the MSAU-Net to a two-stream model (TMSAU-Net) by adding a module with attention mechanism and a temporal stream branch to jointly learn spatial and temporal features. (T)MSAU-Net consistently outperforms existing state-of-the-art solutions on our FG-Emotions and several other datasets, and serves as a strong baseline to drive the future research towards fine-grained facial expression recognition in the wild.
Liqian Liang, Congyan Lang, Yidong Li, Songhe Feng, Jian Zhao 0006
IEEE Trans. Inf. Forensics Secur.2
2021 Cross-Layer Feature Pyramid Network for Salient Object Detection
abstract
Feature pyramid network (FPN) based models, which fuse the semantics and salient details in a progressive manner, have been proven highly effective in salient object detection. However, it is observed that these models often generate saliency maps with incomplete object structures or unclear object boundaries, due to the indirect information propagation among distant layers that makes such fusion structure less effective. In this work, we propose a novel Cross-layer Feature Pyramid Network (CFPN), in which direct cross-layer communication is enabled to improve the progressive fusion in salient object detection. Specifically, the proposed network first aggregates multi-scale features from different layers into feature maps that have access to both the high- and low- level information. Then, it distributes the aggregated features to all the involved layers to gain access to richer context. In this way, the distributed features per layer own both semantics and salient details from all other layers simultaneously, and suffer reduced loss of important information during the progressive feature fusion. At last, CFPN fuses the distributed features of each layer stage-by-stage. This way, the high-level features that contain context useful for locating complete objects are preserved until the final output layer, and the low-level features that contain spatial structure details are embedded into each layer to preserve spatial structural details. Extensive experimental results over six widely used salient object detection benchmarks and with three popular backbones clearly demonstrate that CFPN can accurately locate fairly complete salient regions and effectively segment the object boundaries.
Zun Li 0001, Congyan Lang, Jun Hao Liew, Yidong Li, Qibin Hou, Jiashi Feng
IEEE Trans. Image Process.2
2021 Fine-Grained Semantic Image Synthesis with Object-Attention Generative Adversarial Network
abstract
Semantic image synthesis is a new rising and challenging vision problem accompanied by the recent promising advances in generative adversarial networks. The existing semantic image synthesis methods only consider the global information provided by the semantic segmentation mask, such as class label, global layout, and location, so the generative models cannot capture the rich local fine-grained information of the images (e.g., object structure, contour, and texture). To address this issue, we adopt a multi-scale feature fusion algorithm to refine the generated images by learning the fine-grained information of the local objects. We propose OA-GAN, a novel object-attention generative adversarial network that allows attention-driven, multi-fusion refinement for fine-grained semantic image synthesis. Specifically, the proposed model first generates multi-scale global image features and local object features, respectively, then the local object features are fused into the global image features to improve the correlation between the local and the global. In the process of feature fusion, the global image features and the local object features are fused through the channel-spatial-wise fusion block to learn ‘what’ and ‘where’ to attend in the channel and spatial axes, respectively. The fused features are used to construct correlation filters to obtain feature response maps to determine the locations, contours, and textures of the objects. Extensive quantitative and qualitative experiments on COCO-Stuff, ADE20K and Cityscapes datasets demonstrate that our OA-GAN significantly outperforms the state-of-the-art methods.
Congyan Lang, Liqian Liang, Songhe Feng, Tao Wang 0011, Yutong Gao 0001
ACM Trans. Intell. Syst. Technol.2
2021 GM-PLL: Graph Matching Based Partial Label Learning
abstract
Partial Label Learning (PLL) aims to learn from the data where each training example is associated with a set of candidate labels, among which only one is correct. The key to deal with such problem is to disambiguate the candidate label sets and obtain the correct assignments between instances and their candidate labels. In this paper, we interpret such assignments as instance-to-label matchings, and reformulate the task of PLL as a matching selection problem. To model such problem, we propose a novel Graph Matching based Partial Label Learning (GM-PLL) framework, where Graph Matching (GM) scheme is incorporated owing to its excellent capability of exploiting the instance and label relationship. Meanwhile, since conventional one-to-one GM algorithm does not satisfy the constraint of PLL problem that multiple instances may correspond to the same label, we extend a traditional one-to-one probabilistic matching algorithm to the many-to-one constraint, and make the proposed framework accommodate to the PLL problem. Moreover, we also propose a relaxed matching prediction model, which can improve the prediction accuracy via GM strategy. Extensive experiments on both artificial and real-world data sets demonstrate that the proposed method can achieve superior or comparable performance against the state-of-the-art methods.
Gengyu Lyu, Songhe Feng, Tao Wang 0011, Congyan Lang, Yidong Li
IEEE Trans. Knowl. Data Eng.4
2021 Multi-human Parsing with a Graph-based Generative Adversarial Model
abstract
Human parsing is an important task in human-centric image understanding in computer vision and multimedia systems. However, most existing works on human parsing mainly tackle the single-person scenario, which deviates from real-world applications where multiple persons are present simultaneously with interaction and occlusion. To address such a challenging multi-human parsing problem, we introduce a novel multi-human parsing model named MH-Parser, which uses a graph-based generative adversarial model to address the challenges of close-person interaction and occlusion in multi-human parsing. To validate the effectiveness of the new model, we collect a new dataset named Multi-Human Parsing (MHP), which contains multiple persons with intensive person interaction and entanglement. Experiments on the new MHP dataset and existing datasets demonstrate that the proposed method is effective in addressing the multi-human parsing problem compared with existing solutions in the literature.
Jianshu Li, Jian Zhao 0006, Congyan Lang, Yidong Li, Yunchao Wei, Guodong Guo, Terence Sim, Shuicheng Yan, Jiashi Feng
ACM Trans. Multim. Comput. Commun. Appl.3
2020 Attentive Generative Adversarial Network To Bridge Multi-Domain Gap For Image Synthesis
abstract
Despite the significant progress on text-to-image synthesis, automatically generating realistic images remains a challenging task since the location and specific shape of object are not given in the text descriptions. To address these problems, we propose a novel attentive generative adversarial network with contextual loss (AGAN-CL) algorithm. More specifically, the generative network consists of two sub-networks: a contextual network for generating image contours, and a cycle transformation autoencoder for converting contours to realistic images. Our core idea is the injection of image contours into the generative network, which is the most critical part of our network, since it will guide the whole generative network to focus on object regions. In addition, we also apply contextual loss and cycle-consistent loss to bridge multi-domain gap. Comprehensive results on several challenging datasets demonstrate the advantage of the proposed method over the leading approaches, regarding both visual fidelity and alignment with input descriptions.
Congyan Lang, Liqian Liang, Gengyu Lyu, Songhe Feng, Tao Wang 0011
ICME2
2020 Salient object detection with side information
Qiuning Li, Yidong Li, Congyan Lang
Sci. China Inf. Sci.3
2020 Improving person re-identification via attribute-identity representation and visual attention mechanism
Honglin Quan, Songhe Feng, Congyan Lang, Baifan Chen
Multim. Tools Appl.3
2020 HERA: Partial Label Learning by Combining Heterogeneous Loss with Sparse and Low-Rank Regularization
abstract
Partial label learning (PLL) aims to learn from the data where each training instance is associated with a set of candidate labels, among which only one is correct. Most existing methods deal with this type of problem by either treating each candidate label equally or identifying the ground-truth label iteratively. In this article, we propose a novel PLL approach named HERA, which simultaneously incorporates the HeterogEneous Loss and the SpaRse and Low-rAnk procedure to estimate the labeling confidence for each instance while training the desired model. Specifically, the heterogeneous loss integrates the strengths of both the pairwise ranking loss and the pointwise reconstruction loss to provide informative label ranking and reconstruction information for label identification, whereas the embedded sparse and low-rank scheme constrains the sparsity of ground-truth label matrix and the low rank of noise label matrix to explore the global label relevance among the whole training data, for improving the learning model. Comprehensive ablation study demonstrates the effectiveness of our employed heterogeneous loss, and extensive experiments on both artificial and real-world datasets demonstrate that our method achieves superior or comparable performance against state-of-the-art methods.
Gengyu Lyu, Songhe Feng, Yidong Li, Yi Jin 0001, Guojun Dai, Congyan Lang
ACM Trans. Intell. Syst. Technol.6
2020 End-to-End Text-to-Image Synthesis with Spatial Constrains
abstract
Although the performance of automatically generating high-resolution realistic images from text descriptions has been significantly boosted, many challenging issues in image synthesis have not been fully investigated, due to shapes variations, viewpoint changes, pose changes, and the relations of multiple objects. In this article, we propose a novel end-to-end approach for text-to-image synthesis with spatial constraints by mining object spatial location and shape information. Instead of learning a hierarchical mapping from text to image, our algorithm directly generates multi-object fine-grained images through the guidance of the generated semantic layouts. By fusing text semantic and spatial information into a synthesis module and jointly fine-tuning them with multi-scale semantic layouts generated, the proposed networks show impressive performance in text-to-image synthesis for complex scenes. We evaluate our method both on single-object CUB dataset and multi-object MS-COCO dataset. Comprehensive experimental results demonstrate that our method significantly outperforms the state-of-the-art approaches consistently across different evaluation metrics.
Congyan Lang, Liqian Liang, Songhe Feng, Tao Wang 0011, Yutong Gao 0001
ACM Trans. Intell. Syst. Technol.2
2020 Learning part-alignment feature for person re-identification with spatial-temporal-based re-ranking method
Yi Jin 0001, Yidong Li, Congyan Lang, Songhe Feng, Tao Wang 0011
World Wide Web4
2019 Partial Multi-Label Learning by Low-Rank and Sparse Decomposition
abstract
Multi-Label Learning (MLL) aims to learn from the training data where each example is represented by a single instance while associated with a set of candidate labels. Most existing MLL methods are typically designed to handle the problem of missing labels. However, in many real-world scenarios, the labeling information for multi-label data is always redundant , which can not be solved by classical MLL methods, thus a novel Partial Multi-label Learning (PML) framework is proposed to cope with such problem, i.e. removing the the noisy labels from the multi-label sets. In this paper, in order to further improve the denoising capability of PML framework, we utilize the low-rank and sparse decomposition scheme and propose a novel Partial Multi-label Learning by Low-Rank and Sparse decomposition (PML-LRS) approach. Specifically, we first reformulate the observed label set into a label matrix, and then decompose it into a groundtruth label matrix and an irrelevant label matrix, where the former is constrained to be low rank and the latter is assumed to be sparse. Next, we utilize the feature mapping matrix to explore the label correlations and meanwhile constrain the feature mapping matrix to be low rank to prevent the proposed method from being overfitting. Finally, we obtain the ground-truth labels via minimizing the label loss, where the Augmented Lagrange Multiplier (ALM) algorithm is incorporated to solve the optimization problem. Enormous experimental results demonstrate that PML-LRS can achieve superior or competitive performance against other state-of-the-art methods.
Songhe Feng, Tao Wang 0011, Congyan Lang, Yi Jin 0001
AAAI4
2019 Deformable Surface Tracking by Graph Matching
abstract
This paper addresses the problem of deformable surface tracking from monocular images. Specifically, we propose a graph-based approach that effectively explores the structure information of the surface to enhance tracking performance. Our approach solves simultaneously for feature correspondence, outlier rejection and shape reconstruction by optimizing a single objective function, which is defined by means of pairwise projection errors between graph structures instead of unary projection errors between matched points. Furthermore, an efficient matching algorithm is developed based on soft matching relaxation. For evaluation, our approach is extensively compared to state-of-the-art algorithms on a standard dataset of occluded surfaces, as well as a newly compiled dataset of different surfaces with rich, weak or repetitive texture. Experimental results reveal that our approach achieves robust tracking results for surfaces with different types of texture, and outperforms other algorithms in both accuracy and efficiency.
Tao Wang 0011, Haibin Ling, Congyan Lang, Songhe Feng, Xiaohui Hou
ICCV3
2019 Unsupervised Person Re-identification Based on Clustering and Domain-Invariant Network
Yangru Huang, Yi Jin 0001, Peixi Peng, Congyan Lang, Yidong Li
ICIG (3)4
2019 Vehicle Re-Identification by Multi-Grain Learni
abstract
Vehicle re-identification (re-ID) is to identify the same vehicle captured by different cameras with non-overlapping views, which plays an important role in intelligent transportation system and traffic safety. Compared with face recognition and person re-ID tasks, it is difficult to train an effective vehicle re-ID model since different vehicles of the same vehicle model, such as Mercedes-Benz C300, may exhibit strong inter-class similarity. To handle this difficulty, we propose a multi-grain ranking loss with the auxiliary of vehicle model, which models the vehicle re-ID task as two sub-tasks with different granularities including matching vehicles in the same vehicle model and different vehicle models. In particular, we infer the vehicle model labels online for the unlabeled training samples by clustering. The experimental results on two benchmarks demonstrate the proposed method can achieve state-of-the-art performance.
Xiaoliang Yang, Congyan Lang, Peixi Peng, Junliang Xing
ICIP2
2019 Realtime Human Segmentation in Video
Tairan Zhang, Congyan Lang, Junliang Xing
MMM (2)2
2019 Robust Semi-supervised Multi-label Learning by Triple Low-Rank Regularization
Songhe Feng, Gengyu Lyu, Congyan Lang
PAKDD (2)4
2019 MFAD: A Multi-modality Face Anti-spoofing Dataset
Bingqian Geng, Congyan Lang, Junliang Xing, Songhe Feng, Jun Wu 0007
PRICAI (2)2
2019 Joint Learning of Dictionary and Convolutional Network for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition is to predict the presence of a set of attributes from a given image, and it plays an important role in video surveillance applications. Most existing works model the task as a multi-label classification problem. Although effective, they ignore the existence of correlations among attributes. In this work, to learn multiple attributes jointly, the attributes are modeled as a subspace and a dictionary is introduced to represent the subspace. Furthermore, to extract the convolutional features which are more suitable for attribute prediction, the dictionary is modeled as a network layer which is learned jointly with the convolutional network. Finally, a novel learning algorithm is proposed to optimize the dictionary and the convolutional network corporately. Extensive experimental analyses and evaluations on two largest pedestrian attribute benchmarks PETA and PA-100K demonstrate that the proposed method achieves state-of-the-art performance.
Yan Sha, Congyan Lang, Peixi Peng, Junliang Xing, Danxia Li
VCIP2
2019 General structured sparse learning for human facial age estimation
Yanan Dong, Congyan Lang, Songhe Feng
Multim. Syst.2
2019 Semi-supervised dual low-rank feature mapping for multi-label image annotation
Songhe Feng, Congyan Lang
Multim. Tools Appl.3
2019 Recurrent convolutional network for video-based smoke detection
Mengxia Yin, Congyan Lang, Zun Li 0001, Songhe Feng, Tao Wang 0011
Multim. Tools Appl.2
2019 Co-saliency Detection with Graph Matching
abstract
Recently, co-saliency detection, which aims to automatically discover common and salient objects appeared in several relevant images, has attracted increased interest in the computer vision community. In this article, we present a novel graph-matching based model for co-saliency detection in image pairs. A solution of graph matching is proposed to integrate the visual appearance, saliency coherence, and spatial structural continuity for detecting co-saliency collaboratively. Since the saliency and the visual similarity have been seamlessly integrated, such a joint inference schema is able to produce more accurate and reliable results. More concretely, the proposed model first computes the intra-saliency for each image by aggregating multiple saliency cues. The common and salient regions across multiple images are thus discovered via a graph matching procedure. Then, a graph reconstruction scheme is proposed to refine the intra-saliency iteratively. Compared to existing co-saliency detection methods that only utilize visual appearance cues, our proposed model can effectively exploit both visual appearance and structure information to better guide co-saliency detection. Extensive experiments on several challenging image pair databases demonstrate that our model outperforms state-of-the-art baselines significantly.
Zun Li 0001, Congyan Lang, Jiashi Feng, Yidong Li, Tao Wang 0011, Songhe Feng
ACM Trans. Intell. Syst. Technol.2
2018 Deep Age Estimation Model Stabilization from Images to Videos
abstract
Deep learning models for age estimation from a single image have significantly improved the state-of-the-art. However, when deploying a deep age estimation model from images directly to videos, it often suffers from the fluctuation issue, i.e., the estimated age varies a lot for face frames from the same person. To deal with this problem, this work presents a new deep age estimation model specifically designed for video facial age estimation, which produces very stable and accurate age estimation results. The proposed deep architecture for video facial age estimation incorporates a convolutional neural network with an attention mechanism, where the convolutional neural network extracts the facial features, and an attention block aggregates the facial feature vectors into a single feature representation for final age estimation. The whole model is trained by a novel loss function to guarantee both the accuracy of each frame and the stabilization of age estimation results of all the frames. To evaluate the proposed model for video facial age estimation, a new dataset is collected and annotated. Extensive experimental analyses and comparisons demonstrate the effectiveness of the proposed model and the state-of-the-art performances compared to many competing methods.
Zhipeng Ji, Congyan Lang, Kai Li 0023, Junliang Xing
ICPR2
2018 Constrained Confidence Matching for Planar Object Tracking
abstract
Tracking planar objects has a wide range of applications in robotics. Conventional template tracking algorithms, however, often fail to observe fast object motion or drift significantly after a period of time, due to drastic object appearance change. To address such challenges, we propose a novel constrained confidence matching algorithm for motion estimation and a robust Kalman filter for template updating. Integrated with an accurate occlusion detector, our approach achieves accurate motion estimation in presence of partial occlusion, by excluding occluded pixels from computation of motion parameters. Furthermore, the proposed Kalman filter employs a novel control-input model to handle the object appearance change, which brings our tracker high robustness against sudden illumination change and heavy motion blur. For evaluation, we compare the proposed tracker with several state-of-the-art planar object trackers on two public benchmark datasets. Experimental results show that our algorithm achieves robust tracking results against various environmental variations, and outperforms baseline algorithms remarkably on both datasets.
Tao Wang 0011, Haibin Ling, Congyan Lang, Songhe Feng, Yi Jin 0001, Yidong Li
ICRA3
2018 Hierarchical Discriminant Feature Learning for Heterogeneous Face Recognition
abstract
Heterogeneous Face Recognition (HFR) refers to the problem of recognizing faces across different visual domains and has attached great attention owing to its tremendous potential benefits in practical applications. In this paper, a novel feature learning approach named hierarchical discriminant feature learning (HDFL) has been proposed for HFR. Different from traditional feature learning based HFR approaches, the proposed HDFL aims to learn the most discriminative information via a two-layer hierarchical boosting network (HBN), where the hierarchical discriminative information can be exploited in the learned features and the appearance difference can be effectively reduced, simultaneously. Extensive experiments on three different heterogeneous face databases demonstrate that our approach consistently outperforms the state-of-the-art methods.
Yidong Li, Yi Jin 0001, Congyan Lang, Songhe Feng, Tao Wang 0011
VCIP4
2018 Saliency ranker: A new salient object detection method
Zun Li 0001, Congyan Lang, Songhe Feng, Tao Wang 0011
J. Vis. Commun. Image Represent.2
2018 A novel hypergraph matching algorithm based on tensor refining
Tao Wang 0011, Congyan Lang, Songhe Feng, Yi Jin 0001
J. Vis. Commun. Image Represent.3
2018 Graph Matching with Adaptive and Branching Path Following
abstract
Graph matching aims at establishing correspondences between graph elements, and is widely used in many computer vision tasks. Among recently proposed graph matching algorithms, those utilizing the path following strategy have attracted special research attentions due to their exhibition of state-of-the-art performances. However, the paths computed in these algorithms often contain singular points, which could hurt the matching performance if not dealt properly. To deal with this issue, we propose a novel path following strategy, named branching path following (BPF), to improve graph matching accuracy. In particular, we first propose a singular point detector by solving a KKT system, and then design a branch switching method to seek for better paths at singular points. Moreover, to reduce the computational burden of the BPF strategy, an adaptive path estimation (APE) strategy is integrated into BPF to accelerate the convergence of searching along each path. A new graph matching algorithm named ABPF-G is developed by applying APE and BPF to a recently proposed path following algorithm named GNCCP (Liu & Qiao 2014). Experimental results reveal how our approach consistently outperforms state-of-the-art algorithms for graph matching on five public benchmark datasets.
Tao Wang 0011, Haibin Ling, Congyan Lang, Songhe Feng
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Robust Object Tracking Based on Temporal and Spatial Deep Networks
abstract
Recently deep neural networks have been widely employed to deal with the visual tracking problem. In this work, we present a new deep architecture which incorporates the temporal and spatial information to boost the tracking performance. Our deep architecture contains three networks, a Feature Net, a Temporal Net, and a Spatial Net. The Feature Net extracts general feature representations of the target. With these feature representations, the Temporal Net encodes the trajectory of the target and directly learns temporal correspondences to estimate the object state from a global perspective. Based on the learning results of the Temporal Net, the Spatial Net further refines the object tracking state using local spatial object information. Extensive experiments on four of the largest tracking benchmarks, including VOT2014, VOT2016, OTB50, and OTB100, demonstrate competing performance of the proposed tracker over a number of state-of-the-art algorithms.
Zhu Teng, Junliang Xing, Qiang Wang 0051, Congyan Lang, Songhe Feng, Yi Jin 0001
ICCV4
2017 Multiple path exploration for graph matching
Congyan Lang, Tao Wang 0011
Mach. Vis. Appl.2
2017 Human Facial Age Estimation by Cost-Sensitive Label Ranking and Trace Norm Regularization
abstract
Human facial age estimation has attracted much attention due to its potential applications in forensics, security, and biometrics. In contrast to existing approaches that cast facial age estimation as either a multiclass classification or regression problem, in this work, we propose a novel approach that combines the strength of cost-sensitive label ranking methods with the power of low-rank matrix recovery theories. Instead of having to make a binary decision for each age label, our approach ranks age labels in a descending order in terms of their predicted relevance to the given facial image. In addition, the proposed approach aggregates the linear prediction functions for different ages into a matrix, and introduces the matrix trace norm regularization to explicitly capture the correlations among different age labels and control the model complexity as well. Furthermore, motivated by nonlinear generalization performance of kernel methods, we extend the trace norm regularization from a finite dimensional space to an infinite dimensional space. We also provide theoretical analysis on the efficiency of the proposed kernelized trace normalization, which guarantees the feasibility of the proposed method for solving large-scale prediction problems. Comprehensive experiments on multiple well-known facial image datasets demonstrate the effectiveness of the proposed framework for age estimation compared to the state-of-the-arts.
Songhe Feng, Congyan Lang, Jiashi Feng, Tao Wang 0011, Jiebo Luo 0001
IEEE Trans. Multim.2
2016 Branching Path Following for Graph Matching
Tao Wang 0011, Haibin Ling, Congyan Lang, Jun Wu 0007
ECCV (2)3
2016 A Novel Emotional Saliency Map to Model Emotional Attention Mechanism
Xinmiao Ding, Lulu Huang, Bing Li 0001, Congyan Lang, Zhen Hua
MMM (2)4
2016 Practical anonymity models on protecting private weighted graphs
Yidong Li, Hong Shen 0001, Congyan Lang, Hairong Dong 0001
Neurocomputing3
2016 From sample selection to model update: A robust online visual tracking algorithm against drifting
Zhu Teng, Tao Wang 0011, Feng Liu 0061, Dong-Joong Kang, Congyan Lang, Songhe Feng
Neurocomputing5
2016 Guest Editorial: Learning Multimedia for Real World Applications
Bing-Kun Bao, Congyan Lang, Tao Mei 0001, Alberto Del Bimbo
Multim. Tools Appl.2
2016 Symmetry-aware graph matching
Tao Wang 0011, Haibin Ling, Congyan Lang, Songhe Feng
Pattern Recognit.3
2016 Dual Low-Rank Pursuit: Learning Salient Features for Saliency Detection
abstract
Saliency detection is an important procedure for machines to understand visual world as humans do. In this paper, we consider a specific saliency detection problem of predicting human eye fixations when they freely view natural images, and propose a novel dual low-rank pursuit (DLRP) method. DLRP learns saliency-aware feature transformations by utilizing available supervision information and constructs discriminative bases for effectively detecting human fixation points under the popular low-rank and sparsity-pursuit framework. Benefiting from the embedded high-level information in the supervised learning process, DLRP is able to predict fixations accurately without performing the expensive object segmentation as in the previous works. Comprehensive experiments clearly show the superiority of the proposed DLRP method over the established state-of-the-art methods. We also empirically demonstrate that DLRP provides stronger generalization performance across different data sets and inherits the advantages of both the bottom-up- and top-down-based saliency detection methods.
Congyan Lang, Jiashi Feng, Songhe Feng, Jingdong Wang 0001, Shuicheng Yan
IEEE Trans. Neural Networks Learn. Syst.1
2015 Facial Age Estimation Based on Structured Low-rank Representation
abstract
This paper presents an algorithm based on structured, low- rank representation for facial age estimation. The proposed method learns the discriminative feature representation of images with the constraint of the classwise block-diagonal structure to promote discrimination of representations for robust recognition. A block-sparse regularizer is introduced to exploit the similarity and structure information of class. Based on the new representation, we estimate the accurate age using a regression function. By subtly introducing the structured, low-rank representation, we achieve good age estimation performance. Experimental results on three well-known aging faces datasets have demonstrated that the proposed method is superior to the conventional approaches.
Chenjing Yan, Congyan Lang, Songhe Feng
ACM Multimedia2
2015 An error-tolerant approximate matching algorithm for labeled combinatorial maps
Tao Wang 0011, Congyan Lang, Songhe Feng
Neurocomputing3
2014 Hierarchical sparse representation based Multi-Instance Semi-Supervised Learning with application to image categorization
Songhe Feng, Weihua Xiong, Bing Li 0001, Congyan Lang, Xiankai Huang
Signal Process.4
2014 Touch Saliency: Characteristics and Prediction
abstract
In this work, we propose an alternative ground truth to the eye fixation map in visual attention study, called touch saliency. As it can be directly collected from the recorded data of users' daily browsing behavior on widely used smart phone devices with touch screens, the touch saliency data is easy to obtain. Due to the limited screen size, smart phone users usually move and zoom in the images, and fix the region of interest on the screen when browsing images. Our studies are two-fold. First, we collect and study the characteristics of these touch screen fixation maps (named touch saliency) by comprehensive comparisons with their counterpart, the eye-fixation maps (namely, visual saliency). The comparisons show that the touch saliency is highly correlated with the eye fixations for the same stimuli, which indicates its utility in data collection for visual attention study. Based on the consistency between both touch saliency and visual saliency, our second task is to propose a unified saliency prediction model for both visual and touch saliency detection. This model utilizes middle-level object category features extracted from pre-segmented image superpixels as input to the recently proposed multitask sparsity pursuit (MTSP) framework for saliency prediction. Extensive evaluations show that the proposed middle-level category features can considerably improve the saliency prediction performance when taking both touch saliency and visual saliency as ground truth.
Bingbing Ni, Mengdi Xu, Tam V. Nguyen 0002, Meng Wang 0001, Congyan Lang, ZhongYang Huang, Shuicheng Yan
IEEE Trans. Multim.5
2013 Salient Object Detection via Low-Rank and Structured Sparse Matrix Decomposition
abstract
Salient object detection provides an alternative solution to various image semantic understanding tasks such as object recognition, adaptive compression and image retrieval. Recently, low-rank matrix recovery (LR) theory has been introduced into saliency detection, and achieves impressed results. However, the existing LR-based models neglect the underlying structure of images, and inevitably degrade the associated performance. In this paper, we propose a Low-rank and Structured sparse Matrix Decomposition (LSMD) model for salient object detection. In the model, a tree-structured sparsity-inducing norm regularization is firstly introduced to provide a hierarchical description of the image structure to ensure the completeness of the extracted salient object. The similarity of saliency values within the salient object is then guaranteed by the $\ell _\infty$-norm. Finally, high-level priors are integrated to guide the matrix decomposition and enhance the saliency detection. Experimental results on the largest public benchmark database show that our model outperforms existing LR-based approaches and other state-of-the-art methods, which verifies the effectiveness and robustness of the structure cues in our model.
Houwen Peng, Bing Li 0001, Rongrong Ji, Weiming Hu 0004, Weihua Xiong, Congyan Lang
AAAI6
2013 Contextualizing Tag Ranking and Saliency Detection for Social Images
Congyan Lang, Songhe Feng
MMM (2)2
2013 A novel image tag saliency ranking algorithm based on sparse representation
abstract
As the explosive growth of the web image data, image tag ranking used for image retrieval accurately from mass images is becoming an active research topic. However, the existing ranking approaches are not very ideal, which remains to be improved. This paper proposed a new image tag saliency ranking algorithm based on sparse representation. we firstly propagate labels from image-level to region-level via Multi-instance Learning driven by sparse representation, which means reconstructing the target instance from positive bag via the sparse linear combination of all the instances from training set, instances with nonzero reconstruction coefficients are considered to be similar to the target instance; then visual attention model is used for tag saliency analysis. Comparing with the existing approaches, the proposed method achieves a better effect and shows a better performance.
Zehai Song, Songhe Feng, Congyan Lang, Shuicheng Yan
VCIP4
2013 Adaptive all-season image tag ranking by saliency-driven image pre-classification
Songhe Feng, Congyan Lang, Hongzhe Liu 0001, Xiankai Huang
J. Vis. Commun. Image Represent.2
2013 Supervised sparse patch coding towards misalignment-robust face recognition
Congyan Lang, Songhe Feng, Xiao-Tong Yuan
J. Vis. Commun. Image Represent.1
2013 Improving Bottom-up Saliency Detection by Looking into Neighbors
abstract
Bottom-up saliency detection aims to detect salient areas within natural images usually without learning from labeled images. Typically, the saliency map of an image is inferred by only using the information within this image (referred to as the “current image”). While efficient, such single-image-based methods may fail to obtain reliable results, because the information within a single image may be insufficient for defining saliency. In this paper, we investigate how saliency detection can benefit from the nearest neighbor structure in the image space. First, we show that existing methods can be improved by extending them to include the visual neighborhood information. This verifies the significance of the neighbors. Next, a solution of multitask sparsity pursuit is proposed to integrate the current image and its neighbors to collaboratively detect saliency. The integration is done by first representing each image as a feature matrix, and then seeking the consistently sparse elements from the joint decompositions of multiple matrices into pairs of low-rank and sparse matrices. The computational procedure is formulated as a constrained nuclear norm and ℓ2,1-norm minimization problem, which is convex and can be solved efficiently with the augmented Lagrange multiplier method. Besides the nearest neighbor structure in the visual feature space, the proposed model can also be generalized to handle multiple visual features. Extensive experiments have clearly validated its superiority over other state-of-the-art methods.
Congyan Lang, Jiashi Feng, Guangcan Liu, Jinhui Tang 0001, Shuicheng Yan, Jiebo Luo 0001
IEEE Trans. Circuits Syst. Video Technol.1
2012 Depth Matters: Influence of Depth Cues on Visual Saliency
Congyan Lang, Tam V. Nguyen 0002, Harish Katti, Karthik Yadati, Mohan Kankanhalli, Shuicheng Yan
ECCV (2)1
2012 Towards relevance and saliency ranking of image tags
abstract
Social image tag ranking has emerged as an important research topic recently due to its potential application on web image search. This paper presents an adaptive all-season tag ranking algorithm which can handle the images with and without distinct object(s) using different tag ranking strategies. Firstly, based on saliency map derived from the visual attention model, a linear SVM is trained to pre-classify an image as attentive or non-attentive category by using the gray histogram descriptor on the corresponding saliency map. Then, an image with distinct object is processed by an attention-driven tag saliency ranking algorithm emphasizing distinct object. On the other hand, an image without distinct object is processed by the tag relevance ranking algorithm via the sparse representation based neighbor-voting strategy. Such adaptive ranking strategy can be regarded as taking full advantage of existing tag ranking paradigms. Experiments conducted on well-known image data sets demonstrate the effectiveness and efficiency of the proposed framework.
Songhe Feng, Congyan Lang, Bing Li 0001
ACM Multimedia2
2012 Towards a universal detector by mining concepts with small semantic gaps
Congyan Lang, Jiashi Feng, Yantao Zheng
Expert Syst. Appl.1
2012 A unified supervised codebook learning framework for classification
Congyan Lang, Songhe Feng, Bingbing Ni, Shuicheng Yan
Neurocomputing1
2012 Saliency Detection by Multitask Sparsity Pursuit
abstract
This paper addresses the problem of detecting salient areas within natural images. We shall mainly study the problem under unsupervised setting, i.e., saliency detection without learning from labeled images. A solution of multitask sparsity pursuit is proposed to integrate multiple types of features for detecting saliency collaboratively. Given an image described by multiple features, its saliency map is inferred by seeking the consistently sparse elements from the joint decompositions of multiple-feature matrices into pairs of low-rank and sparse matrices. The inference process is formulated as a constrained nuclear norm and as an l(2, 1)-norm minimization problem, which is convex and can be solved efficiently with an augmented Lagrange multiplier method. Compared with previous methods, which usually make use of multiple features by combining the saliency maps obtained from individual features, the proposed method seamlessly integrates multiple features to produce jointly the saliency map with a single inference step and thus produces more accurate and reliable results. In addition to the unsupervised setting, the proposed method can be also generalized to incorporate the top-down priors obtained from supervised environment. Extensive experiments well validate its superiority over other state-of-the-art methods.
Congyan Lang, Guangcan Liu, Jian Yu 0001, Shuicheng Yan
IEEE Trans. Image Process.1
2011 Supervised Sparse Patch Coding towards Misalignment-Robust Face Recognition
abstract
We address the challenging problem of face recognition under the scenarios where both training and test data are possibly contaminated with spatial misalignments. A supervised sparse coding framework is developed in this paper towards a practical solution to misalignment-robust face recognition. Each given probe face image is then uniformly divided into a set of local patches. We propose to sparsely reconstruct each probe image patch from the patches of all gallery images, and at the same time the reconstructions for all patches of the probe image are regularized by one term towards enforcing sparsity on the subjects of those selected patches. The derived reconstruction coefficients by ℓ1-norm minimization are then utilized to fuse the subject information of the patches for identifying the probe face. Such a supervised sparse coding framework provides a unique solution to face recognition. Extensive face recognition experiments on three benchmark face datasets demonstrate the advantages of the proposed framework over holistic sparse coding and conventional subspace learning based algorithms in terms of robustness to spatial misalignments and image occlusions.
Congyan Lang, Songhe Feng, Xiao-Tong Yuan
ICIG1
2011 Combining visual attention model with multi-instance learning for tag ranking
Songhe Feng, Hong Bao, Congyan Lang, De Xu
Neurocomputing3
2009 Salient region extraction based on Intensity Mapping for image retrieval
abstract
Salient Region Extraction provides an alternative methodology to image description in many applications such as adaptive content delivery and image retrieval. In this paper, we propose a robust approach to extracting the salient region based on bottom-up visual attention. The main contributions are twofold: 1) Instead of the feature parallel integration, the proposed saliencies are derived by serial processing between texture and color feature. 2) A constructive approach is proposed for rendering an image by a non-linear intensity mapping, which can efficiently eliminate high contrast noise regions in the image. And then the salient map can be robustly generated for a variety of nature images. Finally, the salient region extracted by our algorithm is used for image semantic retrieval. Experiments show that the proposed algorithm can characterize the human perception well and achieve satisfied retrieval performance.
Congyan Lang, De Xu, Songhe Feng
SMC1
2005 Perception-Oriented Prominent Region Detection in Video Sequences Using Fuzzy Inference Neural Network
Congyan Lang, De Xu, Xu Yang 0001, Wengang Cheng
ISNN (2)1
2005 Information Theoretic Metrics in Shot Boundary Detection
Wengang Cheng, De Xu, Congyan Lang
KES (3)4
2005 Shot Type Classification in Sports Video Using Fuzzy Information Granular
Congyan Lang, De Xu, Wengang Cheng
KES (2)1
2005 3-DWT Based Motion Suppression for Video Shot Boundary Detection
Xu Yang 0001, De Xu, Guan Tengfei, Aimin Wu, Congyan Lang
KES (2)5