Jianqiang Huang 0001

dblp:207/1901-1 · DBLP profile ↗
← Back
87ranked-venue papers
2as first author
59since 2021 · last 2026
0000-0001-5735-2910ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 58 · 1 first-author · 37 since 2021Artificial intelligence and machine learning · 47 · 1 first-author · 31 since 2021Databases, data management, data science and information retrieval · 12 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Gene Incremental Learning for Single-Cell Transcriptomics
abstract
Classes, as fundamental elements of Computer Vision, have been extensively studied within incremental learning frameworks. In contrast, tokens, which play essential roles in many research fields, exhibit similar characteristics of growth, yet investigations into their incremental learning remain significantly scarce. This research gap primarily stems from the holistic nature of tokens in language, which imposes significant challenges on the design of incremental learning frameworks for them. To overcome this obstacle, in this work, we turn to a type of token, gene, for a large-scale biological dataset—single-cell transcriptomics—to formulate a pipeline for gene incremental learning and establish corresponding evaluations. We found that the forgetting problem also exists in gene incremental learning, thus we adapted existing class incremental learning methods to mitigate the forgetting of genes. Through extensive experiments, we demonstrated the soundness of our framework design and evaluations, as well as the effectiveness of the method adaptations. Finally, we provide a complete benchmark for gene incremental learning in single-cell transcriptomics.
Jiaxin Qi, Jianqiang Huang 0001, Gaogang Xie
AAAI3
2025 A Simple and Comprehensive Benchmark for Single-Cell Transcriptomics
abstract
Single-cell transcriptomics describes complex molecular features at the individual cell level, serving various roles in biological research, such as enhancing gene expression and predicting drug responses. Due to transcriptomic data structurally resembling sequential data, many researchers have trained numerous transformers on extensive transcriptomic datasets. However, they have consistently neglected to explore the intrinsic properties of the data and the appropriateness of their chosen model architecture. In this paper, we carefully investigate the nature of transcriptomics, identifying three overlooked problems: 1) long-tailed data problem, 2) model selection problem, and 3) evaluation problem. Consequently, by applying the weighted sampling strategy, we address the long-tailed data problem and achieve consistent improvement across all settings. By adapting different model structures to transcriptomic data, we discover that transformers are not the only option. By developing three downstream tasks and fair evaluation metrics, we establish a simple and comprehensive benchmark to validate the effectiveness of models for transcriptomics. Through extensive experiments, we clarify the misunderstandings in the traditional methods and provide competitive baselines, thereby paving the way for future research in this field.
Jiaxin Qi, Kailei Guo, Jianqiang Huang 0001, Gaogang Xie
AAAI5
2025 Graph Neural Networks as a Substitute for Transformers in Single-Cell Transcriptomics
abstract
Graph Neural Networks (GNNs) and Transformers share significant similarities in their encoding strategies for interacting with features from nodes of interest, where Transformers use query-key scores and GNNs use edges. Compared to GNNs, which are unable to encode relative positions, Transformers leverage dynamic attention capabilities to better represent relative relationships, thereby becoming the standard backbones in large-scale sequential pre-training. However, the subtle difference prompts us to consider: if positions are no longer crucial, could we substitute Transformers with Graph Neural Networks in some fields such as Single-Cell Transcriptomics? In this paper, we first explore the similarities and differences between GNNs and Transformers, specifically in terms of relative positions. Additionally, we design a synthetic example to illustrate their equivalence where there are no relative positions between tokens in the sample. Finally, we conduct extensive experiments on a large-scale position-agnostic dataset-single-cell transcrip-tomics-finding that GNNs achieve competitive performance compared to Transformers while consuming fewer computation resources. These findings provide novel insights for researchers in the field of single-cell transcriptomics, challenging the prevailing notion that the Transformer is always the optimum choice.
Jiaxin Qi, Jinli Ou, Jianqiang Huang 0001
BIBM4
2025 A Comprehensive Benchmark for Electrocardiogram Time-Series
abstract
Electrocardiogram (ECG), a key bioelectrical time-series signal, is crucial for assessing cardiac health and diagnosing various diseases. Given its time-series format, ECG data is often incorporated into pre-training datasets for large-scale time-series model training. However, existing studies often overlook its unique characteristics and specialized downstream applications, which differ significantly from other time-series data, leading to an incomplete understanding of its properties. In this paper, we present an in-depth investigation of ECG signals and establish a comprehensive benchmark, which includes (1) categorizing its downstream applications into four distinct evaluation tasks, (2) identifying limitations in traditional evaluation metrics for ECG analysis, and introducing a novel metric; (3) benchmarking state-of-the-art time-series models and proposing a new architecture. Extensive experiments demonstrate that our proposed benchmark is comprehensive and robust. The results validate the effectiveness of the proposed metric and model architecture, which establish a solid foundation for advancing research in ECG signal analysis.
Zhijiang Tang, Jiaxin Qi, Yuhua Zheng, Jianqiang Huang 0001
ACM Multimedia4
2024 Criterion-based Heterogeneous Collaborative Filtering for Multi-behavior Implicit Recommendation
abstract
Recent years have witnessed the explosive growth of interaction behaviors in multimedia information systems, where multi-behavior recommender systems have received increasing attention by leveraging data from various auxiliary behaviors such as tip and collect. Among various multi-behavior recommendation methods, non-sampling methods have shown superiority over negative sampling methods. However, two observations are usually ignored in existing state-of-the-art non-sampling methods based on binary regression: (1) users have different preference strengths for different items, so they cannot be measured simply by binary implicit data; (2) the dependency across multiple behaviors varies for different users and items. To tackle the above issue, we propose a novel non-sampling learning framework namedCriterion-guidedHeterogeneousCollaborativeFiltering (CHCF). CHCF introduces both upper and lower thresholds to indicate selection criteria, which will guide user preference learning. Besides, CHCF integrates criterion learning and user preference learning into a unified framework, which can be trained jointly for the interaction prediction of the target behavior. We further theoretically demonstrate that the optimization of Collaborative Metric Learning can be approximately achieved by the CHCF learning framework in a non-sampling form effectively. Extensive experiments on three real-world datasets show the effectiveness of CHCF in heterogeneous scenarios.
Xiao Luo 0001, Daqing Wu, Yiyang Gu, Chong Chen 0002, Luchen Liu, Jinwen Ma, Ming Zhang 0004, Minghua Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ACM Trans. Knowl. Discov. Data9
2024 Networked Time-series Prediction with Incomplete Data via Generative Adversarial Network
abstract
A networked time series (NETS) is a family of time series on a given graph, one for each node. It has a wide range of applications from intelligent transportation to environment monitoring to smart grid management. An important task in such applications is to predict the future values of a NETS based on its historical values and the underlying graph. Most existing methods require complete data for training. However, in real-world scenarios, it is not uncommon to have missing data due to sensor malfunction, incomplete sensing coverage, and so on. In this article, we study the problem of NETS prediction with incomplete data . We propose networked time series Imputation Generative Adversarial Network (NETS-ImpGAN), a novel deep learning framework that can be trained on incomplete data with missing values in both history and future. Furthermore, we propose Graph Temporal Attention Networks , which incorporate the attention mechanism to capture both inter-time series and temporal correlations. We conduct extensive experiments on four real-world datasets under different missing patterns and missing rates. The experimental results show that NETS-ImpGAN outperforms existing methods, reducing the Mean Absolute Error by up to 25%.
Yichen Zhu 0002, Bo Jiang 0003, Haiming Jin, Mengtian Zhang, Jianqiang Huang 0001, Tao Lin 0001, Xinbing Wang
ACM Trans. Knowl. Discov. Data6
2024 A Dual-channel Semi-supervised Learning Framework on Graphs via Knowledge Transfer and Meta-learning
abstract
This article studies the problem of semi-supervised learning on graphs, which aims to incorporate ubiquitous unlabeled knowledge (e.g., graph topology, node attributes) with few-available labeled knowledge (e.g., node class) to alleviate the scarcity issue of supervised information on node classification. While promising results are achieved, existing works for this problem usually suffer from the poor balance of generalization and fitting ability due to the heavy reliance on labels or task-agnostic unsupervised information. To address the challenge, we propose a dual-channel framework for semi-supervised learning on G raphs via K nowledge T ransfer between independent supervised and unsupervised embedding spaces, namely, GKT. Specifically, we devise a dual-channel framework including a supervised model for learning the label probability of nodes and an unsupervised model for extracting information from massive unlabeled graph data. A knowledge transfer head is proposed to bridge the gap between the generalization and fitting capability of the two models. We use the unsupervised information to reconstruct batch-graphs to smooth the label probability distribution on the graphs to improve the generalization of prediction. We also adaptively adjust the reconstructed graphs by encouraging the label-related connections to solidify the fitting ability. Since the optimization of the supervised channel with knowledge transfer contains that of the unsupervised channel as a constraint and vice versa, we then propose a meta-learning-based method to solve the bi-level optimization problem, which avoids the negative transfer and further improves the model’s performance. Finally, extensive experiments validate the effectiveness of our proposed framework by comparing state-of-the-art algorithms.
Ziyue Qiao, Pengyang Wang, Pengfei Wang 0008, Zhiyuan Ning 0001, Yanjie Fu, Yi Du 0010, Yuanchun Zhou, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Hui Xiong 0001
ACM Trans. Web8
2024 Deep Adaptive Graph Clustering via von Mises-Fisher Distributions
abstract
Graph clustering has been a hot research topic and is widely used in many fields, such as community detection in social networks. Lots of works combining auto-encoder and graph neural networks have been applied to clustering tasks by utilizing node attributes and graph structure. These works usually assumed the inherent parameters (i.e., size and variance) of different clusters in the latent embedding space are homogeneous, and hence the assigned probability is monotonous over the Euclidean distance between node embeddings and centroids. Unfortunately, this assumption usually does not hold since the size and concentration of different clusters can be quite different, which limits the clustering accuracy. In addition, the node embeddings in deep graph clustering methods are usually L2 normalized so that it lies on the surface of a unit hyper-sphere. To solve this problem, we proposed D eep A daptive G raph C lustering via von Mises-Fisher distributions, namely DAGC. DAGC assumes the node embeddings H can be drawn from a von Mises-Fisher distribution and each cluster k is associated with cluster inherent parameters ρ k which includes cluster center μ and cluster cohesion degree κ. Then we adopt an EM-like approach (i.e., 𝒫( H | ρ ) and 𝒫( ρ | H ), respectively) to learn the embedding and cluster inherent parameters alternately. Specifically, with the node embeddings, we proposed to update the cluster centers in an attraction-repulsion manner to make the cluster centers more separable. And given the cluster inherent parameters, a likelihood-based loss is proposed to make node embeddings more concentrated around cluster centers. Thus, DAGC can simultaneously improve the intra-cluster compactness and inter-cluster heterogeneity. Finally, extensive experiments conducted on four benchmark datasets have demonstrated that the proposed DAGC consistently outperforms the state-of-the-art methods, especially on imbalanced datasets.
Pengfei Wang 0008, Daqing Wu, Chong Chen 0002, Kunpeng Liu 0001, Yanjie Fu, Jianqiang Huang 0001, Yuanchun Zhou, Jianfeng Zhan, Xian-Sheng Hua 0001
ACM Trans. Web6
2023 ST4ML: Machine Learning Oriented Spatio-Temporal Data Processing at Scale
abstract
Data scientists and researchers utilize enormous spatio-temporal data and build machine learning models to solve practical problems in diverse domains including intelligent transportation, urban planning, epidemic prediction, and many more. Extracting application-specific features from big spatio-temporal data poses system requirements of heterogeneous data support, efficient and scalable computing over spatial and temporal dimensions, as well as a user-friendly programming interface. This paper presents ST4ML, a distributed spatio-temporal data processing system to support scalable machine-learning-oriented applications. We propose a three-stage pipelining computing framework, namely "selection-conversion-extraction" to abstract the distributed computing flow and implement it based on Apache Spark. To the best of our knowledge, ST4ML is the first of its kind to realize our design considerations. Extensive experiments with real-world datasets evidence that ST4ML outperforms straightforward extensions of existing ST data processing systems by up to an order of magnitude. ST4ML is open-sourced at https://github.com/Panrong/st4ml.
Panrong Tong, Mo Li 0001, Jianqiang Huang 0001
Proc. ACM Manag. Data5
2023 Learning comprehensive global features in person re-identification: Ensuring discriminativeness of more local regions
Jiali Xi, Jianqiang Huang 0001, Shibao Zheng, Qin Zhou 0002, Bernt Schiele, Xian-Sheng Hua 0001, Qianru Sun
Pattern Recognit.2
2023 CrowdAtlas: Estimating Crowd Distribution within the Urban Rail Transit System
abstract
While urban rail transit systems are playing an increasingly important role in meeting the transportation demands of people, precise awareness of how the human crowd is distributed within such a system is highly necessary, which serves a range of important applications including emergency response, transit recommendation, and commercial valuation. Most rail transit systems are closed systems where once entered the passengers are free to move around all stations and are difficult to track. In this article, we attempt to estimate the crowd distribution based only on the tap-in and tap-out records of all the rail riders. Specifically, we study Singapore MRT (Mass Rapid Transit) as a vehicle and leverage EZ-Link transit card records to estimate the crowd distribution. Guided by a key observation that the passenger inflows and arrival flows at different MRT stations and time are spatio-temporally correlated due to behavioral consistency of MRT riders, we design and implement a machine learning-based solution, CrowdAtlas, that captures MRT riders’ transition probabilities among stations and across time, and based on that accurately estimates the crowd distribution within the MRT system. Our comprehensive performance evaluations with both trace-driven studies and real-world experiments in MRT disruption cases demonstrate the effectiveness of CrowdAtlas.
Jinlong E, Mo Li 0001, Jianqiang Huang 0001
ACM Trans. Knowl. Discov. Data3
2023 A Survey on Deep Hashing Methods
abstract
Nearest neighbor search aims at obtaining the samples in the database with the smallest distances from them to the queries, which is a basic task in a range of fields, including computer vision and data mining. Hashing is one of the most widely used methods for its computational and storage efficiency. With the development of deep learning, deep hashing methods show more advantages than traditional methods. In this survey, we detailedly investigate current deep hashing algorithms including deep supervised hashing and deep unsupervised hashing. Specifically, we categorize deep supervised hashing methods into pairwise methods, ranking-based methods, pointwise methods as well as quantization according to how measuring the similarities of the learned hash codes. Moreover, deep unsupervised hashing is categorized into similarity reconstruction-based methods, pseudo-label-based methods, and prediction-free self-supervised learning-based methods based on their semantic learning manners. We also introduce three related important topics including semi-supervised deep hashing, domain adaption deep hashing, and multi-modal deep hashing. Meanwhile, we present some commonly used public datasets and the scheme to measure the performance of deep hashing algorithms. Finally, we discuss some potential research directions in conclusion.
Xiao Luo 0001, Haixin Wang 0003, Daqing Wu, Chong Chen 0002, Minghua Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ACM Trans. Knowl. Discov. Data6
2023 Urban Traffic Light Control via Active Multi-Agent Communication and Supply-Demand Modeling
abstract
Urban traffic light control is an important and challenging real-world problem. By regarding intersections as agents, most of the reinforcement learning-based methods generate agents’ actions independently. They can cause action conflict and result in overflow or road resource waste in adjacent intersections. Recently, some collaborative methods have alleviated the above problems by extending the observable surroundings of agents, which can be considered inactive cross-agent communication methods. However, when agents act synchronously in these works, the perceived action value is biased, and the information exchanged is insufficient. In this work, we first propose a novel Multi-agent Communication and Action Rectification (MaCAR) framework. It enables active communication between agents by considering the impact of synchronous actions of agents. Another fundamental problem of traffic light control is the balance between traffic demand and road supply capacity. To fully describe the relation between traffic demand and road supply capacity (Supply-Demand modeling, SD), we further model and forecast the Supply-Demand relation to facilitating the effectiveness of the model’s action. The experiments show that our model outperforms state-of-the-art methods on both synthetic and real-world datasets. Combining the SD with MaCAR, SD-MaCAR can further boost the traffic light control performance even in traffic accident scenarios.
Xin Guo 0006, Zhengxu Yu, Pengfei Wang 0008, Zhongming Jin 0001, Jianqiang Huang 0001, Deng Cai 0001, Xiaofei He 0001, Xian-Sheng Hua 0001
IEEE Trans. Knowl. Data Eng.5
2023 STP-TrellisNets+: Spatial-Temporal Parallel TrellisNets for Multi-Step Metro Station Passenger Flow Prediction
abstract
The drastic increase of metro passengers in recent years inevitably causes the overcrowdedness in the metro systems. Accurately predicting passenger flows at metro stations is critical for efficient metro system management, which helps alleviate such overcrowdedness. Compared to the prevalent next-step prediction, multi-step passenger flow prediction could prominently increase the prediction duration and reveal finer-grained passenger flow variations, which better helps metro system management. Thus, in this paper, we address the problem ofmulti-step metro station passenger (MSP) flow prediction. In light of MSP flows’ unique spatial-temporal characteristics, we proposeSTP-TrellisNets+, which for the first time augments the newly-emerged temporal convolutional frameworkTrellisNetfor multi-step MSP flow prediction. The temporal module of STP-TrellisNets+ (namedCP-TrellisNetsED) employs a Closeness TrellisNet followed by aPeriodicity TrellisNets-based Encoder-Decoder (P-TrellisNetsED)to jointly capture the short- and long-term temporal correlation of MSP flows. In parallel to CP-TrellisNetsED, its spatial module (namedGC-TrellisNetsED) adopts a novel transfer flow-based metric to characterize the spatial correlation among MSP flows, and implements another TrellisNetsED on multiplediffusion graph convolutional networks (DGCNs)in time-series order to capture the dynamics of such spatial correlation. Extensive experiments with two large-scale real-world automated fare collection datasets demonstrate that STP-TrellisNets+ outperforms the state-of-the-art baselines.
Junjie Ou, Yichen Zhu 0002, Haiming Jin, Yijuan Liu, Fan Zhang 0019, Jianqiang Huang 0001, Xinbing Wang
IEEE Trans. Knowl. Data Eng.7
2022 Deconfounded Visual Grounding
abstract
We focus on the confounding bias between language and location in the visual grounding pipeline, where we find that the bias is the major visual reasoning bottleneck. For example, the grounding process is usually a trivial languagelocation association without visual reasoning, e.g., grounding any language query containing sheep to the nearly central regions, due to that most queries about sheep have ground-truth locations at the image center. First, we frame the visual grounding pipeline into a causal graph, which shows the causalities among image, query, target location and underlying confounder. Through the causal graph, we know how to break the grounding bottleneck: deconfounded visual grounding. Second, to tackle the challenge that the confounder is unobserved in general, we propose a confounder-agnostic approach called: Referring Expression Deconfounder (RED), to remove the confounding bias. Third, we implement RED as a simple language attention, which can be applied in any grounding method. On popular benchmarks, RED improves various state-of-the-art grounding methods by a significant margin. Code is available at: https://github.com/JianqiangH/Deconfounded_VG.
Jianqiang Huang 0001, Jiaxin Qi, Qianru Sun, Hanwang Zhang
AAAI1
2022 Homography Loss for Monocular 3D Object Detection
abstract
Monocular 3D object detection is an essential task in autonomous driving. However, most current methods consider each 3D object in the scene as an independent training sample, while ignoring their inherent geometric relations, thus inevitably resulting in a lack of leveraging spatial constraints. In this paper, we propose a novel method that takes all the objects into consideration and explores their mutual relationships to help better estimate the 3D boxes. More-over, since 2D detection is more reliable currently, we also investigate how to use the detected 2D boxes as guidance to globally constrain the optimization of the corresponding predicted 3D boxes. To this end, a differentiable loss function, termed as Homography Loss, is proposed to achieve the goal, which exploits both 2D and 3D information, aiming at balancing the positional relationships between different objects by global constraints, so as to obtain more ac-curately predicted 3D boxes. Thanks to the concise design, our loss function is universal and can be plugged into any mature monocular 3D detector, while significantly boosting the performance over their baseline. Experiments demon-strate that our method yields the best performance (Nov. 2021) compared with the other state-of-the-arts by a large margin on KITTI3D datasets.
Jiaqi Gu 0004, Bojian Wu, Lubin Fan, Jianqiang Huang 0001, Shen Cao, Zhiyu Xiang, Xian-Sheng Hua 0001
CVPR4
2022 Online Convolutional Reparameterization
abstract
Structural re-parameterization has drawn increasing attention in various computer vision tasks. It aims at improving the performance of deep models without introducing any inference-time cost. Though efficient during inference, such models rely heavily on the complicated training-time blocks to achieve high accuracy, leading to large extra training cost. In this paper, we present online convolutional re-parameterization (OREPA), a two-stage pipeline, aiming to reduce the huge training overhead by squeezing the complex training-time block into a single convolution. To achieve this goal, we introduce a linear scaling layer for better optimizing the online blocks. Assisted with the reduced training cost, we also explore some more effective re-param components. Compared with the state-of-the-art re-param models, OREPA is able to save the training-time memory cost by about 70% and accelerate the training speed by around 2×. Meanwhile, equipped with OREPA, the models out-perform previous methods on ImageNet by up to +0.6%. We also conduct experiments on object detection and semantic segmentation and show consistent improvements on the downstream tasks. Codes are available at https://github.com/JUGGHM/OREPA_CVPR2022.
Mu Hu, Junyi Feng, Jiashen Hua, Baisheng Lai, Jianqiang Huang 0001, Xiaojin Gong, Xian-Sheng Hua 0001
CVPR5
2022 Structural and Statistical Texture Knowledge Distillation for Semantic Segmentation
abstract
Existing knowledge distillation works for semantic seg-mentation mainly focus on transfering high-level contextual knowledge from teacher to student. However, low-level texture knowledge is also of vital importance for characterizing the local structural pattern and global statistical prop-erty, such as boundary, smoothness, regularity and color contrast, which may not be well addressed by high-level deep features. In this paper, we are intended to take full advantage of both structural and statistical texture knowledge and propose a novel Structural and Statistical Texture Knowledge Distillation (SSTKD) framework for Semantic Segmentation. Specifically, for structural texture knowledge, we introduce a Contourlet Decomposition Module (CDM) that decomposes low-level features with iterative laplacian pyramid and directional filter bank to mine the structural texture knowledge. For statistical knowledge, we propose a Denoised Texture Intensity Equalization Module (DTIEM) to adaptively extract and enhance statistical texture knowledge through heuristics iterative quantization and denoised operation. Finally, each knowledge learning is supervised by an individual loss function, forcing the student network to mimic the teacher better from a broader perspective. Experiments show that the proposed method achieves state-of-the-art performance on Cityscapes, Pascal VOC 2012 and ADE20K datasets.
Deyi Ji, Mingyuan Tao, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Hongtao Lu 0001
CVPR4
2022 Cloth-Changing Person Re-identification from A Single Image with Gait Prediction and Regularization
abstract
Cloth-Changing person re-identification (CC-ReID) aims at matching the same person across different locations over a long-duration, e.g., over days, and therefore inevitably has cases of changing clothing. In this paper, we focus on handling well the CC-ReID problem under a more challenging setting, i.e., just from a single image, which enables an efficient and latency-free person identity matching for surveillance. Specifically, we introduce Gait recognition as an auxiliary task to drive the Image ReID model to learn cloth-agnostic representations by leveraging personal unique and cloth-independent gait information, we name this framework as GI-ReID. GI-ReID adopts a two-stream architecture that consists of an image ReID-Stream and an auxiliary gait recognition stream (Gait-Stream). The Gait-Stream, that is discarded in the inference for high efficiency, acts as a regulator to encourage the ReID-Stream to capture cloth-invariant biometric motion features during the training. To get temporal continuous motion cues from a single image, we design a Gait Sequence Prediction (GSP) module for Gait-Stream to enrich gait information. Finally, a semantics consistency constraint over two streams is enforced for effective knowledge regularization. Extensive experiments on multiple image-based Cloth-Changing ReID benchmarks, e.g., LTCC, PRCC, Real28, and VC-Clothes, demonstrate that GI-ReID performs favorably against the state-of-the-art methods.
Xin Jin 0014, Tianyu He, Kecheng Zheng, Zhiheng Yin, Xu Shen 0001, Zhen Huang 0007, Ruoyu Feng 0001, Jianqiang Huang 0001, Zhibo Chen 0001, Xian-Sheng Hua 0001
CVPR8
2022 MPC: Multi-view Probabilistic Clustering
abstract
Despite the promising progress having been made, the two challenges of multi-view clustering (MVC) are still waiting for better solutions: i) Most existing methods are either not qualified or require additional steps for incomplete multi-view clustering and ii) noise or outliers might significantly degrade the overall clustering performance. In this paper, we propose a novel unified framework for incomplete and complete MVC named multi-view probabilistic clustering (MPC). MPC equivalently transforms multi-view pairwise posterior matching probability into composition of each view's individual distribution, which tolerates data missing and might extend to any number of views. Then graph-context-aware refinement with path propagation and co-neighbor propagation is used to refine pairwise probability, which alleviates the impact of noise and outliers. Finally, MPC also equivalently transforms probabilistic clustering's objective to avoid complete pairwise computation and adjusts clustering assignments by maximizing joint probability iteratively. Extensive experiments on multiple benchmarks for incomplete and complete MVC show that MPC significantly outperforms previous state-of-the-art methods in both effectiveness and efficiency.
Junjie Liu 0002, Junlong Liu, Shaotian Yan, Rongxin Jiang 0001, Xiang Tian 0002, Boxuan Gu, Yaowu Chen, Chen Shen 0003, Jianqiang Huang 0001
CVPR9
2022 Meta Convolutional Neural Networks for Single Domain Generalization
abstract
In single domain generalization, models trained with data from only one domain are required to perform well on many unseen domains. In this paper, we propose a new model, termed meta convolutional neural network, to solve the single domain generalization problem in image recognition. The key idea is to decompose the convolutional features of images into meta features. Acting as “visual words”, meta features are defined as universal and basic visual elements for image representations (like words for documents in language). Taking meta features as reference, we propose compositional operations to eliminate irrelevant features of local convolutional features by an addressing process and then to reformulate the convolutional feature maps as a composition of related meta features. In this way, images are universally coded without biased information from the unseen domain, which can be processed by following modules trained in the source domain. The compositional operations adopt a regression analysis technique to learn the meta features in an online batch learning manner. Extensive experiments on multiple benchmark datasets verify the superiority of the proposed model in improving single domain generalization ability.
Chaoqun Wan, Xu Shen 0001, Yonggang Zhang 0003, Zhiheng Yin, Xinmei Tian 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001
CVPR7
2022 Balanced and Hierarchical Relation Learning for One-shot Object Detection
abstract
Instance-level feature matching is significantly important to the success of modern one-shot object detectors. Re-cently, the methods based on the metric-learning paradigm have achieved an impressive process. Most of these works only measure the relations between query and target objects on a single level, resulting in suboptimal performance overall. In this paper, we introduce the balanced and hierarchical learning for our detector. The contributions are two-fold: firstly, a novel Instance-level Hierarchical Relation (IHR) module is proposed to encode the contrastive-level, salient-level, and attention-level relations simultane-ously to enhance the query-relevant similarity representation. Secondly, we notice that the batch training of the IHR module is substantially hindered by the positive-negative sample imbalance in the one-shot scenario. We then in-troduce a simple but effective Ratio-Preserving Loss (RPL) to protect the learning of rare positive samples and sup-press the effects of negative samples. Our loss can adjust the weight for each sample adaptively, ensuring the desired positive-negative ratio consistency and boosting query-related IHR learning. Extensive experiments show that our method outperforms the state-of-the-art method by 1.6% and 1.3% on PASCAL VOC and MS COCO datasets for unseen classes, respectively. The code will be available at https://github.com/hero-y/BHRL.
Hanqing Yang 0002, Sijia Cai, Hualian Sheng, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Yu Zhang 0018
CVPR5
2022 On Mitigating Hard Clusters for Face Clustering
Yingjie Chen 0002, Huasong Zhong, Chong Chen 0002, Chen Shen 0003, Jianqiang Huang 0001, Tao Wang 0004, Yun Liang 0001, Qianru Sun
ECCV (12)5
2022 Delving into Details: Synopsis-to-Detail Networks for Video Recognition
Shuxian Liang, Xu Shen 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ECCV (4)3
2022 Rethinking IoU-based Optimization for Single-stage 3D Object Detection
Hualian Sheng, Sijia Cai, Na Zhao 0004, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Minjian Zhao, Gim Hee Lee
ECCV (9)5
2022 Pursuing Knowledge Consistency: Supervised Hierarchical Contrastive Learning for Facial Action Unit Recognition
abstract
With the increasing need for emotion analysis, facial action unit (AU) recognition has attracted much more attention as a fundamental task for affective computing. Although deep learning has boosted the performance of AU recognition to a new level in recent years, it remains challenging to extract subject-consistent representations since the appearance changes caused by AUs are subtle and ambiguous among subjects. We observe that there are three kinds of inherent relations among AUs, which can be treated as strong prior knowledge, and pursuing the consistency of such knowledge is the key to learning subject-consistent representations. To this end, we propose a supervised hierarchical contrastive learning method (SupHCL) for AU recognition to pursue knowledge consistency among different facial images and different AUs, which is orthogonal to methods focusing on network architecture design. Specifically, SupHCL contains three relation consistency modules, i.e., unary, binary, and multivariate relation consistency modules, which take the corresponding kind of inherent relations as extra supervision to encourage knowledge-consistent distributions of both AU-level and image-level representations. Experiments conducted on two commonly used AU benchmark datasets, BP4D and DISFA, demonstrate the effectiveness of each relation consistency module and the superiority of SupHCL.
Yingjie Chen 0002, Chong Chen 0002, Xiao Luo 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Tao Wang 0004, Yun Liang 0001
ACM Multimedia4
2022 Meta Clustering Learning for Large-scale Unsupervised Person Re-identification
abstract
Unsupervised Person Re-identification (U-ReID) with pseudo labeling recently reaches a competitive performance compared to fully-supervised ReID methods based on modern clustering algorithms. However, such clustering-based scheme becomes computationally prohibitive for large-scale datasets, making it infeasible to be applied in real-world application. How to efficiently leverage endless unlabeled data with limited computing resources for better U-ReID is under-explored. In this paper, we make the first attempt to the large-scale U-ReID and propose a "small data for big task" paradigm dubbed Meta Clustering Learning (MCL). MCL only pseudo-labels a subset of the entire unlabeled data via clustering to save computing for the first-phase training. After that, the learned cluster centroids, termed as meta-prototypes in our MCL, are regarded as a proxy annotator to softly annotate the rest unlabeled data for further polishing the model. To alleviate the potential noisy labeling issue in the polishment phase, we enforce two well-designed loss constraints to promise intra-identity consistency and inter-identity strong correlation. For multiple widely-used U-ReID benchmarks, our method significantly saves computational cost while achieving a comparable or even better performance compared to prior works.
Xin Jin 0014, Tianyu He, Xu Shen 0001, Tongliang Liu, Xinchao Wang, Jianqiang Huang 0001, Zhibo Chen 0001, Xian-Sheng Hua 0001
ACM Multimedia6
2022 A topography-aware approach to the automatic generation of urban road networks
abstract
Existing deep-learning tools for road network generation have limited applications in flat urban areas due to their overreliance on the geometric and spatial configurations of street networks and inadequate considerations of topographic information. This paper proposes a new method of street network generation based on a generative adversarial network by designing a pre-positioned geo-extractor module and a geo-merging bypath. The two improvements employ the complementary use of geometric configurations and topographic features to automate street network generation in both flat and hilly urban areas. Our experiments demonstrate that the improved model yields a more realistic prediction of street configurations than conventional image inpainting techniques. The model’s effectiveness is further enhanced when generating streets in hilly areas. Furthermore, the geo-extractor module provides insights from the computer vision perspective in recognizing when topographic information should be considered and which topographic information should receive more attention.
Jiaxin Qi, Lubin Fan, Jianqiang Huang 0001, Ying Jin 0011, Tianren Yang
Int. J. Geogr. Inf. Sci.4
2022 Dynamic supervisor for cross-dataset object detection
Ze Chen 0001, Zhihang Fu, Jianqiang Huang 0001, Mingyuan Tao, Rongxin Jiang 0001, Xiang Tian 0002, Yaowu Chen, Xian-Sheng Hua 0001
Neurocomputing3
2022 Centerness-Aware Network for Temporal Action Proposal
abstract
Temporal action proposal generation aims at localizing the temporal segments containing human actions in a video. This work proposes a centerness-aware network (CAN), which is a novel one-stage approach intended to generate action proposals as keypoint triplets. A keypoint triplet contains two boundary points (starting and ending) and one center point. Specifically, we evaluate the probabilities of each temporal location in the video whether it is at the boundaries or the center region of ground truth action proposals. CAN optimizes the predicted boundary points interactively in a bidirectional adaptation form by exploiting the dependencies among them. Furthermore, to accurately locate the center points of action proposals with different time spans, temporal feature pyramids are utilized to incorporate multi-scale information explicitly. Using the generated three keypoints, CAN efficiently retrieves temporal proposals by grouping keypoints into triplets if they are geometrically aligned. Experiments show that CAN achieves the state-of-the-art performance on the public THUMOS-14 and ActivityNet-1.3 datasets. Moreover, further experiments demonstrate that by applying action classifiers on proposals generated by CAN, our method achieves the state-of-the-art performance in temporal action localization.
Yuan Liu 0017, Jingyuan Chen 0003, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 Progressive Transfer Learning
Zhengxu Yu, Dong Shen 0003, Zhongming Jin 0001, Jianqiang Huang 0001, Deng Cai 0001, Xian-Sheng Hua 0001
IEEE Trans. Image Process.4
2022 Apparel-Invariant Feature Learning for Person Re-Identification
abstract
With the rise of deep learning methods, person Re-Identification (ReID) performance has been improved tremendously in many public datasets. However, most public ReID datasets are collected in a short time window in which persons’ appearance rarely changes. In real-world applications such as in a shopping mall, the same person may change their wearings, and different persons may wear similar apparel. It reveals a critical problem that current ReID models heavily rely on a person’s apparel, resulting in an inconsistent ReID performance. Therefore, it is crucial to learn an apparel-invariant person representation under clothes changing or several persons wearing similar clothes cases. In this work, we tackle this problem from the viewpoint of invariant feature representation learning. The main contributions of this work are as follows. (1) We propose the semi-supervised Apparel-invariant Feature Learning (AIFL) framework to learn an apparel-invariant pedestrian representation using images of the same person wearing different clothes. (2) To obtain images of the same person wearing different clothes, we propose an unsupervised apparel-simulation GAN (AS-GAN) to synthesize cloth-changing images according to the target cloth embedding. It is worth noting that the images used in ReID tasks were cropped from real-world low-quality CCTV videos, making it more challenging to synthesize cloth-changing images. Extensive experiments demonstrate that our proposal can improve the ReID performance of the baseline models.
Zhengxu Yu, Yilun Zhao 0001, Zhongming Jin 0001, Jianqiang Huang 0001, Deng Cai 0001, Xian-Sheng Hua 0001
IEEE Trans. Multim.5
2021 Camera-Aware Proxies for Unsupervised Person Re-Identification
abstract
This paper tackles the purely unsupervised person re-identification (Re-ID) problem that requires no annotations. Some previous methods adopt clustering techniques to generate pseudo labels and use the produced labels to train Re-ID models progressively. These methods are relatively simple but effective. However, most clustering-based methods take each cluster as a pseudo identity class, neglecting the large intra-ID variance caused mainly by the change of camera views. To address this issue, we propose to split each single cluster into multiple proxies and each proxy represents the instances coming from the same camera. These camera-aware proxies enable us to deal with large intra-ID variance and generate more reliable pseudo labels for learning. Based on the camera-aware proxies, we design both intra and inter-camera contrastive learning components for our Re-ID model to effectively learn the ID discrimination ability within and across cameras. Meanwhile, a proxy-balanced sampling strategy is also designed, which facilitates our learning further. Extensive experiments on three large-scale Re-ID datasets show that our proposed approach outperforms most unsupervised methods by a significant margin. Especially, on the challenging MSMT17 dataset, we gain 14.3 percent Rank-1 and 10.2 percent mAP improvements when compared to the second place. Code is available at: https://github.com/Terminator8758/CAP-master.
Menglin Wang 0001, Baisheng Lai, Jianqiang Huang 0001, Xiaojin Gong, Xian-Sheng Hua 0001
AAAI3
2021 Category Dictionary Guided Unsupervised Domain Adaptation for Object Detection
abstract
Unsupervised domain adaption (UDA) is a promising solution to enhance the generalization ability of a model from a source domain to a target domain without manually annotating labels for target data. Recent works in cross-domain object detection mostly resort to adversarial feature adaptation to match the marginal distributions of two domains. However, perfect feature alignment is hard to achieve and is likely to cause negative transfer due to the high complexity of object detection. In this paper, we propose a category dictionary guided (CDG) UDA model for cross-domain object detection, which learns category-specific dictionaries from the source domain to represent the candidate boxes in target domain. The representation residual can be used for not only pseudo label assignment but also quality (e.g., IoU) estimation of the candidate box. A residual weighted self-training paradigm is then developed to implicitly align source and target domains for detection model training. Compared with decision boundary based classifiers such as softmax, the proposed CDG scheme can select more informative and reliable pseudo-boxes. Experimental results on benchmark datasets show that the proposed CDG significantly exceeds the state-of-the-arts in cross-domain object detection.
Shuai Li 0014, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Lei Zhang 0006
AAAI2
2021 Traffic Flow Prediction with Vehicle Trajectories
abstract
This paper proposes a spatiotemporal deep learning framework, Trajectory-based Graph Neural Network (TrGNN), that mines the underlying causality of flows from historical vehicle trajectories and incorporates that into road traffic prediction. The vehicle trajectory transition patterns are studied to explicitly model the spatial traffic demand via graph propagation along the road network; an attention mechanism is designed to learn the temporal dependencies based on neighborhood traffic status; and finally, a fusion of multi-step prediction is integrated into the graph neural network design. The proposed approach is evaluated with a real-world trajectory dataset. Experiment results show that the proposed TrGNN model achieves over 5% error reduction when compared with the state-of-the-art approaches across all metrics for normal traffic, and up to 14% for atypical traffic during peak hours or abnormal events. The advantage of trajectory transitions especially manifest itself in inferring high fluctuation of flows as well as non-recurrent flow patterns.
Mingqian Li, Panrong Tong, Mo Li 0001, Zhongming Jin 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001
AAAI5
2021 Partial Person Re-Identification With Part-Part Correspondence Learning
abstract
Driven by the success of deep learning, the last decade has seen rapid advances in person re-identification (re-ID). Nonetheless, most of approaches assume that the input is given with the fulfillment of expectations, while imperfect input remains rarely explored to date, which is a non-trivial problem since directly apply existing methods without adjustment can cause significant performance degradation. In this paper, we focus on recognizing partial (flawed) input with the assistance of proposed Part-Part Correspondence Learning (PPCL), a self-supervised learning framework that learns correspondence between image patches without any additional part-level supervision. Accordingly, we propose Part-Part Cycle (PP-Cycle) constraint and Part-Part Triplet (PP-Triplet) constraint that exploit the duality and uniqueness between corresponding image patches respectively. We verify our proposed PPCL on several partial person re-ID benchmarks. Experimental results demonstrate that our approach can surpass previous methods in terms of the standard evaluation metric.
Tianyu He, Xu Shen 0001, Jianqiang Huang 0001, Zhibo Chen 0001, Xian-Sheng Hua 0001
CVPR3
2021 Revisiting Knowledge Distillation: An Inheritance and Exploration Framework
abstract
Knowledge Distillation (KD) is a popular technique to transfer knowledge from a teacher model or ensemble to a student model. Its success is generally attributed to the privileged information on similarities/consistency between the class distributions or intermediate feature representations of the teacher model and the student model. However, directly pushing the student model to mimic the probabilities/features of the teacher model to a large extent limits the student model in learning undiscovered knowledge/features. In this paper, we propose a novel inheritance and exploration knowledge distillation framework (IE-KD), in which a student model is split into two parts - inheritance and exploration. The inheritance part is learned with a similarity loss to transfer the existing learned knowledge from the teacher model to the student model, while the exploration part is encouraged to learn representations different from the inherited ones with a dis-similarity loss. Our IE-KD framework is generic and can be easily combined with existing distillation or mutual learning methods for training deep neural networks. Extensive experiments demonstrate that these two parts can jointly push the student model to learn more diversified and effective representations, and our IE-KD can be a general technique to improve the student network to achieve SOTA performance. Furthermore, by applying our IE-KD to the training of two networks, the performance of both can be improved w.r.t. deep mutual learning.
Zhen Huang 0007, Xu Shen 0001, Jun Xing, Tongliang Liu, Xinmei Tian 0001, Houqiang Li, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001
CVPR8
2021 The Blessings of Unlabeled Background in Untrimmed Videos
abstract
Weakly-supervised Temporal Action Localization (WTAL) aims to detect the action segments with only video-level action labels in training. The key challenge is how to distinguish the action of interest segments from the background, which is unlabelled even on the video-level. While previous works treat the background as "curses", we consider it as "blessings". Specifically, we first use causal analysis to point out that the common localization errors are due to the unobserved confounder that resides ubiquitously in visual recognition. Then, we propose a Temporal Smoothing PCA-based (TS-PCA) deconfounder, which exploits the unlabelled background to model an observed substitute for the unobserved confounder, to remove the confounding effect. Note that the proposed deconfounder is model-agnostic and non-intrusive, and hence can be applied in any WTAL method without model re-designs. Through extensive experiments on four state-of-the-art WTAL methods, we show that the deconfounder can improve all of them on the public datasets: THUMOS-14 and ActivityNet-1.31.
Yuan Liu 0002, Jingyuan Chen 0003, Zhenfang Chen, Bing Deng, Jianqiang Huang 0001, Hanwang Zhang
CVPR5
2021 DCT-Mask: Discrete Cosine Transform Mask Representation for Instance Segmentation
abstract
Binary grid mask representation is broadly used in instance segmentation. A representative instantiation is Mask R-CNN which predicts masks on a 28×28 binary grid. Generally, a low-resolution grid is not sufficient to capture the details, while a high-resolution grid dramatically increases the training complexity. In this paper, we propose a new mask representation by applying the discrete cosine transform(DCT) to encode the high-resolution binary grid mask into a compact vector. Our method, termed DCT-Mask, could be easily integrated into most pixel-based instance segmentation methods. Without any bells and whistles, DCT-Mask yields significant gains on different frameworks, backbones, datasets, and training schedules. It does not require any pre-processing or pre-training, and almost no harm to the running speed. Especially, for higher-quality annotations and more complex backbones, our method has a greater improvement. Moreover, we analyze the performance of our method from the perspective of the quality of mask representation. The main reason why DCT-Mask works well is that it obtains a high-quality mask representation with low complexity.
Jirui Yang, Chunbo Wei, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Kewei Liang
CVPR5
2021 Dense Interaction Learning for Video-based Person Re-identification
abstract
Video-based person re-identification (re-ID) aims at matching the same person across video clips. Efficiently exploiting multi-scale fine-grained features while building the structural interaction among them is pivotal for its success. In this paper, we propose a hybrid framework, Dense Interaction Learning (DenseIL), that takes the principal advantages of both CNN-based and Attention-based architectures to tackle video-based person re-ID difficulties. DenseIL contains a CNN encoder and a Dense Interaction (DI) decoder. The CNN encoder is responsible for efficiently extracting discriminative spatial features while the DI decoder is designed to densely model spatial-temporal inherent interaction across frames. Different from previous works, we additionally let the DI decoder densely attends to intermediate fine-grained CNN features and that naturally yields multi-grained spatial-temporal representation for each video clip. Moreover, we introduce Spatio-TEmporal Positional Embedding (STEP-Emb) into the DI decoder to investigate the positional relation among the spatial-temporal inputs. Our experiments consistently and significantly outperform all the state-of-the-art methods on multiple standard video-based person re-ID datasets.
Tianyu He, Xin Jin 0014, Xu Shen 0001, Jianqiang Huang 0001, Zhibo Chen 0001, Xian-Sheng Hua 0001
ICCV4
2021 3D Local Convolutional Neural Networks for Gait Recognition
abstract
The goal of gait recognition is to learn the unique spatiotemporal pattern about the human body shape from its temporal changing characteristics. As different body parts behave differently during walking, it is intuitive to model the spatio-temporal patterns of each part separately. However, existing part-based methods equally divide the feature maps of each frame into fixed horizontal stripes to get local parts. It is obvious that these stripe partition-based methods cannot accurately locate the body parts. First, different body parts can appear at the same stripe (e.g., arms and the torso), and one part can appear at different stripes in different frames (e.g., hands). Second, different body parts possess different scales, and even the same part in different frames can appear at different locations and scales. Third, different parts also exhibit distinct movement patterns (e.g., at which frame the movement starts, the position change frequency, how long it lasts). To overcome these issues, we propose novel 3D local operations as a generic family of building blocks for 3D gait recognition backbones. The proposed 3D local operations support the extraction of local 3D volumes of body parts in a sequence with adaptive spatial and temporal scales, locations and lengths. In this way, the spatio-temporal patterns of the body parts are well learned from the 3D local neighborhood in partspecific scales, locations, frequencies and lengths. Experiments demonstrate that our 3D local convolutional neural networks achieve state-of-the-art performance on popular gait datasets. Code is available at: https://github.com/yellowtownhz/3DLocalCNN.
Zhen Huang 0007, Dixiu Xue, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ICCV6
2021 Video Object Segmentation with Dynamic Memory Networks and Adaptive Object Alignment
abstract
In this paper, we propose a novel solution for object-matching based semi-supervised video object segmentation, where the target object masks in the first frame are provided. Existing object-matching based methods focus on the matching between the raw object features of the current frame and the first/previous frames. However, two issues are still not solved by these object-matching based methods. As the appearance of the video object changes drastically over time, 1) unseen parts/details of the object present in the current frame, resulting in incomplete annotation in the first annotated frame (e.g. view/scale changes). 2) even for the seen parts/details of the object in the current frame, their positions change relatively (e.g. pose changes/camera motion), leading to a misalignment for the object matching. To obtain the complete information of the target object, we propose a novel object-based dynamic memory network that exploits visual contents of all the past frames. To solve the misalignment problem caused by position changes of visual contents, we propose an adaptive object alignment module by incorporating a region translation function that aligns object proposals towards templates in the feature space. Our method achieves state-of-the-art results on latest benchmark datasets DAVIS 2017 ($\mathcal{J}$ of 81.4% and $\mathcal{F}$ of 87.5% on the validation set) and YouTube-VOS (the overall score of 82.7% on the validation set) with a very efficient inference time (0.16 second/frame on DAVIS 2017 validation set). Code is available at: https://github.com/liang4sx/DMN-AOA.
Shuxian Liang, Xu Shen 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ICCV3
2021 Improving 3D Object Detection with Channel-wise Transformer
abstract
Though 3D object detection from point clouds has achieved rapid progress in recent years, the lack of flexible and high-performance proposal refinement remains a great hurdle for existing state-of-the-art two-stage detectors. Previous works on refining 3D proposals have relied on human-designed components such as keypoints sampling, set abstraction and multi-scale feature fusion to produce powerful 3D object representations. Such methods, however, have limited ability to capture rich contextual dependencies among points. In this paper, we leverage the high-quality region proposal network and a Channel-wise Transformer architecture to constitute our two-stage 3D object detection framework (CT3D) with minimal hand-crafted design. The proposed CT3D simultaneously performs proposal-aware embedding and channel-wise context aggregation for the point features within each proposal. Specifically, CT3D uses proposal’s keypoints for spatial contextual modelling and learns attention propagation in the encoding module, mapping the proposal to point embeddings. Next, a new channel-wise decoding module enriches the query-key interaction via channel-wise re-weighting to effectively merge multi-level contexts, which contributes to more accurate object predictions. Extensive experiments demonstrate that our CT3D method has superior performance and excellent scalability. Remarkably, CT3D achieves the AP of 81.77% in the moderate car category on the KITTI test 3D detection benchmark, outperforms state-of-the-art 3D detectors.
Hualian Sheng, Sijia Cai, Yuan Liu 0017, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Minjian Zhao
ICCV5
2021 Graph Contrastive Clustering
abstract
Recently, some contrastive learning methods have been proposed to simultaneously learn representations and clustering assignments, achieving significant improvements. However, these methods do not take the category information and clustering objective into consideration, thus the learned representations are not optimal for clustering and the performance might be limited. Towards this issue, we first propose a novel graph contrastive learning framework, and then apply it to the clustering task, resulting in the Graph Constrastive Clustering (GCC) method. Different from basic contrastive clustering that only assumes an image and its augmentation should share similar representation and clustering assignments, we lift the instance-level consistency to the cluster-level consistency with the assumption that samples in one cluster and their augmentations should all be similar. Specifically, on the one hand, we propose the graph Laplacian based contrastive loss to learn more discriminative and clustering-friendly features. On the other hand, we propose a novel graph-based contrastive learning strategy to learn more compact clustering assignments. Both of them incorporate the latent category information to reduce the intra-cluster variance as well as increase the inter-cluster variance. Experiments on six commonly used datasets demonstrate the superiority of our proposed approach over the state-of-the-art methods.1
Huasong Zhong, Jianlong Wu, Chong Chen 0002, Jianqiang Huang 0001, Minghua Deng, Liqiang Nie, Zhouchen Lin, Xian-Sheng Hua 0001
ICCV4
2021 CrowdAtlas: Estimating Crowd Distribution within the Urban Rail Transit System
abstract
While the urban rail transit systems are playing an increasingly important role in meeting the transportation demands of people, the precise awareness of how the human crowd is distributed within the urban rail transit system is highly necessary, which serves to a range of important applications including emergency response, transit recommendation, commercial valuation, etc. Most urban rail transit systems are closed systems where once entered the travelers are free to move around all stations that are connected into the system and are difficult to track. In this paper, we attempt to estimate the crowd distribution within the urban rail transit system based only on the entrance and exit records of all the rail riders. Specifically, we study Singapore MRT (Mass Rapid Transit) as a vehicle and leverage the tap-in and tap-out records of the EZ-Link transit cards to estimate the crowd distribution. Guided by a key observation that the passenger inflows and arrival flows at various MRT stations are spatio-temporally correlated due to behavioral consistence of MRT riders, we design and implement a machine learning based solution, CrowdAtlas, that accurately estimates the crowd distribution within the MRT system. Our trace-driven performance evaluation demonstrates the effectiveness of CrowdAtlas.
Jinlong E, Mo Li 0001, Jianqiang Huang 0001
ICDE3
2021 CIMON: Towards High-quality Hash Codes
abstract
Recently, hashing is widely used in approximate nearest neighbor search for its storage and computational efficiency. Most of the unsupervised hashing methods learn to map images into semantic similarity-preserving hash codes by constructing local semantic similarity structure from the pre-trained model as the guiding information, i.e., treating each point pair similar if their distance is small in feature space. However, due to the inefficient representation ability of the pre-trained model, many false positives and negatives in local semantic similarity will be introduced and lead to error propagation during the hash code learning. Moreover, few of the methods consider the robustness of models, which will cause instability of hash codes to disturbance. In this paper, we propose a new method named Comprehensive sImilarity Mining and cOnsistency learNing (CIMON). First, we use global refinement and similarity statistical distribution to obtain reliable and smooth guidance. Second, both semantic and contrastive consistency learning are introduced to derive both disturb-invariant and discriminative hash codes. Extensive experiments on several benchmark datasets show that the proposed method outperforms a wide range of state-of-the-art methods in both retrieval performance and robustness.
Xiao Luo 0001, Daqing Wu, Zeyu Ma 0001, Chong Chen 0002, Minghua Deng, Jinwen Ma, Zhongming Jin 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001
IJCAI8
2021 A Statistical Approach to Mining Semantic Similarity for Deep Unsupervised Hashing
abstract
The majority of deep unsupervised hashing methods usually first construct pairwise semantic similarity information and then learn to map images into compact hash codes while preserving the similarity structure, which implies that the quality of hash codes highly depends on the constructed semantic similarity structure. However, since the features of images for each kind of semantics usually scatter in high-dimensional space with unknown distribution, previous methods could introduce a large number of false positives and negatives for boundary points of distributions in the local semantic structure based on pairwise cosine distances. Towards this limitation, we propose a general distribution-based metric to depict the pairwise distance between images. Specifically, each image is characterized by its random augmentations that can be viewed as samples from the corresponding latent semantic distribution. Then we estimate the distances between images by calculating the sample distribution divergence of their semantics. By applying this new metric to deep unsupervised hashing, we come up with Distribution-based similArity sTructure rEconstruction (DATE). DATE can generate more accurate semantic similarity information by using non-parametric ball divergence. Moreover, DATE explores both semantic-preserving learning and contrastive learning to obtain high-quality hash codes. Extensive experiments on several widely-used datasets validate the superiority of our DATE.
Xiao Luo 0001, Daqing Wu, Zeyu Ma 0001, Chong Chen 0002, Minghua Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ACM Multimedia6
2021 Large-scale vehicle trajectory reconstruction with camera sensing network
abstract
Vehicle trajectories provide essential information to understand the urban mobility and benefit a wide range of urban applications. State-of-the-art solutions for vehicle sensing may not build accurate and complete knowledge of all vehicle trajectories. In order to fill the gap, this paper proposes VeTrac, a comprehensive system that employs widely deployed traffic cameras as a sensing network to trace vehicle movements and reconstruct their trajectories in a large scale. VeTrac fuses mobility correlation and vision-based analysis to reduce uncertainties in identifying vehicles. A graph convolution process is employed to maintain the identity consistency across different camera observations, and a self-training process is invoked when aligning with the urban road network to reconstruct vehicle trajectories with confidence. Extensive experiments with real-world data input of over 7 million vehicle snapshots from over one thousand traffic cameras demonstrate that VeTrac achieves 98% accuracy for simple expressway scenario and 89% accuracy for complex urban environment. The achieved accuracy outperforms alternative solutions by 32% for expressway scenario and by 59% for complex urban environment.
Panrong Tong, Mingqian Li, Mo Li 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001
MobiCom4
2021 Self-Supervised Learning Disentangled Group Representation as Feature
abstract
A good visual representation is an inference map from observations (images) to features (vectors) that faithfully reflects the hidden modularized generative factors (semantics). In this paper, we formulate the notion of "good" representation from a group-theoretic view using Higgins' definition of disentangled representation, and show that existing Self-Supervised Learning (SSL) only disentangles simple augmentation features such as rotation and colorization, thus unable to modularize the remaining semantics. To break the limitation, we propose an iterative SSL algorithm: Iterative Partition-based Invariant Risk Minimization (IP-IRM), which successfully grounds the abstract semantics and the group acting on them into concrete contrastive learning. At each iteration, IP-IRM first partitions the training samples into two subsets that correspond to an entangled group element. Then, it minimizes a subset-invariant contrastive loss, where the invariance guarantees to disentangle the group element. We prove that IP-IRM converges to a fully disentangled representation and show its effectiveness on various benchmarks. Codes are available at https://github.com/Wangt-CN/IP-IRM.
Zhongqi Yue, Jianqiang Huang 0001, Qianru Sun, Hanwang Zhang
NeurIPS3
2021 DFOGraph: an I/O- and communication-efficient system for distributed fully-out-of-core graph processing
abstract
With the magnitude of graph-structured data continually increasing, graph processing systems that can scale-out and scale-up are needed to handle extreme-scale datasets. While existing distributed out-of-core solutions have made it possible, they suffer from limited performance due to excessive I/O and communication costs.
Jiping Yu, Xiaowei Zhu 0001, Zhenbo Sun, Jianqiang Huang 0001
PPoPP5
2021 Towards Precise Intra-camera Supervised Person Re-Identification
abstract
Intra-camera supervision (ICS) for person reidentification (Re-ID) assumes that identity labels are independently annotated within each camera view and no inter-camera identity association is labeled. It is a new setting proposed recently to reduce the burden of annotation while expect to maintain desirable Re-ID performance. However, the lack of inter-camera labels makes the ICS Re-ID problem much more challenging than the fully supervised counterpart. By investigating the characteristics of ICS, this paper proposes jointly learned camera-specific non-parametric classifiers, together with a hybrid mining quintuplet loss, to perform intra-camera learning. Then, an inter-camera learning module consisting of a graph-based ID association step and a Re-ID model updating step is conducted. Extensive experiments on three large-scale Re-ID datasets show that our approach outperforms all existing ICS works by a great margin. Our approach performs even comparable to state-of-the-art fully supervised methods in two of the datasets.
Menglin Wang 0001, Baisheng Lai, Jianqiang Huang 0001, Xiaojin Gong, Xian-Sheng Hua 0001
WACV4
2021 Learning to teach and learn for semi-supervised few-shot image classification
Xinzhe Li 0002, Jianqiang Huang 0001, Yaoyao Liu 0001, Qin Zhou 0002, Shibao Zheng, Bernt Schiele, Qianru Sun
Comput. Vis. Image Underst.2
2021 Spatial likelihood voting with self-knowledge distillation for weakly supervised object detection
Ze Chen 0001, Zhihang Fu, Jianqiang Huang 0001, Mingyuan Tao, Rongxin Jiang 0001, Xiang Tian 0002, Yaowu Chen, Xian-Sheng Hua 0001
Image Vis. Comput.3
2021 Unsupervised Discrete Hashing With Affinity Similarity
abstract
In recent years, supervised hashing has been validated to greatly boost the performance of image retrieval. However, the label-hungry property requires massive label collection, making it intractable in practical scenarios. To liberate the model training procedure from laborious manual annotations, some unsupervised methods are proposed. However, the following two factors make unsupervised algorithms inferior to their supervised counterparts: (1) Without manually-defined labels, it is difficult to capture the semantic information across data, which is of crucial importance to guide robust binary code learning. (2) The widely adopted relaxation on binary constraints results in quantization error accumulation in the optimization procedure. To address the above-mentioned problems, in this paper, we propose a novel Unsupervised Discrete Hashing method (UDH). Specifically, to capture the semantic information, we propose a balanced graph-based semantic loss which explores the affinity priors in the original feature space. Then, we propose a novel self-supervised loss, termed orthogonal consistent loss, which can leverage semantic loss of instance and impose independence of codes. Moreover, by integrating the discrete optimization into the proposed unsupervised framework, the binary constraints are consistently preserved, alleviating the influence of quantization errors. Extensive experiments demonstrate that UDH outperforms state-of-the-art unsupervised methods for image retrieval.
Sheng Jin 0002, Hongxun Yao, Qin Zhou 0002, Yao Liu 0014, Jianqiang Huang 0001, Xian-Sheng Hua 0001
IEEE Trans. Image Process.5
2021 Towards Federated Learning in UAV-Enabled Internet of Vehicles: A Multi-Dimensional Contract-Matching Approach
abstract
Coupled with the rise of Deep Learning, the wealth of data and enhanced computation capabilities of Internet of Vehicles (IoV) components enable effective Artificial Intelligence (AI) based models to be built. Beyond ground data sources, Unmanned Aerial Vehicles (UAVs) based service providers for data collection and AI model training, i.e., Drones-as-a-Service (DaaS), is becoming increasingly popular in recent years. However, the stringent regulations governing data privacy potentially impedes data sharing across independently owned UAVs. To this end, we propose the adoption of a Federated Learning (FL) based approach to enable privacy-preserving collaborative Machine Learning across a federation of independent DaaS providers for the development of IoV applications, e.g., for traffic prediction and car park occupancy management. Given the information asymmetry and incentive mismatches between the UAVs and model owners, we leverage on the self-revealing properties of a multi-dimensional contract to ensure truthful reporting of the UAV types, while accounting for the multiple sources of heterogeneity, e.g., in sensing, computation, and transmission costs. Then, we adopt the Gale-Shapley algorithm to match the lowest cost UAV to each subregion. The simulation results validate the incentive compatibility of our contract design, and shows the efficiency of our matching, thus guaranteeing profit maximization for the model owner amid information asymmetry.
Wei Yang Bryan Lim, Jianqiang Huang 0001, Zehui Xiong, Jiawen Kang 0001, Dusit Niyato, Xian-Sheng Hua 0001, Cyril Leung, Chunyan Miao
IEEE Trans. Intell. Transp. Syst.2
2021 Joint Auction-Coalition Formation Framework for Communication-Efficient Federated Learning in UAV-Enabled Internet of Vehicles
abstract
Due to the advanced capabilities of the Internet of Vehicles (IoV) components such as vehicles, Roadside Units (RSUs) and smart devices as well as the increasing amount of data generated, Federated Learning (FL) becomes a promising tool given that it enables privacy-preserving machine learning that can be implemented in the IoV. However, the performance of the FL suffers from the failure of communication links and missing nodes, especially when continuous exchanges of model parameters are required. Therefore, we propose the use of Unmanned Aerial Vehicles (UAVs) as wireless relays to facilitate the communications between the IoV components and the FL server and thus improving the accuracy of the FL. However, a single UAV may not have sufficient resources to provide services for all iterations of the FL process. In this paper, we present a joint auction-coalition formation framework to solve the allocation of UAV coalitions to groups of IoV components. Specifically, the coalition formation game is formulated to maximize the sum of individual profits of the UAVs. The joint auction-coalition formation algorithm is proposed to achieve a stable partition of UAV coalitions in which an auction scheme is applied to solve the allocation of UAV coalitions. The auction scheme is designed to take into account the preferences of IoV components over heterogeneous UAVs. The simulation results show that the grand coalition, where all UAVs join a single coalition, is not always stable due to the profit-maximizing behavior of the UAVs. In addition, we show that as the cooperation cost of the UAVs increases, the UAVs prefer to support the IoV components independently and not to form any coalition.
Jer Shyuan Ng, Wei Yang Bryan Lim, Hongning Dai, Zehui Xiong, Jianqiang Huang 0001, Dusit Niyato, Xian-Sheng Hua 0001, Cyril Leung, Chunyan Miao
IEEE Trans. Intell. Transp. Syst.5
2021 Last-Mile School Shuttle Planning With Crowdsensed Student Trajectories
abstract
By processing a large dataset composed of daily trajectories of thousands of students in Singapore, we find that, instead of simply picking up students from their homes, an optimal school shuttle planning system needs to learn the real transportation usage and plan across all potential pickup locations for every student to generate need-satisfying routes. It is challenging, however, to perform route planning over a large number of students each having multiple potential pickup locations. We develop a graph-based data structure that embeds potential pickup locations of all students with the awareness of real-world constraints and existing public transits. Based on the graph structure, we prove that the optimal last-mile school shuttle planning problem is NP-hard and thereafter design a Tabu-based expansion algorithm to solve the problem, which strikes at a proper balance between the savings of students' commute time and the total cost of operating the shuttle buses. Extensive experiments with large-scale real-world crowdsensed trajectory data demonstrate that our last-mile school shuttles can save the traveling time for most students by over 20% and the savings can be up to 65% for 10% of the students.
Panrong Tong, Wan Du, Mo Li 0001, Jianqiang Huang 0001, Zheng Qin 0004
IEEE Trans. Intell. Transp. Syst.4
2021 Self-Adaptive Neural Module Transformer for Visual Question Answering
abstract
Vision and language understanding is one of the most fundamental and difficult tasks in Multimedia Intelligence. Simultaneously Visual Question Answering (VQA) is even more challenging since it requires complex reasoning steps to the correct answer. To achieve this, Neural Module Network (NMN) and its variants rely on parsing the natural language question into a module layout (i.e., a problem-solving program). In particular, this process follows a feedforward encoder-decoder pipeline: the encoder embeds the question into a static vector and the decoder generates the layout. However, we argue that such conventional encoder-decoder neglects the dynamic nature of question comprehension (i.e., we should attend to different words from step to step) and per-module intermediate results (i.e., we should discard module performing badly) in the reasoning steps. In this paper, we present a novel NMN, called Self-Adaptive Neural Module Transformer (SANMT), which adaptively adjusts both of the question feature encoding and the layout decoding by considering intermediate Q&A results. Specifically, we encode the intermediate results with the given question features by a novel transformer module to generate dynamic question feature embedding which evolves over reasoning steps. Besides, the transformer utilizes the intermediate results from each reasoning step to guide subsequent layout arrangement. Extensive experimental evaluations demonstrate the superiority of the proposed SANMT over NMN and its variants on four challenging benchmarks, including CLEVR, CLEVR-CoGenT, VQAv1.0, and VQAv2.0 (on average the relative improvement over NMN are 1.5, 2.3, 0.7 and 0.5 points with respect to accuracy).
Huasong Zhong, Jingyuan Chen 0003, Chen Shen 0003, Hanwang Zhang, Jianqiang Huang 0001, Xian-Sheng Hua 0001
IEEE Trans. Multim.5
2021 Intuitive and efficient roof modeling for reconstruction and synthesis
abstract
We propose a novel and flexible roof modeling approach that can be used for constructing planar 3D polygon roof meshes. Our method uses a graph structure to encode roof topology and enforces the roof validity by optimizing a simple but effective planarity metric we propose. This approach is significantly more efficient than using general purpose 3D modeling tools such as 3ds Max or SketchUp, and more powerful and expressive than specialized tools such as the straight skeleton. Our optimization-based formulation is also flexible and can accommodate different styles and user preferences for roof modeling. We showcase two applications. The first application is an interactive roof editing framework that can be used for roof design or roof reconstruction from aerial images. We highlight the efficiency and generality of our approach by constructing a mesh-image paired dataset consisting of 2539 roofs. Our second application is a generative model to synthesize new roof meshes from scratch. We use our novel dataset to combine machine learning and our roof optimization techniques, by using transformers and graph convolutional networks to model roof topology, and our roof optimization methods to enforce the planarity constraint.
Jing Ren 0004, Biao Zhang 0005, Bojian Wu, Jianqiang Huang 0001, Lubin Fan, Maks Ovsjanikov, Peter Wonka
ACM Trans. Graph.4
2020 Part-Aware Attention Network for Person Re-identification
Wangmeng Xiang, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Lei Zhang 0006
ACCV (4)2
2020 Second-Order Camera-Aware Color Transformation for Cross-Domain Person Re-identification
Wangmeng Xiang, Hongwei Yong, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Lei Zhang 0006
ACCV (2)3
2020 STP-TrellisNets: Spatial-Temporal Parallel TrellisNets for Metro Station Passenger Flow Prediction
abstract
Recent years have witnessed a drastic increase in the number of urban metro passengers, which inevitably causes the overcrowdedness in the metro systems of many cities. Clearly, an accurate prediction of passenger flows at metro stations is critical for a variety of metro system management operations, such as line scheduling and staff preallocation, that help alleviate such overcrowdedness. Thus, in this paper, we aim to address the problem of accurately predicting metro station passenger (MSP) flows. Similar to other traffic data, such as road traffic volume and highway speed, MSP flows are also spatial-temporal in nature. However, existing methods for other traffic prediction tasks are usually suboptimal to predict MSP flows due to MSP flows' unique spatial-temporal characteristics. As a result, we propose a novel deep learning framework STP-TrellisNets, which for the first time augments the newly-emerged temporal convolutional framework TrellisNet for spatial-temporal prediction. The temporal module of STP-TrellisNets (named CP-TrellisNets) employs two TrellisNets in serial to jointly capture the short- and long-term temporal correlation of MSP flows. In parallel to CP-TrellisNets, its spatial module (named GC-TrellisNet) adopts a novel transfer flow-based metric to characterize the spatial correlation among MSP flows, and implements multiple diffusion graph convolutional networks (DGCNs) in time-series order with their outputs connected to a TrellisNet to capture the dynamics of such spatial correlation. Clearly, GC-TrellisNet essentially integrates TrellisNet with graph convolution, and empowers TrellisNet with the ability to capture dynamic graph-structured correlation. We conduct extensive experiments with two large-scale real-world automated fare collection datasets, which contain respectively about 1.5 billion records in Shenzhen, China and 70 million records in Hangzhou, China. The experimental results demonstrate that STP-TrellisNets outperforms the state-of-the-art baselines.
Junjie Ou, Yichen Zhu 0002, Haiming Jin, Yijuan Liu, Fan Zhang 0019, Jianqiang Huang 0001, Xinbing Wang
CIKM7
2020 Structure Aware Single-Stage 3D Object Detection From Point Cloud
abstract
3D object detection from point cloud data plays an essential role in autonomous driving. Current single-stage detectors are efficient by progressively downscaling the 3D point clouds in a fully convolutional manner. However, the downscaled features inevitably lose spatial information and cannot make full use of the structure information of 3D point cloud, degrading their localization precision. In this work, we propose to improve the localization precision of single-stage detectors by explicitly leveraging the structure information of 3D point cloud. Specifically, we design an auxiliary network which converts the convolutional features in the backbone network back to point-level representations. The auxiliary network is jointly optimized, by two point-level supervisions, to guide the convolutional features in the backbone network to be aware of the object structure. The auxiliary network can be detached after training and therefore introduces no extra computation in the inference stage. Besides, considering that single-stage detectors suffer from the discordance between the predicted bounding boxes and corresponding classification confidences, we develop an efficient part-sensitive warping operation to align the confidences to the predicted bounding boxes. Our proposed detector ranks at the top of KITTI 3D/BEV detection leaderboards and runs at 25 FPS for inference.
Chenhang He, Hui Zeng 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Lei Zhang 0006
CVPR3
2020 Two Causal Principles for Improving Visual Dialog
abstract
This paper unravels the design tricks adopted by us, the champion team MReaL-BDAI, for Visual Dialog Challenge 2019: two causal principles for improving Visual Dialog (VisDial). By "improving", we mean that they can promote almost every existing VisDial model to the state-of-the-art performance on the leader-board. Such a major improvement is only due to our careful inspection on the causality behind the model and data, finding that the community has overlooked two causalities in VisDial. Intuitively, Principle 1 suggests: we should remove the direct input of the dialog history to the answer model, otherwise a harmful shortcut bias will be introduced; Principle 2 says: there is an unobserved confounder for history, question, and answer, leading to spurious correlations from training data. In particular, to remove the confounder suggested in Principle 2, we propose several causal intervention algorithms, which make the training fundamentally different from the traditional likelihood estimation. Note that the two principles are model-agnostic, so they are applicable in any VisDial model.
Jiaxin Qi, Yulei Niu, Jianqiang Huang 0001, Hanwang Zhang
CVPR3
2020 Unbiased Scene Graph Generation From Biased Training
abstract
Today's scene graph generation (SGG) task is still far from practical, mainly due to the severe training bias, e.g., collapsing diverse "human walk on / sit on / lay on beach" into "human on beach". Given such SGG, the down-stream tasks such as VQA can hardly infer better scene structures than merely a bag of objects. However, debiasing in SGG is not trivial because traditional debiasing methods cannot distinguish between the good and bad bias, e.g., good context prior (e.g., "person read book" rather than "eat") and bad long-tailed bias (e.g., "near" dominating "behind / in front of"). In this paper, we present a novel SGG framework based on causal inference but not the conventional likelihood. We first build a causal graph for SGG, and perform traditional biased training with the graph. Then, we propose to draw the counterfactual causality from the trained graph to infer the effect from the bad bias, which should be removed. In particular, we use Total Direct Effect (TDE) as the proposed final predicate score for unbiased SGG. Note that our framework is agnostic to any SGG model and thus can be widely applied in the community who seeks unbiased predictions. By using the proposed Scene Graph Diagnosis toolkit on the SGG benchmark Visual Genome and several prevailing models, we observed significant improvements over the previous state-of-the-art methods.
Kaihua Tang, Yulei Niu, Jianqiang Huang 0001, Jiaxin Shi, Hanwang Zhang
CVPR3
2020 Visual Commonsense R-CNN
abstract
We present a novel unsupervised feature representation learning method, Visual Commonsense Region-based Convolutional Neural Network (VC R-CNN), to serve as an improved visual region encoder for high-level tasks such as captioning and VQA. Given a set of detected object regions in an image (e.g., using Faster R-CNN), like any other unsupervised feature learning methods (e.g., word2vec), the proxy training objective of VC R-CNN is to predict the contextual objects of a region. However, they are fundamentally different: the prediction of VC R-CNN is by using causal intervention: P(Y|do(X)), while others are by using the conventional likelihood: P(Y|X). This is also the core reason why VC R-CNN can learn “sense-making” knowledge like chair can be sat - while not just “common” co-occurrences such as chair is likely to exist if table is observed. We extensively apply VC R-CNN features in prevailing models of three popular tasks: Image Captioning, VQA, and VCR, and observe consistent performance boosts across them, achieving many new state-of-the-arts.
Jianqiang Huang 0001, Hanwang Zhang, Qianru Sun
CVPR2
2020 Gradient Centralization: A New Optimization Technique for Deep Neural Networks
Hongwei Yong, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Lei Zhang 0006
ECCV (1)2
2020 Momentum Batch Normalization for Deep Learning with Small Batch Size
Hongwei Yong, Jianqiang Huang 0001, Deyu Meng, Xian-Sheng Hua 0001, Lei Zhang 0006
ECCV (12)2
2020 Multi-Dimensional Contract-Matching for Federated Learning in UAV-Enabled Internet of Vehicles
abstract
Beyond ground data sources, Unmanned Aerial Vehicles (UAVs) based service providers for data collection and AI model training, i.e., Drones-as-a-Service (DaaS), is increasingly popular in the Internet of Vehicles (IoV) applications in recent years. However, the stringent regulations governing data privacy potentially impedes data sharing across independently owned UAVs. To this end, we propose the adoption of a Federated Learning (FL) based approach to enable privacy-preserving collaborative Machine Learning for the development of IoV applications, e.g., for traffic prediction and car park occupancy management. Given the information asymmetry and incentive mismatches between the UAVs and model owner, we leverage on the self-revealing properties of a multi-dimensional contract to ensure truthful reporting of the UAV types, while accounting for the multiple sources of heterogeneity, e.g., in sensing and transmission costs. Then, we adopt the Gale-Shapley algorithm to match the lowest cost UAV to each subregion. The simulation results validate the incentive compatibility of our contract design and shows the efficiency of our matching.
Wei Yang Bryan Lim, Jianqiang Huang 0001, Zehui Xiong, Jiawen Kang 0001, Dusit Niyato, Xian-Sheng Hua 0001, Cyril Leung, Chunyan Miao
GLOBECOM2
2020 Communication-Efficient Federated Learning in UAV-enabled IoV: A Joint Auction-Coalition Approach
abstract
Due to the advanced capabilities of the Internet of Vehicles (IoV) components such as vehicles, Roadside Units (RSUs) and smart devices as well as the increasing amount of data generated, Federated Learning (FL) becomes a promising tool given that it enables privacy-preserving machine learning. However, the performance of the FL suffers from the failure of communication links and missing nodes. Therefore, we propose the use of Unmanned Aerial Vehicles (UAVs) as wireless relays to facilitate the communications between the IoV components and the FL server and thus improving the accuracy of the FL. However, a single UAV may not have sufficient resources for all iterations of the FL process. In this paper, we present a joint auction-coalition formation framework. The joint auctioncoalition formation algorithm is proposed to achieve a stable partition of UAV coalitions in which an auction scheme is applied. The auction scheme is designed to take into account the preferences of IoV components over heterogeneous UAVs. The simulation results show that the grand coalition, where all UAVs join a single coalition, is not always stable due to the profitmaximizing behavior of the UAVs. In addition, we show that as the cooperation cost of the UAVs increases, the UAVs prefer not to form any coalition.
Jer Shyuan Ng, Wei Yang Bryan Lim, Hongning Dai, Zehui Xiong, Jianqiang Huang 0001, Dusit Niyato, Xian-Sheng Hua 0001, Cyril Leung, Chunyan Miao
GLOBECOM5
2020 A Testbed for Studying COVID-19 Spreading in Ride-Sharing Systems
abstract
Order dispatch is an important area where artificial intelligence (AI) can benefit ride-sharing systems (e.g., Grab, Uber), which has become an integral part of our public transport network. In this paper, we present a multi-agent testbed to study the spread of infectious diseases through such a system. It allows users to vary the parameters of the disease and behaviours to study the interaction effect between technology, disease and people's behaviours in such a complex environment.
Harrison Jun Yong Wong, Zichao Deng, Han Yu 0001, Jianqiang Huang 0001, Cyril Leung, Chunyan Miao
IJCAI4
2020 MaCAR: Urban Traffic Light Control via Active Multi-agent Communication and Action Rectification
abstract
Urban traffic light control is an important and challenging real-world problem. By regarding intersections as agents, most of the Reinforcement Learning (RL) based methods generate actions of agents independently. They can cause action conflict and result in overflow or road resource waste in adjacent intersections. Recently, some collaborative methods have alleviated the above problems by extending the observable surroundings of agents, which can be considered as inactive cross-agent communication methods. However, when agents act synchronously in these works, the perceived action value is biased and the information exchanged is insufficient. In this work, we propose a novel Multi-agent Communication and Action Rectification (MaCAR) framework. It enables active communication between agents by considering the impact of synchronous actions of agents. MaCAR consists of two parts: (1) an active Communication Agent Network (CAN) involving a Message Propagation Graph Neural Network (MPGNN); (2) a Traffic Forecasting Network (TFN) which learns to predict the traffic after agents' synchronous actions and the corresponding action values. By using predicted information, we mitigate the action value bias during training to help rectify agents' future actions. In experiments, we show that our proposal can outperforms state-of-the-art methods on both synthetic and real-world datasets.
Zhengxu Yu, Shuxian Liang, Zhongming Jin 0001, Jianqiang Huang 0001, Deng Cai 0001, Xiaofei He 0001, Xian-Sheng Hua 0001
IJCAI5
2020 Spatio-Temporal Inception Graph Convolutional Networks for Skeleton-Based Action Recognition
abstract
Skeleton-based human action recognition has attracted much attention with the prevalence of accessible depth sensors. Recently, graph convolutional networks (GCNs) have been widely used for this task due to their powerful capability to model graph data. The topology of the adjacency graph is a key factor for modeling the correlations of the input skeletons. Thus, previous methods mainly focus on the design/learning of the graph topology. But once the topology is learned, only a single-scale feature and one transformation exist in each layer of the networks. Many insights, such as multi-scale information and multiple sets of transformations, that have been proven to be very effective in convolutional neural networks (CNNs), have not been investigated in GCNs. The reason is that, due to the gap between graph-structured skeleton data and conventional image/video data, it is very challenging to embed these insights into GCNs. To overcome this gap, we reinvent the split-transform-merge strategy in GCNs for skeleton sequence processing. Specifically, we design a simple and highly modularized graph convolutional network architecture for skeleton-based action recognition. Our network is constructed by repeating a building block that aggregates multi-granularity information from both the spatial and temporal paths. Extensive experiments demonstrate that our network outperforms state-of-the-art methods by a significant margin with only 1/5 of the parameters and 1/10 of the FLOPs.
Zhen Huang 0007, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ACM Multimedia5
2020 PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph Generation
abstract
Today's scene graph generation (SGG) task is largely limited in realistic scenarios, mainly due to the extremely long-tailed bias of predicate annotation distribution. Thus, tackling the class imbalance trouble of SGG is critical and challenging. In this paper, we first discover that when predicate labels have strong correlation with each other, prevalent re-balancing strategies (e.g., re-sampling and re-weighting) will give rise to either over-fitting the tail data (e.g., bench sitting on sidewalk rather than on), or still suffering the adverse effect from the original uneven distribution (e.g., aggregating varied parked on/standing on/sitting on into on). We argue the principal reason is that re-balancing strategies are sensitive to the frequencies of predicates yet blind to their relatedness, which may play a more important role to promote the learning of predicate features. Therefore, we propose a novel Predicate-Correlation Perception Learning (PCPL for short) scheme to adaptively seek out appropriate loss weights by directly perceiving and utilizing the correlation among predicate classes. Moreover, our PCPL framework is further equipped with a graph encoder module to better extract context features. Extensive experiments on the benchmark VG150 dataset show that the proposed PCPL performs markedly better on tail classes while well-preserving the performance on head ones, which significantly outperforms previous state-of-the-art methods.
Shaotian Yan, Chen Shen 0003, Zhongming Jin 0001, Jianqiang Huang 0001, Rongxin Jiang 0001, Yaowu Chen, Xian-Sheng Hua 0001
ACM Multimedia4
2020 Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal Effect
abstract
As the class size grows, maintaining a balanced dataset across many classes is challenging because the data are long-tailed in nature; it is even impossible when the sample-of-interest co-exists with each other in one collectable unit, e.g., multiple visual instances in one image. Therefore, long-tailed classification is the key to deep learning at scale. However, existing methods are mainly based on re-weighting/re-sampling heuristics that lack a fundamental theory. In this paper, we establish a causal inference framework, which not only unravels the whys of previous methods, but also derives a new principled solution. Specifically, our theory shows that the SGD momentum is essentially a confounder in long-tailed classification. On one hand, it has a harmful causal effect that misleads the tail prediction biased towards the head. On the other hand, its induced mediation also benefits the representation learning and head prediction. Our framework elegantly disentangles the paradoxical effects of the momentum, by pursuing the direct causal effect caused by an input sample. In particular, we use causal intervention in training, and counterfactual reasoning in inference, to remove the bad'' while keep thegood''. We achieve new state-of-the-arts on three long-tailed visual recognition benchmarks: Long-tailed CIFAR-10/-100, ImageNet-LT for image classification and LVIS for instance segmentation.
Kaihua Tang, Jianqiang Huang 0001, Hanwang Zhang
NeurIPS2
2020 Incentive Mechanism Design for Federated Learning in the Internet of Vehicles
abstract
In the Internet of Vehicles (IoV) paradigm, a model owner is able to leverage on the enhanced capabilities of Intelligent Connected Vehicles (ICV) to develop promising Artificial Intelligence (AI) based applications, e.g., for traffic efficiency. However, in some cases, a model owner may have insufficient data samples to build an effective AI model. To this end, we propose a Federated Learning (FL) based privacy preserving approach to facilitate collaborative FL among multiple model owners in the IoV. Our system model enables collaborative model training without compromising data privacy given that only the model parameters instead of the raw data are exchanged within the federation. However, there are two main challenges of incentive mismatches between workers and model owners, as well as among model owners. For the former, we leverage on the self-revealing mechanism in contract theory under information asymmetry. For the latter, we use the coalitional game theory approach that rewards model owners based on their marginal contributions. The numerical results validate the performance efficiency of our proposed hierarchical incentive mechanism design.
Wei Yang Bryan Lim, Zehui Xiong, Dusit Niyato, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Chunyan Miao
VTC Fall4
2020 Decouple co-adaptation: Classifier randomization for person re-identification
Zhenyong Wei, Zhongming Jin 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Deng Cai 0001, Xiaofei He 0001
Neurocomputing5
2020 SIF: Self-Inspirited Feature Learning for Person Re-Identification
abstract
The re-identification (ReID) task has received increasing studies in recent years and its performance has gained significant improvement. The progress mainly comes from searching for new network structures to learn person representations. Most of these networks are trained using the classic stochastic gradient descent optimizer. However, limited efforts have been made to explore potential performance of existing ReID networks directly by better training scheme, which leaves a large space for ReID research. In this paper, we propose a Self-Inspirited Feature Learning (SIF) method to enhance performance of given ReID networks from the viewpoint of optimization. We design a simple adversarial learning scheme to encourage a network to learn more discriminative person representation. In our method, an auxiliary branch is added into the network only in the training stage, while the structure of the original network stays unchanged during the testing stage. In summary, SIF has three aspects of advantages: (1) it is designed under general setting; (2) it is compatible with many existing feature learning networks on the ReID task; (3) it is easy to implement and has steady performance. We evaluate the performance of SIF on three public ReID datasets: Market1501, DuckMTMC-reID, and CUHK03(both labeled and detected). The results demonstrate significant improvement in performance brought by SIF. We also apply SIF to obtain state-of-the-art results on all the three datasets. Specifically, mAP / Rank-1 accuracy are: 87.6% / 95.2% (without re-rank) on Market1501, 79.4% / 89.8% on DuckMTMC-reID, 77.0% / 79.5% on CUHK03 (labeled) and 73.9% / 76.6% on CUHK03 (detected), respectively. The code of SIF will be available soon.
Zhenyong Wei, Zhongming Jin 0001, Zhengxu Yu, Jianqiang Huang 0001, Deng Cai 0001, Xiaofei He 0001, Xian-Sheng Hua 0001
IEEE Trans. Image Process.5
2020 Survey of external memory large-scale graph processing on a multi-core system
Jianqiang Huang 0001, Xiaoying Wang 0002
J. Supercomput.1
2020 Concentrated Local Part Discovery With Fine-Grained Part Representation for Person Re-Identification
abstract
The attention mechanism for person re-identification has been widely studied with deep convolutional neural networks. This mechanism works as a good complement to the global features extracted from an image of the entire human body. However, existing works mainly focus on discovering local parts with simple feature representations, such as global average pooling. Moreover, these works either require extra supervision, such as labeling of body joints, or pay little attention to the guidance of part learning, resulting in scattered activation of learned parts. Furthermore, existing works usually extract local features from different body parts via global average pooling and then concatenate them together as good global features. We find that local features acquired in this way contribute little to the overall performance. In this paper, we argue the significance of local part description and explore the attention mechanism from both local part discovery and local part representation aspects. For local part discovery, we propose a new constrained attention module to make the activated regions concentrated and meaningful without extra supervision. For local part representation, we propose a statistical-positional-relational descriptor to represent local parts from a fine-grained viewpoint. Extensive experiments are conducted to validate the overall performance, the effectiveness of each component, and the generalization ability. We achieve a rank-1 accuracy of 95.1% on Market1501, 64.7% on CUHK03, 87.1% on DukeMTMC-ReID, and 79.9% on MSMT17, outperforming state-of-the-art methods.
Chaoqun Wan, Xinmei Tian 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001
IEEE Trans. Multim.4
2019 Quantization Networks
abstract
Although deep neural networks are highly effective, their high computational and memory costs severely hinder their applications to portable devices. As a consequence, lowbit quantization, which converts a full-precision neural network into a low-bitwidth integer version, has been an active and promising research topic. Existing methods formulate the low-bit quantization of networks as an approximation or optimization problem. Approximation-based methods confront the gradient mismatch problem, while optimizationbased methods are only suitable for quantizing weights and can introduce high computational cost during the training stage. In this paper, we provide a simple and uniform way for weights and activations quantization by formulating it as a differentiable non-linear function. The quantization function is represented as a linear combination of several Sigmoid functions with learnable biases and scales that could be learned in a lossless and end-to-end manner via continuous relaxation of the steepness of Sigmoid functions. Extensive experiments on image classification and object detection tasks show that our quantization networks outperform state-of-the-art methods. We believe that the proposed method will shed new lights on the interpretation of neural network quantization.
Jiwei Yang, Xu Shen 0001, Jun Xing, Xinmei Tian 0001, Houqiang Li, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001
CVPR7
2019 Dynamic Anchor Feature Selection for Single-Shot Object Detection
abstract
The design of anchors is critical to the performance of one-stage detectors. Recently, the anchor refinement module (ARM) has been proposed to adjust the initialization of default anchors, providing the detector a better anchor reference. However, this module brings another problem: all pixels at a feature map have the same receptive field while the anchors associated with each pixel have different positions and sizes. This discordance may lead to a less effective detector. In this paper, we present a dynamic feature selection operation to select new pixels in a feature map for each refined anchor received from the ARM. The pixels are selected based on the new anchor position and size so that the receptive filed of these pixels can fit the anchor areas well, which makes the detector, especially the regression part, much easier to optimize. Furthermore, to enhance the representation ability of selected feature pixels, we design a bidirectional feature fusion module by combining features from early and deep layers. Extensive experiments on both PASCAL VOC and COCO demonstrate the effectiveness of our dynamic anchor feature selection (DAFS) operation. For the case of high IoU threshold, our DAFS can improve the mAP by a large margin.
Shuai Li 0014, Lingxiao Yang, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Lei Zhang 0006
ICCV3
2019 Homocentric Hypersphere Feature Embedding for Person Re-Identification
abstract
Triplet loss and softmax loss are two widely used loss functions in Person Re-Identification (Person ReID). However, previous works that try to apply these two loss functions have measure inconsistency during training and testing stage and among different parts of the total loss function, which would cause inferior performance of models. To address this issue, we propose a novel homocentric hypersphere embedding scheme to decouple magnitude and orientation information for both feature and weight vectors, and reformulate the triplet loss and the softmax loss to their angular versions and combine them into an angular discriminative loss. We evaluate our proposed method extensively on the widely used Person ReID benchmarks. Our method demonstrates leading performance on all datasets.
Wangmeng Xiang, Jianqiang Huang 0001, Xianbiao Qi, Xian-Sheng Hua 0001, Lei Zhang 0006
ICIP2
2019 Progressive Transfer Learning for Person Re-identification
abstract
Model fine-tuning is a widely used transfer learning approach in person Re-identification (ReID) applications, which fine-tuning a pre-trained feature extraction model into the target scenario instead of training a model from scratch. It is challenging due to the significant variations inside the target scenario, e.g., different camera viewpoint, illumination changes, and occlusion. These variations result in a gap between the distribution of each mini-batch and the distribution of the whole dataset when using mini-batch training. In this paper, we study model fine-tuning from the perspective of the aggregation and utilization of the global information of the dataset when using mini-batch training. Specifically, we introduce a novel network structure called Batch-related Convolutional Cell (BConv-Cell), which progressively collects the global information of the dataset into a latent state and uses this latent state to rectify the extracted feature. Based on BConv-Cells, we further proposed the Progressive Transfer Learning (PTL) method to facilitate the model fine-tuning process by joint training the BConv-Cells and the pre-trained ReID model. Empirical experiments show that our proposal can improve the performance of the ReID model greatly on MSMT17, Market-1501, CUHK03 and DukeMTMC-reID datasets. The code will be released later on at \url{https://github.com/ZJULearning/PTL}
Zhengxu Yu, Zhongming Jin 0001, Jishun Guo, Jianqiang Huang 0001, Deng Cai 0001, Xiaofei He 0001, Xian-Sheng Hua 0001
IJCAI5
2019 Panoramic Background Image Generation for PTZ Cameras
abstract
Being able to cover a wide range of views, pan-tilt-zoom (PTZ) cameras have been widely deployed in visual surveillance systems. To achieve a global-view perception of a surveillance scene, it is necessary to generate its panoramic background image, which can be used for the subsequent applications such as road segmentation, active tracking, and so on. However, few works have been reported on this problem, partially due to the lack of benchmark dataset and the high complexity of panoramic image generation of PTZ cameras. In this paper, we build, for the first time to our best knowledge, a benchmark PTZ camera dataset with multiple views, and derive a complete set of panoramic transformation formulas for PTZ cameras. We further propose a fast multi-band blending method to address the efficiency issue in panoramic image fusion and mosaicing. Some related panoramic transformations are also developed, such as cylindrical and overlooking transformations. Our proposed approach exhibits impressive accuracy and efficiency in PTZ panorama generation as well as panoramic image mosaicing.
Hongwei Yong, Jianqiang Huang 0001, Wangmeng Xiang, Xian-Sheng Hua 0001, Lei Zhang 0006
IEEE Trans. Image Process.2
2018 Local Convolutional Neural Networks for Person Re-Identification
abstract
Recent works have shown that person re-identification can be substantially improved by introducing attention mechanisms, which allow learning both global and local representations. However, all these works learn global and local features in separate branches. As a consequence, the interaction/boosting of global and local information are not allowed, except in the final feature embedding layer. In this paper, we propose local operations as a generic family of building blocks for synthesizing global and local information in any layer. This building block can be inserted into any convolutional networks with only a small amount of prior knowledge about the approximate locations of local parts. For the task of person re-identification, even with only one local block inserted, our local convolutional neural networks (Local CNN) can outperform state-of-the-art methods consistently on three large-scale benchmarks, including Market-1501, CUHK03, and DukeMTMC-ReID.
Jiwei Yang, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ACM Multimedia5
2017 Stylized Adversarial AutoEncoder for Image Generation
abstract
In this paper, we propose an autoencoder-based generative adversarial network (GAN) for automatic image generation, which is called "stylized adversarial autoencoder". Different from existing generative autoencoders which typically impose a prior distribution over the latent vector, the proposed approach splits the latent variable into two components: style feature and content feature, both encoded from real images. The split of the latent vector enables us adjusting the content and the style of the generated image arbitrarily by choosing different exemplary images. In addition, a multiclass classifier is adopted in the GAN network as the discriminator, which makes the generated images more realistic. We performed experiments on hand-writing digits, scene text and face datasets, in which the stylized adversarial autoencoder achieves superior results for image generation as well as remarkably improves the corresponding supervised recognition task.
Yiru Zhao, Bing Deng, Jianqiang Huang 0001, Hongtao Lu 0001, Xian-Sheng Hua 0001
ACM Multimedia3