Erjin Zhou

dblp:150/4019 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0001-9234-0026ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 IMWA: Iterative Model Weight Averaging benefits class-imbalanced learning
Zitong Huang, Bowen Dong 0001, Chaoqi Liang, Erjin Zhou, Wangmeng Zuo
Pattern Recognit.5
2025 ProtoCLIP: Prototypical Contrastive Language Image Pretraining
abstract
Contrastive language image pretraining (CLIP) has received widespread attention since its learned representations can be transferred well to various downstream tasks. During the training process of the CLIP model, the InfoNCE objective aligns positive image-text pairs and separates negative ones. We show an underlying representation grouping effect during this process: the InfoNCE objective indirectly groups semantically similar representations together via randomly emerged within-modal anchors. Based on this understanding, in this article, prototypical contrastive language image pretraining (ProtoCLIP) is introduced to enhance such grouping by boosting its efficiency and increasing its robustness against the modality gap. Specifically, ProtoCLIP sets up prototype-level discrimination between image and text spaces, which efficiently transfers higher level structural knowledge. Furthermore, prototypical back translation (PBT) is proposed to decouple representation grouping from representation alignment, resulting in effective learning of meaningful representations under a large modality gap. The PBT also enables us to introduce additional external teachers with richer prior language knowledge. ProtoCLIP is trained with an online episodic training strategy, which means it can be scaled up to unlimited amounts of data. We trained our ProtoCLIP on conceptual captions (CCs) and achieved an +5.81% ImageNet linear probing improvement and an +2.01% ImageNet zero-shot classification improvement. On the larger YFCC-15M dataset, ProtoCLIP matches the performance of CLIP with 33% of training time.
Delong Chen, Fan Liu 0003, Zaiquan Yang, Shaoqiu Zheng, Ying Tan 0002, Erjin Zhou
IEEE Trans. Neural Networks Learn. Syst.7
2024 Class Balance Matters to Active Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning has shown remarkable efficacy in efficient learning new concepts with limited annotations. Nevertheless, the heuristic few-shot annotations may not always cover the most informative samples, which largely restricts the capability of incremental learner. We aim to start from a pool of large-scale unlabeled data and then annotate the most informative samples for incremental learning. Based on this premise, Based on this purpose, this paper introduces the Active Class-Incremental Learning (ACIL). The objective of ACIL is to select the most informative samples from the unlabeled pool to effectively train an incremental learner, aiming to maximize the performance of the resulting model. Note that vanilla active learning algorithms suffer from class-imbalanced distribution among annotated samples, which restricts the ability of incremental learning. To achieve both class balance and informativeness in chosen samples, we propose Class-Balanced Selection (CBS) strategy. Specifically, we first cluster the features of all unlabeled images into multiple groups. Then for each cluster, we employ greedy selection strategy to ensure that the Gaussian distribution of the sampled features closely matches the Gaussian distribution of all unlabeled features within the cluster.Our CBS can be plugged and played into those CIL methods which are based on pretrained models with prompts tunning technique.Extensive experiments under ACIL protocol across five diverse datasets demonstrate that CBS outperforms both random selection and other SOTA active learning approaches.
Zitong Huang, Yuanze Li, Bowen Dong 0001, Erjin Zhou, Yong Liu 0026, Rick Siow Mong Goh, Chun-Mei Feng 0001, Wangmeng Zuo
ACM Multimedia5
2022 W2N: Switching from Weak Supervision to Noisy Supervision for Object Detection
Zitong Huang, Yiping Bao, Bowen Dong 0001, Erjin Zhou, Wangmeng Zuo
ECCV (30)4
2021 General Instance Distillation for Object Detection
abstract
In recent years, knowledge distillation has been proved to be an effective solution for model compression. This approach can make lightweight student models acquire the knowledge extracted from cumbersome teacher models. However, previous distillation methods of detection have weak generalization for different detection frameworks and rely heavily on ground truth (GT), ignoring the valuable relation information between instances. Thus, we propose a novel distillation method for detection tasks based on discriminative instances without considering the positive or negative distinguished by GT, which is called general instance distillation (GID). Our approach contains a general instance selection module (GISM) to make full use of feature-based, relation-based and response-based knowledge for distillation. Extensive results demonstrate that the student model achieves significant AP improvement and even outperforms the teacher in various detection frame-works. Specifically, RetinaNet with ResNet-50 achieves 39.1% in mAP with GID on COCO dataset, which surpasses the baseline 36.2% by 2.9%, and even better than the ResNet-101 based teacher model with 38.1% AP.
Xing Dai, Zeren Jiang, Yiping Bao, Zhicheng Wang 0001, Si Liu 0001, Erjin Zhou
CVPR7
2021 Rethinking the Heatmap Regression for Bottom-Up Human Pose Estimation
abstract
Heatmap regression has become the most prevalent choice for nowadays human pose estimation methods. The ground-truth heatmaps are usually constructed via covering all skeletal keypoints by 2D gaussian kernels. The standard deviations of these kernels are fixed. However, for bottom-up methods, which need to handle a large variance of human scales and labeling ambiguities, the current practice seems unreasonable. To better cope with these problems, we propose the scale-adaptive heatmap regression (SAHR) method, which can adaptively adjust the standard deviation for each keypoint. In this way, SAHR is more tolerant of various human scales and labeling ambiguities. However, SAHR may aggravate the imbalance between fore-background samples, which potentially hurts the improvement of SAHR. Thus, we further introduce the weight-adaptive heatmap regression (WAHR) to help balance the fore-background samples. Extensive experiments show that SAHR together with WAHR largely improves the accuracy of bottom-up human pose estimation. As a result, we finally outperform the state-of-the-art model by +1.5AP and achieve 72.0AP on COCO test-dev2017, which is comparable with the performances of most top-down methods. Source codes are available at https://github.com/greatlog/SWAHR-HumanPose.
Zhengxiong Luo 0001, Zhicheng Wang 0001, Yan Huang 0008, Liang Wang 0001, Tieniu Tan, Erjin Zhou
CVPR6
2021 TokenPose: Learning Keypoint Tokens for Human Pose Estimation
abstract
Human pose estimation deeply relies on visual clues and anatomical constraints between parts to locate keypoints. Most existing CNN-based methods do well in visual representation, however, lacking in the ability to explicitly learn the constraint relationships between keypoints. In this paper, we propose a novel approach based on Token representation for human Pose estimation (TokenPose). In detail, each keypoint is explicitly embedded as a token to simultaneously learn constraint relationships and appearance cues from images. Extensive experiments show that the small and large TokenPose models are on par with state-of-the-art CNN-based counterparts while being more lightweight. Specifically, our TokenPose-S and TokenPose-L achieve 72.5 AP and 75.8 AP on COCO validation dataset respectively, with significant reduction in parameters (↓80.6% ; ↓ 56.8%) and GFLOPs (↓ 75.3%; ↓24.7%). Code is publicly available1.
Shoukui Zhang, Zhicheng Wang 0001, Wankou Yang, Shutao Xia, Erjin Zhou
ICCV7
2021 Efficient Human Pose Estimation by Learning Deeply Aggregated Representations
abstract
In this paper, we propose an efficient human pose estimation network (DANet) by learning deeply aggregated representations. Most existing models explore multi-scale infonnation mainly from features with different spatial sizes. Powerful multi-scale representations usually rely on the cascaded pyramid framework. This framework largely boosts the performance but in the meanwhile makes networks very deep and complex. Instead, we focus on exploiting multi-scale information from layers with different receptive-field sizes and then making full of use this infonnation by improving the fusion method. Specifically, we propose an orthogonal attention block (OAB) and a second-order fusion unit (SFU). The OAB learns multi-scale infonnation from different layers and enhances them by encouraging them to be diverse. The SFU adaptively selects and fuses diverse multi-scale infonnation and suppress the redundant ones. With the help of OAB and SFU, our networks could achieve comparable or even better accuracy with much smaller model complexity. Specifically, our DANet-72 achieves 71.0 in AP score on COCO val2017 with only 1.0G FLOPS. Its speed on a CPU platfonn achieves 58 Persons-Per-Second (PPS).
Zhengxiong Luo 0001, Zhicheng Wang 0001, Yuanhao Cai, Guan'an Wang, Liang Wang 0001, Yan Huang 0008, Erjin Zhou, Tieniu Tan, Jian Sun 0001
ICME7
2020 High-Order Information Matters: Learning Relation and Topology for Occluded Person Re-Identification
abstract
Occluded person re-identification (ReID) aims to match occluded person images to holistic ones across dis-joint cameras. In this paper, we propose a novel framework by learning high-order relation and topology information for discriminative features and robust alignment. At first, we use a CNN backbone to learn feature maps and key-points estimation model to extract semantic local features. Even so, occluded images still suffer from occlusion and outliers. Then, we view the extracted local features of an image as nodes of a graph and propose an adaptive direction graph convolutional (ADGC) layer to pass relation information between nodes. The proposed ADGC layer can automatically suppress the message passing of meaningless features by dynamically learning direction and degree of linkage. When aligning two groups of local features, we view it as a graph matching problem and propose a cross-graph embedded-alignment (CGEA) layer to joint learn and embed topology information to local features, and straightly predict similarity score. The proposed CGEA layer can both take full use of alignment learned by graph matching and replace sensitive one-to-one alignment with a robust soft one. Finally, extensive experiments on occluded, partial, and holistic ReID tasks show the effectiveness of our proposed method. Specifically, our framework significantly outperforms state-of-the-art by $6.5\%$ mAP scores on Occluded-Duke dataset.
Guan'an Wang, Shuo Yang 0002, Zhicheng Wang 0001, Yang Yang 0062, Shuliang Wang 0001, Gang Yu 0002, Erjin Zhou, Jian Sun 0001
CVPR8
2020 DPGN: Distribution Propagation Graph Network for Few-Shot Learning
abstract
Most graph-network-based meta-learning approaches model instance-level relation of examples. We extend this idea further to explicitly model the distribution-level relation of one example to all other examples in a 1-vs-N manner. We propose a novel approach named distribution propagation graph network (DPGN) for few-shot learning. It conveys both the distribution-level relations and instance-level relations in each few-shot learning task. To combine the distribution-level relations and instance-level relations for all examples, we construct a dual complete graph network which consists of a point graph and a distribution graph with each node standing for an example. Equipped with dual graph architecture, DPGN propagates label information from labeled examples to unlabeled examples within several update generations. In extensive experiments on few-shot learning benchmarks, DPGN outperforms state-of-the-art results by a large margin in 5%~12% under supervised setting and 7%~13% under semi-supervised setting. Code will be released.
Ling Yang 0006, Zilun Zhang, Erjin Zhou
CVPR5
2020 Learning Delicate Local Representations for Multi-person Pose Estimation
Yuanhao Cai, Zhicheng Wang 0001, Zhengxiong Luo 0001, Binyi Yin, Angang Du, Haoqian Wang, Xiangyu Zhang 0005, Erjin Zhou, Jian Sun 0001
ECCV (3)9
2018 Symmetric Variational Autoencoder and Connections to Adversarial Learning
abstract
A new form of the variational autoencoder (VAE) is proposed, based on the symmetric Kullback- Leibler divergence. It is demonstrated that learn- ing of the resulting symmetric VAE (sVAE) has close connections to previously developed adversarial-learning methods. This relationship helps unify the previously distinct techniques of VAE and adversarially learning, and provides insights that allow us to ameliorate shortcomings with some previously developed adversarial methods. In addition to an analysis that motivates and explains the sVAE, an extensive set of experiments validate the utility of the approach.
Liqun Chen 0001, Shuyang Dai, Yunchen Pu, Erjin Zhou, Chunyuan Li, Qinliang Su, Changyou Chen, Lawrence Carin
AISTATS4
2018 GridFace: Face Rectification via Learning Local Homography Transformations
Erjin Zhou, Zhimin Cao, Jian Sun 0001
ECCV (16)1
2016 Going Deeper with Embedded FPGA Platform for Convolutional Neural Network
abstract
In recent years, convolutional neural network (CNN) based methods have achieved great success in a large number of applications and have been among the most powerful and widely used techniques in computer vision. However, CNN-based methods are com-putational-intensive and resource-consuming, and thus are hard to be integrated into embedded systems such as smart phones, smart glasses, and robots. FPGA is one of the most promising platforms for accelerating CNN, but the limited bandwidth and on-chip memory size limit the performance of FPGA accelerator for CNN.
Jiantao Qiu, Jie Wang 0022, Kaiyuan Guo, Boxun Li, Erjin Zhou, Tianqi Tang 0001, Ningyi Xu, Sen Song, Yu Wang 0002, Huazhong Yang
FPGA6
2016 Approaching human level facial landmark localization by deep learning
Haoqiang Fan, Erjin Zhou
Image Vis. Comput.2
2015 Learning Face Hallucination in the Wild
abstract
Face hallucination method is proposed to generate high-resolution images from low-resolution ones for better visualization. However, conventional hallucination methods are often designed for controlled settings and cannot handle varying conditions of pose, resolution degree, and blur. In this paper, we present a new method of face hallucination, which can consistently improve the resolution of face images even with large appearance variations. Our method is based on a novel network architecture called Bi-channel Convolutional Neural Network (Bi-channel CNN). It extracts robust face representations from raw input by using deep convolutional network, then adaptively integrates two channels of information (the raw input image and face representations) to predict the high-resolution image. Experimental results show our system outperforms the prior state-of-the-art methods.
Erjin Zhou, Haoqiang Fan, Zhimin Cao, Yuning Jiang 0001, Qi Yin
AAAI1
2014 Large scale recurrent neural network on GPU
abstract
Large scale artificial neural networks (ANNs) have been widely used in data processing applications. The recurrent neural network (RNN) is a special type of neural network equipped with additional recurrent connections. Such a unique architecture enables the recurrent neural network to remember the past processed information and makes it an expressive model for nonlinear sequence processing tasks. However, the large computation complexity makes it difficult to effectively train a recurrent neural network and therefore significantly limits the research on the recurrent neural network in the last 20 years. In recent years, the use of graphics processing units (GPUs) becomes a significant advance to speed up the training process of large scale neural networks by taking advantage of the massive parallelism capabilities of GPUs. In this paper, we propose an efficient GPU implementation of the large scale recurrent neural network and demonstrate the power of scaling up the recurrent neural network with GPUs. We first explore the potential parallelism of the recurrent neural network and propose a fine-grained two-stage pipeline implementation. Experiment results show that the proposed GPU implementation can achieve 2 ~ 11 x speed-up compared with the basic CPU implementation with the Intel Math Kernel Library. We then use the proposed GPU implementation to scale up the recurrent neural network and improve its performance. The experiment results of the Microsoft Research Sentence Completion Challenge demonstrate that the large scale recurrent network without class layer is able to beat the traditional class-based modest-size recurrent network and achieve an accuracy of 47%, the best result achieved by a single recurrent neural network on the same dataset.
Boxun Li, Erjin Zhou, Jiayi Duan, Yu Wang 0002, Ningyi Xu, Huazhong Yang
IJCNN2