EDBT 2026 Demo / reviewers in the wild / expert
Annan Li
dblp:91/411
· DBLP profile ↗
47ranked-venue papers
11as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 8 first-author · 18 since 2021Artificial intelligence and machine learning · 17 · 7 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A3Bench: an audience-aligned multilingual benchmark for video audience insights understanding
Yiming Lei 0001, Guozhen Peng, Zeming Liu, Hui Qiu, Haitao Leng, Shaoguo Liu, Tingting Gao, Qingjie Liu 0001, Annan Li, Yunhong Wang 0001 |
Frontiers Comput. Sci. | 9 |
| 2026 | From Gradient Analysis to Norm Control: Rethinking Triplet Loss for Gait RecognitionabstractGait recognition has attracted increasing attention in both academia and industry as a non-intrusive human recognition technology from a distance without requiring cooperation. Triplet loss, which enforces relative distance constraints, is a fundamental component in gait recognition. Recently, several gait-specific triplet losses have been introduced to gait recognition. However, they only focus on sample selection and weighting to enhance constraints without exploring the gradient properties of Cosine/Euclidean metrics, which fundamentally influence the model training efficiency and feature discriminability. In this paper, we theoretically analyze triplet loss gradients combined with weight decay and identify inherent limitations due to inadequate norm-control: Cosine metric triplet loss (Lcos) exhibits excessive gradients resulting from small feature norms, while Euclidean metric triplet loss (Leuc) suffers from a small margin-to-norm ratio due to large feature norms. To address these issues, we propose two norm-control approaches to constrain the feature norm in a stable range: 1) Norm-Variance-Regularized Collaboration. 2) Norm-Based Regularization. Extensive experiments show that our methods outperform state-of-the-art results under both Cosine and Euclidean evaluation metrics on three in-the-wild datasets: Gait3D, GREW, and BUAA-Duke-Gait. The code will be available at https://github.com/bgdpgz/TL-Gait. Guozhen Peng, Yunhong Wang 0001, Zhuguanyu Wu, Shaoxiong Zhang 0001, Ruiyi Zhan, Annan Li |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2026 | Slice-and-Align for Clothes-Irrelevant Features: A Clothes-Changing Person Re-Identification Approach Without Additional InputabstractClothes-changing person re-identification (CC Re-ID) focuses on recognizing pedestrians in a long-term with changes in clothes. Prior arts extract clothes-irrelevant features either by introducing extra modality or clothing labels, having their respective limitations. Instead, we seek to extract clothes-irrelevant features without additional input. We first analyze and find that one impediment to extracting clothes-irrelevant features is the co-occurrence of samples with the same clothes and the same identity. Inspired by this observation, we propose a novel CC Re-ID approach using no additional input. We introduce theSlice-and-Align Framework (SA), which employs a straightforward and intuitive prior: the upper and lower clothes of a person are usually different. SA is a dual-stream framework that slices the original image into upper and lower halves, and then aligns them to extract clothes-irrelevant features. On image CC Re-ID datasets, SA outperforms methods without additional input by a large margin and is comparable to or even better than methods with additional input. Besides, SA also outperforms state-of-the-art on video CC Re-ID task. Guozhen Peng, Annan Li, Yunhong Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?abstractRecent advancements in large language models (LLMs) have showcased impressive code generation capabilities, primarily evaluated through language-to-code benchmarks. However, these benchmarks may not fully capture a model’s code understanding abilities. We introduce CodeJudge-Eval (CJ-Eval), a novel benchmark designed to assess LLMs’ code understanding abilities from the perspective of code judging rather than code generation. CJ-Eval challenges models to determine the correctness of provided code solutions, encompassing various error types and compilation issues. By leveraging a diverse set of problems and a fine-grained judging system, CJ-Eval addresses the limitations of traditional benchmarks, including the potential memorization of solutions. Evaluation of 12 well-known LLMs on CJ-Eval reveals that even state-of-the-art models struggle, highlighting the benchmark’s ability to probe deeper into models’ code understanding abilities. Our benchmark is available at https://github.com/CodeLLM-Research/CodeJudge-Eval . Hongzhan Lin 0001, Weixiang Yan, Annan Li, Jing Ma 0004 |
COLING | 6 |
| 2025 | Saliency Based Data Augmentation for Few-Shot Video Action Recognition
Yongqiang Kong, Yunhong Wang 0001, Annan Li |
MMM (3) | 3 |
| 2025 | RSANet: Relative-sequence quality assessment network for gait recognition in the wild
Guozhen Peng, Yunhong Wang 0001, Shaoxiong Zhang 0001, Annan Li |
Pattern Recognit. | 6 |
| 2024 | GLGait: A Global-Local Temporal Receptive Field Network for Gait Recognition in the WildabstractGait recognition has attracted increasing attention from academia and industry as a human recognition technology from a distance in non-intrusive ways without requiring cooperation. Although advanced methods have achieved impressive success in lab scenarios, most of them perform poorly in the wild. Recently, some Convolution Neural Networks (ConvNets) based methods have been proposed to address the issue of gait recognition in the wild. However, the temporal receptive field obtained by convolution operations is limited for long gait sequences. If directly replacing convolution blocks with visual transformer blocks, the model may not enhance a local temporal receptive field, which is important for covering a complete gait cycle. To address this issue, we design a Global-Local Temporal Receptive Field Network (GLGait). GLGait employs a Global-Local Temporal Module (GLTM) to establish a global-local temporal receptive field, which mainly consists of a Pseudo Global Temporal Self-Attention (PGTA) and a temporal convolution operation. Specifically, PGTA is used to obtain a pseudo global temporal receptive field with less memory and computation complexity compared with a multi-head self-attention (MHSA). The temporal convolution operation is used to enhance the local temporal receptive field. Besides, it can also aggregate pseudo global temporal receptive field to a true holistic temporal receptive field. Furthermore, we also propose a Center-Augmented Triplet Loss (CTL) in GLGait to reduce the intra-class distance and expand the positive samples in the training stage. Extensive experiments show that our method obtains state-of-the-art results on in-the-wild datasets, i.e., Gait3D and GREW. The code is available at https://github.com/bgdpgz/GLGait. Guozhen Peng, Yunhong Wang 0001, Shaoxiong Zhang 0001, Annan Li |
ACM Multimedia | 5 |
| 2024 | Synthesis Pyramid Pooling: A Strong Pooling Method for Gait Recognition in the WildabstractGait recognition has attracted increasing attention from academia and industry as a human recognition technology from a distance in non-intrusive ways without requiring cooperation. Although advanced methods have achieved impressive success in laboratory scenarios, most of them perform poorly in the wild. Prior arts focus on modifying model structure for better extraction of global temporal and partial spatial representations in gait sequences while the aggregation of global spatial and partial temporal information is overlooked. In this paper, we propose a Synthesis Pyramid Pooling framework, named SPP. With no change to the backbone, SPP uses Global Temporal Pooling operation (TP), Horizontal Spatial Pooling operation (HSP), Global Spatial Pooling operation (SP), and Horizontal Temporal Pooling operation (HTP) to extract both global-partial and spatial-temporal gait information. Besides, we propose an Interval Sampling Strategy (ISS) to effectively extract temporal information in HTP. Extensive experiments show that our method obtains state-of-the-art results on two in-the-wild datasets, i.e. Gait3D and GREW, respectively. Guozhen Peng, Annan Li, Yunhong Wang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2024 | EvCap: Element-Aware Video CaptioningabstractVideo captioning is a multi-modal task across computer vision and natural language processing. Previous methods generally follow two paradigms, i.e. template-based and sequence-based. Template-based methods can generate relatively accurate elements (e.g. humans, objects, or actions) to complete a template caption, but with a rather limited vocabulary and syntactic structure; in contrast, sequence-based methods generate more natural descriptions like humans but easily suffer element errors due to their heavy dependence on visual features that often contain much distracting information. In this work, we draw lessons from the element extraction manner in template-based methods and propose a novel Element-aware video Captioning (EvCap) framework that applies linguistic features beyond general visual features to consolidate model awareness of specific elements under the sequence-based paradigm. In particular, we introduce two new linguistic features, i.e. action and object-relevant features, from the upstream encoder of the sequence-based paradigm to encode action and object information (in the forms of phrases and words respectively) that benefits the generation of corresponding elements in the final description. Moreover, to fuse the heterogeneous representations and relieve noise of inaccurate features, we design a post-operation fusion strategy, with semantic interaction and energy weighting to ensure the effective usage of each feature. Experimental results show that our EvCap achieves amazingly promising performance compared with baselines under diverse upstream encoder architectures including CNNs, ViT and CLIP, demonstrating good scalability with respect to encoder choices. Sheng Liu 0009, Annan Li, Yunhong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Transpose and Mask: Simple and Effective Logit-Based Knowledge Distillation for Multi-attribute and Multi-label Classification
Annan Li, Guozhen Peng, Yunhong Wang 0001 |
PRCV (10) | 2 |
| 2023 | Self-Sufficient Feature Enhancing Networks for Video Salient Object DetectionabstractDetecting salient objects in videos is a very challenging task. Current state-of-the-art methods are dominated by motion based deep neural networks, among which optical flow is often leveraged as motion representation. Though with robust performance, these optical flow-based video salient object detection methods face at least two problems that may hinder their generalization and application. First, computing optical flow as a pre-processing step does not support direct end-to-end learning; second, little attention has been given to the quality of visual features due to high computational cost of spatiotemporal feature encoding. In this paper we propose a novel self-sufficient feature enhancing network (SFENet) for video salient object detection, which leverages optical flow estimation as an auxiliary task while being end-to-end trainable. With a joint training scheme of both salient object detection and optical flow estimation, its multi-task architecture can be totally self-sufficient for achieving good performance without any pre-processing. Furthermore, for improving feature quality, we design four lightweight modules in spatial and temporal domains, including cross-layer fusion, multi-level warping, spatial-channel attention and boundary-aware refinement. The proposed method is evaluated through extensive experiments on five video salient object detection datasets. Experimental results show that our SFENet can be easily trained with fast convergence speed. It significantly outperforms previous methods in terms of various evaluation metrics. Moreover, with optical flow estimation and unsupervised video object segmentation as example applications, our method also yields state-of-the-art results on standard datasets. Yongqiang Kong, Yunhong Wang 0001, Annan Li, Qiuyu Huang |
IEEE Trans. Multim. | 3 |
| 2023 | Bidirectional Maximum Entropy Training With Word Co-Occurrence for Video CaptioningabstractVideo captioning aims to generate natural language descriptions for a given video, which is a more challenging task than static image captioning since it requires a more diverse and exhaustive result. Meanwhile, it is also important that the generated captions should be consistent with the language habits of people at a fine granularity. In this work, unlike most recent works enhancing performance with additional data modalities or complex model designs, we focus on optimizing the training process of video captioning models. Firstly, to generate a more diverse video caption, we propose the bidirectional maximum entropy (BME) training, which directly optimizes the probability distribution of neighboring words under a reinforcement learning (RL) framework. Secondly, to search for more human-like captions in the larger search space created by BME, we introduce the word co-occurrence (WCO) weighting. It adaptively guides RL algorithms with co-occurrence statistics in the training corpus. Our method can be deployed on existing captioning models in a plug-and-play manner without introducing any extra parameters. Experimental results show that our method yields up to 5.8% and 7.0% improvements considering the CIDEr score on MSVD and MSR-VTT, respectively. Sheng Liu 0009, Annan Li, Yunhong Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Lagrange Motion Analysis and View Embeddings for Improved Gait RecognitionabstractGait is considered the walking pattern of human body, which includes both shape and motion cues. However, the main-stream appearance-based methods for gait recognition rely on the shape of silhouette. It is unclear whether motion can be explicitly represented in the gait sequence modeling. In this paper, we analyzed human walking using the Lagrange's equation and come to the conclusion that second-order information in the temporal dimension is necessary for identification. We designed a second-order motion extraction module based on the conclusions drawn. Also, a light weight view-embedding module is designed by analyzing the problem that current methods to cross-view task do not take view itself into consideration explicitly. Experiments on CASIA-B and OU-MVLP datasets show the effectiveness of our method and some visualization for extracted motion are done to show the interpretability of our motion extraction module. Tianrui Chai, Annan Li, Shaoxiong Zhang 0001, Yunhong Wang 0001 |
CVPR | 2 |
| 2022 | PACE: Predictive and Contrastive Embedding for Unsupervised Action SegmentationabstractAction segmentation, inferring temporal positions of human actions in an untrimmed video, is an important prerequisite for various video understanding tasks. Recently, unsupervised action segmentation (UAS) has emerged as a more challenging task due to the unavailability of frame-level annotations. Existing clustering- or prediction-based UAS approaches suffer from either over-segmentation or overfitting, leading to unsatisfactory results. To address those problems,we propose Predictive And Contrastive Embedding (PACE), a unified UAS framework leveraging both predictability and similarity information for more accurate action segmentation. On the basis of an auto-regressive transformer encoder, predictive embeddings are learned by exploiting the predictability of video context, while contrastive embeddings are generated by leveraging the similarity of adjacent short video clips. Extensive experiments on three challenging benchmarks demonstrate the superiority of our method, with up to 26.9% improvements in F1-score over the state of the art. Jie Qin 0004, Yunhong Wang 0001, Annan Li |
IJCAI | 4 |
| 2022 | Video Person Re-Identification Using Attribute-Enhanced FeaturesabstractIn this work we propose to boost video-based person re-identification (Re-ID) by using attribute-enhanced feature presentation. To this end, we not only try to use the ID-relevant attributes more effectively, but also for the first time in literature harness the ID-irrelevant attributes to help model training. The former mainly include gender, age, clothing characteristics, etc., which contain rich and supplementary information about the pedestrian; the latter include viewpoint, action, etc., which are seldom used for identification previously. In particular, we use the attributes to enhance the significant areas of the image with a novel Attribute Salient Region Enhance (ASRE) module that can attend more accurately to the body of the pedestrian, so as to better separate the target from the background. Furthermore, we find that many ID-irrelevant but subject-relevant factors, like the view angle and movement of the target pedestrian, have great impact on the two-dimensional appearance of a pedestrian. We then propose to exploit both the ID-relevant and the ID-irrelevant attributes via a novel triplet loss called the Viewpoint and Action-Invariant (VAI) triplet loss. Based on the above, we design an Attribute Salience Assisted Network (ASA-Net) to perform attribute recognition along with identity recognition, and use the attributes for feature enhancement and hard sample mining. Extensive experiments on MARS and DukeMTMC-VideoReID datasets show that our method outperforms the state-of-the-arts. Also, the visualizations of learning results further prove the effectiveness of the proposed method. Tianrui Chai, Annan Li, Jiaxin Chen 0002, Xinyu Mei, Yunhong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Spatiotemporal Saliency Representation Learning for Video Action RecognitionabstractDeep convolutional neural networks (CNNs) have achieved great success in human action recognition, however they are still limited in understanding complex and noisy videos owing to the difficulties of exploiting appearance and motion information. Most existing works have been devoted to designing CNN architectures, which overlook the quality of network inputs that is of great importance. This paper provides an alternative solution of action recognition improvement by focusing on the quality of network inputs. A multi-task video salient object detection approach with object-of-interest segmentation scheme, which takes into account both human and action-relevant cues, is proposed to immunize the input video from background clutter. Further, a simple spatiotemporal residual network architecture is presented, which operates on multiple high-quality inputs for long-term action representation learning. Empirical evaluations on various challenging datasets demonstrate that the proposed framework can perform competitively against state-of-the-art. Besides better performance, learning representations of saliency can help prevent the action recognition model from overfitting and speed up the convergence of training. Yongqiang Kong, Yunhong Wang 0001, Annan Li |
IEEE Trans. Multim. | 3 |
| 2022 | Will You Ever Become Popular? Learning to Predict Virality of Dance ClipsabstractDance challenges are going viral in video communities like TikTok nowadays. Once a challenge becomes popular, thousands of short-form videos will be uploaded within a couple of days. Therefore, virality prediction from dance challenges is of great commercial value and has a wide range of applications, such as smart recommendation and popularity promotion. In this article, a novel multi-modal framework that integrates skeletal, holistic appearance, facial and scenic cues is proposed for comprehensive dance virality prediction. To model body movements, we propose a pyramidal skeleton graph convolutional network (PSGCN) that hierarchically refines spatio-temporal skeleton graphs. Meanwhile, we introduce a relational temporal convolutional network (RTCN) to exploit appearance dynamics with non-local temporal relations. An attentive fusion approach is finally proposed to adaptively aggregate predictions from different modalities. To validate our method, we introduce a large-scale viral dance video (VDV) dataset, which contains over 4,000 dance clips of eight viral dance challenges. Extensive experiments on the VDV dataset well demonstrate the effectiveness of our approach. Furthermore, we show that short video applications such as multi-dimensional recommendation and action feedback can be derived from our model. Yunhong Wang 0001, Nina Weng, Tianrui Chai, Annan Li, Faxi Zhang, Sansi Yu |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2021 | Text2Event: Controllable Sequence-to-Structure Generation for End-to-end Event ExtractionabstractYaojie Lu, Hongyu Lin, Jin Xu, Xianpei Han, Jialong Tang, Annan Li, Le Sun, Meng Liao, Shaoyi Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yaojie Lu 0001, Jin Xu 0014, Xianpei Han, Jialong Tang, Annan Li, Le Sun 0001, Meng Liao, Shaoyi Chen |
ACL/IJCNLP (1) | 6 |
| 2021 | Cross-View Gait Recognition With Deep Universal Linear EmbeddingsabstractGait is considered an attractive biometric identifier for its non-invasive and non-cooperative features compared with other biometric identifiers such as fingerprint and iris. At present, cross-view gait recognition methods always establish representations from various deep convolutional networks for recognition and ignore the potential dynamical information of the gait sequences. If assuming that pedestrians have different walking patterns, gait recognition can be performed by calculating their dynamical features from each view. This paper introduces the Koopman operator theory to gait recognition, which can find an embedding space for a global linear approximation of a nonlinear dynamical system. Furthermore, a novel framework based on convolutional variational autoencoder and deep Koopman embedding is proposed to approximate the Koopman operators, which is used as dynamical features from the linearized embedding space for cross-view gait recognition. It gives solid physical interpretability for a gait recognition system. Experiments on a large public dataset, OU-MVLP, prove the effectiveness of the proposed method. Shaoxiong Zhang 0001, Yunhong Wang 0001, Annan Li |
CVPR | 3 |
| 2021 | Silhouette-Based View-Embeddings for Gait Recognition Under Multiple ViewsabstractGait recognition under multiple views is an important computer vision and pattern recognition task. In the emerging convolutional neural network based approaches, the information of view angle is ignored to some extent. Instead of direct view estimation and training view-specific recognition models, we propose a compatible framework that can embed view information into existing architectures of gait recognition. The embedding is simply achieved by a selective projection layer. Experimental results on two large public datasets show that the proposed framework is very effective. Tianrui Chai, Xinyu Mei, Annan Li, Yunhong Wang 0001 |
ICIP | 3 |
| 2021 | Semantically-Guided Disentangled Representation for Robust Gait RecognitionabstractGait is an important biometric that can recognize people at a distance. Recently, Disentangled Representation Learning (DRL) has been introduced for distinguishing identity-irrelevant covariate features from identity features for better recognition performance. However, such a simple gait energy image (GEI) pairing operation inevitably brings in over-disentanglement effects that degrade the performance. To address this issue, we proposed a covariate feature control gate module that compensates for the discriminative feature loss by using additional semantic labels. Furthermore, a shared attention module, which allows the identity and covariate part to pay attention to different spatial regions, is also proposed for better spatial disentanglement. Experimental results show that our method outperforms the state-of-the-art and well-explain the mechanism of how the improvement is achieved. The code is available at https://github.com/ctrasd/GA-ICDNet. Tianrui Chai, Xinyu Mei, Annan Li, Yunhong Wang 0001 |
ICME | 3 |
| 2021 | Few-shot Fine-Grained Action Recognition via Bidirectional Attention and Contrastive Meta-LearningabstractFine-grained action recognition is attracting increasing attention due to the emerging demand of specific action understanding in real-world applications, whereas the data of rare fine-grained categories is very limited. Therefore, we propose the few-shot fine-grained action recognition problem, aiming to recognize novel fine-grained actions with only few samples given for each class. Although progress has been made in coarse-grained actions, existing few-shot recognition methods encounter two issues handling fine-grained actions: the inability to capture subtle action details and the inadequacy in learning from data with low inter-class variance. To tackle the first issue, a human vision inspired bidirectional attention module (BAM) is proposed. Combining top-down task-driven signals with bottom-up salient stimuli, BAM captures subtle action details by accurately highlighting informative spatio-temporal regions. To address the second issue, we introduce contrastive meta-learning (CML). Compared with the widely adopted ProtoNet-based method, CML generates more discriminative video representations for low inter-class variance data, since it makes full use of potential contrastive pairs in each training episode. Furthermore, to fairly compare different models, we establish specific benchmark protocols on two large-scale fine-grained action recognition datasets. Extensive experiments show that our method consistently achieves state-of-the-art performance across evaluated tasks. Yunhong Wang 0001, Sheng Liu 0009, Annan Li |
ACM Multimedia | 4 |
| 2021 | Group Activity Recognition by Exploiting Position Distribution and Appearance Relation
Duoxuan Pei, Annan Li, Yunhong Wang 0001 |
MMM (1) | 2 |
| 2020 | Two-Stream Temporal Convolutional Network for Dynamic Facial Attractiveness PredictionabstractIn the field of facial attractiveness prediction, while deep models using static pictures have shown promising results, little attention is paid to dynamic facial information, which is proven to be influential by psychological studies. Meanwhile, the increasing popularity of short video apps creates an enormous demand for facial attractiveness prediction from short video clips. In this paper, we target on the dynamic facial attractiveness prediction problem. To begin with, a large-scale video-based facial attractiveness prediction dataset (VFAP) with more than one thousand clips from TikTok is collected. A two-stream temporal convolutional network (2S-TCN) is then proposed to capture dynamic attractiveness features from both facial appearance and landmarks. We employ attentive feature enhancement along with specially designed modality and temporal fusion strategies to better explore the temporal dynamics. Extensive experiments on the proposed VFAP dataset demonstrate that 2S-TCN has a distinct advantage over the state-of-the-art static prediction methods. Nina Weng, Annan Li, Yunhong Wang 0001 |
ICPR | 3 |
| 2020 | Assessing Action Quality via Attentive Spatio-Temporal Convolutional Networks
Zhengyin Du, Annan Li, Yunhong Wang 0001 |
PRCV (2) | 3 |
| 2020 | Sports Video Captioning via Attentive Motion Representation and Group Relationship ModelingabstractSports video captioning refers to the task of automatically generating a textual description for sports events (football, basketball, or volleyball games). Although a great deal of previous work has shown promising performance in producing a coarse and a general description of a video but lack of professional sports knowledge, it is still quite challenging to caption a sports video with multiple fine-grained player's actions and complex group relationship between players. In this paper, we present a novel hierarchical recurrent neural network-based framework with an attention mechanism for sports video captioning, in which a motion representation module is proposed to capture individual pose attribute and dynamical trajectory cluster information with extra professional sports knowledge, and a group relationship module is employed to design a scene graph for modeling players' interaction by a gated graph convolutional network. Moreover, we introduce a new dataset called sports video captioning dataset-volleyball for evaluation. The proposed model is evaluated on three widely adopted public datasets and our collected new dataset, on which the effectiveness of our method is well demonstrated. Mengshi Qi, Yunhong Wang 0001, Annan Li, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | stagNet: An Attentive Semantic RNN for Group Activity and Individual Action RecognitionabstractIn real life, group activity recognition plays a significant and fundamental role in a variety of applications, e.g. sports video analysis, abnormal behavior detection, and intelligent surveillance. In a complex dynamic scene, a crucial yet challenging issue is how to better model the spatio-temporal contextual information and inter-person relationship. In this paper, we present a novel attentive semantic recurrent neural network (RNN), namely, stagNet, for understanding group activities and individual actions in videos, by combining the spatio-temporal attention mechanism and semantic graph modeling. Specifically, a structured semantic graph is explicitly modeled to express the spatial contextual content of the whole scene, which is further incorporated with the temporal factor through structural-RNN. By virtue of the “factor sharing” and “message passing” mechanisms, our stagNet is capable of extracting discriminative and informative spatio-temporal representations and capturing inter-person relationships. Moreover, we adopt a spatio-temporal attention model to focus on key persons/frames for improved recognition performance. Besides, a body-region attention and a global-part feature pooling strategy are devised for individual action recognition. In experiments, four widely-used public datasets are adopted for performance evaluation, and the extensive results demonstrate the superiority and effectiveness of our method. Mengshi Qi, Yunhong Wang 0001, Jie Qin 0004, Annan Li, Jiebo Luo 0001, Luc Van Gool |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | STC-GAN: Spatio-Temporally Coupled Generative Adversarial Networks for Predictive Scene ParsingabstractPredictive scene parsing is a task of assigning pixellevel semantic labels to a future frame of a video. It has many applications in vision-based artificial intelligent systems, e.g., autonomous driving and robot navigation. Although previous work has shown its promising performance in semantic segmentation of images and videos, it is still quite challenging to anticipate future scene parsing with limited annotated training data. In this paper, we propose a novel model called STC-GAN, Spatio-Temporally Coupled Generative Adversarial Networks for predictive scene parsing, which employ both convolutional neural networks and convolutional long short-term memory (LSTM) in the encoderdecoder architecture. By virtue of STC-GAN, both spatial layout and semantic context can be captured by the spatial encoder effectively, while motion dynamics are extracted by the temporal encoder accurately. Furthermore, a coupled architecture is presented for establishing joint adversarial training where the weights are shared and features are transformed in an adaptive fashion between the future frame generation model and predictive scene parsing model. Consequently, the proposed STCGAN is able to learn valuable features from unlabeled video data. We evaluate our proposed STC-GAN on two public datasets, i.e., Cityscapes and CamVid. Experimental results demonstrate that our method outperforms the state-of-the-art. Mengshi Qi, Yunhong Wang 0001, Annan Li, Jiebo Luo 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | KE-GAN: Knowledge Embedded Generative Adversarial Networks for Semi-Supervised Scene ParsingabstractIn recent years, scene parsing has captured increasing attention in computer vision. Previous works have demonstrated promising performance in this task. However, they mainly utilize holistic features, whilst neglecting the rich semantic knowledge and inter-object relationships in the scene. In addition, these methods usually require a large number of pixel-level annotations, which is too expensive in practice. In this paper, we propose a novel Knowledge Embedded Generative Adversarial Networks, dubbed as KE-GAN, to tackle the challenging problem in a semi-supervised fashion. KE-GAN captures semantic consistencies of different categories by devising a Knowledge Graph from the large-scale text corpus. In addition to readily-available unlabeled data, we generate synthetic images to unveil rich structural information underlying the images. Moreover, a pyramid architecture is incorporated into the discriminator to acquire multi-scale contextual information for better parsing results. Extensive experimental results on four standard benchmarks demonstrate that KE-GAN is capable of improving semantic consistencies and learning better representations for scene parsing, resulting in the state-of-the-art performance. Mengshi Qi, Yunhong Wang 0001, Jie Qin 0004, Annan Li |
CVPR | 4 |
| 2019 | Atrous Temporal Convolutional Network for Video Action SegmentationabstractFine-grained temporal human action segmentation in untrimmed videos is receiving increasing attention due to its extensive applications in surveillance, robotics, and beyond. It is crucial for an action segmentation system to be robust to the temporal scale of different actions since in practical applications the duration of an action can vary from less than a second to tens of minutes. In this paper, we introduce a novel atrous temporal convolutional network (AT-Net), which explicitly generates multiscale video contextual representations by utilizing atrous temporal pyramid pooling (ATPP) and has an architecture of encoder-decoder fully convolutional network. In the decoding stage, AT-Net combines multiscale contextual features with low-level local features to generate high-quality action segmentation results. Experiments on the 50 Salads, GTEA and JIGSAWS benchmarks demonstrate that AT-Net achieves improvement over the state of the art. Zhengyin Du, Annan Li, Yunhong Wang 0001 |
ICIP | 3 |
| 2019 | Adversarial Binary Coding for Efficient Person Re-IdentificationabstractPerson re-identification (ReID) aims at associating persons with the same identity across different views/scenes. Most existing methods improve matching accuracy by proposing high-dimensional real-valued features to represent person images comprehensively. However, considering the increasing data scale in real-world applications, the storage and matching efficiencies should be paid attention to as well. In this paper, we propose a binary coding approach for efficient ReID, inspired by the recent advances in adversarial learning. Specifically, the proposed Adversarial Binary Coding (ABC) implicitly fits the feature distribution to the expected binary one by optimizing the Wasserstein distance. To further enhance the semantic discriminability of binary codes, we seamlessly embed the ABC into a similarity measuring deep neural network. By end-to-end learning the framework, compact and discriminative binary features are generated for efficient and accurate ReID. Extensive experiments on large-scale benchmarks demonstrate the superiority of our approach over the state-of-the-art methods in both efficiency and accuracy. Zheng Liu 0014, Jie Qin 0004, Annan Li, Yunhong Wang 0001, Luc Van Gool |
ICME | 3 |
| 2019 | A Temporal Attentive Approach for Video-Based Pedestrian Attribute Recognition
Annan Li, Yunhong Wang 0001 |
PRCV (2) | 2 |
| 2019 | Hierarchical Integration of Rich Features for Video-Based Person Re-IdentificationabstractPerson re-identification (ReID) aims to associate the identity of pedestrians captured by cameras across non-overlapped areas. Video-based ReID plays an important role in intelligent video surveillance systems and has attracted growing attention in recent years. In this paper, we propose an end-to-end video-based ReID framework based on the convolutional neural network (CNN) for efficient spatio-temporal modeling and enhanced similarity measuring. Specifically, we build our descriptor of sequences by basic mathematical calculations on the semantic mid-level image features, which avoids the time consuming computations and the loss of spatial correlations. We further hierarchically extract image features from multiple intermediate CNN stages to build multi-level sequence descriptors. For a descriptor at one stage, we design an effective auxiliary pairwise loss which is jointly optimized with a triplet loss. To integrate hierarchical representation, we propose an intuitive yet effective summation-based similarity integration scheme to match identities more accurately. Furthermore, we extend our framework by a multi-model ensemble strategy, which effectively assembles three popular CNN models to represent walking sequences more comprehensively and improve the performance. Extensive experiments on three video-based ReID datasets show that the proposed framework outperforms the state-of-the-art methods. Zheng Liu 0014, Yunhong Wang 0001, Annan Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | BUAA-PRO: A Tracking Dataset with Pixel-Level Annotation
Annan Li, Yunhong Wang 0001 |
BMVC | 1 |
| 2018 | stagNet: An Attentive Semantic RNN for Group Activity Recognition
Mengshi Qi, Jie Qin 0004, Annan Li, Yunhong Wang 0001, Jiebo Luo 0001, Luc Van Gool |
ECCV (10) | 3 |
| 2018 | Combining Multiple Deep Features for Glaucoma ClassificationabstractGlaucoma is one of the leading cause of blindness. Although there is still no cure, early detection can prevent serious vision loss. Therefore automated glaucoma detection/classification is an important issue. In the past decade, segmentation based approach such as those based on cup-to-disc-ratio are popular, but single indicator limit its performance. Recently, convolutional neural network based image classification approaches that can use more image cues achieve good performance. In this paper, we propose a new glaucoma classification by combining multiple features extracted by different convolutional neural networks. Its effectiveness is clearly demonstrated on the publicly available Origa [1] dataset. It achieves an area under the receiver operating characteristic curve of 0.8483, which better than the 0.838 given by on manual marked cup-to-disc-ratio. To our knowledge, it is the first approach surpass human in glaucoma classification. Annan Li, Yunhong Wang 0001, Jun Cheng 0003, Jiang Liu 0001 |
ICASSP | 1 |
| 2018 | Learning supervised descent directions for optic disc segmentation
Annan Li, Zhiheng Niu, Jun Cheng 0003, Fengshou Yin, Damon Wing Kee Wong, Shuicheng Yan, Jiang Liu 0001 |
Neurocomputing | 1 |
| 2017 | Online Cross-Modal Scene Retrieval by Binary Representation and Semantic GraphabstractIn recent years, cross-modal scene retrieval has attracted more attention. However, most existing approaches neglect the semantic relationship between objects in a scene together with the embedded spatial layouts. Moreover, these methods mostly apply the batch learning strategy, which is not suitable for processing streaming data. To address the aforementioned problems, we propose a new framework for online cross-modal scene retrieval based on binary representations and semantic graph. Specially, we adopt the cross-modal hashing based on the quantization loss of different modalities. By introducing the semantic graph, we are able to extract wealthy semantics and measure their correlation across different modalities. Further more, we propose a two-step optimization procedure based on stochastic gradient descent for online update. Experimental results on four datasets show the superiority of our approach over the state-of-the-art. Mengshi Qi, Yunhong Wang 0001, Annan Li |
ACM Multimedia | 3 |
| 2016 | NUS-PRO: A New Visual Tracking ChallengeabstractNumerous approaches on object tracking have been proposed during the past decade with demonstrated success. However, most tracking algorithms are evaluated on limited video sequences and annotations. For thorough performance evaluation, we propose a large-scale database which contains 365 challenging image sequences of pedestrians and rigid objects. The database covers 12 kinds of objects, and most of the sequences are captured from moving cameras. Each sequence is annotated with target location and occlusion level for evaluation. A thorough experimental evaluation of 20 state-of-the-art tracking algorithms is presented with detailed analysis using different metrics. The database is publicly available and evaluation can be carried out online for fair assessments of visual tracking algorithms. Annan Li, Yi Wu 0001, Ming-Hsuan Yang 0001, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Clothing Attributes Assisted Person ReidentificationabstractPerson reidentification across nonoverlapping camera views is a rather challenging task. Due to the difficulties in obtaining identifiable faces, clothing appearance becomes the main cue for identification purposes. In this paper, we present a comprehensive study on clothing attributes assisted person reidentification. First, the body parts and their local features are extracted for alleviating the pose-misalignment issue. A latent support vector machine (LSVM)-based person reidentification approach is proposed to describe the relations among the low-level part features, middle-level clothing attributes, and high-level reidentification labels of person pairs. Motivated by the uncertainties of clothing attributes, we treat them as real-value variables instead of using them as discrete variables. Moreover, a large-scale real-world dataset with 10 camera views and about 200 subjects is collected and thoroughly annotated for this paper. The extensive experiments on this dataset show: 1) part features are more effective than features extracted from the holistic human bounding boxes; 2) the clothing attributes embedded in the LSVM model may further boost reidentification performance compared with support vector machine without clothing attributes; and 3) treating clothing attributes as real-value variables is more effective than using them as discrete variables in person reidentification. Annan Li, Luoqi Liu, Kang Wang 0002, Si Liu 0001, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | CovGa: A novel descriptor based on symmetry of regions for head pose estimation
Bingpeng Ma, Annan Li, Xiujuan Chai, Shiguang Shan |
Neurocomputing | 2 |
| 2014 | Object Tracking With Only Background CuesabstractBackground cues mainly play a supplementary or accompanying role in most previous approaches for object tracking. If object tracking is treated as a binary target/background classification problem, then the similarity with the target and the difference from the background can be considered equally informative. This leads to an interesting question: is it possible to perform object tracking using background cues only? To answer this question, we propose an object tracking approach that utilizes background cues only. The extensive experimental results positively validate the possibility of performing object tracking using background cues only. The results revealed in this paper also provide a new reference in designing future object tracking methods. Annan Li, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | Coupled Bias-Variance Tradeoff for Cross-Pose Face RecognitionabstractSubspace-based face representation can be looked as a regression problem. From this viewpoint, we first revisited the problem of recognizing faces across pose differences, which is a bottleneck in face recognition. Then, we propose a new approach for cross-pose face recognition using a regressor with a coupled bias-variance tradeoff. We found that striking a coupled balance between bias and variance in regression for different poses could improve the regressor-based cross-pose face representation, i.e., the regressor can be more stable against a pose difference. With the basic idea, ridge regression and lasso regression are explored. Experimental results on CMU PIE, the FERET, and the Multi-PIE face databases show that the proposed bias-variance tradeoff can achieve considerable reinforcement in recognition performance. Annan Li, Shiguang Shan, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2011 | Face recognition based on non-corresponding region matchingabstractIn previous works of face recognition, similarity between faces is measured by comparing corresponding face regions. That is to say, matching eyes with eyes and mouths with mouths etc.. In this paper, we propose that face can be also recognized by matching non-corresponding facial regions. In another word face can be recognized by matching eyes with mouths, for example. Specifically, the problem we study in this paper can be formulated as how to measure the possibility whether two non-corresponding face regions belong to the same face. We propose that the possibility can be measured via canonical correlation analysis. Experimental results show that it is feasible to recognize face via non-corresponding region matching. The proposed method provides an alternative and more flexible way to recognize faces. Annan Li, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
ICCV | 1 |
| 2011 | Cross-pose face recognition based on partial least squares
Annan Li, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
Pattern Recognit. Lett. | 1 |
| 2009 | Maximizing intra-individual correlations for face recognition across pose differencesabstractThe variations of pose lead to significant performance decline in face recognition systems, which is a bottleneck in face recognition. A key problem is how to measure the similarity between two image vectors of unequal length that viewed from different pose. In this paper, we propose a novel approach for pose robust face recognition, in which the similarity is measured by correlations in a media subspace between different poses on patch level. The media subspace is constructed by Canonical Correlation Analysis, such that the intra-individual correlations are maximized. Based on the media subspace two recognition approaches are developed. In the first, we transform non-frontal face into frontal for recognition. And in the second, we perform recognition in the media subspace with probabilistic modeling. The experimental results on FERET database demonstrate the efficiency of our approach. Annan Li, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
CVPR | 1 |
| 2008 | Recovering 3D facial shape via coupled 2D/3D space learningabstractThis paper presents a method for recovering 3D facial shape from single image via learning the relationship between the 2D intensity images and the 3D facial shapes. With a coupled training set, the intensity images and their corresponding facial shapes make up two vector spaces respectively. But only the correlated components in both spaces are useful for inference, so there must be embedded hidden subspaces in each space which preserve the inter-space correlation information. Thus by learning the projection onto hidden subspaces based on maximum correlation criteria and optimizing the linear transform between the hidden spaces, 3D facial shape is inferred from the intensity image. The effectiveness of the method is demonstrated on both synthesized and real world data. Annan Li, Shiguang Shan, Xilin Chen 0001, Xiujuan Chai, Wen Gao 0001 |
FG | 1 |