Longteng Kong

dblp:180/2842 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-9772-4579ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Sample Synthesis and Decoupled Distillation for Black-Box Attack
Longteng Kong
ICPR (10)2
2025 CASIA-PR-V1: A Multi-Ethnic, Multi-Device and Cross-Spectral Dataset and a Multiscale Disentangled Model for Periocular Recognition
abstract
Periocular recognition is regarded as an alternative trait for biometric recognition that can effectively solve the identification problem under large occlusions. However, few datasets are tailored for periocular recognition. For most compromises, iris datasets at near-infrared wavelengths, miss information about the eyebrows or eyelids. In this paper, a challenging dataset for real scenarios named CASIA-PR-V1 with evaluation protocols is released for periocular recognition. It is collected from multiple types of mobile devices with different resolutions or wavelengths. A rich set of attributes, e.g., ethnicities, is tagged to support fine-grained classification tasks. Moreover, we consider a wide range of noisy data in unconstrained environment, especially for glasses. Superior to its counterparts, this periocular dataset is highly valuable for studying cross-device and cross-spectral periocular recognition with occlusions, as well as fine-grained attribute classification. Additionally, a multiscale disentangled model is proposed to extract discriminating representations for periocular recognition with severe occlusions. Extensive experiments are conducted on CASIA-PR-V1, and the results indicate the superiority of our model for unconstraint periocular recognition.
Yiwei Ru, Yushan Han, Longteng Kong, Zijian Wang 0009, Yong He 0009, Zhenan Sun
IEEE Trans. Multim.4
2025 Multi-Scale Reconstruction and Relation Decomposition Modeling for Group Activity Recognition
abstract
Group activity recognition (GAR) is a challenging task in computer vision, which needs to comprehensively model the spatiotemporal relations among actors. However, most previous methods tend to only model unitary actor relations and directly aggregate actor features to form group representation at a single scale. To address these issues, we propose a novel GAR approach termed multi-scale cross-distance transformer (MSCD-Former), capable of capturing diverse actor relation contexts in multiple spatiotemporal scales. A cross-distance attentive block (CDA-Block) is designed to decompose the actor relations into local and distant ones, diversifying the relation features in rearranged groups. The multi-scale group descriptors are then enhanced by deploying stacked CDA-Blocks to cascaded stages and tightening the sampling scales accordingly. Moreover, we introduce a multi-scale reconstructive learning measure (MSR-Learning) between adjacent scales of CDA-Blocks. Via the reconstruction of actor relational features from lower scales to upper scales, MSR-Learning can enforce semantic consistency in multiple spatiotemporal scales. Consequently, our MSCD-Former boosts GAR by fusing such discriminative relation features of different scales. We extensively evaluate the proposed approach on the VolleyTactic, Volleyball, Collective Activity, NBA, and JRDB-PAR datasets, and the experimental results demonstrate its superiority.
Longteng Kong, Yongjian Huai, Jie Qin 0004
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Contextualized Relation Predictive Model for Self-Supervised Group Activity Representation Learning
abstract
Group activity analysis has attracted remarkable attention recently due to the widespread applications in security, entertainment and military. This article targets at learning group activity representations with self-supervision, which differs from the majorities relying heavily on manually annotated labels. Moreover, existing Self-Supervised Learning (SSL) methods for videos are sub-optimal to generate such representations because of the complex context dynamics in group activities. In this article, an end-to-end framework termed Contextualized Relation Predictive Model (Con-RPM) is proposed for self-supervised group activity representation learning with predictive coding. It involves the Serial-Parallel Transformer Encoder (SPTrans-Encoder) to model the context of spatial interactions and temporal variations, and the Hybrid Context Transformer Decoder (HConTrans-Decoder) to predict the future spatio-temporal relations guided by holistic scene context. Additionally, to improve the discriminability and consistency of prediction, we introduce a united loss integrating group-wise and person-wise contrastive losses in frame-level as well as the adversarial loss in global sequence-level. Consequently, our Con-RPM learns robust group representations via describing temporal evolutions of individual relationships and scene semantics explicitly. Extensive experimental results on downstream tasks indicate the effectiveness and generalization of our model in self-supervised learning, and present state-of-the-art performance on the Volleyball, Collective Activity, VolleyTactic, and Choi's New datasets.
Longteng Kong, Yushan Han, Jie Qin 0004, Zhenan Sun
IEEE Trans. Multim.2
2023 Group Activity Representation Learning With Long-Short States Predictive Transformer
abstract
The research goal of this paper is to learn the group activity representations in a self-supervised fashion instead of through the use of conventional methods that rely on manually annotated labels. It is essential for this task to better describe the complex group states and their future transitions. To this end, we propose a long-short state predictive Transformer (LSSPT), which mines the meaningful spatiotemporal features of group activities by predicting the future group states with long- and short-term historical state dynamics. LSSPT consists of an encoder that models diverse spatiotemporal state representations in the observation, together with a decoder that exploits rich dynamic patterns by attending to both the short-term spatial context and long-term history state evolutions to predict future group states. Furthermore, we consider the distinguishability and consistency of the predicted states and introduce a joint learning mechanism to optimize the models, enabling LSSPT to describe more reliable state transitions. Finally, extensive experiments are carried out to evaluate the learned representation on downstream tasks on the Volleyball, Collective Activity and VolleyTactic datasets, which showcases the method’s state-of-the-art performance over the existing self-supervised learning approaches.
Longteng Kong, Duoxuan Pei, Zhaofeng He 0001, Di Huang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Key Role Guided Transformer for Group Activity Recognition
abstract
Group Activity Recognition (GAR) is a challenging task, where modeling spatio-temporal relationships among participants plays a fundamental role. To address this issue, we propose a novel end-to-end trainable network, termed Key Role Guided Transformer (KRGFormer). Different from current methods that concurrently take all individuals into account for global reasoning, it captures crucial contextual information by emphasizing a set of key individuals in a coarse-to-fine manner considering that group activities are usually dominated by them. A Key Individual-aware Block (KIaBlock) is designed to select relevant individuals and enhance their relationships with the reservation of global dependencies of the entire group. The representations are then iteratively refined by deploying multiple stacked KIaBlocks, leading to a stronger discriminative power to distinguish group activities. Moreover, along with general data augmentation schemes, several “actor-centric” ones are presented to relieve the over-fitting risk, which further boost the performance. We extensively evaluate the proposed approach on the Volleyball, VolleyTactic and NBA datasets, and the experimental results demonstrate its superiority.
Duoxuan Pei, Di Huang 0001, Longteng Kong, Yunhong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Group Activity Representation Learning with Self-supervised Predictive Coding
Longteng Kong, Zhaofeng He 0001, Man Zhang 0005, Yunzhi Xue
PRCV (3)1
2022 Hierarchical Long-Short Transformer for Group Activity Recognition
Zhaofeng He 0001, Longteng Kong
PRCV (3)3
2022 Spatio-Temporal Player Relation Modeling for Tactic Recognition in Sports Videos
abstract
Tactic recognition in sports videos is a challenging task. To address this, we present a novel spatio-temporal relation modeling approach, which captures both detailed player interactions and long-range group dynamics in tactics. In spatial modeling, we propose an Adaptive Graph Convolutional Network (A-GCN), and it represents individual and common patterns of data through local and global graphs to learn diverse player interactions. In temporal modeling, we propose an Attentive Temporal Convolutional Network (A-TCN) and with spatial configurations as input, it builds group dynamics and is robust to redundant content by considering sequence dependencies. Due to adaptive interaction and attentive dynamics modeling, our approach is able to comprehensively describe team cooperation over time in a tactic. We extensively evaluate the proposed approach on the Volleyball dataset and a newly collected VolleyTactic dataset, and the experimental results show its advantage.
Longteng Kong, Duoxuan Pei, Di Huang 0001, Yunhong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Towards Practical Compressed Video Action Recognition: A Temporal Enhanced Multi-Stream Network
abstract
Current compressed video action recognition methods are mainly based on complete data. However, in a real transmission scenario, the compressed video packets are usually disorderly received and even lost due to network jitters or congestion. To recognize actions in early phases with limited packets, e.g. for quickly forecasting possible potential risks, in this paper, we propose a Temporal Enhanced Multi-Stream Network (TEMSN) towards practical compressed video action recognition. First, we make use of three modalities in the compressed domain as complementary cues and build a multi-stream network to capture rich information from compressed video packets. Second, we design a temporal enhanced module based on an Encoder-Decoder structure, which is applied to each stream to infer missing packets, generating more accurate action dynamics. Thanks to the multiple modalities and their temporal enhancement, our approach better models actions with partial available compressed video packets. Experiments on the HMDB-51 and UCF-101 datasets validate its effectiveness and efficiency.
Longteng Kong, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001, Yunhong Wang 0001
ICPR2
2020 A Joint Framework for Athlete Tracking and Action Recognition in Sports Videos
abstract
Sports video analysis has received increasing attention in recent years. Athlete tracking and action recognition are its two major issues that are highly related to each other; however, they are individually considered and processed in the existing studies. In this paper, we propose a joint framework for athlete tracking and action recognition in sports videos. In athlete tracking, we propose a scaling and occlusion robust tracker, named scaling and occlusion robust compressive tracking (CT), to localize the position of specific athlete in each frame. It follows the approach of CT but extends it in two aspects, i.e., scale refinement as well as occlusion recovery. For the former, an objectness method, edge box, is adopted to generate proposals, which replace the fixed sampling boxes in CT and better fit the scales of the candidate objects. For the latter, a candidate obstruction-based solution is presented, which brings in additional trackers to detect possible obstructions and to relocate the target as occlusion ends. Regarding action recognition, we propose a long-term recurrent region-guided convolutional network, which recognizes pre-defined actions by modeling discriminative temporal cues of the tracking results. We employ SPP-net to extract the robust feature of the tracked region of each frame. The features of all the frames are then fed into a stack of recurrent sequence models to capture the long-term region-level information. We extensively evaluate the proposed approach on a newly collected sports video benchmark and on the off-the-shelf UIUC2 dataset, and the experimental results clearly show its effectiveness.
Longteng Kong, Di Huang 0001, Jie Qin 0004, Yunhong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Long-Term Action Dependence-Based Hierarchical Deep Association for Multi-Athlete Tracking in Sports Videos
abstract
Tracking multiple athletes in sports videos is a very challenging Multi-Object Tracking (MOT) task, as athletes generally share high similarity in appearance with large deformations. In this paper, unlike the existing hand-crafted solutions, we propose a novel and effective approach to this issue, which hierarchically associates detections of the same identity through discriminative and robust deep features. First, in detection association, we make use of athlete appearances and poses instead of traditional position cues to generate short tracklets for better initialization. Second, in tracklet association, a new deep architecture, namely Siamese Tracklet Affinity Networks (STAN), is presented, which is able to bi-directionally simulate the unseen dynamics of actions, comprehensively models the long-term action dependences, and sequentially estimates their affinity. Such hierarchical association is finally solved as a minimum-cost network flow problem. We extensively evaluate the proposed approach on the APIDIS, NCAA Basketball and VolleyTrack (newly collected) databases, and the experimental results show its advantages.
Longteng Kong, Di Huang 0001, Yunhong Wang 0001
IEEE Trans. Image Process.1
2019 A Robust Multi-Athlete Tracking Algorithm by Exploiting Discriminant Features and Long-Term Dependencies
Nan Ran, Longteng Kong, Yunhong Wang 0001, Qingjie Liu 0001
MMM (1)2
2018 Hierarchical Attention and Context Modeling for Group Activity Recognition
abstract
Group activity recognition in videos is a challenging task, with two major issues, i.e. attending to those persons and their body parts that contribute significantly to the activity, and modeling contextual person structures in the group. Most previous approaches fail to provide a practical solution to jointly address both issues, however. In this paper, we propose to simultaneously deal with both issues via a hierarchical attention and context modeling framework based on Long Short-Term Memory (LSTM) networks. For the former, we propose `Hierarchical Attention Networks' applied at the part/person level, capable of attending distinctively to different persons and their body parts. For the latter, we build `Hierarchical Context Networks' that take the attentively pooled person-level features as input and recurrently model intra/inter-group contextual structures. The attentive and contextual representations are concatenated and fed into another LSTM to generate high-level discriminative temporal representations for group activity recognition. Extensive experiments on two widely-used group activity datasets demonstrate the effectiveness and superiority of the proposed framework.
Longteng Kong, Jie Qin 0004, Di Huang 0001, Yunhong Wang 0001, Luc Van Gool
ICASSP1
2016 Scaling and occlusion robust athlete tracking in sports videos
abstract
This paper proposes a novel approach to athlete tracking in sports videos. It follows the framework of Compressive Tracking (CT), but extends it by two manners, i.e. scale refinement as well as occlusion recovery. For the former, an objectness method, namely Edge Box (EB) is adopted to generate proposals, replacing the fixed sampling box in CT, which better fits the scales of the candidate objects. For the latter, a candidate obstruction based solution is presented, which makes use of additional trackers to detect possible obstructions especially the ones possessing highly similar appearances as the target one, and relocate the target as occlusion ends. Therefore, the proposed method inherits the advantage of CT in robust object modelling and fast processing speed, and embodies the tolerance to occlusion and scaling. We evaluate the proposed method on a collection of videos of beach volleyball games, and the experimental results and the comparison with recent advanced trackers highlight its effectiveness.
Jianghu Lu, Di Huang 0001, Yunhong Wang 0001, Longteng Kong
ICASSP4