VLDB 2026 Research / reviewers in the wild / expert
Chunhui Liu 0002
dblp:20/5393-2
· DBLP profile ↗
10ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0001-6707-5187ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data ExpansionabstractTransformer-based multimodal models are widely used in industrialscale recommendation, search, and advertising systems for content understanding and relevance ranking.Enhancing labeled training data quality and cross-modal fusion significantly improves model performance, influencing key metrics such as quality view rates and ad revenue.High-quality annotations are crucial for advancing content modeling, yet traditional statistical-based active learning (AL) methods face limitations: they struggle to detect overconfident misclassifications and are less effective in distinguishing semantically similar items in deep neural networks.Additionally, audio information plays an increasing role, especially in short-video platforms, yet most pretrained multimodal architectures primarily focus on text and images.While training from scratch across all three modalities is possible, it sacrifices the benefits of leveraging existing pretrained visual-language (VL) and audio models.To address these challenges, we propose kNN-based Latent Space Broadening (LSB) to enhance AL efficiency, achieving an up to 9% recall improvement at 80% precision on proprietary datasets.Additionally, we introduce Vision-Language Modeling with Audio Enhancement (VLMAE), a mid-fusion approach integrating audio into VL models, yielding up * Author corresponded for this research. Yu Sun 0088, Ruixiao Sun, Chunhui Liu 0002, Fangming Zhou, Ze Jin, Xiang Shen 0001, Zhuolin Hao, Hongyu Xiong |
KDD (2) | 4 |
| 2022 | TubeR: Tubelet Transformer for Video Action DetectionabstractWe propose TubeR: a simple solution for spatio-temporal video action detection. Different from existing methods that depend on either an offline actor detector or hand-designed actor-positional hypotheses like proposals or anchors, we propose to directly detect an action tubelet in a video by simultaneously performing action localization and recognition from a single representation. TubeR learns a set of tubelet-queries and utilizes a tubelet-attention module to model the dynamic spatio-temporal nature of a video clip, which effectively reinforces the model capacity compared to using actor-positional hypotheses in the spatio-temporal space. For videos containing transitional states or scene changes, we propose a context aware classification head to utilize short-term and long-term context to strengthen action classification, and an action switch regression head for detecting the precise temporal action extent. TubeR directly produces action tubelets with variable lengths and even maintains good results for long video clips. TubeR outperforms the previous state-of-the-art on commonly used action detection datasets AVA, UCF101-24 and JHMDB51-21. Code will be available on GluonCV(https://cv.gluon.ai/). Jiaojiao Zhao, Yanyi Zhang, Xinyu Li 0003, Hao Chen 0024, Bing Shuai, Chunhui Liu 0002, Kaustav Kundu, Yuanjun Xiong, Davide Modolo, Ivan Marsic, Cees Snoek, Joseph Tighe |
CVPR | 7 |
| 2022 | NUTA: Non-uniform Temporal Aggregation for Action RecognitionabstractIn the world of action recognition research, one primary focus has been on how to construct and train networks to model the spatial-temporal volume of an input video. These methods typically uniformly sample a segment of an input clip (along the temporal dimension). However, not all parts of a video are equally important to determine the action in the clip. In this work, we focus instead on learning where to extract features, so as to focus on the most informative parts of the video. We propose a method called the non-uniform temporal aggregation (NUTA), which aggregates features only from informative temporal segments. We also introduce a synchronization method that allows our NUTA features to be temporally aligned with traditional uniformly sampled video features, so that both local and clip-level features can be combined. Our model has achieved state-of-the-art performance on four widely used large-scale action-recognition datasets (Kinetics400, Kinetics700, Something-something V2 and Charades). In addition, we have created a visualization to illustrate how the proposed NUTA method selects only the most relevant parts of a video clip. Xinyu Li 0003, Chunhui Liu 0002, Bing Shuai, Yi Zhu 0001, Hao Chen 0024, Joseph Tighe |
WACV | 2 |
| 2022 | SSCAP: Self-supervised Co-occurrence Action Parsing for Unsupervised Temporal Action SegmentationabstractTemporal action segmentation is a task to classify each frame in the video with an action label. However, it is quite expensive to annotate every frame in a large corpus of videos to construct a comprehensive supervised training dataset. Thus in this work we propose an unsupervised method, namely SSCAP, that operates on a corpus of unlabeled videos and predicts a likely set of temporal segments across the videos. SSCAP leverages Self-Supervised learning to extract distinguishable features and then applies a novel Co-occurrence Action Parsing algorithm to not only capture the correlation among sub-actions under-lying the structure of activities, but also estimate the temporal path of the sub-actions in an accurate and general way. We evaluate on both classic datasets (Breakfast, 50Sal-ads) and the emerging fine-grained action dataset (Fine-Gym) with more complex activity structures and similar sub-actions. Results show that SSCAP achieves state-of-the-art performance on all datasets and can even outperform some weakly-supervised approaches, demonstrating its effectiveness and generalizability. Zhe Wang 0013, Hao Chen 0024, Xinyu Li 0003, Chunhui Liu 0002, Yuanjun Xiong, Joseph Tighe, Charless C. Fowlkes |
WACV | 4 |
| 2021 | Selective Feature Compression for Efficient Activity Recognition InferenceabstractMost action recognition solutions rely on dense sampling to precisely cover the informative temporal clip. Extensively searching temporal region is expensive for a real-world application. In this work, we focus on improving the inference efficiency of current action recognition backbones on trimmed videos, and illustrate that an action model can accurately classify an action with a single pass over the video unlike the multi-clip sampling common with SOTA by learning to drop non-informative features. We present Selective Feature Compression (SFC), an action recognition inference strategy that greatly increases model inference efficiency without compromising accuracy. Different from previous works that compress kernel size and decrease the channel dimension, we propose to compress features along the spatio-temporal dimensions without the need to change backbone parameters. Our experiments on Kinetics-400, UCF101 and ActivityNet show that SFC is able to reduce inference speed by 6-7x and memory usage by 5-6x compared with the commonly used 30 crop dense sampling procedure, while also slightly improving Top1 Accuracy. We perform thorough quantitative and qualitative evaluation and show how our SFC learns to attend to important video regions for the task of action recognition. Chunhui Liu 0002, Xinyu Li 0003, Hao Chen 0024, Davide Modolo, Joseph Tighe |
ICCV | 1 |
| 2021 | VidTr: Video Transformer Without ConvolutionsabstractWe introduce Video Transformer (VidTr) with separable-attention for video classification. Comparing with commonly used 3D networks, VidTr is able to aggregate spatiotemporal information via stacked attentions and provide better performance with higher efficiency. We first introduce the vanilla video transformer and show that transformer module is able to perform spatio-temporal modeling from raw pixels, but with heavy memory usage. We then present VidTr which reduces the memory cost by 3.3× while keeping the same performance. To further optimize the model, we propose the standard deviation based topK pooling for attention (pooltopK_std), which reduces the computation by dropping non-informative features along temporal dimension. VidTr achieves state-of-the-art performance on five commonly used datasets with lower computational requirement, showing both the efficiency and effectiveness of our design. Finally, error analysis and visualization show that VidTr is especially good at predicting actions that require long-term temporal reasoning. Yanyi Zhang, Xinyu Li 0003, Chunhui Liu 0002, Bing Shuai, Yi Zhu 0001, Biagio Brattoli, Hao Chen 0024, Ivan Marsic, Joseph Tighe |
ICCV | 3 |
| 2020 | Application of Multi-Object Tracking with Siamese Track-RCNN to the Human in Events DatasetabstractMulti-object tracking systems often consist of a combination of a detector, a short term linker, a re-identification feature extractor and a solver that takes the output from these separate components and makes a final prediction. Differently, this work aims to unify all these in a single tracking system. Towards this, we propose Siamese Track-RCNN, a two stage detect-and-track framework which consists of three functional branches: (1) the detection branch localizes object instances; (2) the Siamese-based track branch estimates the object motion and (3) the object re-identification branch re-activates the previously terminated tracks when they re-emerge. We used this design and apply it to the Human in Events dataset. Bing Shuai, Andrew G. Berneshawi, Manchen Wang, Chunhui Liu 0002, Davide Modolo, Xinyu Li 0003, Joseph Tighe |
ACM Multimedia | 4 |
| 2020 | A Benchmark Dataset and Comparison Study for Multi-modal Human Action AnalyticsabstractLarge-scale benchmarks provide a solid foundation for the development of action analytics. Most of the previous activity benchmarks focus on analyzing actions in RGB videos. There is a lack of large-scale and high-quality benchmarks for multi-modal action analytics. In this article, we introduce PKU Multi-Modal Dataset (PKU-MMD), a new large-scale benchmark for multi-modal human action analytics. It consists of about 28,000 action instances and 6.2 million frames in total and provides high-quality multi-modal data sources, including RGB, depth, infrared radiation (IR), and skeletons. To make PKU-MMD more practical, our dataset comprises two subsets under different settings for action understanding, namely Part I and Part II. Part I contains 1,076 untrimmed video sequences with 51 action classes performed by 66 subjects, while Part II contains 1,009 untrimmed video sequences with 41 action classes performed by 13 subjects. Compared to Part I, Part II is more challenging due to short action intervals, concurrent actions and heavy occlusion. PKU-MMD can be leveraged in two scenarios: action recognition with trimmed video clips and action detection with untrimmed video sequences. For each scenario, we provide benchmark performance on both subsets by conducting different methods with different modalities under two evaluation protocols, respectively. Experimental results show that PKU-MMD is a significant challenge to many state-of-the-art methods. We further illustrate that the features learned on PKU-MMD can be well transferred to other datasets. We believe this large-scale dataset will boost the research in the field of action analytics for the community. Jiaying Liu 0001, Sijie Song, Chunhui Liu 0002, Yanghao Li, Yueyu Hu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2017 | Temporal Perceptive Network for Skeleton-Based Action Recognition
Yueyu Hu, Chunhui Liu 0002, Yanghao Li, Jiaying Liu 0001 |
BMVC | 2 |
| 2017 | Online action detection and forecast via Multitask deep Recurrent Neural NetworksabstractOnline human action detection and forecast on untrimmed 3D skeleton sequences is a novel task based on traditional action recognition and has not been fully studied. Its aim is to localize and recognize one action in a long sequence while doing forecasting task at the same time. In this paper, we propose an online detection algorithm featuring Multi-Task Recurrent Neural Network to solve this problem. First, a deep Long Short Term Memory (LSTM) network is designed for feature extraction and temporal dynamic modeling. Then we utilize a classification subnetwork to classify one action, and predict the status of it at the same time. To forecast the occurrence of actions and estimate the accurate time of occurrence, we incorporate a regression subnetwork to our model. Then we split the action classes to three stages and train the model by optimizing a joint classification regression objective function. Experimental results show that the proposed model achieves satisfactory results on online action detection and forecast. Chunhui Liu 0002, Yanghao Li, Yueyu Hu, Jiaying Liu 0001 |
ICASSP | 1 |