VLDB 2026 Research / reviewers in the wild / expert
Qianwen Cao
dblp:286/4863
· DBLP profile ↗
12ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Sparse Transformer Inference on GPUabstractLarge language models (LLMs) are popular around the world due to their powerful understanding capabilities. As the core component of LLMs, accelerating Transformer through parallelization has gradually become a hot research topic. Mask layers introduce sparsity into Transformer to reduce calculations. However, previous works rarely focus on the performance optimization of sparse Transformer. In addition, current static operator fusion schemes fail to adapt to diverse application scenarios. To address the above problems, we propose STOF, a framework that incorporates optimizations for Sparse Transformer that enables flexible masking and Operator Fusion on GPU. For multi-head attention (MHA) structure, STOF maps the computation to row-wise or block-wise kernels with unique storage formats according to analytical modeling. For downstream operators, STOF maps the fusion scheme to compilation templates and determines the optimal running configuration through two-stage searching. The experimental results show that compared to the state-of-the-art work, STOF achieves maximum speedups of 1.6× in MHA computation and 1.4× in end-to-end inference. Wenhao Dai, Haodong Deng, Mengfei Rong, Fangxin Liu, Hailong Yang 0002, Qianwen Cao, Qingxiao Sun |
PPoPP | 8 |
| 2026 | Efficient video captioning annotation: a semi-automated framework via LLM-driven heuristics
Qianwen Cao, Guixin Liao |
Multim. Syst. | 1 |
| 2025 | QueryNILM: Query based Non-Intrusive Load Monitoring with Appliance EmbeddingsabstractNon-Intrusive Load Monitoring (NILM), also known as energy disaggregation, aims to infer appliance-level energy consumption from aggregated readings obtained from a single meter measuring a household’s total electricity demand. The increasing availability of large-scale and fine-grained household electricity usage datasets has facilitated the adoption of deep learning techniques for NILM, leading to state-of-the-art performance. However, a key limitation of current approaches is the requirement to train separate models for each appliance, which impedes both model scalability and the ability to integrate new appliances. To address this limitation, we propose a query-based NILM model that leverages appliance embeddings as queries, enabling joint training for multi-appliance disaggregation. By encoding appliance energy usage patterns as embeddings, the proposed model can be queried with these embeddings to extract the energy consumption of corresponding appliances from the aggregated readings and facilitate zero-shot disaggregation. Experimental results on the UK-DALE dataset demonstrate the efficacy of the proposed query-based NILM model. Jie Jiang 0011, Qiuqiang Kong, Qianwen Cao |
IJCNN | 5 |
| 2025 | From Skeleton to Flesh: Aggregated Relational Transformer Towards Controllable Video Captioning with Two-Step DecodingabstractVideo captioning is the task of automatically producing natural language descriptions for a given video. The great success of the Transformer architecture in NLP and CV has inspired recent attempts at constructing an end-to-end Transformer-based model for video captioning. Although achieving impressive performance, the one-step framework relies on brute force to learn highly compressed semantics from a large number of video patches, lacking the review process, which is a common behavior of human beings for understanding videos/papers/books. Taking the rough but significant impression captured at first sight as the skeleton, additional review can verify and investigate specific information as flesh to deepen precise recognition. In this work, we introduce the review process in the Transformer-based framework and propose a novel network, Aggregated Relational Transformer (ART) to conduct two-step decoding for video captioning. Since the relation triplet concisely summarizes the main structure, it is assigned as the objective of the first-pass decoding. Then the relations are utilized as a skeleton and guide the review process to potentially obtain better captions by looking into semantic components with a global view, where relations can also play the role of the prompt signal for the controllable caption generation at the second decoding pass. Extensive experiments show that our method achieves the SOTA performance on MSVD, MSRVTT, and VATEX datasets for video captioning, and is capable of controlling captions to respond to different semantic contexts. Qianwen Cao, Heyan Huang, Boran Wang |
ICMR | 1 |
| 2024 | Screening through a broad pool: Towards better diversity for lexically constrained text generation
Changsen Yuan, Heyan Huang, Yixin Cao 0002, Qianwen Cao |
Inf. Process. Manag. | 4 |
| 2023 | Ada-SwinBERT: Adaptive Token Selection for Efficient Video Captioning with Online Self-DistillationabstractVideo captioning aims at producing textual descriptions for the given video. Benefiting from the self-attention mechanism for capturing long-distance dependencies between video patches and language sentences, the fully Transformer-based models achieve promising performance recently. However, due to continuous temporal information, there exists a large amount of redundant and unimportant visual content. Indiscriminate use of video patches results in expensive computation and inefficient use of resources. To tackle this issue, we propose Ada-SwinBERT, a novel approach that adaptively selects salient video tokens to achieve a balance between efficiency and performance for video captioning. Moreover, we devise a training strategy with online self-distillation to make up for the information loss caused by discarding video tokens. Video-text alignment knowledge distilled from the teacher leads to a robust training process. By pruning 78.1% input tokens hierarchically, our approach greatly reduces 62.0% FLOPs compared with the base model while achieving competitive performance with SOTA methods. Qianwen Cao, Heyan Huang, Minpeng Liao, Xianling Mao |
ICME | 1 |
| 2023 | Concept-Enhanced Relation Network for Video Visual Relation InferenceabstractVideo visual relation inference aims at extracting the relation triplets in the form of$>$in videos. With the development of deep learning, existing approaches are designed based on data-driven neural networks. But the datasets are always biased in terms of objects and relation triplets, which make relation inference challenging. Existing approaches often describe the relationships from visual, spatial, and semantic characteristics. The semantic description plays a key role to indicate the potential linguistic connections between objects, that are crucial to transfer knowledge across relationships, especially for the determination of novel relations. However, in these works, the semantic features are not emphasized, but simply obtained by mapping object labels, which can not reflect sufficient linguistic meanings. To alleviate the above issues, we propose a novel network, termed Concept-Enhanced Relation Network (CERN), to facilitate video visual relation inference. Thanks to the attributes and linguistic contexts implied in concepts, the semantic representations aggregated with related concept knowledge of objects are of benefit to relation inference. To this end, we incorporate retrieved concepts with local semantics of objects via the gating mechanism to generate the concept-enhanced semantic representations. Extensive experimental results show that our approach has achieved state-of-the-art performance on two public datasets: ImageNet-VidVRD and VidOR. Qianwen Cao, Heyan Huang, Mucheng Ren, Changsen Yuan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Video Visual Relation Detection With Contextual Knowledge EmbeddingabstractVideo visual relation detection (VidVRD) aims at abstracting structured relations in the form of$< $subject-predicate-object$>$from videos. The triple formation makes the search space extremely huge and the distribution unbalanced. Usually, existing works predict the relationships from visual, spatial, and semantic cues. Among them, semantic cues are responsible for exploring the semantic connections between objects, which is crucial to transfer knowledge across relations. However, most of these works extract semantic cues via simply mapping the object labels to classified features, which ignore the contextual surroundings, resulting in poor performance for low-frequency relations. To alleviate these issues, we propose a novel network, termed Contextual Knowledge Embedded Relation Network (CKERN), to facilitate VidVRD through establishing contextual knowledge embeddings for detected object pairs in relations from two aspects: commonsense attributes and prior linguistic dependencies. Specifically, we take the pair as a query to extract relational facts in the commonsense knowledge base, then encode them to explicitly construct semantic surroundings for relations. In addition, the statistics of object pairs with different predicates distilled from large-scale visual relations are taken into account to represent the linguistic regularity of relations. Extensive experimental results on benchmark datasets demonstrate the effectiveness and robustness of our proposed model. Qianwen Cao, Heyan Huang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Piecewise graph convolutional network with edge-level attention for relation extraction
Changsen Yuan, Heyan Huang, Chong Feng 0001, Qianwen Cao |
Neural Comput. Appl. | 4 |
| 2022 | VSRN: Visual-Semantic Relation Network for Video Visual Relation InferenceabstractVideo visual relation inference refers to the task of automatically detecting the relation triplets between the observed objects in videos with the form of${ < subject, predicate, object>}$, which requires correctly labeling each detected object and their interaction predicates. Despite the recent advances in image visual relation detection using deep learning techniques, relation inference in videos remains a challenging topic. On one hand, since the introduction of temporal information, it needs to model the rich spatio-temporal visual information for objects and videos. On the other hand, wild videos are often annotated with incomplete relation triplet tags and a few of them are semantically overlapped. However, previous methods adopt hand-crafted visual features extracted from the trajectories, describing local appearance characteristics of isolated objects. And they treat the problem as a multi-class classification task, which makes the relation tags mutually exclusive. To address the above issues, we propose a novel model, termed Visual-Semantic Relation Network (VSRN). In this network, we leverage three-dimensional convolution kernel to capture spatio-temporal features, and encode global visual features in videos through pooling operation on each time slice. Moreover, the semantic collocations between objects are also incorporated so as to obtain comprehensive representations of the relationships. For relation classification, we treat the problem as a multi-label classification task and regard each tag to be independent to predict various relationships. Additionally, we modify commonly used evaluation metric, video-wise recall, to a pair-wise metric (Roop) for testing the performance of models in predicting multiple relationships for the object pairs, Extensive experimental results on two large-scale datasets demonstrate the effectiveness of our proposed model which significantly outperforms the previous works. Qianwen Cao, Heyan Huang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Attention Guided Relation Detection Approach for Video Visual Relation DetectionabstractVideo Visual Relation Detection (VidVRD) aims at detecting the relation instances between two observed objects in the form of$< $subject-predicate-object$>$. Unlike image visual relation detection, due to the introduction of the time dimensions, the various predicates and spatial-temporal locations are both required to be tackled, making the task challenging. To balance these challenges, most existing works perform this task in two phases: first predicting relationships in segmented clips to capture the motions, and then associating them into the relation instances with proper locations in videos. These works detect different relationships by collecting the cues from multi-aspects, but treat them equally without distinction. Furthermore, due to the dynamic scenes and drifting problem in object tracking, the rigid spatial overlap used to determine the association in previous works is insufficient, which leads to missing associations. To address the problems, in this paper, we propose a novel attention guided relation detection approach for VidVRD. In order to model the distinction among different cues and strengthen the salient characteristics, we assign these cues the attention weights for relationship prediction and association decision-making. In addition, to comprehensively measure whether merging the relationships, we put forward a customized network to take both visual appearance and geometric location into account. Extensive experiment results on ImageNet-VidVRD dataset and VidOR dataset demonstrate the effectiveness of our proposed approach. And abundant ablation studies verify the component designed in the approach is essential. Qianwen Cao, Heyan Huang |
IEEE Trans. Multim. | 1 |
| 2021 | 3-D Relation Network for visual relation recognition in videos
Qianwen Cao, Heyan Huang, Xindi Shang, Boran Wang, Tat-Seng Chua |
Neurocomputing | 1 |