VLDB 2026 Research / reviewers in the wild / expert
Kun Wang 0057
dblp:05/1958-57
· DBLP profile ↗
12ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-6735-7667ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond 1 × 1 Convolutions: A Dynamic Select-and-Fuse Channel Sampling Strategy for On-Device Human Activity RecognitionabstractThe proliferation of low-cost, portable sensors has made wearable human activity recognition (HAR) a cornerstone for real-time health monitoring and behavior analysis. However, deploying accurate yet lightweight deep learning models on resource-constrained wearable devices poses a significant challenge for on-device activity recognition. While channel pruning is a common solution to accelerate deep Convolutional Neural Networks (CNNs), existing works often require specialized implementations or pre-trained models, which potentially degrade performance by simply removing an entire channel, limiting their ability to handle complex multimodal sensor inputs. Moreover, lightweight CNN design, particularly the heavy use of 1×1 convolution layers for channel squeezing, remain inefficient for sensor-based HAR, which consume resources without expanding the receptive field due to their pointwise nature. To address these issues, we propose a novel dynamic channel sampling module, Select-and-Fuse (SaF), specifically designed for sensor-based HAR. SaF divides channels into subsets and performs a dynamic, input-dependent selection from them, with the picking decision being made per-time-step based on the input sensor signal activations, allowing for fine-grained feature adaptation to multi-modal sensor signals. While integrated into compact backbones, SaF significantly reduces model size and inference latency while maintaining high accuracy. Extensive evaluations on public UCI-HAR, OPPORTUNITY, WISDM, and UniMiB-SHAR benchmarks confirm a favorable performance-cost trade-off. Crucially, we measure actual inference latency on a Raspberry Pi, proving its practicality for resource-constrained HAR applications. Code will be released. Guangjie Chen, Xin Liu 0176, Lei Zhang 0130, Qifan Sun, Kun Wang 0057, Hao Wu 0010, Aiguo Song |
IEEE Internet Things J. | 6 |
| 2026 | TSA-Former: Linear Transformer With Taylor Series Attention for Sensor-Based Human Activity RecognitionabstractTransformer models have demonstrated superior capability in capturing long-range temporal dependencies crucial for Sensor-Based Human Activity Recognition (HAR). However, the quadratic computational complexity inherent to the Softmax-Attention mechanism significantly impedes their deployment on resource-constrained wearable devices and real-time streaming tasks. To address this, we propose a novel Linear Transformer with Taylor Series Attention specifically tailored for the HAR domain, named TSA-Former. It leverages the first-order Taylor expansion to approximate the Softmax-Attention and utilizes the norm-preserving mapping to approximate the high-order non-linear information, resulting in a linear computational complexity. In addition, TSA-Former integrates a multi-branch architecture featuring multi-scale patch embedding, which enables the model to dynamically capture multi-scale temporal features while minimizing overhead. Experimental results across four public HAR benchmarks, namely UniMiB-SHAR, UCI-HAR, WISDM, and OPPORTUNITY, demonstrate that TSA-Former achieves state-of-the-art (SOTA) accuracy and efficiency, outperforming conventional Transformers and existing linear-attention models. Deployment experiments conducted on the Raspberry Pi 5 platform further validate the model’s superior low-latency and minimal power consumption profile, confirming its robust suitability for real-world embedded HAR applications. Code will be released. Qifan Sun, Kun Wang 0057, Zenan Fu, Guangjie Chen, Lei Zhang 0130, Hao Wu 0010, Aiguo Song |
IEEE Internet Things J. | 3 |
| 2025 | SKL-CLIP: Learning Skeleton-Based Action Representations via Language SupervisionabstractCLIP, widely used in multimodal learning, excels due to its large-scale image-text pretraining. However, applying CLIP-like architectures to skeleton-based action representation learning presents challenges due to the incompatible non-visual data structure and the limited scale of skeleton datasets, which hinder robust generalization. To address these issues, we propose SKL-CLIP, a framework that incorporates Supervised Self-Contrastive Learning to mitigate overfitting and enhance transferable representation learning, Knowledge Distillation from the textual encoder of pretrained CLIP models to preserve generalization while adapting to skeleton-text scenarios, and Multi-Domain Parallel Training to leverage diverse support datasets, improving cross-dataset and zero-shot recognition. Extensive experiments on NTU and PKU datasets demonstrate that SKL-CLIP significantly advances skeleton-based action representation, achieving state-of-the-art performance across fully-supervised, cross-data unsupervised domain adaptation, and zero-shot tasks. Kun Wang 0057, Jiuxin Cao, Jiawei Ge 0002, Chang Liu 0113, Bo Liu 0004 |
ICME | 1 |
| 2025 | Gen4Track: A Tuning-free Data Augmentation Framework via Self-correcting Diffusion Model for Vision-Language TrackingabstractThe performance of current Vision-Language Tracking (VLT) models is constrained by the limited diversity and quantity of labeled data. Compared to constructing large-scale datasets, data augmentation offers a more cost-saving strategy for VLT by synthesizing new samples from existing data, rather than generating them from scratch. However, conventional techniques like rotation and flipping may disrupt scene composition, causing conflicts between visual layouts and textual annotations. Recent advances in generative models have inspired the use of synthetic videos for data augmentation. Yet, existing approaches fail to address the core concerns of data augmentation in VLT (shown in Fig. 1)-target location accuracy, text-video consistency, and video content coherency. To bridge the gap, we propose Gen4Track, a tuning-free data augmentation framework that leverages the self-correcting mechanism to dynamically generate high-quality video data with annotations. Our approach involves (1) optimizing the attention calculations in a frozen text-to-image diffusion model to synthesize coherent videos that satisfy specific conditions (e.g., spatial location, category, color, and style), and (2) implementing a self-correcting mechanism based on a Large Language Model (LLM) to improve text-video consistency. During video augmentation, we propose content-coherent self-attention and location-enhanced cross-attention mechanisms, ensuring that image-level editings are accurately and coherently propagated throughout the video. Then, with the goal of maximizing text-video consistency, we iteratively refine the augmentation instruction with our designed self-correcting mechanism for a more aligned video. Extensive experiments validate that Gen4Track significantly boosts the performance of SOTA VLT models (achieving improvements of up to 3.2% in SUC and 3.5% in PRE), opening a new chapter of training Vision-Language trackers with synthetic videos rather than manually annotated data. Jiawei Ge 0002, Xin-Yu Zhang 0027, Jiuxin Cao, Xuelin Zhu, Qingqing Gao, Biwei Cao, Kun Wang 0057, Chang Liu 0113, Bo Liu 0004, Chen Feng 0028, Ioannis Patras |
ACM Multimedia | 8 |
| 2025 | Causality-inspired representation learning for weakly supervised skeleton-based action recognition
Kun Wang 0057, Jiuxin Cao, Jiawei Ge 0002, Chang Liu 0113, Bo Liu 0004 |
Knowl. Based Syst. | 1 |
| 2025 | Text-to-face synthesis based on facial landmarks prediction
Kun Wang 0057, Biwei Cao, Bo Liu 0004, Jiuxin Cao |
Mach. Vis. Appl. | 1 |
| 2025 | Beyond Visual Cues: Synchronously Exploring Target-Centric Semantics for Vision-Language TrackingabstractSingle object tracking aims to locate one specific target in video sequences, given its initial state. Classical trackers rely solely on visual cues, restricting their ability to handle challenges such as appearance variations, ambiguity, and distractions. Hence, Vision-Language Tracking (VLT) has emerged as a promising approach, incorporating language descriptions to directly provide high-level semantics and enhance tracking performance. However, current Vision-Language (VL) trackers have not fully exploited the power of multi-modal learning, as they suffer from limitations such as heavily relying on off-the-shelf backbones for feature extraction, ineffective asynchronous fusion designs, and the absence of VL-related loss functions for optimizing multi-modal representation. Consequently, we present a novel tracker that progressively explores target-centric semantics for VLT. Specifically, we propose the first Synchronous Learning Backbone (SLB) for VLT, which consists of two novel modules: the Target Enhance Module (TEM) and the Semantic-Aware Module (SAM). These modules together ensure the multi-modal feature extraction and interaction at the same pace, facilitating the tracker to synchronously perceive target-related semantics from both visual and textual modalities. Moreover, we devise the dense matching loss to further strengthen multi-modal representation learning. Extensive experiments on VLT datasets demonstrate the superiority and effectiveness of our methods. Jiawei Ge 0002, Jiuxin Cao, Xiangmei Chen, Xuelin Zhu, Chang Liu 0113, Kun Wang 0057, Bo Liu 0004 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | Dual-Domain Triple Contrast for Cross-Dataset Skeleton-Based Action RecognitionabstractSkeleton-based Action Recognition (SAR) is widely recognized for its robustness and efficiency in human action analysis, but its performance in cross-dataset tasks has been limited due to domain shifts between different datasets. To address this challenge, current methods typically approach cross-dataset SAR as an Unsupervised Domain Adaptation (UDA) task, which is tackled using domain adaptation or self-supervised learning strategies. In this article, we propose a Dual-Domain Triple Contrast (D2TC) framework for cross-dataset SAR under the UDA setting. Unlike existing UDA methods that either focus on a single strategy or superficially combine strategies, our D2TC leverages contrastive learning to integrate both strategies into a unified framework. It performs three types of contrastive learning: Self-Supervised Contrastive Learning, Supervised Contrastive Learning, and UDA with Contrastive Learning, across both source and target domains. The triple contrasts go beyond mere summation, effectively bridging the domain gap and enhancing the model’s representational capacity. Additionally, we introduce multi-modal ensemble contrast and extreme skeleton augmentation methods to further enhance the skeleton-based representation learning. Extensive experiments on six cross-dataset settings validate the superiority of our D2TC framework over state-of-the-art methods, demonstrating its effectiveness in reducing domain discrepancies and improving cross-dataset SAR performance. The codes are available on https://github.com/KennCoder7/DualDomainTripleContrast . Kun Wang 0057, Jiuxin Cao, Jiawei Ge 0002, Chang Liu 0113, Bo Liu 0004 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning for Visual Story SynthesisabstractThe excellent text-to-image synthesis capability of diffusion models has driven progress in synthesizing coherent visual stories. The current state-of-the-art method combines the features of historical captions, historical frames, and the current captions as conditions for generating the current frame. However, this method treats each historical frame and caption as the same contribution. It connects them in order with equal weights, ignoring that not all historical conditions are associated with the generation of the current frame. To address this issue, we propose Causal-Story. This model incorporates a local causal attention mechanism that considers the causal relationship between previous captions, frames, and current captions. By assigning weights based on this relationship, Causal-Story generates the current frame, thereby improving the global consistency of story generation. We evaluated our model on the PororoSV and FlintstonesSV datasets and obtained state-of-the-art FID scores, and the generated frames also demonstrate better storytelling in visuals. Tianyi Song, Jiuxin Cao, Kun Wang 0057, Bo Liu 0004 |
ICASSP | 3 |
| 2024 | Consistencies are All You Need for Semi-supervised Vision-Language TrackingabstractVision-Language Tracking (VLT) requires locating a specific target in video sequences, given a natural language prompt and an initial object box. Despite recent advancements, existing approaches heavily rely on expensive and time-consuming human annotations. To mitigate this limitation, directly generating pseudo labels from raw videos seems to be a straightforward solution; however, it inevitably introduces undesirable noise during the training process. Moreover, we insist that an efficient tracker should excel in tracking the target, regardless of the temporal direction. Building upon these insights, we propose the pioneering semi-supervised learning scheme for VLT task, representing a crucial step towards reducing the dependency on high-quality yet costly labeled data. Specifically, drawing inspiration from the natural attributes of a video (i.e., space, time, and semantics), our approach progressively leverages inherent consistencies from these aspects: (1) Spatially, each frame and any object cropped from it naturally form an image-bbox (bounding box) pair for self-training; (2) Temporally, bidirectional tracking trajectories should exhibit minimal differences; (3) Semantically, the correlation between visual and textual features is expected to remain consistent. Furthermore, the framework is validated with a simple yet effective tracker we devised, named ATTracker (Asymmetrical Transformer Tracker). It modifies the self-attention operation in an asymmetrical way, striving to enhance target-related features while suppressing noise. Extensive experiments confirm that our ATTracker serves as a robust baseline, outperforming fully supervised base trackers. By unveiling the potential of learning with limited annotations, this study aims to attract attention and pave the way for Semi-supervised Vision-Language Tracking (SS-VLT). Jiawei Ge 0002, Jiuxin Cao, Xuelin Zhu, Xin-Yu Zhang 0027, Chang Liu 0113, Kun Wang 0057, Bo Liu 0004 |
ACM Multimedia | 6 |
| 2024 | EnsCLR: Unsupervised skeleton-based action recognition via ensemble contrastive learning of representation
Kun Wang 0057, Jiuxin Cao, Biwei Cao, Bo Liu 0004 |
Comput. Vis. Image Underst. | 1 |
| 2021 | Sequential Weakly Labeled Multiactivity Localization and Recognition on Wearable Sensors Using Recurrent Attention NetworksabstractWith the popularity and development of the wearable devices such as smartphones, human activity recognition (HAR) based on sensors has become as a key research area in human computer interaction and ubiquitous computing. The emergence of deep learning leads to a recent shift in the research of HAR, which requires massive strictly labeled data in supervised learning scenario. In comparison with video data, activity data recorded from accelerometer or gyroscope are often more difficult to interpret and segment. Recently, several attention mechanisms are proposed to handle the weakly labeled human activity data, which do not require accurate data annotation. However, these attention-based models can only handle the weakly labeled dataset whose sample includes one target activity, as a result it limits efficiency and practicality. In the article, we propose a recurrent attention networks (RAN) to handle sequential weakly labeled multiactivity recognition and location tasks. The model can repeatedly perform steps of attention on multiple activities of one sample and each step is corresponding to the current focused activity. The effectiveness of the RAN model is validated on a collected sequential weakly labeled multiactivity dataset and the other two public datasets. The experiment results show that our RAN model can simultaneously infer multiactivity types from the coarse-grained sequential weak labels and determine specific locations of every target activity with only knowledge of which types of activities contained in the long sequence. It will greatly reduce the burden of manual labeling.1 Kun Wang 0057, Jun He 0006, Lei Zhang 0130 |
IEEE Trans. Hum. Mach. Syst. | 1 |