Kaiwen Dong

dblp:301/7629 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning
abstract
Xiang Gao, Yuguang Yao, Qi Zhang, Kaiwen Dong, Avinash Baidya, Ruocheng Guo, Hilaf Hasson, Kamalika Das. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiang Gao 0011, Yuguang Yao, Kaiwen Dong, Avinash Baidya, Ruocheng Guo, Hilaf Hasson, Kamalika Das
ACL (1)4
2026 Unsupervised video stitching via spatiotemporal coupling with temporal propagation and dynamic memory
abstract
Video stitching poses a fundamental challenge: while video frames exhibit strong spatiotemporal coupling, prevailing decoupled paradigms artificially separate spatial and temporal information, leading to limited precision and severe efficiency bottlenecks. Consequently, achieving a joint spatiotemporal optimization remains a critical, unresolved objective in this field. To overcome this, we propose an unsupervised video stitching framework that fundamentally reconceptualizes the task as a joint spatiotemporal state estimation problem, driven by temporal propagation and a dynamic memory mechanism. Building upon spatiotemporally decoupled multi-view inputs, we design a Spatially-Aware Gated Recurrent Unit (SA-GRU) structure that mines intrinsic spatiotemporal feature correlations to establish cross-frame temporal propagation and spatiotemporal mapping, thereby compensating for missing temporal information and improving alignment. To overcome the perceptual constraints of sliding windows, we introduce a Temporally-Gated Recurrent Unit (T-GRU) structure integrated with a Spatiotemporal Attention Fusion (STA-Fusion) module, which adaptively integrates long-term historical context with current spatiotemporal constraints to smooth stitching trajectories and enable long-term stabilization with a minimal memory footprint. Evaluations on the Stabstitch-D dataset demonstrate that our method outperforms existing approaches in alignment and geometric fidelity, achieving a PSNR of 31.18, an SSIM of 0.907, and a distortion measure of 0.011. Compared to traditional sliding-window approaches, our framework drastically reduces the memory footprint.
Xiao Yun, Fanfei Yu, Kaiwen Dong, Yanjing Sun
Knowl. Based Syst.3
2026 Temporal Consistency and Variation-Guided Spatio-Temporal Aggregation for Few-Shot Action Recognition
abstract
Few-shot Action Recognition (FSAR) aims to recognize novel actions from only a few labeled examples, posing challenges due to limited supervision and complex temporal dynamics. Existing methods often adopt a unified motion modeling strategy for both short- and long-term dynamics, overlooking the need to adapt motion pattern extraction to the specific temporal properties inherent to different timescales. This forces models to hedge against multi-scale relevance through exhaustive searches over temporal tuples, followed by heavy spatio-temporal fusion, which substantially increases parameters and computation and ultimately limits efficiency. To this end, we propose the efficient Temporal Consistency and Variation-Guided Spatio-Temporal Aggregation Network (TCV-STA), which comprises four key components: the Temporal Consistency Module (TCM), the Temporal Variation Module (TVM), the Spatio-Temporal Aggregation attention (STA), and the Shifted Window Temporal Attention (SWTA). The TCM captures stable motion patterns to suppress short-term perturbations and enhance temporal consistency for robust motion representation, while the TVM models dynamic motion patterns to highlight long-term variations that improve inter-class discriminability and facilitate intra-class alignment. Built upon these complementary motion cues, the STA selectively aggregates spatial and temporal representations under the guidance of the learned stable and dynamic motion patterns, avoiding global dense fusion. Finally, to address the limited receptive field and discontinuous modeling caused by frame grouping in TCM and TVM, we adapt a SWTA to capture longer-range temporal dependencies and ensure smooth transitions across subaction segments for few-shot action recognition. Experiments demonstrate that TCV-STA achieves competitive accuracy across four widely-used FSAR benchmarks while reducing parameters by up to 27.9% and computational cost by 21.3%, striking a favorable balance between accuracy and efficiency for deployment in resource-constrained scenarios.
Kaiwen Dong, Quanyi Li, Yanjing Sun, Xiao Yun, Yu Zhou 0009, Kévin Riou, Xiaofeng Hou, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.1
2026 CORE: Data Augmentation for Link Prediction via Information Bottleneck
abstract
Link Prediction (LP) is a fundamental task in graph representation learning, with numerous applications in diverse domains. However, the generalizability of LP models is often compromised due to the presence of noisy or spurious information in graphs and the inherent incompleteness of graph data. To address these challenges, we draw inspiration from the Information Bottleneck principle and propose a novel data augmentation method, COmplete and REduce (CORE) to learn compact and predictive augmentations for LP models. In particular, CORE aims to recover missing edges in graphs while simultaneously removing noise from the graph structures, thereby enhancing the model’s robustness and performance. Extensive experiments on multiple benchmark datasets demonstrate the applicability and superiority of CORE over state-of-the-art methods, showcasing its potential as a leading approach for robust LP in graph representation learning.
Kaiwen Dong, Zhichun Guo, Nitesh V. Chawla
ACM Trans. Knowl. Discov. Data1
2025 Transaction Categorization with Relational Deep Learning in QuickBooks
Kaiwen Dong, Padmaja Jonnalagedda, Xiang Gao 0011, Ayan Acharya, Maria Kissa, Mauricio Flores, Nitesh V. Chawla, Kamalika Das
ECML/PKDD (9)1
2025 QoI-Aware Configuration Adaptation and Heterogeneous Resource Allocation for Edge Video Analytics
abstract
Edge video analytics enables agile responses of machine-centric applications by streaming videos from end devices to edge servers (ESs) for resource-intensive deep neural network (DNN) inference. Quality of Inference (QoI), reflected by end-to-end analytics delay and inference accuracy, is crucial for the high-quality real-time decision-making. Due to the limited ES resources and dynamic network conditions, video configuration adaptation and fine-grained resource allocation are necessary to enhance the QoI of video streams, which involves a delicate balance between delay and accuracy. In this article, we propose an edge-enabled multivideo analytics framework, in which performing DNN inference relies on heterogeneous resources, including CPU, GPU, and memory. Considering the different impacts of heterogeneous resources on QoI and double-queue backlogs, we formulate the joint video configuration adaptation, CPU-GPU resource allocation, and batch size selection (JCCGB) problem based on semi-Markov decision process to maximize the long-term average QoI. Then, Lyapunov optimization is applied to transform the original problem into one that minimizes the Lyapunov drift-plus-penalty upper bound, thereby ensuring the stability of both the transmission and computation queues. To tackle the potential multimodality of the optimal configuration adaptation and resource allocation policy, we propose the diffusion-deep-reinforcement-learning-based JCCGB (DD-JCCGB) scheme. This approach effectively enhances the policy performance and accelerates the training process. Experiments driven by real-world network traces demonstrate that the DD-JCCGB algorithm improves the long-term average QoI and outperforms baseline schemes.
Yanjing Sun, Beibei Zhang 0001, Kaiwen Dong, Bowen Wang 0004, Hongli Xu 0001
IEEE Internet Things J.4
2025 Heterogeneous modal collaborative training network for human action recognition
Xiao Yun, Yanjing Sun, Kaiwen Dong
Knowl. Based Syst.4
2024 Behavioral Recognition of Skeletal Data Based on Targeted Dual Fusion Strategy
abstract
The deployment of multi-stream fusion strategy on behavioral recognition from skeletal data can extract complementary features from different information streams and improve the recognition accuracy, but suffers from high model complexity and a large number of parameters. Besides, existing multi-stream methods using a fixed adjacency matrix homogenizes the model’s discrimination process across diverse actions, causing reduction of the actual lift for the multi-stream model. Finally, attention mechanisms are commonly applied to the multi-dimensional features, including spatial, temporal and channel dimensions. But their attention scores are typically fused in a concatenated manner, leading to the ignorance of the interrelation between joints in complex actions. To alleviate these issues, the Front-Rear dual Fusion Graph Convolutional Network (FRF-GCN) is proposed to provide a lightweight model based on skeletal data. Targeted adjacency matrices are also designed for different front fusion streams, allowing the model to focus on actions of varying magnitudes. Simultaneously, the mechanism of Spatial-Temporal-Channel Parallel Attention (STC-P), which processes attention in parallel and places greater emphasis on useful information, is proposed to further improve model’s performance. FRF-GCN demonstrates significant competitiveness compared to the current state-of-the-art methods on the NTU RGB+D, NTU RGB+D 120 and Kinetics-Skeleton 400 datasets. Our code is available at: https://github.com/sunbeam-kkt/FRF-GCN-master.
Xiao Yun, Kévin Riou, Kaiwen Dong, Yanjing Sun, Song Li 0001, Kévin Subrin, Patrick Le Callet
AAAI4
2024 Evaluating 3D Human Pose Estimation in Occluded Multi-Sensor Scenarios: Dataset and Annotation Approach
abstract
Obtaining ground truth annotations for 3D pose estimation (3D HPE) typically depends on motion capture equipment (Mocap), which is not only expensive but impractical for widespread deployment. In contrast, triangulation can reconstruct 3D poses solely from multi-view 2D poses with known camera parameters, eliminating the need for Mocap. However, inherent noise in 2D pose predictions introduces uncertainties, compromising the reliability of the results. To obtain more reliable annotations with noisy input, we introduce an annotation approach for the 3D HPE task, driven by prior knowledge of the skeletal configuration. We split our approach into two steps: first a parametric model is designed to enhance confidence predictions. Then, a differentiable weighted triangulation is employed to estimate the 3D pose in world space, leveraging the predicted confidence scores as weights. The pipeline is trained using a bone length loss. Moreover, we collect a multi-view dataset for 3D HPE and annotate it using our proposed annotation tool. This dataset is characterized by more construction scenarios, including heavier occlusion cases, diverse viewing directions, and the integration of various optical sensors, setting it apart from existing datasets. Experiments on both our dataset and Human3.6M demonstrate the effectiveness of our method.
Kévin Riou, Kaiwen Dong, Kévin Subrin, Patrick Le Callet, Yanjing Sun
ICIP2
2024 Pure Message Passing Can Estimate Common Neighbor for Link Prediction
abstract
Message Passing Neural Networks (MPNNs) have emerged as the {\em de facto} standard in graph representation learning. However, when it comes to link prediction, they are not always superior to simple heuristics such as Common Neighbor (CN). This discrepancy stems from a fundamental limitation: while MPNNs excel in node-level representation, they stumble with encoding the joint structural features essential to link prediction, like CN. To bridge this gap, we posit that, by harnessing the orthogonality of input vectors, pure message-passing can indeed capture joint structural features. Specifically, we study the proficiency of MPNNs in approximating CN heuristics. Based on our findings, we introduce the Message Passing Link Predictor (MPLP), a novel link prediction model. MPLP taps into quasi-orthogonal vectors to estimate link-level structural features, all while preserving the node-level complexities. We conduct experiments on benchmark datasets from various domains, where our method consistently outperforms the baseline methods, establishing new state-of-the-arts.
Kaiwen Dong, Zhichun Guo, Nitesh V. Chawla
NeurIPS1
2024 Asymmetric network pseudo labels mutual refinement for unsupervised domain adaptation person re-identification
Xiao Yun, Kaiwen Dong, Yanjing Sun
Multim. Tools Appl.4
2023 Heterogeneous Graph Masked Autoencoders
abstract
Generative self-supervised learning (SSL), especially masked autoencoders, has become one of the most exciting learning paradigms and has shown great potential in handling graph data. However, real-world graphs are always heterogeneous, which poses three critical challenges that existing methods ignore: 1) how to capture complex graph structure? 2) how to incorporate various node attributes? and 3) how to encode different node positions? In light of this, we study the problem of generative SSL on heterogeneous graphs and propose HGMAE, a novel heterogeneous graph masked autoencoder model to address these challenges. HGMAE captures comprehensive graph information via two innovative masking techniques and three unique training strategies. In particular, we first develop metapath masking and adaptive attribute masking with dynamic mask rate to enable effective and stable learning on heterogeneous graphs. We then design several training strategies including metapath-based edge reconstruction to adopt complex structural information, target attribute restoration to incorporate various node attributes, and positional feature prediction to encode node positional information. Extensive experiments demonstrate that HGMAE outperforms both contrastive and generative state-of-the-art baselines on several tasks across multiple datasets. Codes are available at https://github.com/meettyj/HGMAE.
Yijun Tian 0001, Kaiwen Dong, Chuxu Zhang, Nitesh V. Chawla
AAAI2
2023 From Temporal-Evolving to Spatial-Fixing: A Keypoints-Based Learning Paradigm for Visual Robotic Manipulation
abstract
The current learning pipelines for robotics manipulation infer movement primitives sequentially along the temporal-evolving axis, which can result in an accumulation of prediction errors and subsequently cause the visual observations to fall out of the training distribution. This paper proposes a novel hierarchical behavior cloning approach which tries to dissociate standard behaviour cloning (BC) pipeline to two stages. The intuition of this approach is to eliminate accumu-lation errors using a fixed spatial representation. At first stage, a high-level planner will be employed to translate the initial observation of the scene into task-specific spatial waypoints. Then, a low-level robotic path planner takes over the task of guiding the robot by executing a set of pre-defined elementary movements or actions known as primitives, with the goal of reaching the previously predicted waypoints. Our hierarchical keypoints-based paradigm aims to simplify existing temporal-evolving approach to a more simple way: directly spatialize the whole sequential primitives as a set of 8D waypoints only from the very first observation. Plentiful experiments demon-strate that our paradigm can achieve comparable results with Reinforcement Learning (RL) and outperforms existing offline BC approaches, with only a single-shot inference from the initial observation. Code and models are available at: https://github.com/KevinRiou22/spatial-fixing-il
Kévin Riou, Kaiwen Dong, Kévin Subrin, Yanjing Sun, Patrick Le Callet
IROS2
2023 Kinetic particles : from human pose estimation to an immersive and interactive piece of art questionning thought-movement relationships
abstract
Digital tools offer extensive solutions to explore novel interactive-art paradigms, by relying on various sensors to create installations and performances where the human activity can be captured, analysed and used to generate visual and sound universes in real-time. Deep learning approaches, including human detection and human pose estimation, constitute ideal human-art interaction mediums, as they allow automatic human gesture analysis, which can be directly used to produce the interactive piece of art. In this context, this paper presents an interactive work of art that explores the relationship between thought and movement by combining dance, philosophy, numerical arts, and deep learning. We present a novel system that combines a multi-camera setup to capture human movement, state-of-the-art human pose estimation models to automatically analyze this movement, and an immersive 180° projection system that projects a dynamic textual content that intuitively responds to the users’ behaviors. The demonstration being proposed consists of two parts. Firstly, a professional dancer will utilize the proposed setup to deliver a conference-show. Secondly, the audience will be given the opportunity to experiment and discover the potential of the proposed setup, which has been transformed into an interactive installation. This allows multiple spectators to engage simultaneously with clusters of words and letters extracted from the conference text.
Mickael Lafontaine, Julie Cloarec-Michaud, Kévin Riou, Kaiwen Dong, Patrick Le Callet
IMX5
2023 Combining detailed appearance and multi-scale representation: a structure-context complementary network for human pose estimation
Kaiwen Dong, Yanjing Sun, Xiaozhou Cheng
Appl. Intell.1
2022 Triplet attention multiple spacetime-semantic graph convolutional network for skeleton-based action recognition
Yanjing Sun, Xiao Yun, Kaiwen Dong
Appl. Intell.5
2022 Knowledge self-distillation for visible-infrared cross-modality person re-identification
Yu Zhou 0009, Yanjing Sun, Kaiwen Dong, Song Li 0001
Appl. Intell.4
2021 An Optimized NL2SQL System for Enterprise Data Mart
Kaiwen Dong, David A. Cieslak, Nitesh V. Chawla
ECML/PKDD (5)1