EDBT 2026 Demo / reviewers in the wild / expert
Dongkai Wang
dblp:263/3255
· DBLP profile ↗
15ranked-venue papers
10as first author
14since 2021 · last 2026
0000-0002-0266-4340ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 9 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MRDNet: Multivariable Relational Decomposition Network for Multivariate Time Series Forecasting
Ao Hu, Liangjian Wen, Yong Dai 0001, Dongkai Wang, Jun Wang 0089, Jiang Duan |
Knowl. Based Syst. | 5 |
| 2026 | TimeCNN: Refining inscross-variable interaction on time point for time series forecasting
Ao Hu, Liangjian Wen, Yong Dai 0001, Shiyi Qi, Jun Wang 0089, Xun Zhou 0001, Dongkai Wang, Zenglin Xu, Jiang Duan |
Neural Networks | 8 |
| 2026 | Spike Camera Optical Flow Estimation Based on Continuous Spike StreamsabstractSpike camera is an emerging bio-inspired vision sensor with ultra-high temporal resolution. It records scenes by accumulating photons and outputting binary spike streams. Optical flow estimation aims to estimate pixel-level correspondences between different moments, describing motion information along time, which is a key task of spike camera. High-quality optical flow is important since motion information is a foundation for analyzing spikes. However, extracting stable light-intensity information from spikes is difficult due to the randomness of binary spikes. Besides, the continuity of spikes can offer contextual information for optical flow. In this paper, we propose a network Spike2Flow++ to estimate optical flow for spike camera. In Spike2Flow++, we propose a differential of spike firing time (DSFT) to represent information in binary spikes. Moreover, we propose a dual DSFT representation and a dual correlation construction to extract stable light-intensity information for reliable correlations. To use the continuity of spikes as motion contextual information, we propose a joint correlation decoding (JCD) that jointly estimates a series of flow fields. To adaptively fuse different motions in JCD, we propose a global motion bank aggregation to construct an information bank for all motions and adaptively extract contexts from the bank for each iteration during recurrent decoding of each motion. To train and evaluate our network, we construct a real scene with spikes and flow++ (RSSF++) based on real-world scenes. Experiments demonstrate that our Spike2Flow++ achieves state-of-the-art performance on RSSF++, photo-realistic high-speed motion (PHM), and real-captured data. Rui Zhao 0010, Ruiqin Xiong, Dongkai Wang, Shiyu Xuan, Jian Zhang 0018, Xiaopeng Fan 0001, Tiejun Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | FDNet: High-frequency disentanglement network with information-theoretic guidance for multivariate time series forecasting
Ao Hu, Liangjian Wen, Jiang Duan, Yong Dai 0001, Dongkai Wang, Shudong Huang, Jun Wang 0089, Zenglin Xu |
Pattern Recognit. | 5 |
| 2025 | Generalizable Object Keypoint Localization from Generative PriorsabstractGeneralizable object keypoint localization is a fundamental computer vision task in understanding the object structure. It is challenging for existing keypoint localization methods because their limited training data cannot provide generalizable shape and semantic cues, leading to inferior performance and generalization capability. Instead of relying on large scale training data, this work tackles this challenge by exploiting the rich priors from large generative models. We propose a data-efficient generalizable localization method named GenLoc. GenLoc extracts the generative priors from a pre-trained image generation model by calculating the correlation map between image latent feature and condition embedding. Those priors are hence optimized with our proposed heatmap expectation loss to perform object keypoint localization. Benefited by the rich knowledge of generative priors in understanding of object semantics and structures, GenLoc achieves superior performance on various object keypoint localization benchmarks. It shows more substantial performance enhancements in cross-domain, few-shot and zero-shot evaluation settings, e.g., getting 20%+ AP enhancement over CLAMP [43] in various zero-shot settings. Dongkai Wang, Jiang Duan, Liangjian Wen, Shiyu Xuan, Hao Chen 0061, Shiliang Zhang |
CVPR | 1 |
| 2025 | InfMasking: Unleashing Synergistic Information by Contrastive Multimodal InteractionsabstractIn multimodal representation learning, synergistic interactions between modalities not only provide complementary information but also create unique outcomes through specific interaction patterns that no single modality could achieve alone. Existing methods may struggle to effectively capture the full spectrum of synergistic information, leading to suboptimal performance in tasks where such interactions are critical. This is particularly problematic because synergistic information constitutes the fundamental value proposition of multimodal representation. To address this challenge, we introduce InfMasking, a contrastive synergistic information extraction method designed to enhance synergistic information through an Infinite Masking strategy. InfMasking stochastically occludes most features from each modality during fusion, preserving only partial information to create representations with varied synergistic patterns. Unmasked fused representations are then aligned with masked ones through mutual information maximization to encode comprehensive synergistic information. This infinite masking strategy enables capturing richer interactions by exposing the model to diverse partial modality combinations during training. As computing mutual information estimates with infinite masking is computationally prohibitive, we derive an InfMasking loss to approximate this calculation. Through controlled experiments, we demonstrate that InfMasking effectively enhances synergistic information between modalities. In evaluations on large-scale real-world datasets, InfMasking achieves state-of-the-art performance across seven benchmarks. Code is released at https://github.com/brightest66/InfMasking. Liangjian Wen, Qun Dai, Jianzhuang Liu, Jiangtao Zheng, Yong Dai 0001, Dongkai Wang, Zhao Kang 0001, Jun Wang 0089, Zenglin Xu, Jiang Duan |
NeurIPS | 6 |
| 2024 | LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language ModelabstractThe capacity of existing human keypoint localization models is limited by keypoint priors provided by the training data. To alleviate this restriction and pursue more gen-eral model, this work studies keypoint localization from a different perspective by reasoning locations based on key-piont clues in text descriptions. We propose LocLLM, the first Large-Language Model (LLM) based keypoint local-ization model that takes images and text instructions as in-puts and outputs the desired keypoint coordinates. LocLLM leverages the strong reasoning capability of LLM and clues of keypoint type, location, and relationship in textual de-scriptions for keypoint localization. To effectively tune Lo-cLLM, we construct localization-based instruction conver-sations to connect keypoint description with corresponding coordinates in input image, and fine-tune the whole model in a parameter-efficient training pipeline. LocLLM shows remarkable performance on standard 2D/3D keypoint lo-calization benchmarks. Moreover, incorporating language clues into the localization makes LocLLM show superior flexibility and generalizable capability in cross dataset key-point localization, and even detecting novel type of key-points unseen during training††Project page: https://github.com/kennethwdk/LocLLM. Dongkai Wang, Shiyu Xuan, Shiliang Zhang |
CVPR | 1 |
| 2024 | Spatial-Aware Regression for Keypoint LocalizationabstractRegression-based keypoint localization shows advan-tages of high efficiency and better robustness to quantization errors than heatmap-based methods. However, existing regression-based methods discard the spatial location prior in input image with a global pooling, leading to in-ferior accuracy and are limited to single instance localization tasks. We study the regression-based keypoint localization from a new perspective by leveraging the spatiallocation prior. Instead of regressing on the pooled feature, the proposed Spatial-Aware Regression (SAR) maintains the spatial location map and outputs spatial coordinates and confidence score for each grid, which are optimized with a unified objective. Benefited by the location prior, these spatial-aware outputs can be efficiently optimized, resulting in better localization performance. Moreover, incorporating spatial prior makes SAR more general and can be applied into various keypoint localization tasks. We test the proposed method in 4 keypoint localization tasks including single/multi-person 2D/3D pose estimation, and the whole-body pose estimation. Extensive experiments demonstrate its promising performance, e.g., consistently outperforming recent regressions-based methods††project pagn: https://github.com/kennethwdk/SAR. Dongkai Wang, Shiliang Zhang |
CVPR | 1 |
| 2023 | 3D Human Mesh Recovery with Sequentially Global Rotation EstimationabstractModel-based 3D human mesh recovery aims to reconstruct a 3D human body mesh by estimating its parameters from monocular RGB images. Most of recent works adopt the Skinned Multi-Person Linear (SMPL) model to regress relative rotations for each body joint along the kinematics chain. This pipeline needs to transform each relative rotation matrix into a global rotation matrix to articulate the canonical mesh, and suffers from accumulated errors along the kinematics chain. This paper proposes to directly estimate the global rotation of each joint to avoid error accumulation and pursue better accuracy. The proposed Sequentially Global Rotation Estimation (SGRE) directly predicts the global rotation matrix of each joint on the kinematics chain. SGRE features a residual learning module to leverage complementary features and previously predicted rotations of parent joints to guide the estimation of subsequent child joints. Thanks to this global estimation pipeline and residual learning module, SGRE alleviates error accumulation and produces more accurate 3D human mesh. It can be flexibly integrated into existing regression-based methods and achieves superior performance on various benchmarks. For example, it improves the latest method 3DCrowdNet by 3.3 mm MPJPE and 5.0 mm PVE on 3DPW dataset and 3.0 AP on COCO dataset, respectively†. Dongkai Wang, Shiliang Zhang |
ICCV | 1 |
| 2023 | HumVis: Human-Centric Visual Analysis SystemabstractHuman-centric visual analysis is a fundamental task for many multimedia and computer vision applications, such as self-driving, multimedia retrieval, and augmented reality, etc. Based on our recent research efforts on fine-grained human visual analysis, we develop a robust and efficient human-centric visual analysis system named as HumVis. HumVis is built on a simple yet efficient contextual instance decoupling (CID) module, which can effectively separate different persons in an input image and output corresponding person structure information for visual analysis. Based on CID, HumVis achieves accurate multi-person pose estimation, multi-person foreground segmentation, multi-person part segmentation and 3D human mesh recovery for user-uploaded images/videos and support live stream presentation. Dongkai Wang, Shiliang Zhang, Yaowei Wang 0001, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001 |
ACM Multimedia | 1 |
| 2023 | Contextual Instance Decoupling for Instance-Level Human AnalysisabstractOne fundamental challenge of instance-level human analysis is to decouple instances in crowded scenes, where multiple persons are overlapped with each other. This paper proposes the Contextual Instance Decoupling (CID), which presents a new pipeline of decoupling persons for multi-person instance-level analysis. Instead of relying on person bounding boxes to spatially differentiate persons, CID decouples persons in an image into multiple instance-aware feature maps. Each of those feature maps is hence adopted to infer instance-level cues for a specific person, e.g., keypoints, instance mask or part segmentation masks. Compared with bounding box detection, CID is differentiable and robust to detection errors. Decoupling persons into different feature maps also allows to isolate distractions from other persons, and explore context cues at scales larger than the bounding box size. Extensive experiments on various tasks including multi-person pose estimation, person foreground segmentation, and part segmentation, show that CID consistently outperforms previous methods in both accuracy and efficiency. For instance, it achieves 71.3% AP on CrowdPose in multi-person pose estimation, outperforming the recent single-stage DEKR by 5.6%, the bottom-up CenterAttention by 3.7%, and the top-down JC-SPPE by 5.3%. This advantage sustains on multi-person segmentation and part segmentation tasks. Dongkai Wang, Shiliang Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Contextual Instance Decoupling for Robust Multi-Person Pose EstimationabstractCrowded scenes make it challenging to differentiate persons and locate their pose keypoints. This paper proposes the Contextual Instance Decoupling (CID), which presents a new pipeline for multi-person pose estimation. Instead of relying on person bounding boxes to spatially differentiate persons, CID decouples persons in an image into multiple instance-aware feature maps. Each of those feature maps is hence adopted to infer keypoints for a specific person. Compared with bounding box detection, CID is differentiable and robust to detection errors. Decoupling persons into different feature maps allows to isolate distractions from other persons, and explore context cues at scales larger than the bounding box size. Experiments show that CID outperforms previous multi-person pose estimation pipelines on crowded scenes pose estimation benchmarks in both accuracy and efficiency. For instance, it achieves 71.3% AP on CrowdPose, outperforming the recent single-stage DEKR by 5.6%, the bottom-up CenterAttention by 3.7%, and the top-down JC-SPPE by 5.3%. This advantage sustains on the commonly used COCO benchmark††Code is available at https://github.com/kennethwdk/CID. Dongkai Wang, Shiliang Zhang |
CVPR | 1 |
| 2022 | Unsupervised Person Re-Identification via Multi-Label Classification
Dongkai Wang, Shiliang Zhang |
Int. J. Comput. Vis. | 1 |
| 2021 | Robust Pose Estimation in Crowded Scenes with Direct Pose-Level InferenceabstractMulti-person pose estimation in crowded scenes is challenging because overlapping and occlusions make it difficult to detect person bounding boxes and infer pose cues from individual keypoints. To address those issues, this paper proposes a direct pose-level inference strategy that is free of bounding box detection and keypoint grouping. Instead of inferring individual keypoints, the Pose-level Inference Network (PINet) directly infers the complete pose cues for a person from his/her visible body parts. PINet first applies the Part-based Pose Generation (PPG) to infer multiple coarse poses for each person from his/her body parts. Those coarse poses are refined by the Pose Refinement module through incorporating pose priors, and finally are fused in the Pose Fusion module. PINet relies on discriminative body parts to differentiate overlapped persons, and applies visual body cues to infer the global pose cues. Experiments on several crowded scenes pose estimation benchmarks demonstrate the superiority of PINet. For instance, it achieves 59.8% AP on the OCHuman dataset, outperforming the recent works by a large margin. Dongkai Wang, Shiliang Zhang, Gang Hua 0001 |
NeurIPS | 1 |
| 2020 | Unsupervised Person Re-Identification via Multi-Label ClassificationabstractThe challenge of unsupervised person re-identification (ReID) lies in learning discriminative features without true labels. This paper formulates unsupervised person ReID as a multi-label classification task to progressively seek true labels. Our method starts by assigning each person image with a single-class label, then evolves to multi-label classification by leveraging the updated ReID model for label prediction. The label prediction comprises similarity computation and cycle consistency to ensure the quality of predicted labels. To boost the ReID model training efficiency in multi-label classification, we further propose the memory-based multi-label classification loss (MMCL). MMCL works with memory-based non-parametric classifier and integrates multi-label classification and single-label classification in an unified framework. Our label prediction and MMCL work iteratively and substantially boost the ReID performance. Experiments on several large-scale person ReID datasets demonstrate the superiority of our method in unsupervised person ReID. Our method also allows to use labeled person images in other domains. Under this transfer learning setting, our method also achieves state-of-the-art performance. Dongkai Wang, Shiliang Zhang |
CVPR | 1 |