Kaiwen Du

dblp:292/9285 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0002-9352-6927ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2026 Vision-Language Enhancement Network Based on Decoupling-Joint Adaptation for Few-Shot Action Recognition
abstract
Learning robust and generalizable feature extractors to generate discriminative prototypes is crucial for few-shot action recognition. However, most existing methods rely on fine-tuning large pre-trained image models, easily leading to transferability and overfitting issues. In this paper, we propose a novel vision-language enhancement network based on decoupling-joint adaptation (VEDA) for few-shot action recognition, which decouples visual features into temporal and spatial branches, followed by a joint operation that integrates these two branches using an adapter-tuning paradigm. VEDA can gradually equip the model with spatio-temporal reasoning capabilities. Since relying exclusively on local frame feature matching results in inaccurate performance, we design a video-level relation module (VLR) to enhance video context awareness through global feature matching. In addition, we design a vision-language fusion module (VLF) that introduces multimodal information to alleviate the data scarcity issue. Simultaneously, we apply adapter-tuning to both visual and textual branches to enhance the generalization ability. Based on the proposed components above, our network can extract both informative and discriminative prototypes, resulting in excellent recognition performance. Experimental results on five challenging benchmarks demonstrate the effectiveness of the proposed VEDA. The code will be released soon at https://github.com/ReverseSuzhou/VEDA.
Suzhou Que, Hanyu Guo, Kaiwen Du, Yan Yan 0001, Yanwei Pang, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.3
2025 Edge Guided Network With Motion Enhancement for Few-Shot Action Recognition
abstract
Existing state-of-the-art methods for few-shot action recognition (FSAR) achieve promising performance by spatial and temporal modeling. However, most current methods ignore the importance of edge information and motion cues, leading to inferior performance. For the few-shot task, it is important to effectively explore limited data. Additionally, effectively utilizing edge information is beneficial for exploring motion cues, and vice versa. In this paper, we propose a novel edge guided network with motion enhancement (EGME) for FSAR. To the best of our knowledge, this is the first work to utilize the edge information as guidance in the FSAR task. Our EGME contains two crucial components, including an edge information extractor (EIE) and a motion enhancement module (ME). Specifically, EIE is used to obtain edge information on video frames. Afterward, the edge information is used as guidance to fuse with the frame features. In addition, ME can adaptively capture motion-sensitive features of videos. It adopts a self-gating mechanism to highlight motion-sensitive regions in videos from a large temporal receptive field. Based on the above designed components, EGME can capture edge information and motion cues, resulting in superior recognition performance. Experimental results on four challenging benchmarks show that EGME performs favorably against recent advanced methods.
Kaiwen Du, Weirong Ye, Hanyu Guo, Yan Yan 0001, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.1
2024 Spatio-Temporal Correlation Learning for Multiple Object Tracking
abstract
Multi-object tracking (MOT) has gained remarkable progress in recent years, while due to the complexity of real-world environments, there are still many challenges that remain unsolved, such as object occlusion and deformation. To effectively alleviate this problem, we propose a simple yet effective Transformer-based tracker, named CLNet, consisting of an Instance-Aware Localization (IAL) module and a Temporal Context Aggregation (TCA) module. Specifically, the former learns the correlation of object positions for potential location estimation, and the latter learns the correlation of background contexts to obtain robust re-ID features for data association. Experimental results show that CLNet outperforms the baseline method by +2.2 MOTA and +2.3 IDF1 on MOT17 and +6.3 MOTA and +3.6 IDF1 on MOT20 respectively, which demonstrate the effectiveness of the proposed method.
Yajun Jian, Chihui Zhuang, Wenyan He, Kaiwen Du, Yang Lu 0009, Hanzi Wang
ICASSP4
2022 Egnet: A Novel Edge Guided Network for Instance Segmentation
abstract
Edge information plays a significant role in instance segmentation. However, many instance segmentation methods directly perform pixel-wise classification via fully convolutional networks, which may ignore object edges. In this paper, we propose a novel Edge Guided Network (EGNet), which exploits edge information to improve the mask accuracy, for instance segmentation. Specifically, we propose an edge branch to extract edge information. Then, we use edge information as guidance and fuse it with mask features, in order to enrich the mask features. Furthermore, we propose a Spatial Attention (SA) module and add it to the backbone of our EGNet, enabling the network to focus more on foreground objects. In addition, we incorporate a Semantic Enhancement (SE) module into the edge branch, aiming to obtain additional global context information. Experimental results on the COCO 2017 dataset show the effectiveness of the proposed EGNet.
Kaiwen Du, Xiao Wang 0072, Yan Yan 0009, Yang Lu 0009, Hanzi Wang
ICIP1
2022 Dualfeat: Dual Feature Aggregation for Video Object Detection
abstract
Video object detection aims to detect and track each object in a given video. However, due to the problem of appearance deterioration in the video, it is still challenging to obtain good results when we apply traditional image object detection methods to videos. In this paper, we propose a new feature aggregation method, called Dual Feature Aggregation (DualFeat) for video object detection. By effectively combining the temporal and spatial attention mechanisms, we make full use of the temporal and spatial information in videos. Meanwhile, we leverage a real-time tracker to track detected objects in video frames, where features are aggregated again with previously obtained features. Such a way helps to obtain more comprehensive and richer features, greatly improving the accuracy of video object detection. We perform experiments on the ILSVRC2017 dataset, and the experimental results also verify the effectiveness of our method.
Kaiwen Du, Yan Yan 0001, Hanzi Wang
ICIP2
2022 Dual Selection Network for Video Object Detection
abstract
Some off-the-shelf video object detection methods usually enhance the degraded proposal features of target frames by aggregating the proposal features from support frames. However, the proposals generated by region proposal network may not be accurate, resulting in inaccurate proposal features and limited performance. To mitigate this, we propose a novel dual selection network (DSNet) for video object detection, which contains two successive stages: selecting proposals that fit objects more closely, and selecting proposal features that are more conducive to feature aggregation. Correspondingly, the proposal selection module (PSM) aims to select better proposals by exploiting their boundary information, and the selective aggregation module (SAM) aims to select better proposal features for aggregation. Consequently, DSNet can generate more robust proposal features through the novel dual selection mechanism implemented by PSM and SAM. Extensive experiments show that our DSNet obtains 83.7% mAP and achieves superior performance over several state-of-the-art methods.
Tianxiang Hou, Qiang Qi, Yang Lu 0009, Kaiwen Du, Hanzi Wang
ICME4
2020 Learning intra-inter semantic aggregation for video object detection
abstract
Video object detection is a challenging task due to the appearance deterioration problems in video frames. Thus, object features extracted from different frames of a video are usually deteriorated in varying degrees. Currently, some state-of-the-art methods enhance the deteriorated object features in a reference frame by aggregating the undeteriorated object features extracted from other frames, simply based on their learned appearance relation among object features. In this paper, we propose a novel intra-inter semantic aggregation method (ISA) to learn more effective intra and inter relations for semantically aggregating object features. Specifically, in the proposed ISA, we first introduce an intra semantic aggregation module (Intra-SAM) to enhance the deteriorated spatial features based on the learned intra relation among the features at different positions of an individual object. Then, we present an inter semantic aggregation module (Inter-SAM) to enhance the deteriorated object features in the temporal domain based on the learned inter relation among object features. As a result, by leveraging Intra-SAM and Inter-SAM, the proposed ISA can generate discriminative features from the novel perspective of intra-inter semantic aggregation for robust video object detection. We conduct extensive experiments on the ImageNet VID dataset to evaluate ISA. The proposed ISA obtains 84.5% mAP and 85.2% mAP with ResNet-101 and ResNeXt-101, and it achieves superior performance compared with several state-of-the-art video object detectors.
Haosheng Chen 0001, Kaiwen Du, Yan Yan 0001, Hanzi Wang
MMAsia3
2020 Robust visual tracking via scale-aware localization and peak response strength
abstract
Existing regression-based deep trackers usually localize a target based on a response map, where the highest peak response corresponds to the predicted target location. Nevertheless, when the background distractors appear or the target scale changes frequently, the response map is prone to produce multiple sub-peak responses to interfere with model prediction. In this paper, we propose a robust online tracking method via Scale-Aware localization and Peak Response strength (SAPR), which can learn a discriminative model predictor to estimate a target state accurately. Specifically, to cope with large scale variations, we propose a Scale-Aware Localization (SAL) module to provide multi-scale response maps based on the scale pyramid scheme. Furthermore, to focus on the target response, we propose a simple yet effective Peak Response Strength (PRS) module to fuse the multi-scale response maps and the response maps generated by a correlation filter. According to the response map with the maximum classification score, the model predictor iteratively updates its filter weights for accurate target state estimation. Experimental results on three benchmark datasets, including OTB100, VOT2018 and LaSOT, demonstrate that the proposed SAPR accurately estimates the target state, achieving the favorable performance against several state-of-the-art trackers.
Luo Xiong, Kaiwen Du, Yan Yan 0001, Hanzi Wang
MMAsia3