EDBT 2026 Demo / reviewers in the wild / expert
Zhiwen Wang 0001
dblp:08/5974-1 · also Zhi-Wen Wang 0001
· DBLP profile ↗
39ranked-venue papers
3as first author
33since 2021 · last 2026
0000-0003-2309-7282ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 1 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DSA mamba: A model for advanced medical image classification
Zhiwen Wang 0001, Haoyu Yin, Jilin Yu, Mengsi Gong, Qiaoqiao Chen, Zhenlin He, Danying Wang |
Expert Syst. Appl. | 1 |
| 2026 | Adaptive confidence-driven learning and cross-modal hard sample mining for unsupervised visible-infrared person re-identification
Canlong Zhang, Haifei Ma, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei |
Inf. Process. Manag. | 5 |
| 2026 | GKC-Net: Gated KAN with Channel-Position Attention Mechanism for Image Deraining
Mengsi Gong, Jilin Yu, Zhiwen Wang 0001, Senlin Chi |
Pattern Recognit. | 3 |
| 2026 | CMAG: Cross-Modal Attention and Graph-Enhanced Memory for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identification (USL-VI-ReID) has garnered widespread attention due to its surveillance application value in complex environments. However, it faces four key challenges: modality discrepancy, batch training limitations, pseudo-label noise, and camera view bias. This paper proposes the CMAG (Cross-Modal Attention and Graph-enhanced Memory) framework, which innovatively combines circular topology structure with cross-modal attention mechanisms to address these challenges. CMAG introduces four core innovations: (1) applying circular topology structure to provide pseudo-label verification through detecting circular paths in feature space, effectively addressing the pseudo-label noise problem; (2) designing a cross-modal attention mechanism for Vision Transformers with residual fusion to balance modality-specific and shared information, solving the modality discrepancy issue; (3) constructing a graph-structured memory enhancement module with adaptive graph construction and multi-layer feature propagation to overcome batch training limitations; and (4) integrating camera-specific clustering with circular structure constraints to reduce camera background bias. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate the effectiveness of CMAG, achieving approximately 3.5% improvement in Rank-1 accuracy and 2.8% in mAP on average compared to state-of-the-art methods, validating our approach’s advantages in addressing key challenges in unsupervised cross-modal person re-identification.Code is available at https://github.com/hurryup186/CMAG. Canlong Zhang, Junwei Tian, Haifei Ma, Zhixin Li 0001, Zhiwen Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Uncertainty-Aware Prototype Semantic Decoupling for Text-Based Person Search in Full Images
Zengli Luo, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei |
KSEM (2) | 4 |
| 2025 | DiffusionLight:a multi-agent reinforcement learning approach for traffic signal control based on shortcut-diffusion model
Jilin Yu, Zhiwen Wang 0001 |
Appl. Intell. | 2 |
| 2025 | Max-Min Pooling and Squeeze Excitation Lightweight Bidirectional Mamba for image classification
Senlin Chi, Zhiwen Wang 0001, Lianyuan Jang, Mengsi Gong |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Recursively learning fine-grained spatial-temporal features for video-based person Re-identification
Haifei Ma, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Multi-scale Feature Refinement via Perspective Scaling and Adaptive Regularization for text-based person search
Sheng Xie, Canlong Zhang, Runcong Ma, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Unsupervised infrared-visible person re-identification by multi-level Dual-Stream Contrastive Learning
Canlong Zhang, Haifei Ma, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei |
Neurocomputing | 5 |
| 2025 | Dynamic feature projection and grouped contrastive learning for text-to-image person re-identification
Shun He, Canlong Zhang, Xiaochun Lu, Zhixin Li 0001, Zhiwen Wang 0001 |
Knowl. Based Syst. | 5 |
| 2025 | Joint feature augmentation and posture label for cloth-changing person re-identification
Liman Jiang, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei |
Multim. Syst. | 5 |
| 2025 | Learning Simultaneous and Sequential Decisions in Multi-Agent Systems With Application to Traffic Signal ControlabstractEfficient traffic signal control (TSC) has been one of the most useful ways for reducing urban road congestion. By modeling each intersection as an autonomous agent, multiagent reinforcement learning (MARL) shows remarkable performance in solving dynamic TSC. However, most of TSC methods based on MARL suffer from a non-stationarity problem since agents update their policies simultaneously. To resolve this issue, this paper considers multi-intersection TSC as a multi-agent sequential decision-making process with policy online learning. We utilize a sequential model such as Transformer architecture to learn the multi-agent joint policy. By carefully designing the advantage function of each agent, the monotonic improvement property can be guaranteed. Moreover, to fully exploit the advantages of both simultaneous and sequential MARL, we further propose a novel MARL network selection algorithm (MARL-NS) which selectively employs simultaneous MARL only at states that sequential MARL might fall into local optimum. Our theory proves that MARL-NS preserves cooperative MARL converge properties. Finally, we validate the proposed MARL-NS method on a unified TSC benchmark, LibSignal. Experimental results show that our method can outperform the baseline methods in network-level and arterial coordination. Haipeng Zhang 0005, Zhiwen Wang 0001, Jilin Yu, Caoqing Jiang, Gongkun Luo, Weiwei Wu 0001, Wanyuan Wang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Gradually Spatio-Temporal Feature Activation for Target TrackingabstractMost existing transformer-based trackers use ViT [1] as the backbone to extract and fuse feature tokens of target templates and search region. Since both the target template and the search region contain background information, their tokens are prone to background interference in interaction that affects tracking performance. We propose a GFATrack that combines spatiotemporal information with prominent target features. The tracker mainly consists of the Dynamic Template Refinement branch and the Search Feature Enhancement branch. The former activates the target features in the dynamic template and provides temporal information. The latter enhances search features by aggregating spatial information from the initial template and temporal information from the dynamic template to achieve precise tracking. Our proposed Feature Activation Module can effectively fuse refined features with reference features and highlight high-similarity features among them. At the same time, we propose a concise and effective dynamic threshold update strategy to capture time context update dynamic templates from historical prediction results. Many experiments have verified the effectiveness and latest performance of the proposed method. Yanfang Deng, Canlong Zhang, Zhixin Li 0001, Chunrong Wei, Zhiwen Wang 0001, Shuqi Pan |
ICASSP | 5 |
| 2024 | Mask-guided Salient Feature Mining for Cloth-Changing Person Re-identificationabstractCloth-changing person re-identification (CC-ReID) aims at retrieving pedestrians with changing clothes across multiple cameras. Most methods often focus on exploiting discriminative biometric features for resisting the clothing changes. Nevertheless, simply concatenating various features not only increases the computations, but also introduces redundant information. In this paper, we propose a Mask-guided Salient Feature Mining (MSFM) to learn cloth-irrelevant features. Specifically, we introduce human parsing results in the data pre-processing stage, and exploit region-specific cues by performing pixel-level mask on the parsing results. Besides, a Multi-local Attention (MLA) is proposed, where the model can focus on local cues in horizontal direction and obtain robust identity-related representations. Meanwhile, we introduce a part loss supervised by selective masking regions for capturing fine-grained features and constraining clothing features. Extensive experiments on two public cloth-changing datasets demonstrate our proposed MSFM can achieve superior performance over existing state-of-the-art methods. Liman Jiang, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei |
ICME | 5 |
| 2024 | Person Re-identification utilizing Text to Search VideoabstractPrevious research in pedestrian re-identification can be broadly categorized into three classes: image-to-image, video-to-video, and text-to-image pedestrian re-identification. However, these paradigms exhibit certain limitations in practical applications. Hence, this paper introduces a novel task: utilizing natural language to retrieve pedestrians in videos. Specifically, given a textual description of a person, the model’s objective is to retrieve the pedestrian from a video dataset that best matches the provided text. Due to the absence of datasets specifically designed for text-to-video pedestrian re-identification, we undertook manual annotations on the video-based pedestrian re-identification dataset, MARS, to establish the T-MARS dataset. In this paper, we propose the Implicit Alignment Framework with Local Capture Attention Module(IAFL) for text-to-video pedestrian re-identification. The Local Capture Attention Module incorporates local priors while modeling inter-frame complementary relationships, thereby effectively transferring knowledge from text-to-image models to the task of text-to-video pedestrian re-identification. A plethora of text-to-video models are evaluated and compared on this dataset. We observe that the proposed Implicit Alignment Framework is more suitable for achieving cross-modal granularity alignment in pedestrian re-identification tasks. Additionally, IAFL establishes the state-of-the-art performance in pedestrian search. Shunkai Zhou, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei |
ICME | 4 |
| 2024 | Detection of pneumonia in chest X-ray image based on separable convolutional neural networkabstractSummary Pneumonia disease progress rapidly leading to missed diagnosis and diagnosis error. However, it is significant to be diagnosed accurately as well as adopted timely treatments, particularly for children and elderly people over 65. In the first place, we propose a novel method for data augmentation to expand non‐case samples in order to solve the problem of imbalance of the raw data. Then we have designed a neural network based on separable convolution layers, using deep learning to classify between Pneumonia cases and normal cases from chest X‐ray images. Finally, well‐known performance measures are utilized to evaluate our model, obtaining the following results: Accuracy (0.97), AUC (0.99), Sensitivity (0.98), Specificity (0.98), Precision (0.95), F1‐Score (0.97), and Youden Index (0.90). Simultaneously, we highlighted the proposed neural network recognition area via Class Activation Method for the purpose of providing an explainable diagnosis. Compared with the‐state‐of‐art models, our model performed better, and it may be regarded as an alternative in countries or regions lack of the relevant equipment and professional radiologists. Gongkun Luo, Zhiwen Wang 0001, Kuangquan Wang |
Concurr. Comput. Pract. Exp. | 2 |
| 2024 | Full-view salient feature mining and alignment for text-based person search
Sheng Xie, Canlong Zhang, Enhao Ning, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei |
Expert Syst. Appl. | 5 |
| 2024 | A review on video person re-identification based on deep learning
Haifei Ma, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei |
Neurocomputing | 5 |
| 2024 | Pedestrian detection using RetinaNet with multi-branch structure and double pooling attention mechanism
Lincai Huang, Zhiwen Wang 0001, Xiaobiao Fu |
Multim. Tools Appl. | 2 |
| 2024 | Text-based person search by non-saliency enhancing and dynamic label smoothing
Yonghua Pang, Canlong Zhang, Zhixin Li 0001, Chunrong Wei, Zhiwen Wang 0001 |
Neural Comput. Appl. | 5 |
| 2024 | Target-Aware Tracking With Spatial-Temporal Context AttentionabstractCurrent trackers only rely on a fixed target template to localize the target in each frame, which is however prone to fail in case of fast appearance changes or the presence of distractor objects. Having some historical knowledge about the tracked targets as well as their surrounding scenes can be highly beneficial for robust tracking. This historical information can be propagated through the sequence and used to timely perceive the change in target appearance and explicitly avoid distractor objects. In this work, we propose a Spatial-Temporal Context Attention (STCA) model which utilizes the appearance and state information of previously tracked targets as well as their surrounding scenes to more accurately localize the real target in the current frame. We embed an improved position encoder into the STCA, which enables the target template, context template and search patch to perform extensive interactional fusion through simultaneously self-attention and cross-attention calculation. By embedding the STCA module into Transformer, we construct a target-aware based online tracking network (named TATrack) that has a backbone to extract features better suited to the tracking task, a neck to further suppress distractors and highlight target, and a classification-regression head to make the tracking scores consistently reflect the quality of the bounding boxes. In addition, we also design a simple yet effective online updating approach to select high-quality context templates. Our tracker reaches the latest level on several benchmarks, including LaSOT, TrackingNet, GOT10k, OTB100 and UAV123. The code and trained models are available at https://github.com/hekaijie123/TATrack. Kaijie He, Canlong Zhang, Sheng Xie, Zhixin Li 0001, Zhiwen Wang 0001, Rui-Guo Qin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Target-Aware Tracking with Long-Term Context AttentionabstractMost deep trackers still follow the guidance of the siamese paradigms and use a template that contains only the target without any contextual information, which makes it difficult for the tracker to cope with large appearance changes, rapid target movement, and attraction from similar objects. To alleviate the above problem, we propose a long-term context attention (LCA) module that can perform extensive information fusion on the target and its context from long-term frames, and calculate the target correlation while enhancing target features. The complete contextual information contains the location of the target as well as the state around the target. LCA uses the target state from the previous frame to exclude the interference of similar objects and complex backgrounds, thus accurately locating the target and enabling the tracker to obtain higher robustness and regression accuracy. By embedding the LCA module in Transformer, we build a powerful online tracker with a target-aware backbone, termed as TATrack. In addition, we propose a dynamic online update algorithm based on the classification confidence of historical information without additional calculation burden. Our tracker achieves state-of-the-art performance on multiple benchmarks, with 71.1% AUC, 89.3% NP, and 73.0% AO on LaSOT, TrackingNet, and GOT-10k. The code and trained models are available on https://github.com/hekaijie123/TATrack. Kaijie He, Canlong Zhang, Sheng Xie, Zhixin Li 0001, Zhiwen Wang 0001 |
AAAI | 5 |
| 2023 | Reinforced domain adaptation with attention and adversarial learning for unsupervised person Re-ID
Peiyi Wei, Canlong Zhang, Yanping Tang, Zhixin Li 0001, Zhiwen Wang 0001 |
Appl. Intell. | 5 |
| 2023 | Multi-scene image enhancement based on multi-channel illumination estimation
Runxing Zhao, Zhiwen Wang 0001, Wuyuan Guo, Canlong Zhang |
Expert Syst. Appl. | 2 |
| 2023 | Discriminative feature mining with relation regularization for person re-identification
Jing Yang 0046, Canlong Zhang, Zhixin Li 0001, Yanping Tang, Zhiwen Wang 0001 |
Inf. Process. Manag. | 5 |
| 2022 | SATNet: Captioning with Semantic Alignment and Feature Enhancement
Wenhui Bai, Canlong Zhang, Zhixin Li 0001, Peiyi Wei, Zhiwen Wang 0001 |
ICONIP (3) | 5 |
| 2022 | Multi-level Network Based on Text Attention and Pose-Guided for Person Re-ID
Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001 |
ICONIP (7) | 4 |
| 2022 | Graph Structure Guided Transformer for Semantic SegmentationabstractSegmentation is an essential operation of image processing, and utilizing long-range context information is the key for pixel-wise prediction tasks such as semantic segmentation. Convolutional Neural Networks (CNNs) are good at modeling local relationships through convolutional operations, but they are often inefficient in capturing global relationships between distant regions and require stacking multiple convolutional lay-ers. Utilizing the advantages of transformer in modeling long-range dependency, this paper proposes a novel Graph Structure Guided Transformer (GSGT) to realize semantic segmentation. Different from the previous methods that hard-divide the image in a regular grid manner, our graph projection method maps the two-dimensional feature map into a graph structure according to certain semantic relevance, so as to meet the data structure form required by the transformer. Meanwhile, to fully utilize the graph structure information, we also propose a graph embedding attention module, which utilizes the local topology of the graph structure to complement the global context of transformer. Moreover, GSGT is easy to be incorporated with various CNN backbones and transformer model variants to significantly improve the segmentation accuracy and convergence speed. Experiments on Cityscapes, VOC and ADE20K datasets demonstrate that the proposed method performs well in semantic seamentation task. Luyang Qian, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001 |
ICTAI | 4 |
| 2022 | Image Captioning According to User's Intention and StyleabstractExciting image captioning models are usually individuality-agnostic, and they cannot generate individual description according to the user's intention and language style. To address above problem, this paper proposes a personalized image captioning model by using fine-grained scene control graph and the style control factor to respectively represent user's intention and speaking style. More specifically, we first construct a scene control graph that consists of the object, its attribute and the relationship between it and other objects, and employe the graph flow attention to control the focus of the description. Secondly, propose a style generator to extract user's style pattern from his profile generated based on his gender, age, education level and other information. Finally, enter style factor in the style control module before generating sentences, which enables the language decoder to output a personalized caption for image. The experimental results on MSCOCO and FlickrStyle datasets show that the proposed method can generate personalized and diverse image caption sentences. Canlong Zhang, Zhiwen Wang 0001, Zhixin Li 0001 |
IJCNN | 3 |
| 2022 | Pedestrian detection in infrared image based on depth transfer learning
Zhiwen Wang 0001 |
Multim. Tools Appl. | 1 |
| 2021 | Domain-Adaptation Person Re-Identification via Style Translation and Clustering
Peiyi Wei, Canlong Zhang, Zhixin Li 0001, Yanping Tang, Zhiwen Wang 0001 |
ICONIP (1) | 5 |
| 2021 | Low illumination color image enhancement based on Gabor filtering and Retinex theory
Zhiwen Wang 0001, Dong Lv, Chanlong Zhang |
Multim. Tools Appl. | 2 |
| 2020 | Multi-level Visual Fusion Networks for Image CaptioningabstractImage captioning is a multi-modal complex task in machine learning. Traditional methods focus only on entities in visual strategy networks, and can't reason about the relationship between entities and attributes. There are problems of exposure bias and error accumulation in language strategy networks. To this end, this paper proposes a multi-level visual fusion network model based on reinforcement learning. In the visual strategy network, multi-level neural network modules are used to transform visual features into feature sets of visual knowledge. The fusion network generates function words that make the description more fluent, and is used for the interaction between the visual strategy network and the language strategy network. The self-criticism strategy gradient algorithm based on reinforcement learning in language strategy networks is used to achieve end-to-end optimization of visual fusion networks. We evaluated our model on the Flickr 30K and MS-COCO datasets, and verified the accuracy of the model and the diversity of model learning subtitles through experiments. Our model achieves better performance over state-of-the-art methods. Dongming Zhou 0003, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001 |
IJCNN | 4 |
| 2019 | Sparse High-Level Attention Networks for Person Re-IdentificationabstractWhen extracting convolutional features from person images with low resolution, a large amount of available information will be lost due to the pooling, which will lead to the reduction of the accuracy of person classification models. This paper proposes a new classification model, which can effectively to reduce the loss of important information about the convolutional neural works. Firstly, the SE module in the Squeeze-and-Excitation Networks (SENet) is extracted and normalized to generate the Normalized Squeeze-and-Excitation (NSE) module. Then, 4 NSE modules are applied to the convolutional layers of ResNet. Finally, a Sparse Normalized Squeeze-and-Excitation Network (SNSENet) is constructed by adding 4 shortcut connections between the convolutional layers. The experimental results of Market-1501 show that the rank-1 of SNSE-ResNet-50 is 3.7% and 4.2% higher than that of SE-ResNet-50 and ResNet-50 respectively, it has done well in other person re-identification datasets. Sheng Xie, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001 |
ICTAI | 4 |
| 2019 | Joint spatiograms for multi-modality tracking with online update
Canlong Zhang, Yanping Tang, Zhixin Li 0001, Zhiwen Wang 0001 |
Pattern Recognit. Lett. | 4 |
| 2018 | Parallel Connecting Deep and Shallow CNNs for Simultaneous Detection of Big and Small Objects
Canlong Zhang, Dongcheng He, Zhixin Li 0001, Zhiwen Wang 0001 |
PRCV (4) | 4 |
| 2018 | Joint compressive representation for multi-feature tracking
Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001 |
Neurocomputing | 3 |
| 2017 | Analysis of Influencing Factors of Shooting Rate Based on Trajectory Prediction of the BasketballabstractThe main factors affecting the shooting rate are the angle of shooting, the speed of shot, the height of the shot and the trajectory of the basketball under the influence of direction and speed of wind. In this paper, we propose the method to analyze the influence factors of shooting rate based on trajectory prediction. It can accurately get the motion trajectory of basketball through the forecast. We use forecast analysis with mechanical theory and motion trajectory to get shot height, angle and the release speed of the shooting of the impact, providing theoretical basis for shooting training. The experimental results show that the shooting parameters can be greatly improved by shooting the relevant parameters of the shooting point. Zhiwen Wang 0001, Lianyuan Jiang, Canlong Zhang, Zhenghuan Hu |
WISA | 1 |