VLDB 2026 Research / reviewers in the wild / expert
Wenzhe Wang
dblp:150/1794
· DBLP profile ↗
22ranked-venue papers
7as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SEM-RRT*: Fast Risk Assessment and Path Planning in Uneven Terrain using Statistical Elevation MapabstractPath planning in uneven terrain scenarios is one of the core capabilities of intelligent off-road robots and vehicles. The complex terrain undulations often cause bumpy motion and sharp turns along the planned path, making smooth and safe path planning challenging. Most path planning methods in this community rely on dense point cloud maps as direct inputs, which inevitably incur high computational overhead for map representation and terrain assessment. To address these problems, we propose a novel path planning method toward uneven terrains, SEM-RRT*, which balances both planning quality and computational efficiency. First, we propose a map representation namely Statistical Elevation Map (SEM), which is lightweight to store and compute. Then, to enable fast terrain risk assessment, a terrain risk filter with omnidirectional and multi-scale characteristics is designed. Finally, we incorporate multi-objective cost evaluation, backward search, and rolling optimization strategies into the Informed RRT* framework, leveraging its path optimality on large scale map. Extensive experiments in challenging terrain scenarios, such as hills, canyons, and volcanic landscapes, show that SEM-RRT* outperforms existing methods in both path quality and computational time. Xudong Dong 0004, Wenzhe Wang |
IROS | 4 |
| 2024 | SinLane: Siamese Visual Transformer via Pyramid Feature Integration for Lane DetectionabstractLane detection is an important yet challenging task in autonomous driving systems. Based on the development of the Visual Transformer, early Transformer-based lane detection studies have achieved promising results in some scenarios. However, for complex road conditions such as uneven illumination intensity and heavy traffic, the performance of these methods remains limited and may even be worse than that of contemporaneous CNN-based methods. In this paper, we propose a novel Transformer-based end-to-end network, called SinLane, that attains the attention weights focusing on the sparse yet meaningful locations and improves the accuracy of lane detection in complex environments. SinLane is composed of a novel Siamese Visual Transformer structure and a novel Feature Pyramid Network (FPN) structure called Pyramid Feature Integration (PFI). We utilize the proposed PFI to better integrate global semantics and finer-scale features and to promote the optimization of the Transformer. Moreover, the designed Siamese Visual Transformer is combined with multiple levels of the PFI and is employed to refine the multi-scale lane line features output from the PFI. Extensive experiments on three benchmark datasets of lane detection demonstrate that our SinLane achieves state-of-the-art results with high accuracy and efficiency. Specifically, our SinLane improves the accuracy by over 3% compared to the current best-performing Transformer-based method for lane detection on CULane. Zinan Lv, Wenzhe Wang, Danny Ziyi Chen |
ECAI | 3 |
| 2024 | Handling the Non-smooth Challenge in Tensor SVD: A Multi-objective Tensor Recovery Framework
Wanglong Lu, Wenzhe Wang, Yankai Cao, Xiaoqin Zhang 0002, Xianta Jiang |
ECCV (14) | 3 |
| 2024 | A Siamese Transformer with Hierarchical Refinement for Lane DetectionabstractLane detection is an important yet challenging task in autonomous driving systems. Existing lane detection methods mainly rely on finer-scale information to identify key points of lane lines. Since local information in realistic road environments is frequently obscured by other vehicles or affected by poor outdoor lighting conditions, these methods struggle with the regression of such key points. In this paper, we propose a novel Siamese Transformer with hierarchical refinement for lane detection to improve the detection accuracy in complex road environments. Specifically, we propose a high-to-low hierarchical refinement Transformer structure, called LAne TRansformer (LATR), to refine the key points of lane lines, which integrates global semantics information and finer-scale features. Moreover, exploiting the thin and long characteristics of lane lines, we propose a novel Curve-IoU loss to supervise the fit of lane lines. Extensive experiments on three benchmark datasets of lane detection demonstrate that our proposed new method achieves state-of-the-art results with high accuracy and efficiency. Specifically, our method achieves improved F1 scores on the OpenLane dataset, surpassing the current best-performing method by 5.0 points. Zinan Lv, Wenzhe Wang, Danny Ziyi Chen |
NeurIPS | 3 |
| 2024 | Self-triggered adjustable prescribed performance control for stochastic multiagent systems with communication faults
Wenzhe Wang, Yingnan Pan, Hongru Ren |
Appl. Intell. | 1 |
| 2024 | Emotional agents enabled bilateral negotiation: Persuasion strategies generated by agents' affect infusion and preference
Jinghua Wu, Wenzhe Wang, Yan Li 0104 |
Expert Syst. Appl. | 2 |
| 2024 | Synchronous MDADT-Based Fuzzy Adaptive Tracking Control for Switched Multiagent Systems via Modified Self-Triggered MechanismabstractIn this paper, a self-triggered fuzzy adaptive switched control strategy is proposed to address the synchronous tracking issue in switched stochastic multiagent systems (MASs) based on mode-dependent average dwell-time (MDADT) method. Firstly, a synchronous slow switching mechanism is considered in switched stochastic MASs and realized through a class of designed switching signals under MDADT property. By utilizing the information of both specific agents under switching dynamics and observers with switching features, the synchronous switching signals are designed, which reduces the design complexity. Then, a switched state observer via a switching-related output mask is proposed. The information of agents and their preserved neighbors is utilized to construct the observer and the observation performance of states is improved. Moreover, a modified self-triggered mechanism is designed to improve control performance via proposing auxiliary function. Finally, by analysing the relationship between the synchronous switching problem and the different switching features of the followers, the synchronous slow switching mechanism based on MDADT is obtained. Meanwhile, the designed self-triggered controller can guarantee that all signals of the closed-loop system are ultimately bounded under the switching signals. The effectiveness of the designed control method can be verified by some simulation results. Hongjing Liang, Wenzhe Wang, Yingnan Pan, Hak-Keung Lam, Jiayue Sun |
IEEE Trans. Fuzzy Syst. | 2 |
| 2022 | Handling Slice Permutations Variability in Tensor RecoveryabstractThis work studies the influence of slice permutations on tensor recovery, which is derived from a reasonable assumption about algorithm, i.e. changing data order should not affect the effectiveness of the algorithm. However, as we will discussed in this paper, this assumption is not satisfied by tensor recovery under some cases. We call this interesting problem as Slice Permutations Variability (SPV) in tensor recovery. In this paper, we discuss SPV of several key tensor recovery problems theoretically and experimentally. The obtained results show that there is a huge gap between results by tensor recovery using tensor with different slices sequences. To overcome SPV in tensor recovery, we develop a novel tensor recovery algorithm by Minimum Hamiltonian Circle for SPV (TRSPV) which exploits a low dimensional subspace structures within data tensor more exactly. To the best of our knowledge, this is the first work to discuss and effectively solve the SPV problem in tensor recovery. The experimental results demonstrate the effectiveness of the proposed algorithm in eliminating SPV in tensor recovery. Xiaoqin Zhang 0002, Wenzhe Wang, Xianta Jiang |
AAAI | 3 |
| 2022 | Discriminative Cervical Lesion Detection in Colposcopic Images With Global Class Activation and Local Bin ExcitationabstractAccurate cervical lesion detection (CLD) methods using colposcopic images are highly demanded in computer-aided diagnosis (CAD) for automatic diagnosis of High-grade Squamous Intraepithelial Lesions (HSIL). However, compared to natural scene images, the specific characteristics of colposcopic images, such as low contrast, visual similarity, and ambiguous lesion boundaries, pose difficulties to accurately locating HSIL regions and also significantly impede the performance improvement of existing CLD approaches. To tackle these difficulties and better capture cervical lesions, we develop novel feature enhancing mechanisms from both global and local perspectives, and propose a new discriminative CLD framework, called CervixNet, with a Global Class Activation (GCA) module and a Local Bin Excitation (LBE) module. Specifically, the GCA module learns discriminative features by introducing an auxiliary classifier, and guides our model to focus on HSIL regions while ignoring noisy regions. It globally facilitates the feature extraction process and helps boost feature discriminability. Further, our LBE module excites lesion features in a local manner, and allows the lesion regions to be more fine-grained enhanced by explicitly modelling the inter-dependencies among bins of proposal feature. Extensive experiments on a number of 9888 clinical colposcopic images verify the superiority of our method (AP$_{.75}$= 20.45) over state-of-the-art models on four widely used metrics. Tingting Chen 0002, Xuechen Liu 0004, Ruiwei Feng, Wenzhe Wang, Chunnv Yuan, Weiguo Lu, Haizhen He, Honghao Gao, Haochao Ying, Danny Ziyi Chen, Jian Wu 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | PR-Net: Preference Reasoning for Personalized Video Highlight DetectionabstractPersonalized video highlight detection aims to shorten a long video to interesting moments according to a user’s preference, which has recently raised the community’s attention. Current methods regard the user’s history as holistic information to predict the user’s preference but negating the inherent diversity of the user’s interests, resulting in vague preference representation. In this paper, we propose a simple yet efficient preference reasoning framework (PR-Net) to explicitly take the diverse interests into account for frame-level highlight prediction. Specifically, distinct user-specific preferences for each input query frame are produced, presented as the similarity weighted sum of history highlights to the corresponding query frame. Next, distinct comprehensive preferences are formed by the user-specific preferences and a learnable generic preference for more overall highlight measurement. Lastly, the degree of highlight and non-highlight for each query frame is calculated as semantic similarity to its comprehensive and non-highlight preferences, respectively. Besides, to alleviate the ambiguity due to the incomplete annotation, a new bidirectional contrastive loss is proposed to ensure a compact and differentiable metric space. In this way, our method significantly outperforms state-of-the-art methods with a relative improvement of 12% in mean accuracy precision. Runnan Chen, Penghao Zhou, Wenzhe Wang, Nenglun Chen, Xing Sun 0001, Wenping Wang 0001 |
ICCV | 3 |
| 2021 | Dig into Multi-modal Cues for Video Retrieval with Hierarchical AlignmentabstractMulti-modal cues presented in videos are usually beneficial for the challenging video-text retrieval task on internet-scale datasets. Recent video retrieval methods take advantage of multi-modal cues by aggregating them to holistic high-level semantics for matching with text representations in a global view. In contrast to this global alignment, the local alignment of detailed semantics encoded within both multi-modal cues and distinct phrases is still not well conducted. Thus, in this paper, we leverage the hierarchical video-text alignment to fully explore the detailed diverse characteristics in multi-modal cues for fine-grained alignment with local semantics from phrases, as well as to capture a high-level semantic correspondence. Specifically, multi-step attention is learned for progressively comprehensive local alignment and a holistic transformer is utilized to summarize multi-modal cues for global alignment. With hierarchical alignment, our model outperforms state-of-the-art methods on three public video retrieval datasets. Wenzhe Wang, Mengdan Zhang, Runnan Chen, Guanyu Cai, Penghao Zhou, Xing Sun 0001 |
IJCAI | 1 |
| 2021 | Frame Aggregation and Multi-modal Fusion Framework for Video-Based Person Recognition
Fangtao Li, Wenzhe Wang, Chenghao Yan, Bin Wu 0001 |
MMM (1) | 2 |
| 2021 | Multi-modality fusion learning for the automatic diagnosis of optic neuropathy
Zheng Cao 0005, Chuanbin Sun, Wenzhe Wang, Xiangshang Zheng, Jian Wu 0001, Honghao Gao |
Pattern Recognit. Lett. | 3 |
| 2021 | Interactive Few-Shot Learning: Limited Supervision, Better Medical Image SegmentationabstractMany known supervised deep learning methods for medical image segmentation suffer an expensive burden of data annotation for model training. Recently, few-shot segmentation methods were proposed to alleviate this burden, but such methods often showed poor adaptability to the target tasks. By prudently introducing interactive learning into the few-shot learning strategy, we develop a novel few-shot segmentation approach called Interactive Few-shot Learning (IFSL), which not only addresses the annotation burden of medical image segmentation models but also tackles the common issues of the known few-shot segmentation methods. First, we design a new few-shot segmentation structure, called Medical Prior-based Few-shot Learning Network (MPrNet), which uses only a few annotated samples (e.g., 10 samples) as support images to guide the segmentation of query images without any pre-training. Then, we propose an Interactive Learning-based Test Time Optimization Algorithm (IL-TTOA) to strengthen our MPrNet on the fly for the target task in an interactive fashion. To our best knowledge, our IFSL approach is the first to allow few-shot segmentation models to be optimized and strengthened on the target tasks in an interactive and controllable manner. Experiments on four few-shot segmentation tasks show that our IFSL approach outperforms the state-of-the-art methods by more than 20% in the DSC metric. Specifically, the interactive optimization algorithm (IL-TTOA) further contributes ~10% DSC improvement for the few-shot segmentation models. Ruiwei Feng, Xiangshang Zheng, Tianxiang Gao, Jintai Chen, Wenzhe Wang, Danny Ziyi Chen, Jian Wu 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Dual-Level Selective Transfer Learning for Intrahepatic Cholangiocarcinoma Segmentation in Non-enhanced Abdominal CT
Wenzhe Wang, Qingyu Song 0004, Jiarong Zhou, Ruiwei Feng, Tingting Chen 0002, Wenhao Ge, Danny Ziyi Chen, Shaohua Kevin Zhou, Jian Wu 0001 |
MICCAI (1) | 1 |
| 2020 | Multi-Cue and Temporal Attention for Person Recognition in Videos
Wenzhe Wang, Bin Wu 0001, Fangtao Li |
PRCV (2) | 1 |
| 2019 | Multi-view Learning with Feature Level Fusion for Cervical Dysplasia Diagnosis
Tingting Chen 0002, Xinjun Ma, Xuechen Liu 0004, Wenzhe Wang, Ruiwei Feng, Jintai Chen, Chunnv Yuan, Weiguo Lu, Danny Ziyi Chen, Jian Wu 0001 |
MICCAI (1) | 4 |
| 2019 | LSRC: A Long-Short Range Context-Fusing Framework for Automatic 3D Vertebra Localization
Jintai Chen, Ruoqian Guo, Bohan Yu, Tingting Chen 0002, Wenzhe Wang, Ruiwei Feng, Danny Ziyi Chen, Jian Wu 0001 |
MICCAI (6) | 6 |
| 2019 | Regular and Small Target Detection
Wenzhe Wang, Bin Wu 0001, Jinna Lv, Pilin Dai |
MMM (2) | 1 |
| 2018 | Road Damage Detection and Classification with Faster R-CNNabstractThis technical paper presents the method that we use in the Road Damage Detection and Classification Challenge, which is designed to detect damages contained in road images photographed by a vehicle-mounted smartphone. In this task, we apply Faster R-CNN to detect and classify damaged roads. Through analyses of aspect ratios and sizes of the damaged areas in the training dataset, we adjust relevant parameters of the model. In order to solve the problem of unbalanced data distribution of different classes, we introduce some data augmentation techniques (contrast transformation, brightness adjustment, and Gaussian blur) before training. Experimental results demonstrate that our method can achieve a Mean F1-Score of 0.6255 in the competition. The source code and model are publicly available at https://github.com/zhezheey/tf-faster-rcnn-rddc. Wenzhe Wang, Bin Wu 0001, Sixiong Yang |
IEEE BigData | 1 |
| 2018 | A Framework for Identifying Diabetic Retinopathy Based on Anti-noise Detection and Attention-Based Fusion
Zhiwen Lin, Ruoqian Guo, Tingting Chen 0002, Wenzhe Wang, Danny Ziyi Chen, Jian Wu 0001 |
MICCAI (2) | 6 |
| 2018 | Deep Active Self-paced Learning for Accurate Pulmonary Nodule Segmentation
Wenzhe Wang, Tingting Chen 0002, Danny Ziyi Chen, Jian Wu 0001 |
MICCAI (2) | 1 |