VLDB 2026 Research / reviewers in the wild / expert
Wei Zhao 0022
dblp:z/WeiZhao22
· DBLP profile ↗
16ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-6060-1022ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Robust Multi-Oriented License Plate Detector and A Derived End-to-End License Plate RecognizerabstractABSTRACT Automatic license plate recognition (ALPR) systems critically depend on the robust and efficient detection of LPs under unconstrained environmental conditions, including significant viewpoint variations and complex backgrounds. To address these challenges, this paper introduces CAD‐Net, a novel corner‐aware LP detection architecture that combines a computationally efficient ResNet‐18 encoder with an efficient multi‐scale feature decoder for accurate LP corner localization. The decoder aggregates and refines features through group dilated convolutions, coordinate attention, and context gated attention, enabling enhanced focus on semantically salient regions while capturing intricate spatial dependencies. The detected LP corner points enable a polygonal region‐of‐interest alignment strategy for geometric rectification of LP features, which is integrated into an end‐to‐end LP recognition framework named CAR‐Net. Comprehensive experiments demonstrate the efficacy of our method. For LP detection, CAD‐Net attains LP detection rates of 99.9% on CCPD‐Base and 100.0% on AOLP‐RP, with a processing speed of 105 frames per second. For end‐to‐end LP recognition, CAR‐Net achieves state‐of‐the‐art performance on multiple benchmarks, CCPD (98.9%), AOLP‐RP (99.2%), PKUdata (98.5%), CLPD (82.3%), and OpenALPR‐BR (99.1%), while maintaining a real‐time inference speed of 72 frames per second. These results confirm practical viability for deployment in real‐world ALPR systems. Xudong Fan, Wei Zhao 0022 |
IET Image Process. | 2 |
| 2026 | A unified framework for image anomaly detection via reconstruction, segmentation and spatial relationship modeling
Ziniu Zhang, Wei Zhao 0022, Enrico Zio |
Neurocomputing | 2 |
| 2025 | Multiple Object Tracking in Video SAR: A Benchmark and Tracking BaselineabstractIn the context of multiobject tracking using video synthetic aperture radar (Video SAR), Doppler shifts induced by target motion result in artifacts that are easily mistaken for shadows caused by static occlusions. Moreover, appearance changes of the target caused by Doppler mismatch may lead to association failures and disrupt trajectory continuity. A major limitation in this field is the lack of public benchmark datasets for standardized algorithm evaluation. To address the above challenges, we collected and annotated 45 video SAR sequences containing moving targets, and named the video SAR MOT benchmark (VSMB). Specifically, to mitigate the effects of trailing and defocusing in moving targets, we introduce a line feature enhancement mechanism that emphasizes the positive role of motion shadows and reduces false alarms induced by static occlusions. In addition, to mitigate the adverse effects of target appearance variations, we propose a motion-aware clue discarding mechanism that substantially improves tracking robustness in video SAR. The proposed model achieves state-of-the-art performance on the VSMB, and the dataset and model are released athttps://github.com/softwarePupil/VSMB Haoxiang Chen 0008, Wei Zhao 0022, Rufei Zhang, Dongjin Li |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | PIFTrack: Point-of-Interest Flows for Multiobject Tracking in Satellite VideosabstractData annotation is extremely difficult due to the satellite imaging conditions, under which the targets are usually small, obscured, and scattered. Therefore, satellite video data annotation inevitably has noise and errors. Moreover, the irregular acceleration, sudden turns, and stops of the target make prediction and trajectory maintenance highly challenging. In this study, we propose the Points of Interest Flows Track (PIFTrack) to address the aforementioned challenges. PIFTrack improves tracking accuracy by modeling target uncertainty distributions and nonlinear motion patterns, while leveraging the spatial inclusion relationships of points of interest (PoIs) across consecutive frames. Specifically, we eliminate the rigid Dirac-based labeling assumption by employing a set of PoIs to model the spatial probability distribution of the target. PoIs enable the model to infer optimal outputs in the vicinity of annotations, thereby improving robustness to annotation errors. Secondly, to capture the real motion transfer patterns of targets in the data, we introduce a diffusion-based ordinary differential equation (ODE) model. Ultimately, we alleviate the impact of tiny object localization drifts on association results by exploiting the inclusion relationship between PoIs. PIFTrack has been extensively evaluated on the VISO, AIR-MOT, CGSTL, and VSMB datasets, exhibiting competitive performance relative to contemporary studies. Our code is open-source and available at https://github.com/softwarePupil/PIFTrack. Haoxiang Chen 0008, Wei Zhao 0022, Xudong Fan, Xiping Shang, Rufei Zhang, Dongjin Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | CCLDet: A Cross-Modality and Cross-Domain Low-Light DetectorabstractVehicle detection based on remote sensing images is widely used in urban traffic management and disaster rescue. RGB images, which are used more, lead to poor detection performance in low light conditions due to the imaging mechanism. At present, the main solution is to improve the detection performance in low light by fusing with infrared images. However, the current methods often overlook the impact of illumination changes on RGB images, and ignore the important role of high-frequency information for object detection, especially for low-light target detection. In this paper, we propose a Cross-modality and Cross-domain Low-light Detector (CCLDet) for low-light vehicle detection, including three improvements. First, an object illumination-aware module (OIAM) is proposed, which can adjust adaptively the weight of different modalities according to the object illumination intensity in the training phase and enables the detector to adapt to different lighting conditions. Second, we propose a visibility loss, which converts the position deviation into the illumination intensity deviation of each point in the object area. Compared with relying only on semantic information for object localization, the illumination makes the information that can be used for localization more abundant. Third, we design a cross-domain feature fusion module (CDFFM), which can enhance high-frequency features and enrich target information when low-frequency features are lost due to low light pollution. Extensive experiments on three challenging RGB-infrared objects detection datasets demonstrated the mAP and the parameter quantities of CCLDet over popular object detectors. Xiping Shang, Dongjin Li, Jianwei Lv, Wei Zhao 0022, Rufei Zhang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | SR-Stereo & DAPE: Stepwise Regression and Pre-Trained Edges for Practical Stereo MatchingabstractDue to the difficulty in obtaining real samples and ground truth, the generalization performance and domain adaptation performance are critical for the feasibility of stereo matching methods in practical applications. However, there are significant distributional discrepancies among different domains, which pose challenges for generalization and domain adaptation of the model. Inspired by the iteration-based methods, we propose a novel stepwise regression architecture. This architecture regresses the disparity error through multiple range-controlled clips, which effectively overcomes domain discrepancies. We implement this architecture based on the iterative-based methods, and refer to this new stereo method as SR-Stereo. Specifically, a new stepwise regression unit is proposed to replace the original update unit in order to control the range of output. Meanwhile, a regression objective segment is proposed to set the supervision individually for each stepwise regression unit. In addition, to enhance the edge awareness of models adapting new domains with sparse ground truth, we propose Domain Adaptation based on Pre-trained Edges (DAPE). In DAPE, a pre-trained stereo model and an edge estimator are used to estimate the edge maps of the target domain images, which along with the sparse ground truth disparity are used to fine-tune the stereo model. The proposed SR-Stereo and DAPE are extensively evaluated on SceneFlow, KITTI, Middbury 2014 and ETH3D. Compared with the SOTA methods and generalized methods, the proposed SR-Stereo achieves competitive in-domain and cross-domain performances. Meanwhile, the proposed DAPE significantly improves the performance of the fine-tuned model, especially in the texture-less and detailed regions. The code is available at https://github.com/zhuxing0/SR-Stereov1-DAPE Weiqing Xiao, Wei Zhao 0022 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Multiple Object Tracking in Satellite Video With Graph-Based Multiclue Fusion TrackerabstractWith the rapid advancement of satellite technology, satellite video has emerged as a key method for acquiring dynamic terrestrial information, facilitating multiple object tracking (MOT). Satellites are capable of surveying vast urban landscapes, yet the observed objects are small and dispersed among complex interference from the background, heightening the challenges in detection and association tasks for object tracking. However, current trackers often dissociate the classification task from the localization task, leading to drift in tiny object detection (TOD), and rely on prior knowledge for clue ranking, limiting model robustness. In this article, we introduce the graph-based multiclue fusion tracker (GMFTracker). Initially, we introduce a sparse sampling-based feature map correction approach to rectify the misalignment between the classification and localization feature maps. Furthermore, we developed graph neural networks (GNNs) for object relationship modeling, free from presuppositions, to tackle association challenges using relational features. GMFTracker was rigorously tested on VISO, CGSTL, and TinyPerson datasets, demonstrating its competitive performance relative to contemporary studies. Haoxiang Chen 0008, Dongjin Li, Jianwei Lv, Wei Zhao 0022, Rufei Zhang, Jingyu Xu 0005 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Trigonometric-Coded Refined Detector for High Precision Oriented Object DetectionabstractOriented object detection in aerial images is a crucial link in earth observation. As a special parameter in oriented object representation, angle is the key to achieving high-precision detection. However, the widely-used regression-based methods suffer from boundary discontinuity problem due to the periodicity of angle. To address this issue, we proposed a novel angle prediction method called Fixed Step Trigonometric Coder (FSTC). Exploiting the innate periodicity of trigonometric functions, FSTC can encode angles cyclically in a succinct, continuous, and uniform manner. Based on FSTC, we designed a single-shot oriented object detector, namely, Trigonometric-coded Refined Detector (TRDet), for high-precision object detection in real-time. TRDet consists of two modules: the Angle Optimization Module (AOM) and the Object Detection Module (ODM). AOM employs FSTC to generate high-quality rotated anchors. In ODM, a Dynamically Weighted Loss (DWL) was proposed to make the model focus on hard samples with higher angle deviation. Extensive experiments on DOTA and HRSC2016 show that both FSTC and TRDet can achieve competitive performance compared with peer works. Rufei Zhang, Sheng Shen 0013, Wei Zhao 0022, Zhiliang Zeng, Dongjin Li |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | APS-Net: An Adaptive Point Set Network for Optical Remote-Sensing Object DetectionabstractOriented object detection in optical remote-sensing images has been a challenging task due to arbitrary orientations and densely packed distribution of objects. Specifically, most existing methods lack adaptivity when regressing objects with different shapes and orientations. Although the point set representation is relatively flexible, the initial distribution of the point set is fixed in advance. In addition, some models based on the point set cannot get high location precision of points, affecting the bounding box generation. In this letter, we propose an Adaptive Point Set Network (APS-Net) for optical remote-sensing object detection, including three improvements. First, we propose the initial distribution learner (IDL) to learn the optimal initial aspect ratio, which helps the point set fit the object’s shape well. Second, we design the uncertainty measurement module (UMM), which considers the uncertainty of point location to improve location precision. Third, we introduce the local outlier factor (LOF) in the loss to punish outlier points more reasonably. Extensive experiments demonstrate that our proposed model achieves state-of-the-art performance on three commonly used datasets (i.e., DOTA-v1.0, UCAS-AOD, and HRSC2016) in the remote-sensing field. Junfeng Zhou, Rufei Zhang, Wei Zhao 0022, Sheng Shen 0013 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Improving Robustness of License Plates Automatic Recognition in Natural ScenesabstractAutomatic license plate recognition plays an important role in intelligent transportation systems and is of great significance. However, at present, most current approaches are only concerned of license plate recognition under restrictive conditions, where the license plates are shot in a frontal view and under good light conditions. These approaches are not robust enough in real-world complex capture scenarios, such as uneven light condition or oblique shooting angle. In order to improve the robustness of recognizing license plates under complex capture scenarios, a robust license plate detection network (CA-CenterNet) is proposed in this paper, together with a segmentation-free network (CNNG) for the recognition of license plate characters. CA-CenterNet can detect not only the center of each license plate, but also four vectors pointing to the four corners of the corresponding license plate, regardless of the rotation and distortion of the license plates, which gives us the possibility to rectify the distorted license plates in the source images. Then, CNNG can accurately identify the characters in the detected license plates without character segmentation. Experimental results prove that our automatic license plate recognition system has good performance in real-world complex capture scenarios and outperforms current license plate recognition models. Xudong Fan, Wei Zhao 0022 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Encoder- and Decoder-Based Networks Using Multiscale Feature Fusion and Nonlocal Block for Remote Sensing Image Semantic SegmentationabstractWith the development of convolutional neural networks, the semantic segmentation of remote sensing images has been widely developed, but there are still some unsolved problems in this field due to the lack of multiscale information and the feature mismatch at the upsampling process. To solve these problems, we propose a network called multiscale feature fusion and alignment network (MFANet). MFANet is composed of an encoder and a decoder. The encoder contains a fully convolutional network, a multilevel feature fusion block (MLFFB), and a multiscale feature pyramid (MSFP). These subnetworks can obtain fine-grained feature maps that are full of multiscale and global features and improve segmentation results at multiple object scales. Moreover, MFANet uses a light convolution subnetwork, called decoder, to upsample the segmentation map stage by stage. Combining three scales of features, the decoder can promote the feature alignment at the upsampling stage. Along with the decoder, MFANet utilizes a multistage supervision loss to enhance the localization performance and boundary regression ability. Benefitting from the encoder and decoder structure and the innovative components inside encoder, MFANet is very powerful for the semantic segmentation of remote sensing images and can suit the complicated environment. We evaluate our MFANet on the Vaihingen and Potsdam data sets, and it outperforms the state-of-art methods both in the metric and visual effect. Zhaochen Sun, Wei Zhao 0022 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | Innovative CFAR detector with effective parameter estimation method for generalised gamma distribution and iterative sliding window strategyabstractFor high‐resolution synthetic aperture radar (SAR), constant false alarm rate (CFAR) detectors are widely used to separate targets from background and it is proved that generalised gamma distribution ( ) can model the non‐homogeneous clutter appropriately. However, CFAR detectors based on possess two problems. First, the methods for solving highly non‐linear equations of resort to numerical iterative algorithms, which are computationally expensive. To avoid this, the authors present a novel analytic solution for parameter estimation of by exploiting a third‐order approximation of the polygamma function. This novel analytic solution can result in more accurate parameter estimation and can fit a wider range of the parameters. The second problem is that in a multi‐target SAR image, some target pixels may be classified as clutter pixels. To select pixels of interest, an iterative sliding window approach is often used in CFAR. They analyse the relationship between the detection probability and the number of iterations, and prove that this strategy can reduce the miss rate and false detection rate effectively. On the basis of the aforementioned parameter estimation method and the iterative sliding window approach, an innovative and effective CFAR detector is proposed in this study and its superiority is demonstrated by experiments. Wei Zhao 0022, Xueqing Yang, Qichao Peng |
IET Image Process. | 1 |
| 2018 | A Constrained Optimization Approach for Image Gradient EnhancementabstractThe human visual system is not very sensitive to the absolute luminance of an image, but rather responds to local luminance changes, i.e., the gradient of an image. In this paper, we propose a constrained optimization approach for image gradient enhancement. The gradient strength of the enhanced image can be controlled directly using a target gradient strength parameter in the cost function. To suppress artifacts and ensure that contrast improves, a novel constraint is included in the optimization. Due to the number of variables in optimization-based image enhancement techniques being equal to the number of gray scales, we quantize the image using a k -means clustering-based histogram mergence (KCHM) method before enhancement. KCHM can significantly reduce the number of image gray scales while effectively preserving the subjective quality. This is useful considering the reduction of variables is good for solving optimization and reducing computation cost. Experimental results demonstrate that the proposed method can significantly improve the subjective image quality by enhancing both the contrast and the image gradient. Lidong Huang, Wei Zhao 0022, Besma Roui-Abidi, Mongi A. Abidi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | A Novel Hybrid Method of Parameters Tuning in Support Vector Regression for Reliability Prediction: Particle Swarm Optimization Combined With Analytical SelectionabstractSupport vector regression (SVR) is a widely used technique for reliability prediction. The key issue for high prediction accuracy is the selection of SVR parameters, which is essentially an optimization problem. As one of the most effective evolutionary optimization methods, particle swarm optimization (PSO) has been successfully applied to tune SVR parameters and is shown to perform well. However, the inherent drawbacks of PSO, including slow convergence and local optima, have hindered its further application in practical reliability prediction problems. To overcome these drawbacks, many improvement strategies are being developed on the mechanisms of PSO, whereas there is little research exploring a priori information about historical data to improve the PSO performance in the SVR parameter selection task. In this paper, a novel method controlling the inertial weight of PSO is proposed to accelerate its convergence and guide the evolution out of local optima, by utilizing the analytical selection (AS) method based on a priori knowledge about SVR parameters. Experimental results show that the proposed ASPSO method is almost as accurate as the traditional PSO and outperforms it in convergence speed and ability in tuning SVR parameters. Therefore, the proposed ASPSO-SVR shows promising results for practical reliability prediction tasks. Wei Zhao 0022, Enrico Zio |
IEEE Trans. Reliab. | 1 |
| 2015 | Entropy maximisation histogram modification scheme for image enhancementabstractContrast enhancement plays an important role in image processing applications. The global histogram equalisation (GHE)‐based techniques are very popular for their simpleness. In the author's study, the authors originally divide the GHE techniques into two steps, that is, the pixel populations mergence (PPM) step and the grey‐levels distribution (GLD) step. In the PPM step, the pixel populations of adjoining grey scales to be mapped to the same grey scale are merged firstly in input histogram. Then, the new grey scales are redistributed according to a corresponding transformation function in the GLD step. This division is meaningful because the entropy of enhanced image is only determined by pixel populations regardless of grey levels. Then, they prove the entropy of enhanced image is reduced because of mergence. Inspired by GHE, they propose a novel entropy maximisation histogram modification scheme, which also consists of PPM and GLD steps. However, the entropy is maximised, that is, the reduction of entropy is minimised under originally presented entropy maximisation rule in their PPM step. In the GLD step, they redistribute the grey scales in the merged histogram using a log‐based distribution function to control the enhancement level. Experimental results demonstrate the proposed method is effective. Wei Zhao 0022, Lidong Huang, Zebin Sun |
IET Image Process. | 1 |
| 2015 | An advanced gradient histogram and its application for contrast and gradient enhancement
Lidong Huang, Wei Zhao 0022, Zebin Sun |
J. Vis. Commun. Image Represent. | 2 |