VLDB 2026 Research / reviewers in the wild / expert
Donghai Zhai
dblp:142/6734
· DBLP profile ↗
13ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0001-8396-5710ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-category spatiotemporal consensus and discriminative networks for weakly-supervised temporal action localization
Kunlun Wu, Donghai Zhai |
Neural Networks | 2 |
| 2025 | Monocular 3D Vehicle Detection Based on Video Point Tracking: A Novel Approach Without Prior InformationabstractMonocular vision systems have emerged as a promising solution for 3D object detection due to their cost-effectiveness and deployment simplicity. However, existing methods heavily rely on prior information such as camera parameters and complex 3D dataset annotations. Most current approaches focus on static RGB images, struggling to recover 3D information from 2D inputs. We propose a novel monocular 3D vehicle detection method based on point tracking from a roadside perspective, leveraging dynamic inter-frame associations in video data. Our method operates without prior data like vehicle dimensions or camera parameters, and eliminates the need for complex 3D dataset construction and annotation. It integrates object detection and semantic segmentation for feature extraction, high-precision inter-frame data association, and trajectory-based state analysis. 3D reconstruction is achieved through reference point determination and coordinate calibration under geometric constraints. Experimental results on the DAIR-V2X-I dataset demonstrate superior performance over baseline models, with error reductions ranging from 1.3% to 7.3% across various metrics (A3DS, A2DS, AGS, ALS, APS, AW). This approach presents a practical solution for monocular 3D object detection while eliminating dependency on prior information and complex datasets. Da Yang 0004, Donghai Zhai, Meng-Si Yu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Daily Schedule Recommendation in Urban Life Based on Deep Reinforcement LearningabstractIn our daily lives, people frequently consider daily schedule to meet their needs, such as going to a barbershop for a haircut, then eating in a restaurant, and finally shopping in a supermarket. Reasonable activity location [or point-of-interest (POI)] and activity sequencing will help people save a lot of time and get better services. In this article, we propose a reinforcement learning-based deep activity factor balancing model to recommend a reasonable daily schedule according to user's current location and needs. The proposed model consists of a deep activity factor balancing network (DAFB) and a reinforcement learning framework. First, the DAFB is proposed to fuse multiple factors that affect daily schedule recommendation (DSR). Then, a reinforcement learning framework based on policy gradient is used to learn the parameters of the DAFB. Further, on the feature storage based on the matrix method, we compress the feature storage space of the candidate POIs. Finally, the proposed method is compared with seven benchmark methods using two real-world datasets. Experimental results show that the proposed method is adaptive and effective. Jia Liu 0033, Donghai Zhai, Wei Huang 0037, Shenggong Ji, Junbo Zhang 0004, Tianrui Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Boundary-Aware Axial Attention Network for High-Quality Pavement Crack DetectionabstractPavement crack detection is a practical and challenging task that has the ability to significantly reduce the burden of manual building and road maintenance in intelligent transportation systems. Existing methods mainly focus on addressing common crack diseases and are poor in generalizing to other conditions of crack detection due to diverse environmental factors (e.g., illumination), topology complexity, and intensity in-homogeneity. Moreover, the samples suffer from the severe foreground-background imbalance and the model is easily prone to overfitting on trained anomalies, resulting in unsatisfactory performance. To tackle the aforementioned challenges and achieve high-quality pavement crack detection, we propose an innovative approach termed boundary-aware axial attention network (BAAN), which is composed of multiple position-guided axial attention (PAA) modules in a hierarchical encoder-decoder architecture. Specifically, it learns efficient contextual information via decomposed multidimensional position-guided attention to capture more precise spatial structures, and the proposed boundary regularization module (BRM) mines more discriminative foreground-background relationships to regularize the ambiguous details between diverse spatial regions. Moreover, we propose a novel boundary refinement loss (BRL) to alleviate the challenges associated with regional losses (e.g., pixel-wise cross-entropy loss) in the context of heavily imbalanced crack detection problems. The proposed BAAN is evaluated on four crack datasets and experimental results indicate that the BAAN consistently outperforms the state-of-the-art methods with fewer computational requirements. Kunlun Wu, Bo Peng 0006, Donghai Zhai |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Multitask-Guided Deep Clustering With Boundary AdaptationabstractMultitask learning uses external knowledge to improve internal clustering and single-task learning. Existing multitask learning algorithms mostly use shallow-level correlation to aid judgment, and the boundary factors on high-dimensional datasets often lead algorithms to poor performance. The initial parameters of these algorithms cause the border samples to fall into a local optimal solution. In this study, a multitask-guided deep clustering (DC) with boundary adaptation (MTDC-BA) based on a convolutional neural network autoencoder (CNN-AE) is proposed. In the first stage, dubbed multitask pretraining (M-train), we construct an autoencoder (AE) named CNN-AE using the DenseNet-like structure, which performs deep feature extraction and stores captured multitask knowledge into model parameters. In the second phase, the parameters of the M-train are shared for CNN-AE, and clustering results are obtained by deep features, which is termed as single-task fitting (S-fit). To eliminate the boundary effect, we use data augmentation and improved self-paced learning to construct the boundary adaptation. We integrate boundary adaptors into the M-train and S-fit stages appropriately. The interpretability of MTDC-BA is accomplished by data transformation. The model relies on the principle that features become important as the reconfiguration loss decreases. Experiments on a series of typical datasets confirm the performance of the proposed MTDC-BA. Compared with other traditional clustering methods, including single-task DC algorithms and the latest multitask clustering algorithms, our MTDC-BA achieves better clustering performance with higher computational efficiency. Deep features clustering results demonstrate the stability of MTDC-BA by visualization and convergence verification. Through the visualization experiment, we explain and analyze the whole model data input and the middle characteristic layer. Further understanding of the principle of MTDC-BA. Through additional experiments, we know that the proposed MTDC-BA is efficient in the use of multitask knowledge. Finally, we carry out sensitivity experiments on the hyper-parameters to verify their optimal performance. Xiaole Zhao, Dengmin Wen, Donghai Zhai |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | YOLOv5s-BSS: A Novel Deep Neural Network for Crack Detection of Road DamageabstractCracks are one of the most common and significant types of road surface damage, posing a threat to the safety of pedestrians and vehicles. If left untreated, cracks can lead to severe consequences such as road and bridge collapse. Therefore, it is essential to develop an efficient road crack detection method. Traditional crack identification methods have the problem of being largely affected by the environment and having low recognition accuracy. In this paper, we propose a road crack detection model based on an improved You Only Look Once version 5 (YOLOv5) model that addresses the limitations of existing state-of-the-art crack detection methods in terms of accuracy and detection speed. First, we replace the intersection over union (IoU) loss function with the SCYLLA-IoU (SIoU) loss function for better accuracy. Second, to enhance detection performance, we replace the feature pyramid network (FPN) with a bi-directional feature pyramid network (BiFPN). Finally, to better extract spatial feature information of different sizes, we modify the original Spatial Pyramid Pooling-Fast (SPPF) module of YOLOv5 by using Spatial Pyramid Pooling Cross-Stage Partial Connections (SPPCSPC). We evaluated our YOLOv5s-BiFPN-SPPCSPC-SIoU (YOLOv5s-BSS) method on the dataset from the IEEE 2020 Global Road Damage Detection Challenge (GRDDC) and achieved promising results on road damage datasets from China, Japan, and the United States. The [email protected] of different cracks in three datasets reached 84.9%, 54.6%, and 71%. Our method outperforms related methods, with an increase of 0.7%, 0.7%, and 2.8% over YOLOv5s. Conghua Wei, Qianjun Zhang, Yan Yang 0001, Jixin Zhang, Donghai Zhai |
IEEE Big Data | 6 |
| 2023 | Enhancing diversity and robustness of clustering ensemble via reliability weighted measure
Panpan Ni, Donghai Zhai, Tianrui Li 0001 |
Appl. Intell. | 3 |
| 2023 | A hybrid deep learning pavement crack semantic segmentation
Zaid Al-Huda, Bo Peng 0006, Riyadh Nazar Ali Algburi, Mugahed A. Al-antari, Rabea Al-Jarazi, Donghai Zhai |
Eng. Appl. Artif. Intell. | 6 |
| 2022 | An overview of edge and object contour detection
Daipeng Yang, Bo Peng 0006, Zaid Al-Huda, Asad Malik 0002, Donghai Zhai |
Neurocomputing | 5 |
| 2022 | ASS-GAN: Asymmetric semi-supervised GAN for breast ultrasound image segmentation
Donghai Zhai, Bijie Hu, Haipeng Zou |
Neurocomputing | 1 |
| 2022 | Self-Supervised Robust Deep Matrix Factorization for Hyperspectral UnmixingabstractHyperspectral unmixing is a critical step to process hyperspectral images (HSIs). Nonnegative matrix factorization (NMF) has drawn extensive attention in remotely sensed hyperspectral unmixing since it does not require prior knowledge about the pure spectral constituents (endmembers) in the scene. However, this approach is normally implemented as a single-layer procedure, which does not allow for a refinement of the obtained endmember abundances. In addition, HSIs suffer from the interference of sparse noise (besides Gaussian noise), which brings challenges when pursuing efficient hyperspectral unmixing. To address these issues, we propose a new self-supervised robust deep matrix factorization (SSRDMF) model for hyperspectral unmixing, which consists of two parts:encoderanddecoder. In theencoder, a multilayer nonlinear structure is designed to directly map the observed HSI data to the corresponding abundances. The abundances are then decoded by thedecoder, in which the connected weights are treated as the extracted endmembers. By modeling the sparse noise explicitly, the proposed method can reduce the effect caused by both Gaussian and sparse noise. Furthermore, a self-supervised constraint is included for exploring the spectral information, which is beneficial to further improve unmixing performance. To validate our method, we have conducted extensive experiments on both synthetic and real datasets. Our experiments reveal that our newly developed SSRDMF achieves superior unmixing performance compared to other state-of-the-art methods. Heng-Chao Li 0001, Xin-Ru Feng, Donghai Zhai, Qian Du 0001, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Optimal Scale of Hierarchical Image Segmentation with Scribbles Guidance for Weakly Supervised Semantic SegmentationabstractDeep convolutional neural networks (DCNNs) trained on the pixel-level annotated images have achieved improvements in semantic segmentation. Due to the high cost of labeling training data, their applications may have great limitation. However, weakly supervised segmentation approaches can significantly reduce human labeling efforts. In this paper, we introduce a new framework to generate high-quality initial pixel-level annotations. By using a hierarchical image segmentation algorithm to predict the boundary map, we select the optimal scale of high-quality hierarchies. In the initialization step, scribble annotations and the saliency map are combined to construct a graphic model over the optimal scale segmentation. By solving the minimal cut problem, it can spread information from scribbles to unmarked regions. In the training process, the segmentation network is trained by using the initial pixel-level annotations. To iteratively optimize the segmentation, we use a graphical model to refine segmentation masks and retrain the segmentation network to get more precise pixel-level annotations. The experimental results on Pascal VOC 2012 dataset demonstrate that the proposed framework outperforms most of weakly supervised semantic segmentation methods and achieves the state-of-the-art performance, which is [Formula: see text] mIoU. Zaid Al-Huda, Donghai Zhai, Yan Yang 0001, Riyadh Nazar Ali Algburi |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2021 | Local2Global: Unsupervised multi-view deep graph representation learning with Nearest Neighbor Constraint
Yan Yang 0001, Donghai Zhai, Tianrui Li 0001, Jielei Chu, Hao Wang 0068 |
Knowl. Based Syst. | 3 |