VLDB 2026 Research / reviewers in the wild / expert
Borui Zhang
dblp:230/7918
· DBLP profile ↗
16ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Alignment-Invertibility Regularization for Explainable Neural NetworksabstractDeep learning has profoundly impacted society, yet the inherent nature of deep neural networks hinders further application to high-reliability industries. To demystify these closed-boxes, numerous works attempt to improve the explainability by observing or impacting internal variables of the models. However, existing methods rely on heuristics without rigorous theoretical foundations, often requiring intricate model modifications or redesigns. This work first formalizes two fundamental properties of explainability: alignment and invertibility, serving as theoretical pillars for rigorous interpretability analysis. Building on these, we introduce Bort, a plug-and-play optimizer that enforces Boundedness and orthogonality constraints on model parameters to improve explainability. These constraints are theoretically derived from the alignment and invertibility principles. Considering conventional optimizers can not leverage data features for precise attribution, we present a data-aware extension, termed DBort, which integrates an auxiliary loss term. Intriguingly, in the linear case, DBort converges to Principal Component Analysis (PCA). Our in-depth analysis of penalty term design reveals that $l_{1}$l1-based penalties provide a more stringent adherence to the imposed constraints compared to their $l_{2}$l2 counterparts. Our experiments involve reconstructing and backtracking through the optimized model representations, which reveal a marked enhancement in explainability. Furthermore, leveraging Bort, we successfully synthesize explainable adversarial examples without additional training. Notably, Bort consistently improves the classification accuracy across diverse architectures, including ResNet and DeiT, on benchmark datasets such as MNIST, CIFAR-10, and ImageNet. Borui Zhang, Qihang Rao, Jie Zhou 0001, Jiwen Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Task Allocation and Trajectory Optimization for Multi-UAV Cargo Systems with Cellular-Connected ConstraintsabstractABSTRACT This paper investigates a multi‐UAV cargo delivery scenario, where each UAV picks up goods from one location and delivers them to another destination while maintaining connectivity with the ground cellular network. Optimizing task assignment and UAV trajectory design to minimize completion time under the constraints is a significant challenge. To address this, the approach is structured into two principal phases. First, Dijkstra's algorithm is utilized to derive the shortest paths between points while ensuring communication connectivity meets specific quality constraints. Second, these paths are integrated with a novel hybrid optimization algorithm fusing a genetic algorithm and an ant colony algorithm to solve the coupled task assignment and route planning problem subject to communication and payload limitations. The hybrid approach efficiently balances exploration and exploitation, leading to superior task allocation and route planning. Numerical results show that our proposed method is effective in balancing task allocation and reducing overall completion time by comparing it with other integrated optimization techniques. Borui Zhang, Kui Huang, Dingcheng Yang |
IET Commun. | 1 |
| 2024 | SelfOcc: Self-Supervised Vision-Based 3D Occupancy Predictionabstract3D occupancy prediction is an important task for the robustness of vision-centric autonomous driving, which aims to predict whether each point is occupied in the surrounding 3D space. Existing methods usually require 3D occupancy labels to produce meaningful results. However, it is very laborious to annotate the occupancy status of each voxel. In this paper, we propose SelfOcc to explore a self-supervised way to learn 3D occupancy using only video sequences. We first transform the images into the 3D space (e.g., bird's eye view) to obtain 3D representation of the scene. We directly impose constraints on the 3D representations by treating them as signed distance fields. We can then render 2D images of previous and future frames as self-supervision signals to learn the 3D representations. We propose an MVS-embedded strategy to directly optimize the SDF-induced weights with multiple depth proposals. Our SelfOcc out-performs the previous best method SceneRF by 58.7% using a single frame as input on SemanticKITTI and is the first self-supervised work that produces reasonable 3D occupancy for surround cameras on nuScenes. SelfOcc produces high-quality depth and achieves state-of-the-art results on novel depth synthesis, monocular depth estimation, and surround-view depth estimation on the SemanticKITTI, KITTI-2015, and nuScenes, respectively. Code: https://github.com/huang-yh/SelfOcc. Yuanhui Huang 0002, Wenzhao Zheng, Borui Zhang, Jie Zhou 0001, Jiwen Lu |
CVPR | 3 |
| 2024 | LowRankOcc: Tensor Decomposition and Low-Rank Recovery for Vision-Based 3D Semantic Occupancy PredictionabstractIn this paper, we present a tensor decomposition and low-rank recovery approach (LowRankOcc) for vision-based 3D semantic occupancy prediction. Conventional methods model outdoor scenes with fine-grained 3D grids, but the sparsity of non-empty voxels introduces consider-able spatial redundancy, leading to potential overfitting risks. In contrast, our approach leverages the intrinsic low-rank property of 3D occupancy data, factorizing voxel representations into low-rank components to efficiently mitigate spatial redundancy without sacrificing performance. Specifically, we present the Vertical-Horizontal (VH) de-composition block factorizes 3D tensors into vertical vectors and horizontal matrices. With our “decomposition-encoding-recovery” framework, we encode 3D contexts with only 1/2D convolutions and poolings, and subsequently recover the encoded compact yet informative context features back to voxel representations. Experimental results demonstrate that LowRankOcc achieves state-of-the-art performances in semantic scene completion on the Se-manticKITTI dataset and 3D occupancy prediction on the nuScenes dataset. Linqing Zhao, Xiuwei Xu, Ziwei Wang 0010, Borui Zhang, Wenzhao Zheng, Dalong Du, Jie Zhou 0001, Jiwen Lu |
CVPR | 5 |
| 2024 | OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
Wenzhao Zheng, Yuanhui Huang 0002, Borui Zhang, Yueqi Duan, Jiwen Lu |
ECCV (13) | 4 |
| 2024 | Path Choice Matters for Clear Attributions in Path MethodsabstractRigorousness and clarity are both essential for interpretations of DNNs to engender human trust. Path methods are commonly employed to generate rigorous attributions that satisfy three axioms. However, the meaning of attributions remains ambiguous due to distinct path choices. To address the ambiguity, we introduce Concentration Principle, which centrally allocates high attributions to indispensable features, thereby endowing aesthetic and sparsity. We then present SAMP, a model-agnostic interpreter, which efficiently searches the near-optimal path from a pre-defined set of manipulation paths. Moreover, we propose the infinitesimal constraint (IC) and momentum strategy (MS) to improve the rigorousness and optimality. Visualizations show that SAMP can precisely reveal DNNs by pinpointing salient image pixels.
We also perform quantitative experiments and observe that our method significantly outperforms the counterparts. Borui Zhang, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
ICLR | 1 |
| 2024 | GERWR: Identifying the Key Pathogenicity- Associated sRNAs of Magnaporthe Oryzae Infection in Rice Based on Graph Embedding and Random Walk With RestartabstractRice blast, caused by Magnaporthe oryzae(M.oryzae), is a destructive rice disease that reduces rice yield by 10% to 30% annually. It also affects other cereal crops such as barley, wheat, rye, millet, sorghum, and maize. Small RNAs (sRNAs) play an essential regulatory role in fungus-plant interaction during the fungal invasion, but studies on pathogenic sRNAs during the fungal invasion of plants based on multi-omics data integration are rare. This paper proposes a novel approach called Graph Embedding combined with Random Walk with Restart (GERWR) to identify pathogenic sRNAs based on multi-omics data integration during M.oryzae invasion. By constructing a multi-omics network (MRMO), we identified 29 pathogenic sRNAs of rice blast fungus. Further analysis revealed that these sRNAs regulate rice genes in a many-to-many relationship, playing a significant regulatory role in the pathogenesis of rice blast disease. This paper explores the pathogenic factors of rice blast disease from the perspective of multi-omics data analysis, revealing the inherent connection between pathogenic factors of different omics. It has essential scientific significance for studying the pathogenic mechanism of rice blast fungus, the rice blast fungus-rice model system, and the pathogen-host interaction in related fields. Hao Zhang 0064, Tianheng Zhao, Enshuang Zhao, Lanhui Li, Guihua Li, Borui Zhang, Qing-Ming Qin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2024 | Toward Integrity and Detail With Ensemble Learning for Salient Object Detection in Optical Remote-Sensing ImagesabstractOptical remote sensing image salient object detection (ORSI-SOD) poses significant challenges due to complicated object variances and interfering surroundings. Although existing methods have achieved impressive performance, they encounter difficulties in balancing deep and shallow features, leading to limitations in preserving object integrity and edge detail. To address this, we propose the Integrated and Detailed Ensemble Learning (IDEL) framework, which incorporates hierarchical branches with deep supervision. By divide-and-conquer, each branch captures information with a specific granularity, while the fusion module combines all outputs to generate the final saliency maps. To ensure the effectiveness of ensemble learning, IDEL is designed to satisfy two necessary conditions: the weak learner property and branch independence. Firstly, we utilize the Transformer blocks with a global receptive field and purify intermediate features with the Deep Supervision Module (DSM) to enhance the performance of each branch. Secondly, we disentangle multiple branches through hardness-aware weights and hierarchical supervision labels, allowing them to learn distinct features. Qualitative visualizations demonstrate the effectiveness of each module, and extensive experimental results conducted on three popular ORSI datasets confirm the superiority of IDEL compared to other state-of-the-art (SOTA) counterparts. Kangjie Liu, Borui Zhang, Jiwen Lu, Haibin Yan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Interpretable Traffic Accident Prediction: Attention Spatial-Temporal Multi-Graph Traffic Stream Learning ApproachabstractTraffic accident prediction plays a vital role in Intelligent Transportation Systems (ITS), where a large number of traffic streaming data are generated on a daily basis for spatiotemporal big data analysis. The rarity of accidents and the absent interconnection information make it hard for spatiotemporal modeling. Moreover, the inherent characteristic of the black box predictive model makes it difficult to interpret the reliability and effectiveness of the deep learning model. To address these issues, a novel self-explanatory spatial-temporal deep learning model–Attention Spatial-Temporal Multi-Graph Convolutional Network (ASTMGCN) is proposed for traffic accident prediction. The original recorded rare accident data is formulated as a multivariate irregularly interval-aligned dataset, and the temporal discretization method is used to transfer into regularly sampled time series. Multiple graphs are defined to construct edge features and represent spatial relationships when node-related information is missing. Multi-graph convolutional operators and attention mechanisms are integrated into a Sequence-to-Sequence (Seq2Seq) framework to effectively capture dynamic spatial and temporal features and correlations in multi-step prediction. Comparative experiments and interpretability analysis are conducted on a real-world data set, and results indicate that our model can not only yield superior prediction performance but also has the advantage of interpretability. Chaojie Li, Borui Zhang, Zeyu Wang 0011, Yin Yang 0001, Xiaojun Zhou 0001, Shirui Pan, Xinghuo Yu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Bort: Towards Explainable Neural Networks with Bounded Orthogonal Constraint
Borui Zhang, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
ICLR | 1 |
| 2023 | Graph Reinforcement Learning for Securing Critical Loads by E-Mobility
Borui Zhang, Chaojie Li, Boyang Hu, Xiangyu Li 0008, Zhao Yang Dong |
ICONIP (7) | 1 |
| 2022 | Attributable Visual Similarity LearningabstractThis paper proposes an attributable visual similarity learning (AVSL) framework for a more accurate and ex-plainable similarity measure between images. Most existing similarity learning methods exacerbate the unexplain-ability by mapping each sample to a single point in the em-bedding space with a distance metric (e.g., Mahalanobis distance, Euclidean distance). Motivated by the human se-mantic similarity cognition, we propose a generalized simi-larity learning paradigm to represent the similarity between two images with a graph and then infer the overall simi-larity accordingly. Furthermore, we establish a bottom-up similarity construction and top-down similarity inference framework to infer the similarity based on semantic hier-archy consistency. We first identify unreliable higher-level similarity nodes and then correct them using the most co-herent adjacent lower-level similarity nodes, which simulta-neously preserve traces for similarity attribution. Extensive experiments on the CUB-200-2011, Cars196, and Stanford Online Products datasets demonstrate significant improve-ments over existing deep similarity learning methods and verify the interpretability of our framework.11Code: https://github.com/zbr17/AVSL. Borui Zhang, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
CVPR | 1 |
| 2022 | Dynamic Metric Learning with Cross-Level Concept Distillation
Wenzhao Zheng, Yuan Huang 0002, Borui Zhang, Jie Zhou 0001, Jiwen Lu |
ECCV (24) | 3 |
| 2022 | Asynchronous Autoregressive Prediction for Satellite Anomaly DetectionabstractThis paper proposes an ASynchronous Autoregressive Prediction (ASAP) method for satellite anomaly detection. We empirically observe that a single classification model can hardly detect unknown anomalous situations and neglect the Markov nature of temporal satellite data. To address this, we adopt an autoregressive model to deal with the prediction of unknown anomaly for satellite data. We further propose a non-uniform temporal encoding method for asynchronous data and a median filtering method for more accurate detection. To reduce the effect of outliers, we employ an adaptive threshold selection method to achieve a more robust classification boundary. Experiments on real satellite data demonstrate that the proposed ASAP method outperforms the baseline classification method by 55.79%. Haopeng Zhang 0014, Lifang Yuan, Borui Zhang, Chengkun Wang |
VCIP | 4 |
| 2021 | Deep Relational Metric LearningabstractThis paper presents a deep relational metric learning (DRML) framework for image clustering and retrieval. Most existing deep metric learning methods learn an embedding space with a general objective of increasing interclass distances and decreasing intraclass distances. However, the conventional losses of metric learning usually suppress intraclass variations which might be helpful to identify samples of unseen classes. To address this problem, we propose to adaptively learn an ensemble of features that characterizes an image from different aspects to model both interclass and intraclass distributions. We further employ a relational module to capture the correlations among each feature in the ensemble and construct a graph to represent an image. We then perform relational inference on the graph to integrate the ensemble and obtain a relation-aware embedding to measure the similarities. Extensive experiments on the widely-used CUB-200-2011, Cars196, and Stanford Online Products datasets demonstrate that our framework improves existing deep metric learning methods and achieves very competitive results.1 Wenzhao Zheng, Borui Zhang, Jiwen Lu, Jie Zhou 0001 |
ICCV | 2 |
| 2021 | GAEBic: A Novel Biclustering Analysis Method for miRNA-Targeted Gene Data Based on Graph Autoencoder
Hao Zhang 0064, Haowu Chang, Qing-Ming Qin, Borui Zhang, Xue-Qing Li, Tianheng Zhao |
J. Comput. Sci. Technol. | 5 |