EDBT 2026 Demo / reviewers in the wild / expert
Zhenbo Lu
dblp:42/501
· DBLP profile ↗
24ranked-venue papers
3as first author
18since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging Visual Captions for Enhanced Zero-Shot HOI DetectionabstractZero-shot Human-Object Interaction (HOI) detection aims to identify both seen and unseen HOI categories in an image. Most existing methods rely on semantic knowledge distilled from CLIP to find novel interactions but fail to fully exploit the powerful generalization ability of vision-language models, leading to impaired transferability. In this paper, we introduce a novel framework for zero-shot HOI detection. We first utilize vision-language models (VLMs) to generate visual captions from multiple perspectives, including humans, objects, and environments, to enhance interaction understanding. Then, we propose a multi-modal fusion encoder to fully leverage these visual captions. Additionally, to equip the HOI detector with a thorough consideration of contextual information in the image, we design a novel multi-branch HOI network that aggregates features at the instance, union, and global levels. Experiments on prevalent benchmarks demonstrate that our model achieves promising performance under a variety of zero-shot settings. The source codes are available at https://github.com/aqingcv/VC-HOI. Yanqing Zeng, Yunyao Mao, Zhenbo Lu, Wengang Zhou 0001, Houqiang Li |
ICASSP | 3 |
| 2025 | $\hbox {I}^2$MD: 3D Action Representation Learning with Inter- and Intra-Modal Mutual Distillation
Yunyao Mao, Jiajun Deng, Wengang Zhou 0001, Zhenbo Lu, Wanli Ouyang, Houqiang Li |
Int. J. Comput. Vis. | 4 |
| 2025 | Bayesian Deep Learning Approach for Real-Time Lane-Based Arrival Curve Reconstruction at Intersection Using License Plate Recognition DataabstractThe acquisition of real-time and accurate traffic arrival information is of vital importance for proactive traffic control systems, especially in partially connected vehicle environments. License plate recognition (LPR) data that record both vehicle departures and identities are proven to be desirable in reconstructing lane-based arrival curves in previous works. Existing LPR data-based methods are predominantly designed for reconstructing historical arrival curves. For real-time reconstruction of multi-lane urban roads, it is pivotal to determine the lane choice of real-time link-based arrivals, which has not been exploited in previous studies. In this study, we propose a Bayesian deep learning approach for real-time lane-based arrival curve reconstruction, in which the lane choice patterns and uncertainties of link-based arrivals are both characterized. Specifically, the learning process is designed to effectively capture the relationship between partially observed link-based arrivals and lane-based arrivals, which can be physically interpreted as lane choice proportion. Moreover, the lane choice uncertainties are characterized using Bayesian parameter inference techniques, minimizing arrival curve reconstruction uncertainties, especially in low LPR data matching rate conditions. Real-world experiment results conducted in multiple matching rate scenarios demonstrate the superiority and necessity of lane choice modeling in reconstructing arrival curves. Chengchuan An, Yao-Jan Wu, Zhenbo Lu, Jingxin Xia |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Effective Adversarial Attack Approach to Assess the Vulnerability of Autonomous Vehicle Trajectory Prediction ModelsabstractTrajectory prediction is crucial for autonomous vehicle (AV) trajectory planning. The deep learning based trajectory prediction models are easily manipulated by cyber attack such as adversarial attack or confidential information tampering. Current research in adversarial attack typically relies on vehicle physical motion boundaries to conduct linear search, which limits the diversity of samples and covers up the vulnerabilities of model. Moreover, the reckless driving behaviors underlying the generated trajectory samples can be easily detected and smoothed. In this study, a dual constraint optimization framework for adversarial attack is developed. The proposed framework integrates hard constraint of physical boundary with soft constraint of driving risk map to simulate the actual vehicles interaction. Subsequently, Stochastic Gradient Descent (SGD) incorporates Hard-Soft constraint to increase the search space of local optimal solution. The high-precision vehicle trajectory data (sampling interval 0.1s) from the Next Generation Simulation (NGSIM) dataset supports microscopic traffic flow analysis and is used for validating our methods. The vulnerability of the prediction model is revealed from number of attack frames and input features. Results show that our proposed method increases the Average Displacement Errors (ADE) by 42.04% and Final Displacement Error (FDE) by 24.19% compared to the state-of-the-art method. Chengchuan An, Jingxin Xia, Zhenbo Lu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Diffusion With Reinforcement Learning for Pedestrian Trajectory PredictionabstractThe trajectories of pedestrian movements involve uncertainty, requiring a predictive probability model capable of modeling the underlying multimodality. To predict the trajectory of pedestrian, most existing methods try to learn the probability distribution of real pedestrian trajectories and then independently sample multiple times from this distribution to obtain a set of possible future paths. However, naively learning the distribution of real-world trajectories leads to sub-optimal results. In this paper, we design a model-agnostic reinforcement learning-based framework for pedestrian trajectory prediction. This framework models pedestrian trajectory generation as a denoising process, which is further formulated as a multi-step decision-making process. In our framework, we subtly design a reward function, which is used to optimize the diffusion model with policy-based reinforcement learning. We make evaluation on multiple benchmark datasets, including ETH/UCY and SDD datasets, where our approach achieves promising results. Our source code will be released at:https://github.com/ustc-yaojinchen/DRL-for-PTP Jinchen Yao, Zhenbo Lu, Yunyao Mao, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Remember the Past for Better Future: Memory-Augmented Offline RLabstractAs a foundation of human intelligence, memory has been found to be critical for human attention and decision making. However, it is usually underutilized in current reinforcement learning literature, primarily serving as training data. Researchers have rarely noticed the use of memory in other perspectives. To explore the potential of memory architectures, we focus on the offline reinforcement learning setting, where a fixed memory buffer is provided, and propose a novel framework to exploit it. Specifically, an attention-based architecture is designed to adaptively utilize past memories in learned environment dynamic models, providing reliable references for the estimation of future states. Such memory-augmented environment dynamic models are then applied to boost the training of RL policies. While demonstrating superior empirical performance, our method is highly extendable to most of offline model-based RL algorithms without any change in the pipelines or theoretical conclusions. Yaodong Yang 0001, Zhenbo Lu, Wengang Zhou 0001, Houqiang Li |
IJCNN | 3 |
| 2024 | Coordinate-aligned multi-camera collaboration for active multi-object tracking
Zeyu Fang, Jian Zhao 0018, Mingyu Yang 0003, Zhenbo Lu, Wengang Zhou 0001, Houqiang Li |
Multim. Syst. | 4 |
| 2024 | Efficient and Robust Freeway Traffic Speed Estimation Under Oblique Grid Using Vehicle Trajectory DataabstractAccurately estimating spatiotemporal traffic states on freeways is a significant challenge due to limited sensor deployment and potential data corruption. In this study, we propose an efficient and robust low-rank model for precise spatiotemporal traffic speed state estimation (TSE) using low-penetration vehicle trajectory data. Leveraging traffic wave priors, an oblique grid-based matrix is first designed to transform the inherent dependencies of spatiotemporal traffic states into the algebraic low-rankness of a matrix. Then, with the enhanced traffic state low-rankness in the oblique matrix, a low-rank matrix completion method is tailored to explicitly capture spatiotemporal traffic propagation characteristics and precisely reconstruct traffic states. In addition, an anomaly-tolerant module based on a sparse matrix is developed to accommodate corrupted data input and thereby improve the TSE model robustness. Notably, driven by the understanding of traffic waves, the computational complexity of the proposed efficient method is only correlated with the problem size itself, not with dataset size and hyperparameter selection prevalent in existing studies. Extensive experiments demonstrate the effectiveness, robustness, and efficiency of the proposed model. The performance of the proposed method achieves up to a 12% improvement in Root Mean Squared Error (RMSE) in the TSE scenarios and an 18% improvement in RMSE in the robust TSE scenarios, and it runs more than 20 times faster than the state-of-the-art (SOTA) methods. Chengchuan An, Yuheng Jia, Jiachao Liu, Zhenbo Lu, Jingxin Xia |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | An Integrated Intra-View and Inter-View Framework for Multiple Traffic Variable Data Simultaneous RecoveryabstractRapid advancements in traffic monitoring and sensing technologies have permitted the multiplex and democratized gathering of numerous traffic data (e.g. speed, volume), depicting identical traffic dynamics from various but complementary views. Incomplete values are ubiquitous in these data, which undermines their utility in subsequent applications. In order to manage and enhance traffic data quality, most existing methods recover single traffic variable data independently based on intra-view spatiotemporal correlations, while the inter-view complementarities are ignored. In this paper, we leverage both intra-view and inter-view correlations for multiple traffic variable data simultaneous recovery. To explore the inter-view relationships, a multi-view subspace consistency learning module is developed to bridge connections and activate complementarities among multi-view traffic data. Specifically, the latent subspace features of each data view are extracted and organized as a multi-view subspace tensor with low-rank regularization. The multi-view low-rank tensor captures the consistent subspace structure across multiple data views while reserving unique features within each data view. To characterize the intra-view dependencies, a tensor-based low-rank representation is presented to explore the distinct spatiotemporal patterns within single-view traffic data. For model validation, we additionally design a nonrandom missing pattern to simulate sensor permanent failure cases in practice. Extensive experiments implemented on three real-world multi-view traffic datasets demonstrate the effectiveness and robustness of the proposed model. Yuheng Jia, Yunqing Jia, Chengchuan An, Zhenbo Lu, Jingxin Xia |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Asymmetric Feature Fusion for Image RetrievalabstractIn asymmetric retrieval systems, models with different capacities are deployed on platforms with different computational and storage resources. Despite the great progress, existing approaches still suffer from a dilemma between retrieval efficiency and asymmetric accuracy due to the limited capacity of the lightweight query model. In this work, we propose an Asymmetric Feature Fusion (AFF) paradigm, which advances existing asymmetric retrieval systems by considering the complementarity among different features just at the gallery side. Specifically, it first embeds each gallery image into various features, e.g., local features and global features. Then, a dynamic mixer is introduced to aggregate these features into compact embedding for efficient search. On the query side, only a single lightweight model is deployed for feature extraction. The query model and dynamic mixer are jointly trained by sharing a momentum-updated classifier. Notably, the proposed paradigm boosts the accuracy of asymmetric retrieval without introducing any extra overhead to the query side. Exhaustive experiments on various landmark retrieval datasets demonstrate the superiority of our paradigm. Min Wang 0019, Wengang Zhou 0001, Zhenbo Lu, Houqiang Li |
CVPR | 4 |
| 2023 | Text-Only Training for Visual StorytellingabstractVisual storytelling aims to generate a narrative based on a sequence of images, necessitating both vision-language alignment and coherent story generation. Most existing solutions predominantly depend on paired image-text training data, which can be costly to collect and challenging to scale. To address this, we formulate visual storytelling as a visual-conditioned story generation problem and propose a text-only training method that separates the learning of cross-modality alignment and story generation. Our approach specifically leverages the cross-modality pre-trained CLIP model to integrate visual control into a story generator, trained exclusively on text data. Moreover, we devise a training-free visual condition planner that accounts for the temporal structure of the input image sequence while balancing global and local visual content. The distinctive advantage of requiring only text data for training enables our method to learn from external text story data, enhancing the generalization capability of visual storytelling. We conduct extensive experiments on the VIST benchmark, showcasing the effectiveness of our approach in both in-domain and cross-domain settings. Further evaluations on expression diversity and human assessment underscore the superiority of our method in terms of informativeness and robustness. Yuechen Wang, Wengang Zhou 0001, Zhenbo Lu, Houqiang Li |
ACM Multimedia | 3 |
| 2023 | Hierarchical Multi-Agent Skill DiscoveryabstractSkill discovery has shown significant progress in unsupervised reinforcement learning. This approach enables the discovery of a wide range of skills without any extrinsic reward, which can be effectively combined to tackle complex tasks. However, such unsupervised skill learning has not been well applied to multi-agent reinforcement learning (MARL) due to two primary challenges. One is how to learn skills not only for the individual agents but also for the entire team, and the other is how to coordinate the skills of different agents to accomplish multi-agent tasks. To address these challenges, we present Hierarchical Multi-Agent Skill Discovery (HMASD), a two-level hierarchical algorithm for discovering both team and individual skills in MARL. The high-level policy employs a transformer structure to realize sequential skill assignment, while the low-level policy learns to discover valuable team and individual skills. We evaluate HMASD on sparse reward multi-agent benchmarks, and the results show that HMASD achieves significant performance improvements compared to strong MARL baselines. Mingyu Yang 0003, Yaodong Yang 0001, Zhenbo Lu, Wengang Zhou 0001, Houqiang Li |
NeurIPS | 3 |
| 2023 | Multi-Agent First Order Constrained Optimization in Policy SpaceabstractIn the realm of multi-agent reinforcement learning (MARL), achieving high performance is crucial for a successful multi-agent system.
Meanwhile, the ability to avoid unsafe actions is becoming an urgent and imperative problem to solve for real-life applications.
Whereas, it is still challenging to develop a safety-aware method for multi-agent systems in MARL. In this work, we introduce a novel approach called Multi-Agent First Order Constrained Optimization in Policy Space (MAFOCOPS), which effectively addresses the dual objectives of attaining satisfactory performance and enforcing safety constraints. Using data generated from the current policy, MAFOCOPS first finds the optimal update policy by solving a constrained optimization problem in the nonparameterized policy space. Then, the update policy is projected back into the parametric policy space to achieve a feasible policy. Notably, our method is first-order in nature, ensuring the ease of implementation, and exhibits an approximate upper bound on the worst-case constraint violation. Empirical results show that our approach achieves remarkable performance while satisfying safe constraints on several safe MARL benchmarks. Youpeng Zhao 0001, Yaodong Yang 0001, Zhenbo Lu, Wengang Zhou 0001, Houqiang Li |
NeurIPS | 3 |
| 2023 | Improving Person Re-Identification With Multi-Cue Similarity Embedding and PropagationabstractMost existing person re-identification (Re-ID) methods rely on the visual appearance of the human body. However, face cues are rarely explored in the Re-ID community despite the face that it is an important biometric identifier for human beings. In this work, we propose a Similarity Ensemble Framework (SEF) that uses multi-cue similarity embedding and propagation to effectively fuse body and face information for person re-identification. Specifically, for each query, we first perform standard pedestrian retrieval using body and face cues, respectively, to obtain some candidate results with high confidence. Next, the body and face similarities are combined and embedded into a shared space as node features, and two graphs with the same nodes and different edges with respect to body and face affinities are constructed. Then, the similarity features are propagated in both body and face graphs using graph convolution to capture the relationship among the candidates using different cues. Lastly, the refined features are used to compute the final similarities with the query. The proposed method not only combines the similarities of body and face, but also takes into account the relationship among all the other candidate samples under different cues. Extensive experiments demonstrate that the use of face cues effectively improves the performance of person Re-ID even if the performance obtained by the face alone is much lower than that of the body, suggesting that our approach is able to capture valuable information beyond body from weaker face cues in person Re-ID scenarios. Qiaokang Xie, Zhenbo Lu, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Multim. | 2 |
| 2023 | Weakly Supervised Hashing with Reconstructive Cross-modal AttentionabstractOn many popular social websites, images are usually associated with some meta-data such as textual tags, which involve semantic information relevant to the image and can be used to supervise the representation learning for image retrieval. However, these user-provided tags are usually polluted by noise, therefore the main challenge lies in mining the potential useful information from those noisy tags. Many previous works simply treat different tags equally to generate supervision, which will inevitably distract the network learning. To this end, we propose a new framework, termed as Weakly Supervised Hashing with Reconstructive Cross-modal Attention (WSHRCA), to learn compact visual-semantic representation with more reliable supervision for retrieval task. Specifically, for each image-tag pair, the weak supervision from tags is refined by cross-modal attention, which takes image feature as query to aggregate the most content-relevant tags. Therefore, tags with relevant content will be more prominent while noisy tags will be suppressed, which provides more accurate supervisory information. To improve the effectiveness of hash learning, the image embedding in WSHRCA is reconstructed from hash code, which is further optimized by cross-modal constraint and explicitly improves hash learning. The experiments on two widely-used datasets demonstrate the effectiveness of our proposed method for weakly-supervised image retrieval. The code is available at https://github.com/duyc168/weakly-supervised-hashing . Yongchao Du, Min Wang 0019, Zhenbo Lu, Wengang Zhou 0001, Houqiang Li |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | CMD: Self-supervised 3D Action Representation Learning with Cross-Modal Mutual Distillation
Yunyao Mao, Wengang Zhou 0001, Zhenbo Lu, Jiajun Deng, Houqiang Li |
ECCV (3) | 3 |
| 2022 | UDoc-GAN: Unpaired Document Illumination Correction with Background Light PriorabstractDocument images captured by mobile devices are usually degraded by uncontrollable illumination, which hampers the clarity of document content. Recently, a series of research efforts have been devoted to correcting the uneven document illumination. However, existing methods rarely consider the use of ambient light information, and usually rely on paired samples including degraded and the corrected ground-truth images which are not always accessible. To this end, we propose UDoc-GAN, the first framework to address the problem of document illumination correction under the unpaired setting. Specifically, we first predict the ambient light features of the document. Then, according to the characteristics of different level of ambient lights, we re-formulate the cycle consistency constraint to learn the underlying relationship between normal and abnormal illumination domains. To prove the effectiveness of our approach, we conduct extensive experiments on DocProj dataset under the unpaired setting. Compared with the state-of-the-art approaches, our method demonstrates promising performance in terms of character error rate (CER) and edit distance (ED), together with better qualitative results for textual detail preservation. The source code is now publicly available at \urlhttps://github.com/harrytea/UDoc-GAN. Wengang Zhou 0001, Zhenbo Lu, Houqiang Li |
ACM Multimedia | 3 |
| 2022 | Hidden Mixture Vehicle Discharge State Inference at Signalized Intersection Using Vehicle Travel Time and Discharge Headway DataabstractAccurate and reliable traffic state identification is crucial to developing responsive and proactive traffic management applications. In this study, the problem of vehicle discharge state identification at signalized intersections is investigated, which focuses on the vehicle discharge process during the green interval. Instead of using detailed vehicle trajectory data and treating the observations of vehicles independently, this study formulates the vehicle discharge process in a Hidden Markov Model (HMM) framework using sequential observations of vehicle travel time and discharge headway as inputs. Three vehicle discharge states (i.e., overflow, single stop, and free arrival) are encoded as latent states, and a restricted left-to-right state transition matrix is imposed to respect the nature of the vehicle discharge process in the real world. The standard HMM is further extended to incorporate two informative covariates to parameterize the probabilities of the initial states and state transitions. The proposed models have been validated on the Next Generation Simulation (NGSIM) dataset. Compared to a benchmark model, the proposed models show their strength in correctly inferring the vehicle discharge state and are more reliable to use in presence of random missing observations. The effectiveness of covariate incorporation is also investigated, and several extended applications of the proposed models are discussed. Chengchuan An, Haoliang Shen, Yueru Xu, Zhenbo Lu, Jingxin Xia |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2017 | Video restoration based on a novel second order nonlocal total variation model
Zhenbo Lu, Qing Ling 0001, Houqiang Li, Weiping Li 0003 |
Signal Process. | 1 |
| 2015 | A Bayesian adaptive weighted total generalized variation model for image restorationabstractIn recent years, the Total Generalized Variation (TGV) model has received lots of attention in image processing community. Though this model can restore image with natural intensity transitions, its spatial identical parameter setting limits its performance. In this paper, we propose a novel Adaptive Weighted Total Generalized Variation model for image restoration. We analyze the TGV model from Bayesian Probability view and derive a novel adaptive parameter calculation scheme for it, exploiting the image's self-similarity. Experiment results on image deblurring and reconstruction show that by adapting the parameters in TGV model to image contents, the proposed model can restore image's edges and details well and achieve significant improvement over state of the art variational based models. Zhenbo Lu, Houqiang Li, Weiping Li 0003 |
ICIP | 1 |
| 2015 | Image deblocking via group sparsity optimizationabstractBlock-wise compressed image often suffers from the blocking artifacts. In this paper, we propose a novel deblocking scheme for compressed image, by combining image's sparse property and its self-similarity together, called group sparsity optimization. Instead of processing each image patch individually, in the proposed scheme, similar patches in one group are required to be well-represented on learned dictionary collaboratively, using group sparsity regularization. The group sparsity not only imposes every patch's representation to be sparse, bus also requires patches' coefficients in the group share the similar pattern. The experiment results on standard test images demonstrate that our scheme can improve the PSNR of the compressed images by an average of 1.25 dB, and outperform state of the art deblocking approaches. Zhenbo Lu, Houqiang Li, Weiping Li 0003 |
ISCAS | 1 |
| 2014 | A new non-local video denoising scheme using low-rank representation and total variation regularizationabstractIn this article, we present a novel non-local video denoising scheme using low-rank representation and total variation regularization. The proposed scheme attempts to make full use of the intrinsic properties that the grouping similar patches not only lie in a low-rank subspace but are also sparse in total variation (TV) domain. For a group of similar patches, we formulate video denoising problem into a concise model that combines nuclear norm, TV regularization and l1norm. The experiments demonstrate that the proposed scheme is capable of handling multi-type noise including dense Gaussian noise and random-valued sparse noise, while maintaining the texture information meantime. The results show that our scheme achieves noticeable performance improvement over the state-of-the-art video denoising methods. Qingbo Lu, Zhenbo Lu, Xiaoqing Tao, Houqiang Li |
ISCAS | 2 |
| 2013 | Noise reduction for hyperspectral images based on structural sparse and low-rank matrix decompositionabstractIn this paper, a noise reduction approach for hyperspectral images (HSIs) is presented. Due to the assorted noise sources of HSIs, it seems difficult to describe the noise in a concise manner. Commonly, noise reduction algorithms are dedicated to a certain kind of noise, such as random or striping noise. Most of them in addition have somewhat idealized hypotheses. For example, the random noise is white or signal-independent, or the observed scene is spatially homogeneous or quasi-homogeneous. Thus a practically efficient and universal denoising method is preferred. Thanks to the low-rank characteristic of HSI signal, and the structural sparsity of HSI noise, we draw inspiration from low-rank matrix decomposition and the emerging mixed norm, to propose a method dealing with various patterns of noise simultaneously. Both simulated and real data experiments show the effectiveness of the proposed approach. Zhenbo Lu, Qingbo Lu, Houqiang Li, Weiping Li 0003 |
IGARSS | 2 |
| 2013 | Detection of Blotch and Scratch in Video Based on Video DecompositionabstractIn old video restoration, automatic detection of common defects, e.g., scratches and blotches, has always been emphasized. While prior thoughts mainly focus on detecting blotches and linear, vertical scratches separately, this paper contributes to a more generalized and challenging issue: simultaneous detection of blotches and complex scratches in video, with much less knowledge of them. We investigate the characteristics of blotches and scratches in space and time domain, and propose a novel detection method based on two main steps: cartoon-texture decomposition in the space domain and content-defect separation in the time domain. We then formulate it into convex optimization problems and develop corresponding algorithms. The experiment results demonstrate that the proposed method is of high detection accuracy, verifying the effectiveness of our detection via a video decomposition method. Houqiang Li, Zhenbo Lu, Zhangyang Wang, Qing Ling 0001, Weiping Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |