EDBT 2026 Demo / reviewers in the wild / expert
Bencheng Liao
dblp:289/0295
· DBLP profile ↗
15ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Better early detector for high-performance detection transformer
Bin Hu 0020, Bencheng Liao, Jiyang Qi, Shusheng Yang, Wenyu Liu 0001 |
Image Vis. Comput. | 2 |
| 2025 | ViG: Linear-complexity Visual Sequence Learning with Gated Linear AttentionabstractRecently, linear complexity sequence modeling networks have achieved modeling capabilities similar to Vision Transformers on a variety of computer vision tasks, while using fewer FLOPs and less memory. However, their advantage in terms of actual runtime speed is not significant. To address this issue, we introduce Gated Linear Attention (GLA) for vision, leveraging its superior hardware-awareness and efficiency. We propose direction-wise gating to capture 1D global context through bidirectional modeling and a 2D gating locality injection to adaptively inject 2D local details into 1D global context. Our hardware-aware implementation further merges forward and backward scanning into a single kernel, enhancing parallelism and reducing memory cost and latency. The proposed model, ViG, offers a favorable trade-off in accuracy, parameters, and FLOPs on ImageNet and downstream tasks, outperforming popular Transformer and CNN-based models. Bencheng Liao, Xinggang Wang, Lianghui Zhu, Chang Huang |
AAAI | 1 |
| 2025 | DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous DrivingabstractRecently, the diffusion model has emerged as a powerful generative technique for robotic policy learning, capable of modeling multi-mode action distributions. Leveraging its capability for end-to-end autonomous driving is a promising direction. However, the numerous denoising steps in the robotic diffusion policy and the more dynamic, open-world nature of traffic scenes pose substantial challenges for generating diverse driving actions at a real-time speed. To address these challenges, we propose a novel truncated diffusion policy that incorporates prior multi-mode anchors and truncates the diffusion schedule, enabling the model to learn denoising from anchored Gaussian distribution to the multi-mode driving action distribution. Additionally, we design an efficient cascade diffusion decoder for enhanced interaction with conditional scene context. The proposed model, DiffusionDrive, demonstrates 10× reduction in denoising steps compared to vanilla diffusion policy, delivering superior diversity and quality in just 2 steps. On the planning-oriented NAVSIM dataset, with aligned ResNet-34 backbone, DiffusionDrive achieves 88.1 PDMS without bells and whistles, setting a new record, while running at a real-time speed of 45 FPS on an NVIDIA 4090. Qualitative results on challenging scenarios further confirm that DiffusionDrive can robustly generate diverse plausible driving actions. Bencheng Liao, Shaoyu Chen, Bo Jiang 0011, Cheng Wang 0048, Sixu Yan, Xinbang Zhang, Qian Zhang 0009, Xinggang Wang |
CVPR | 1 |
| 2025 | DiG: Scalable and Efficient Diffusion Models with Gated Linear AttentionabstractDiffusion models with large-scale pre-training have achieved significant success in the field of visual content generation, particularly exemplified by Diffusion Transformers (DiT). However, DiT models have faced challenges with quadratic complexity efficiency, especially when handling long sequences. In this paper, we aim to incorporate the sub-quadratic modeling capability of Gated Linear Attention (GLA) into the 2D diffusion backbone. Specifically, we introduce Diffusion Gated Linear Attention Transformers (DiG), a simple, adoptable solution with minimal parameter overhead. We offer two variants, i,e, a plain and U-shape architecture, showing superior efficiency and competitive effectiveness. In addition to superior performance to DiT and other sub-quadratic-time diffusion models at 256 × 256 resolution, DiG demonstrates greater efficiency than these methods starting from a 512 resolution. Specifically, DiG-S/2 is 2.5× faster and saves 75.7% GPU memory compared to DiT-S/2 at a 1792 resolution. Additionally, DiG-XL/2 is 4.2× faster than the Mamba-based model at a 1024 resolution and 1.8× faster than DiT with FlashAttention-2 at a 2048 resolution. Lianghui Zhu, Bencheng Liao, Jun Hao Liew, Hanshu Yan, Jiashi Feng, Xinggang Wang |
CVPR | 3 |
| 2025 | MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
Yingyue Li, Bencheng Liao, Wenyu Liu 0001, Xinggang Wang |
ICCV | 2 |
| 2025 | RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement LearningabstractExisting end-to-end autonomous driving (AD) algorithms typically follow the Imitation Learning (IL) paradigm, which faces challenges such as causal confusion and an open-loop gap. In this work, we propose RAD, a 3DGS-based closed-loop Reinforcement Learning (RL) framework for end-to-end Autonomous Driving. By leveraging 3DGS techniques, we construct a photorealistic digital replica of the real physical world, enabling the AD policy to extensively explore the state space and learn to handle out-of-distribution scenarios through large-scale trial and error. To enhance safety, we design specialized rewards to guide the policy in effectively responding to safety-critical events and understanding real-world causal relationships. To better align with human driving behavior, we incorporate IL into RL training as a regularization term. We introduce a closed-loop evaluation benchmark consisting of diverse, previously unseen 3DGS environments. Compared to IL-based methods, RAD achieves stronger performance in most closed-loop metrics, particularly exhibiting a 3× lower collision rate. Abundant closed-loop results are presented in the supplementary material. Code is available at https://github.com/hustvl/RAD for facilitating future research. Shaoyu Chen, Bo Jiang 0011, Bencheng Liao, Yiang Shi, Yuechuan Pu, Xinbang Zhang, Wenyu Liu 0001, Qian Zhang 0001, Xinggang Wang |
NeurIPS | 4 |
| 2025 | MapTRv2: An End-to-End Framework for Online Vectorized HD Map Construction
Bencheng Liao, Shaoyu Chen, Yunchi Zhang, Bo Jiang 0011, Qian Zhang 0009, Wenyu Liu 0001, Chang Huang, Xinggang Wang |
Int. J. Comput. Vis. | 1 |
| 2025 | MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
Jialv Zou, Bencheng Liao, Wenyu Liu 0001, Xinggang Wang |
Int. J. Comput. Vis. | 2 |
| 2024 | Lane Graph as Path: Continuity-Preserving Path-Wise Modeling for Online Lane Graph Construction
Bencheng Liao, Shaoyu Chen, Bo Jiang 0011, Tianheng Cheng, Qian Zhang 0009, Wenyu Liu 0001, Chang Huang, Xinggang Wang |
ECCV (44) | 1 |
| 2024 | Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelabstractRecently the state space models (SSMs) with efficient hardware-aware designs, i.e., the Mamba deep learning model, have shown great potential for long sequence modeling. Meanwhile building efficient and generic vision backbones purely upon SSMs is an appealing direction. However, representing visual data is challenging for SSMs due to the position-sensitivity of visual data and the requirement of global context for visual understanding. In this paper, we show that the reliance on self-attention for visual representation learning is not necessary and propose a new generic vision backbone with bidirectional Mamba blocks (Vim), which marks the image sequences with position embeddings and compresses the visual representation with bidirectional state space models. On ImageNet classification, COCO object detection, and ADE20k semantic segmentation tasks, Vim achieves higher performance compared to well-established vision transformers like DeiT, while also demonstrating significantly improved computation & memory efficiency. For example, Vim is 2.8x faster than DeiT and saves 86.8% GPU memory when performing batch inference to extract features on images with a resolution of 1248x1248. The results demonstrate that Vim is capable of overcoming the computation & memory constraints on performing Transformer-style understanding for high-resolution images and it has great potential to be the next-generation backbone for vision foundation models. Lianghui Zhu, Bencheng Liao, Qian Zhang 0009, Wenyu Liu 0001, Xinggang Wang |
ICML | 2 |
| 2024 | Learning accurate monocular 3D voxel representation via bilateral voxel transformer
Tianheng Cheng, Haoyi Jiang, Shaoyu Chen, Bencheng Liao, Qian Zhang 0009, Wenyu Liu 0001, Xinggang Wang |
Image Vis. Comput. | 4 |
| 2023 | VAD: Vectorized Scene Representation for Efficient Autonomous DrivingabstractAutonomous driving requires a comprehensive understanding of the surrounding environment for reliable trajectory planning. Previous works rely on dense rasterized scene representation (e.g., agent occupancy and semantic map) to perform planning, which is computationally intensive and misses the instance-level structure information. In this paper, we propose VAD, an end-to-end vectorized paradigm for autonomous driving, which models the driving scene as a fully vectorized representation. The proposed vectorized paradigm has two significant advantages. On one hand, VAD exploits the vectorized agent motion and map elements as explicit instance-level planning constraints which effectively improves planning safety. On the other hand, VAD runs much faster than previous end-to-end planning methods by getting rid of computation-intensive rasterized representation and hand-designed post-processing steps. VAD achieves state-of-the-art end-to-end planning performance on the nuScenes dataset, outperforming the previous best method by a large margin. Our base model, VAD-Base, greatly reduces the average collision rate by 29.0% and runs 2.5× faster. Besides, a lightweight variant, VAD-Tiny, greatly improves the inference speed (up to 9.3×) while achieving comparable planning performance. We believe the excellent performance and the high efficiency of VAD are critical for the real-world deployment of an autonomous driving system. Code and models are available at https://github.com/hustvl/VAD for facilitating future research. Bo Jiang 0011, Shaoyu Chen, Bencheng Liao, Helong Zhou, Qian Zhang 0009, Wenyu Liu 0001, Chang Huang, Xinggang Wang |
ICCV | 4 |
| 2023 | MapTR: Structured Modeling and Learning for Online Vectorized HD Map Construction
Bencheng Liao, Shaoyu Chen, Xinggang Wang, Tianheng Cheng, Qian Zhang 0009, Wenyu Liu 0001, Chang Huang |
ICLR | 1 |
| 2021 | You Only Look at One Sequence: Rethinking Transformer in Vision through Object DetectionabstractCan Transformer perform $2\mathrm{D}$ object- and region-level recognition from a pure sequence-to-sequence perspective with minimal knowledge about the $2\mathrm{D}$ spatial structure? To answer this question, we present You Only Look at One Sequence (YOLOS), a series of object detection models based on the vanilla Vision Transformer with the fewest possible modifications, region priors, as well as inductive biases of the target task. We find that YOLOS pre-trained on the mid-sized ImageNet-$1k$ dataset only can already achieve quite competitive performance on the challenging COCO object detection benchmark, e.g., YOLOS-Base directly adopted from BERT-Base architecture can obtain $42.0$ box AP on COCO val. We also discuss the impacts as well as limitations of current pre-train schemes and model scaling strategies for Transformer in vision through YOLOS. Code and pre-trained models are available at https://github.com/hustvl/YOLOS. Bencheng Liao, Xinggang Wang, Jiemin Fang, Jiyang Qi, Rui Wu 0018, Jianwei Niu 0004, Wenyu Liu 0001 |
NeurIPS | 2 |
| 2021 | Real-time and accurate object detection in compressed video by long short-term feature aggregation
Xinggang Wang, Zhaojin Huang, Bencheng Liao, Lichao Huang, Yongchao Gong, Chang Huang |
Comput. Vis. Image Underst. | 3 |