Zhenxin Li

dblp:209/9403 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DriveSuprim: Towards Precise Trajectory Selection for End-to-End Planning
abstract
Autonomous vehicles must navigate safely in complex driving environments. Imitating a single expert trajectory, as in regression-based approaches, usually does not explicitly assess the safety of the predicted trajectory. Selection-based methods address this by generating and scoring multiple trajectory candidates and predicting the safety score for each. However, they face optimization challenges in precisely selecting the best option from thousands of candidates and distinguishing subtle but safety-critical differences, especially in rare and challenging scenarios. We propose DriveSuprim to overcome these challenges and advance the selection-based paradigm through a coarse-to-fine paradigm for progressive candidate filtering, a rotation-based augmentation method to improve robustness in out-of-distribution scenarios, and a self-distillation framework to stabilize training. DriveSuprim achieves state-of-the-art performance, reaching 93.5% PDMS in NAVSIM v1 and 87.1% EPDMS in NAVSIM v2 without extra data, with 83.02 Driving Score and 60.00 Success Rate on Bench2Drive, demonstrating superior planning capabilities in various driving scenarios.
Wenhao Yao, Zhenxin Li, Shiyi Lan, Xinglong Sun, José M. Álvarez 0004, Zuxuan Wu
AAAI2
2026 Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
Zihan Chang, Sheng Xiao, Shuibing He, Xuechen Zhang 0001, Siling Yang, Zhenxin Li, Weijian Chen 0002
Euro-Par (2)6
2025 Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop Training
Zhenxin Li, Shiyi Lan, Zhiding Yu, Zuxuan Wu, José M. Álvarez 0004
ICCV1
2025 Enhancing Autonomous Driving Safety with Collision Scenario Integration
abstract
Autonomous vehicle safety is crucial for the successful deployment of self-driving cars. However, most existing planning methods rely heavily on imitation learning, which limits their ability to leverage collision data effectively. Moreover, collecting collision or near-collision data is inherently challenging, as it involves risks and raises ethical and practical concerns. In this paper, we propose SafeFusion, a training framework to learn from collision data. Instead of over-relying on imitation learning, SafeFusion integrates safety-oriented metrics during training to enable collision avoidance learning. In addition, to address the scarcity of collision data, we propose CollisionGen, a scalable data generation pipeline to generate diverse, high-quality scenarios using natural language prompts, generative models, and rule-based filtering. Experimental results show that our approach improves planning performance in collision-prone scenarios by 56% over previous state-of-the-art planners while maintaining effectiveness in regular driving situations. Our work provides a scalable and effective solution for advancing the safety of autonomous driving systems.
Shiyi Lan, Xinglong Sun, Nadine Chang, Zhenxin Li, Zhiding Yu, José M. Álvarez 0004
IROS5
2025 Fighting Malicious Media Data: A Survey on Tampering Detection and Deepfake Detection
abstract
Online media data, in the form of images and videos, are becoming mainstream communication channels. However, recent advances in deep learning (DL), particularly deep generative models, open the doors for producing perceptually convincing images and videos at a low cost, which not only poses a serious threat to the trustworthiness of digital information but also has severe societal implications. This motivates a growing interest in research in media tampering detection (TD), i.e., using DL techniques to examine whether media data have been maliciously manipulated. Depending on the content of the targeted images, media forgery could be divided into image tampering and Deepfake techniques. The former typically moves or erases the visual elements in ordinary images, while the latter manipulates the expressions and even the identity of human faces. Accordingly, the means of defense include image TD and Deepfake detection (DFD), which share a wide variety of properties. In this article, we provide a comprehensive review of the current media TD approaches and discuss the challenges and trends in this field for future research.
Zhenxin Li, Chao Zhang 0001, Jingjing Chen 0001, Zuxuan Wu, Larry Davis 0001, Yu-Gang Jiang 0001
Proc. IEEE2
2024 BEVNeXt: Reviving Dense BEV Frameworks for 3D Object Detection
abstract
Recently, the rise of query-based Transformer decoders is reshaping camera-based 3D object detection. These query-based decoders are surpassing the traditional dense BEV (Bird's Eye View)-based methods. However, we argue that dense BEV frameworks remain important due to their out-standing abilities in depth estimation and object localization, depicting 3D scenes accurately and comprehensively. This paper aims to address the drawbacks of the existing dense BEV-based 3D object detectors by introducing our proposed enhanced components, including a CRF-modulated depth estimation module enforcing object-level consistencies, a long-term temporal aggregation module with extended receptive fields, and a two-stage object decoder combining perspective techniques with CRF-modulated depth embedding. These enhancements lead to a “modernized” dense BEV framework dubbed BEVNeXt. On the nuScenes benchmark, BEVNeXt outperforms both BEV-based and query-based frameworks under various settings, achieving a state-of-the-art result of 64.2 NDS on the nuScenes test set.
Zhenxin Li, Shiyi Lan, José M. Álvarez 0004, Zuxuan Wu
CVPR1
2024 CCL-BTree: A Crash-Consistent Locality-Aware B+-Tree for Reducing XPBuffer-Induced Write Amplification in Persistent Memory
abstract
In persistent B+ -Tree, random updates of small key-value (KV) pairs will cause severe XPBuffer-induced write amplification (XBI-amplification) because CPU cacheline size is smaller than media access granularity in persistent memory (PM). We observe that XBI-amplification directly determines the application performance when the PM bandwidth is exhausted in multi-thread scenarios. However, none of the existing work can efficiently address the XBI-amplification issue while maintaining superior range query performance.
Zhenxin Li, Shuibing He, Zheng Dang, Peiyi Hong, Xuechen Zhang 0001, Rui Wang 0076, Fei Wu 0001
EuroSys1
2024 Efficient Large Graph Processing with Chunk-Based Graph Representation Model
Rui Wang 0076, Weixu Zong, Shuibing He, Zhenxin Li, Zheng Dang
USENIX ATC5
2024 PMAlloc: A Holistic Approach to Improving Persistent Memory Allocation
abstract
Persistent memory allocation is a fundamental building block for developing high-performance and in-memory applications. Existing persistent memory allocators suffer from many performance issues. First, they may introduce repeated cache line flushes and small random accesses in persistent memory for their poor heap metadata management. Second, they use static slab segregation resulting in a dramatic increase in memory consumption when allocation request size is changed. Third, they are not aware of NUMA effect, leading to remote persistent memory accesses in memory allocation and deallocation processes. In this article, we design a novel allocator, named PMAlloc, to solve the above issues simultaneously. (1) PMAlloc eliminates cache line reflushes by mapping contiguous data blocks in slabs to interleaved metadata entries stored in different cache lines. (2) It writes small metadata units to a persistent bookkeeping log in a sequential pattern to remove random heap metadata accesses in persistent memory. (3) Instead of using static slab segregation, it supports slab morphing, which allows slabs to be transformed between size classes to significantly improve slab usage. (4) It uses a local-first allocation policy to avoid allocating remote memory blocks. And it supports a two-phase deallocation mechanism including recording and synchronization to minimize the number of remote memory access in the deallocation. PMAlloc is complementary to the existing consistency models. Results on six benchmarks demonstrate that PMAlloc improves the performance of state-of-the-art persistent memory allocators by up to 6.4× and 57× for small and large allocations, respectively. PMAlloc with NUMA optimizations brings a 2.9× speedup in multi-socket evaluation and is up to 36× faster than other persistent memory allocators. Using PMAlloc reduces memory usage by up to 57.8%. Besides, we integrate PMAlloc in a persistent FPTree. Compared to the state-of-the-art allocators, PMAlloc improves the performance of this application by up to 3.1×.
Zheng Dang, Shuibing He, Xuechen Zhang 0001, Peiyi Hong, Zhenxin Li, Haozhe Song, Xian-He Sun, Gang Chen 0001
ACM Trans. Comput. Syst.5
2023 A New Pareto Discrete NSGAII Algorithm for Disassembly Line Balance Problem
abstract
With the increasing variety and quantity of end‐of‐life (EOL) products, the traditional disassembly process has become inefficient. In response to this phenomenon, this article proposes a random multiproduct U‐shaped mixed‐flow incomplete disassembly line balancing problem (MUPDLBP). MUPDLBP introduces a mixed disassembly method for multiple products and incomplete disassembly method into the traditional DLBP, while considering the characteristics of U‐shaped disassembly lines and the uncertainty of the disassembly process. First, mixed‐flow disassembly can improve the efficiency of disassembly lines, reducing factory construction and maintenance costs. Second, by utilizing the characteristics of incomplete disassembly to reduce the number of dismantled components and the flexibility and efficiency of U‐shaped disassembly lines in allocating disassembly tasks, further improvement in disassembly efficiency can be achieved. In addition, this paper also addresses the characteristics of EOL products with heavy weight and high rigidity. While retaining the basic settings of MUPDLBP, the stability of the assembly during the disassembly process is considered, and a new problem called MUPDLBP_S, which takes into account the disassembly stability, is further proposed. The corresponding mathematical model is provided. To obtain high‐quality disassembly plans, a new and improved algorithm called INSGAII is proposed. The INSGAII algorithm uses the initialization method based on Monte Carlo tree simulation (MCTI) and the Group Global Crowd Degree Comparison (GCDC) operator to replace the initialization method and crowding distance comparison operator in the NSGAII algorithm, effectively improving the coverage of the initial population individuals in the entire solution space and the evenness and spread of the Pareto front. Finally, INSGAII’s effectiveness has been affirmed by tackling both current disassembly line balancing problems and the proposed MUPDLBP and MUPDLBP_S. Importantly, INSGAII outshines six comparison algorithms with a top rank of 1 in the Friedman test, highlighting its superior performance.
Zhenxin Li, Yixin Zou
Int. J. Intell. Syst.3
2022 NVAlloc: rethinking heap metadata management in persistent memory allocators
abstract
Persistent memory allocation is a fundamental building block for developing high-performance and in-memory applications. Existing persistent memory allocators suffer from suboptimal heap organizations that introduce repeated cache line flushes and small random accesses in persistent memory. Worse, many allocators use static slab segregation resulting in a dramatic increase in memory consumption when allocation request size is changed. In this paper, we design a novel allocator, named NVAlloc, to solve the above issues simultaneously. First, NVAlloc eliminates cache line reflushes by mapping contiguous data blocks in slabs to interleaved metadata entries stored in different cache lines. Second, it writes small metadata units to a persistent bookkeeping log in a sequential pattern to remove random heap metadata accesses in persistent memory. Third, instead of using static slab segregation, it supports slab morphing, which allows slabs to be transformed between size classes to significantly improve slab usage. NVAlloc is complementary to the existing consistency models. Results on 6 benchmarks demonstrate that NVAlloc improves the performance of state-of-the-art persistent memory allocators by up to 6.4x and 57x for small and large allocations, respectively. Using NVAlloc reduces memory usage by up to 57.8%. Besides, we integrate NVAlloc in a persistent FPTree. Compared to the state-of-the-art allocators, NVAlloc improves the performance of this application by up to 3.1x.
Zheng Dang, Shuibing He, Peiyi Hong, Zhenxin Li, Xuechen Zhang 0001, Xian-He Sun, Gang Chen 0001
ASPLOS4
2022 PhaST: Hierarchical Concurrent Log-Free Skip List for Persistent Memory
abstract
Skip list (skiplist) is a competitive index structure that offers superior concurrency and excellent performance but with high memory overhead and low access locality. Emerging persistent memory (PM) technologies present an opportunity to mitigate the capacity constraint of DRAM. However, data consistency on PM typically results in excessive write overhead. In addition, fast concurrent access to an index is critical to the throughput on high-end contemporary computer systems. In this article, we propose a Partitioned HierArchical SkiplisT calledPhaST, which can simultaneously reduce the skiplist height and improve its access locality, through its hierarchy of component structures, while enabling fast parallel recovery in case of failure. To ensure high concurrency and fast data consistency, we also have developed writelock-free concurrent insert and log-free atomic split. Furthermore, we have developed a durable lock-free concurrent search that can discern transient structural inconsistencies and deliver highly concurrent read operations. We have conducted an extensive evaluation ofPhaSTcompared to state-of-the-art studies such as NV-Skiplist, wB+-Tree, FPTree, and FAST-FAIR. Our evaluation results showPhaSToutperforms other indexing structures by up to 4.05× and 2.87× in single-threaded inserts and searches, and 1.56× and 2.62× in concurrent inserts and searches.
Zhenxin Li, Bing Jiao, Shuibing He, Weikuan Yu
IEEE Trans. Parallel Distributed Syst.1