VLDB 2026 Research / reviewers in the wild / expert
Lancheng Zou
dblp:321/2579
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
0009-0004-6820-7064ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Analytical FFN-to-MoE Restructuring via Activation Pattern AnalysisabstractZehua Pei, Hui-Ling Zhen, Lancheng Zou, Xianzhi Yu, Wulong Liu, Sinno Jialin Pan, Mingxuan Yuan, Bei Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zehua Pei, Hui-Ling Zhen, Lancheng Zou, Xianzhi Yu, Wulong Liu, Sinno Jialin Pan, Mingxuan Yuan, Bei Yu 0001 |
ACL (1) | 3 |
| 2026 | IncreMacro-3D: Incremental Macro Placement for Face-to-Face Stacked Memory-on-Logic 3D ICsabstractFace-to-face stacked 3D ICs, such as memory-on-logic (MoL) architectures, have emerged as a promising solution to overcome the limitations of traditional 2D integration by offering enhanced performance, power efficiency, and density. Given the increasing design complexity of modern system-on-chips (SoCs), achieving high-quality macro placement is critical, as it plays a decisive role in determining the final performance, power, and area (PPA) metrics. However, existing RTL-to-GDS 3D physical design flows for MoL 3D ICs rely heavily on manual macro placement, which becomes increasingly challenging and time-consuming for modern SoCs with a vast number of macros. In this paper, we introduce an innovative macro placement algorithm, IncreMacro-3D, which employs graph neural network-based macro repartitioning and 3D macro position refinement, thereby facilitating subsequent steps in 3D physical design flow. The experimental results on several benchmark circuits demonstrate that the proposed approach can reduce the routed wirelength, worst negative slack (WNS), total negative slack (TNS), and total power consumption by 6.1%, 44.2%, 62.8%, and 0.6% compared to state-of-the-art analytical placer for MoL 3D ICs. Lancheng Zou, Sing Sen Ye, Yuan Pu 0001, Jiaxi Jiang, Siting Liu 0002, Yuxuan Zhao 0001, Bei Yu 0001 |
DATE | 1 |
| 2026 | Oiso: Outlier-Isolated Data Format for Low-Bit Large Language Model QuantizationabstractThe scale of large language models (LLMs) has steadily increased over time, leading to enhanced performance in multi-modal understanding and complex reasoning, but with significant execution overhead on hardware. Quantization is a promising approach to reduce computation and memory overhead for LLM deployment. However, maintaining accuracy and efficiency simultaneously is challenging due to the presence of outliers. Moreover, low-bit quantization tends to deteriorate accuracy due to its limited precision. Existing outlier-aware quantization/hardware co-design methods split the sparse outliers from the normal values with dedicated encoding schemes. However, such separation produces a non-uniform data format for normal values and outliers, leading to additional hardware design and inefficient memory access. This paper presents an outlier-isolated data format for low-bit LLM quantization called Oiso. Oiso is a unified representation for both outliers and normal values. It isolates the normal values from the outliers, which can reduce the impact of outliers on the normal values during the quantization process. Taking advantage of the uniform format, Oiso arithmetic can be performed using a homogeneous computational unit, and Oiso values can be stored in a standardized format. Hierarchical block encoding with a subblock alignment scheme is introduced to reduce the encoding cost and the hardware overhead. We introduce the Oiso architecture, equipped with Oiso processing elements and encoders tailored for Oiso arithmetic, realizing efficient low-bit LLM inference. Oiso quantization can push the limits of low-bit LLM quantization, and the Oiso accelerator outperforms the state-of-the-art outlieraware accelerator design with 1.26× performance improvement and 25% energy reduction. Lancheng Zou, Mingzi Wang, Wenqian Zhao 0002, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | Invited: Physical Design for Advanced 3D ICs: Challenges and SolutionsabstractAs technology scaling predicted by Moore's law slows down, 3D integrated circuits (3D ICs) have emerged as a promising alternative to enhance performance while maintaining cost-effectiveness. With the advancement of fabrication and bonding technologies, wafer-level 3D integration enables fine-grain 3D interconnects that maximize the benefits in power, performance, and area (PPA). However, a multitude of challenges have obstructed traditional electronic design automation (EDA) methodologies for 3D IC implementations. This paper surveys the major challenges in the physical design of advanced 3D ICs. We provide a comprehensive review of existing solutions, analyzing their advantages and disadvantages in depth. Finally, we discuss open problems and research opportunities in the development of native 3D EDA tools. Yuxuan Zhao 0001, Lancheng Zou, Bei Yu 0001 |
ISPD | 2 |
| 2025 | PermLLM: Learnable Channel Permutation for N: M Sparse Large Language ModelsabstractChannel permutation is a powerful technique for enhancing the accuracy of N:M sparse models by reordering the channels of weight matrices to prioritize the retention of important weights.
However, traditional channel permutation methods rely on handcrafted quality metrics, which often fail to accurately capture the true impact of pruning on model performance.
To address this limitation, we propose PermLLM, a novel post-training pruning framework that introduces learnable channel permutation (LCP) for N:M sparsity.
LCP leverages Sinkhorn normalization to transform discrete permutation matrices into differentiable soft permutation matrices, enabling end-to-end optimization.
Additionally, PermLLM incorporates an efficient block-wise channel permutation strategy, which significantly reduces the number of learnable parameters and computational complexity.
PermLLM seamlessly integrates with existing one-shot pruning methods to adaptively optimize channel permutations, effectively mitigating pruning-induced errors.
Extensive experiments on the LLaMA series, Qwen, and OPT models demonstrate that PermLLM achieves superior performance in optimizing N:M sparse models. Lancheng Zou, Zehua Pei, Tsung-Yi Ho, Farzan Farnia, Bei Yu 0001 |
NeurIPS | 1 |
| 2025 | Lay-Net: Grafting Netlist Knowledge on Layout-Based Congestion PredictionabstractCongestion modeling is crucial for enhancing the routability of VLSI placement solutions. The underutilization of netlist information constrains the efficacy of existing layout-based congestion modeling techniques. We devise a novel approach that grafts netlist-based message passing into a layout-based model, thereby achieving a better knowledge fusion between layout and netlist to improve congestion prediction performance. The innovative heterogeneous message-passing paradigm more effectively incorporates routing demand into the model by considering connections between cells, overlaps of nets, and interactions between cells and nets. Leveraging multi-scale features, the proposed model effectively captures connection information across various ranges, addressing the issue of inadequate global information present in existing models. Using contrastive learning and mini-Gnet techniques allows the model to learn and represent features more effectively, boosting its capabilities and achieving superior performance. Extensive experiments demonstrate a notable performance enhancement of the proposed model compared to existing methods.Our code is available at: https://github.com/lanchengzou/congPred. Lancheng Zou, Su Zheng, Peng Xu 0052, Siting Liu 0002, Bei Yu 0001, Martin D. F. Wong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | HAPE: Hardware-Aware LLM Pruning For Efficient On-Device Inference OptimizationabstractOver the past few years, large language models (LLMs) have demonstrated remarkable performance and versatility across a variety of complex tasks. However, their deployment has been challenged by their substantial model size and computational requirements. Pruning is a effective approach to make the model parameters sparse, thereby acquire inference acceleration. While not everyone requires training or fine-tuning large models, the diverse range of applications necessitates the deployment of LLMs on different devices. Model pruning and compression have emerged as areas of deep research interest to address these challenges. In consideration of versatility and practicality, we have designed a hardware-aware pruning process for general-purpose hardware/edge devices to enable efficient deployment and inference of LLMs. Instead of considering sparse ratio alone, we are motivated to design a pruning framework that incorporates genuine inference speed-up sensitivity from each pruning structure. Moreover, our framework breaks the layer-by-layer pruning setting and fuse several layers into one pruning stage to allow cross-layer optimization. Apart from that, we hold pragmatism by conducting compilation optimization during pruning. This step is critical because most sparsity patterns barely show distinct speed acceleration with corresponding dataflow and memory optimization. Our process operates within a post-training framework, obviating the need for additional training and thereby reducing resource requirements, while ensuring diverse inference speed and accuracy requirements on hardware. Wenqian Zhao 0002, Lancheng Zou, Zixiao Wang 0001, Xufeng Yao, Bei Yu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2024 | PDRC: Package Design Rule Checking via GPU-Accelerated Geometric Intersection Algorithms for Non-Manhattan GeometryabstractWith the emergence of chiplet technology, the scale of IC packaging design has been steadily increasing, making the utilization of traditional design rule checking (DRC) methods more time-consuming. In this paper, we propose PDRC, a package-level design rule checker for non-manhattan geometry with GPU acceleration. PDRC employs hierarchical interval lists within an iterative parallel sweepline framework to implement the geometric intersection algorithm, thereby finishing design rule checking tasks. Experimental results have demonstrated 30 - 50 times speedup achieved by PDRC compared with two CPU-based checkers. Jiaxi Jiang, Lancheng Zou, Wenqian Zhao 0002, Zhuolun He, Tinghuan Chen, Bei Yu 0001 |
DAC | 2 |
| 2024 | BiE: Bi-Exponent Block Floating-Point for Large Language Models QuantizationabstractNowadays, Large Language Models (LLMs) mostly possess billions of parameters, bringing significant challenges to hardware platforms. Although quantization is an efficient approach to reduce computation and memory overhead for inference optimization, we stress the challenge that mainstream low-bit quantization approaches still suffer from either various data distribution outliers or a lack of hardware efficiency. We also find that low-bit data format has further potential expressiveness to cover the atypical language data distribution. In this paper, we propose a novel numerical representation, Bi-Exponent Block Floating Point (BiE), and a new quantization flow. BiE quantization shows accuracy superiority and hardware friendliness on various models and benchmarks. Lancheng Zou, Wenqian Zhao 0002, Qi Sun 0002, Bei Yu 0001 |
ICML | 1 |
| 2023 | Mitigating Distribution Shift for Congestion Optimization in Global PlacementabstractThe placement and routing (PnR) flow plays a critical role in physical design. Poor routing congestion is a possible problem causing severe routing detours, which can lead to deteriorated timing performance or even routing failure. Deep-learning-based congestion prediction model is designed to guide the global placement process in previous work. However, the distribution shift problem in this method limits its performance. In this paper, we mitigate the distribution shift problem with a look-ahead mechanism inspired by optical flow prediction and an invariant feature space learning technique. With the proposed method, we can achieve better congestion prediction performance and less-congested placement results. Su Zheng, Lancheng Zou, Siting Liu 0002, Yibo Lin, Bei Yu 0001, Martin D. F. Wong |
DAC | 2 |
| 2023 | Lay-Net: Grafting Netlist Knowledge on Layout-Based Congestion PredictionabstractCongestion modeling is a key point for improving the routability of VLSI placement solutions. The underuti-lization of netlist information limits the performance of ex-isting layout-based congestion modeling methods. Combining the knowledge from netlist and layout, we graft netlist-based message passing on a layout-based model to achieve better congestion prediction performance. The novel heterogeneous message-passing paradigm better embeds the routing demand into the model by considering both connections between cells and overlaps of nets. With the help of multi-scale features, the proposed model can effectively capture connection information across different ranges, overcoming the problem of insufficient global information in existing models. Based on the advancements, the proposed model achieves significant improvement compared with existing methods. Su Zheng, Lancheng Zou, Peng Xu 0052, Siting Liu 0002, Bei Yu 0001, Martin D. F. Wong |
ICCAD | 2 |