VLDB 2026 Research / reviewers in the wild / expert
Zhixiong Di
dblp:165/4598
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0001-7323-5052ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LEGALM 2.0: A Versatile Augmented Lagrangian Method-Based Methodology for Mixed-Cell-Height LegalizationabstractLegalization is a crucial step in VLSI physical design, ensuring design rule compliance while minimizing disruptions to global placement. With the rise of multi-row-height cells in advanced nodes, mixed-cell-height legalization poses significant challenges due to complex cell shapes and design constraints. In this work, we present LEGALM 2.0, a versatile legalization methodology that efficiently handles routability and hybrid region constraints. Our approach introduces a linearized augmented Lagrangian formulation, a scanline-based initial legalization algorithm, and a connectivity-based local optimum escape strategy to enhance convergence. Additionally, we propose a block gradient descent method and a GPU-optimized triplefold partitioning strategy for improved parallelism. Experimental results show that LEGALM 2.0 outperforms state-of-the-art legalizers, achieving 6-36% better quality scores on ICCAD2017 benchmarks and 1.61-4.30W speedup on large-scale designs. For hybrid region constraints, it reduces displacement by 25% and wirelength perturbation by 21%, demonstrating its effectiveness in modern physical design. Jing Mai, Chunyuan Zhao, Zuodong Zhang, Zhixiong Di, Runsheng Wang, Yibo Lin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | LEGALM: Efficient Legalization for Mixed-Cell-Height Circuits with Linearized Augmented Lagrangian MethodabstractAdvanced technologies increasingly adopt mixed-cell-height circuits due to their superior power efficiency, compact area usage, enhanced routability, and improved performance. However, the complex constraints of modern circuit design, including routing challenges and fence region constraints, increase the difficulty of mixed-cell-height legalization. In this paper, we introduce LEGALM, a state-of-the-art mixed-cell-height legalizer that can address routability and fence region constraints more efficiently. We propose an augmented Lagrangian formulation coupled with a block gradient descent method that offers a novel analytical perspective on the mixed-cell-height legalization problem. To further enhance efficiency, we develop a series of GPU-accelerated kernels and a triplefold partitioning technique with minor quality overhead. Experimental results on ICCAD-2017 and modified ISPD-2015 benchmarks show that our approach significantly outperforms current state-of-the-art legalization algorithms in both quality and efficiency. Jing Mai, Chunyuan Zhao, Zuodong Zhang, Zhixiong Di, Yibo Lin, Runsheng Wang, Ru Huang 0001 |
ISPD | 4 |
| 2025 | An r-DFA-Based Layout Pattern Match Method Supporting Fuzzy MatchingabstractAs chip manufacturing approaches physical limits, the probability of defects due to specific chip layout structures has significantly increased. These defect-prone structures are known as lithographic hotspots. Pattern matching method is widely used in hotspot detection algorithms due to its efficiency and accuracy. However, traditional pattern matching algorithms face major challenges in both solution efficiency and flexibility for fuzzy matching. To overcome these limitations, an integer range-based deterministic finite automaton (r-DFA)-based layout pattern matching method supporting parallelization and fuzzy matching is proposed. Manhattan polygons in the layout are represented as multiple path strings, thereby transforming the pattern matching problem into a string regular expression search problem. To simplifies the construction of large integer range elements in fuzzy matching, the r-DFA is employed, enhancing construction efficiency and enabling the algorithm to achieve linear time complexity. Moreover, this approach focuses most of the matching process within each individual layout polygon, enabling parallelized matching across diverse layout polygons. Compared to the state-of-the-art, our approach supports fuzzy matching, and shows an average efficiency improvement of 1.23 times. Qianxi Chen, Yujiao Deng, Qiang Wu 0014, Zhixiong Di |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | A Robust FPGA Router With Optimization of High-Fanout Nets and Intra-CLB ConnectionsabstractRouting is the most time-consuming step in the implementation flow of field programmable gate array (FPGA) designs. With the advance in transistor scaling and system integration, hardware resources in FPGA devices are growing in a larger quantity and diversity. The routing architecture is designed to be more complicated for mapping RTL designs to FPGA devices correctly, which brings significant challenges for current FPGA routing algorithms. The key challenges for routing algorithms lie in large solution space and heavy congestion, especially for high-fanout nets (HFNets). We propose a partition-based algorithm to accelerate the routing of HFNets, which decomposes the global routing guide to shrink the search space for connecting each sink. Meanwhile, the congestions existing inside configurable logic block (CLB) are hard to handle by traditional sequential negotiation-based algorithms, because the industrial routing architecture is quite complex. We propose a concurrent intra-CLB rerouting algorithm to effectively resolve routing congestion inside a CLB tile induced by connections between intra-CLB logic pins, e.g., logic elements and switch boxes. Experimental results on modified ISPD2016 benchmarks demonstrate that our framework can achieve 100% routability in 9.8% less wirelength and$11\times $less runtime, while the state-of-the-art VTR 8.0 routing algorithm fails at 7 of 12 benchmarks. Xun Jiang 0002, Jing Mai, Zhixiong Di, Yibo Lin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | Efficient Schemes and Architectures for Check Node Update in Shuffled Min-Sum Decoding of LDPC CodesabstractBy dividing Variable Nodes (VNs) into groups and processing groups sequentially, the Shuffled Min-Sum (SMS) decoding of Low-Density Parity-Check (LDPC) codes achieves a good trade-off between hardware resource and throughput. However, under the grouping strategy where at least one Check Node (CN) has multiple VN neighbors in one group, the CN update of State-Of-The-Art (SOTA) SMS decoding has long latency and high complexity. To address the issue, this paper presents two efficient CN update schemes. The first one uses the first two minimum Variable-to-Check (V2C) message magnitudes in each group and a size-λ Monotone Double-ended Queue (λMDQ), leading to an Improved λMDQ (I-λMDQ) scheme. The second one called the Modified λMDQ (M-λMDQ) scheme uses only the minimum in each group. As a result, our schemes reduce the number of Selection Modules (SMs) required by the CN update to only one. Simulation results show that both of our schemes have comparable error-correction performance to the SOTA schemes. Furthermore, we simplify the SM via a decomposition design and also present the detailed hardware architectures for our schemes. Analysis using the 90nm CMOS technology library shows that, compared to the SOTA schemes, our I-λMDQ scheme reduces the area and latency by up to 75% and 72%, respectively, while improving the throughput by up to 254%. Similarly, our M-λMDQ scheme reduces the area and latency by up to 79% and 70%, respectively, and improves the throughput by up to 229%. Lisha Luo, Qin Du, Kaining Han, Zhixiong Di, Xiaohu Tang 0004 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | A Remote FPGA-based Experimental Teaching System Design Supporting Single-board Multi-user and Multi-board Single-user Operations in MOOCsabstractUnlike traditional offline courses, where the enrollment number is generally fixed, massive open online courses (MOOCs) often exhibit significant disparities in the enrollment number across different sections. Namely, this number can vary by several or even hundreds of times depending on the MOOC sections. To this end, this study proposes a remote FPGA-based experimental teaching system with two main innovations. First, a software-hardware co-work framework is designed to divide a single physical FPGA into multiple independently virtual FPGAs, allowing for multiple users to use the same physical FPGA concurrently. Second, an "X86 CPU+Multi-PYNQ" collaborative computing system supporting up to 16 PYNQs for parallel computing is designed. This system uses an X86 CPU as a main processor to distribute computing tasks to the PYNQ-cluster for scheduling FPGA parallel computations, which enables a single user to use multiple FPGAs concurrently. Therefore, the proposed system can effectively address the problem of underutilized boards in MOOCs where the enrollment number is smaller than the number of available FPGA boards. In summary, the system proposed in this paper can effectively mitigate the conflict between the number of students and the number of FPGA boards in MOOCs. Zhixiong Di, Xufeng Wei, Yiduo Chen, Shuanglong Wu, Peihao Sun, Qiang Wu 0014 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | Multielectrostatic FPGA Placement Considering SLICEL-SLICEM Heterogeneity, Clock Feasibility, and Timing OptimizationabstractWhen modern FPGA architecture becomes increasingly complicated, modern FPGA placement is a mixed optimization problem with multiple objectives, including wirelength, routability, timing closure, and clock feasibility. Typical FPGA devices nowadays consist of heterogeneous SLICEs like SLICEL and SLICEM. The resources of a SLICE can be configured to {LUT, FF, distributed RAM, SHIFT, CARRY}. Besides such heterogeneity, advanced FPGA architectures also bring complicated constraints like timing, clock routing, carry chain alignment, etc. The above heterogeneity and constraints impose increasing challenges to FPGA placement algorithms. In this work, we propose a multielectrostatic FPGA placer considering the aforementioned SLICEL–SLICEM heterogeneity under timing, clock routing and carry chain alignment constraints. We first propose an effective SLICEL–SLICEM heterogeneity model with a novel electrostatic-based density formulation. We also design a dynamically adjusted preconditioning and carry chain alignment technique to stabilize the optimization convergence. We then propose a timing-driven net weighting scheme to incorporate timing optimization. Finally, we put forward a nested Lagrangian relaxation-based placement framework to incorporate the optimization objectives of wirelength, routability, timing, and clock feasibility. Experimental results on both academic and industrial benchmarks demonstrate that our placer outperforms the state-of-the-art placers in quality and efficiency. Jing Mai, Zhixiong Di, Yibo Lin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | LEAPS: Topological-Layout-Adaptable Multi-Die FPGA Placement for Super Long Line MinimizationabstractMulti-die FPGAs are crucial components in modern computing systems, particularly for high-performance applications such as artificial intelligence and data centers. Super long lines (SLLs) provide interconnections between super logic regions (SLRs) for a multi-die FPGA on a silicon interposer. They have significantly higher delay compared to regular interconnects, which need to be minimized. With the increase in design complexity, the growth of SLLs gives rise to challenges in timing and power closure. Existing placement algorithms focus on optimizing the number of SLLs but often face limitations due to specific topologies of SLRs. Furthermore, they fall short of achieving continuous optimization of SLLs throughout the entire placement process. This highlights the necessity for more advanced and adaptable solutions. In this paper, we propose LEAPS, a comprehensive, systematic, and adaptable multi-die FPGA placement algorithm for SLL minimization. Our contributions are threefold: 1) proposing a high-performance global placement algorithm for multi-die FPGAs that optimizes the number of SLLs while addressing other essential design constraints such as wirelength, routability, and clock routing; 2) introducing a versatile method for more complex SLR topologies of multi-die FPGAs, surpassing the limitations of existing approaches; and 3) executing continuous optimization of SLL counts across the whole placement stages, including global placement (GP), legalization (LG), and detailed placement (DP). Experimental results demonstrate the effectiveness of LEAPS in reducing SLLs and enhancing circuit performance. Compared with the most recent state-of-the-art (SOTA) method, LEAPS achieves an average reduction of 43.08% in SLL counts and 9.99% in HPWL while exhibiting a notable 34.34$\times$improvement in runtime. Zhixiong Di, Runzhe Tao, Jing Mai, Yibo Lin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | A Robust FPGA Router with Concurrent Intra-CLB ReroutingabstractRouting is the most time-consuming step in the FPGA design flow with increasingly complicated FPGA architectures and design scales. The growing complexity of connections between logic pins inside CLBs of FPGAs challenges the efficiency and quality of FPGA routers. Existing negotiation-based rip-up and reroute schemes will result in a large number of iterations when generating paths inside CLBs. In this work, we propose a robust routing framework for FPGAs with complex connections between logic elements and switch boxes. We propose a concurrent intra-CLB rerouting algorithm that can effectively resolve routing congestion inside a CLB tile. Experimental results on modified ISPD 2016 benchmarks demonstrate that our framework can achieve 100% routability in less wirelength and runtime, while the state-of-the-art VTR 8.0 routing algorithm fails at 4 of 12 benchmarks. Jing Mai, Zhixiong Di, Yibo Lin |
ASP-DAC | 3 |
| 2023 | A High Precision CV Control Scheme for Low Power AC-DC BUCK Converter ControllerabstractThis article presents a constant voltage control scheme to improve the transfer efficiency and output voltage accuracy for single-stage AC-DC converters. The proposed scheme also has the advantage of low cost because of the simple application and few peripheral devices. To realize above features, the peak current and frequency curves with high efficiency are selected by analyzing the power losses of the applied topology. A voltage self-compensation circuit is designed to improve the output accuracy, especially the output voltage below 5V. Furthermore, a multi-stage startup circuit is also proposed to reduce the output voltage overshoot. To verify the feasibility of the proposed constant voltage control scheme, a BUCK controller adopting the proposed control scheme has been designed and fabricated with$0.18~\mu \text{m}$BCD process. Using the proposed controller, the peak efficiency of the prototype can reach as high as 66.9%, 78.9% and 82.8% under 115Vac input and 3.3V@300mA, 5V@300mA and 12V@300mA respectively. The load regulation is counted to be within ±1%. Qiang Wu 0014, Linjun Wu, Yongyuan Li, Zhixiong Di, Shubin Liu 0001, Zhangming Zhu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | Multi-electrostatic FPGA placement considering SLICEL-SLICEM heterogeneity and clock feasibilityabstractModern field-programmable gate arrays (FPGAs) contain heterogeneous resources, including CLB, DSP, BRAM, IO, etc. A Configurable Logic Block (CLB) slice is further categorized to SLICEL and SLICEM, which can be configured as specific combinations of instances in {LUT, FF, distributed RAM, SHIFT, CARRY}. Such kind of heterogeneity challenges the existing FPGA placement algorithms. Meanwhile, limited clock routing resources also lead to complicated clock constraints, causing difficulties in achieving clock feasible placement solutions. In this work, we propose a heterogeneous FPGA placement framework considering SLICEL-SLICEM heterogeneity and clock feasibility based on a multi-electrostatic formulation. We support a comprehensive set of the aforementioned instance types with a uniform algorithm for wirelength, routability, and clock optimization. Experimental results on both academic and industrial benchmarks demonstrate that we outperform the state-of-the-art placers in both quality and efficiency. Jing Mai, Yibai Meng, Zhixiong Di, Yibo Lin |
DAC | 3 |
| 2022 | Learned Compression Framework With Pyramidal Features and Quality Enhancement for SAR ImagesabstractCurrent image compression algorithms based on transforms can achieve ideal performance for natural images, but do not do well with synthetic aperture radar (SAR) images. We propose a learned compression framework with pyramidal features and quality enhancement to fully exploit the redundancy among image pixels and to improve the compression bitrate and reconstruction quality. Based on the variational autoencoder (VAE) architecture, pyramidal decomposition is performed at the first autoencoder to extract both global and coarse feature maps. The latent distribution is modeled by the second hyperprior autoencoder with a single-Gaussian model for more accurate and flexible entropy estimation. Universal quantization is applied to consolidate the entropy estimation accuracy of the hyperprior network. To further improve reconstruction quality, a residual dense network (RDN) is adopted to fully capture local and global features. Experimental results demonstrate that the proposed framework provides a better rate-distortion tradeoff than standard codecs such as JPEG, JPEG2000, and learning-based methods on both the Sandia and ICEYE datasets. Zhixiong Di, Qiang Wu 0014, Jiangyi Shi, Quanyuan Feng, Yibo Fan |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Synthetic Aperture Radar Image Compression Based on a Variational AutoencoderabstractGiven the uniqueness of synthetic aperture radar (SAR) images, traditional optical image compression algorithms cannot fully exploit their redundant information. To improve SAR image compression in terms of rate–distortion performance and visual perception, an end-to-end SAR image compression convolutional neural network (CNN) model based on a variational autoencoder is proposed. The proposed CNN model consists of a main autoencoder and a hyper autoencoder. To reduce dependencies in latent space, a joint transform of linear CNN and nonlinear generalized divisive normalization (GDN) activation is applied in the main autoencoder. Moreover, residual blocks are combined with the transforms to boost the efficiency of feature learning and make use of subpixels to improve the quality of reconstructed images. Instead of a fixed entropy model, a conditioned entropy model that works with a hyperprior network is used to learn the distribution of latents, which helps to further improve the compression quality. During training, the model is optimized by evaluating the rate–distortion performance. The experimental results show that the proposed method can achieve better distortion performance than JPEG, JPEG2000, and the available CNN-based method in terms of objective evaluation criteria and human vision perception quality. Qihan Xu, Yunfan Xiang, Zhixiong Di, Yibo Fan, Quanyuan Feng, Qiang Wu 0014, Jiangyi Shi |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | NBLG: A Robust Legalizer for Mixed-Cell-Height Modern DesignabstractWith the increasing complexity of modern design, mixed-cell-height designs have become more popular, which makes the legalization problem more challenging. In this article, a robust negotiation-based legalizer (NBLG) is proposed to reduce average displacement and maximum displacement for mixed-cell-height circuits with considering the fence region and technology constraints. By dissecting the main components of the negotiation-based method, we divide the placement grid in terms of placement sites and reformulate the legalization problem as a resource allocation task. We then allow all movable cells to gradually remove overlaps in a surrounding window with two individual techniques: 1) isolation point and 2) adaptive penalty function. We also adopt a deterministic multithreading technique to accelerate the convergence of our algorithm. Experimental results show that our legalizer achieved the minimal average displacement and maximum displacement in a reasonable runtime compared with state-of-the-art methods. Jinwei Chen 0005, Zhixiong Di, Jiangyi Shi, Quanyuan Feng, Qiang Wu 0014 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | A High Throughput and Energy Efficient Lepton Hardware Encoder With Hash-Based Memory OptimizationabstractAlthough it has been surpassed by many subsequent coding standards, JPEG occupies a large storage share of the current data hosting service. To reduce the storage costs, DropBox proposed a secondary lossless compression algorithm, Lepton, to further improve the compression rate of JPEG images. However, the bloated probability models defined by Lepton severely restrict its throughput and energy efficiency. To solve this problem, we construct an access probability-based hash function for the probability models, and then propose a hardware-friendly memory optimization method by combining the proposed hash function with N-way Set-Associative unit. Besides, we also propose a synchronization mechanism for the serial accessing probability model, so that the syntax elements can be processed in parallel without changing the resulting bitstream. After that, we implement a high throughput, high energy efficiency, and low-cost Lepton hardware encoder. To the best of our knowledge, this is the first hardware implementation of transparent image recompression. The synthesis result shows that the proposed hardware structure reduces the total area of the probability models by 70.97%. Compared with DropBox’s software solution, the throughput and the energy efficiency of the proposed Lepton hardware encoder are increased by 63 and 1398 times on average. In terms of manufacturing cost, the proposed Lepton hardware encoder is also lower than the general-purpose CPU used by DropBox. Xiao Yan 0006, Zhixiong Di, Minjiang Li, Xiaoyang Zeng, Yibo Fan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | ASIC Design Principle Course with Combination of Online-MOOC and Offline-Inexpensive FPGA BoardabstractASIC Design Principle (ASICDP) is a compulsory course for undergraduate majors in microelectronics and integrated circuits, and the focus of this paper is the teaching methods of online theoretical teaching and offline experimental teaching of this course. As is well known, in order to prevent and control COVID-19, the use of online platforms to carry out online teaching has attracted worldwide attention. In this paper, the teaching strategy "Online-MOOC + Offline Inexpensive FPGA Board" in ASICDP in the Spring 2020 semester is demonstrated, in where MOOC means Massive Open Online Course. The theoretical teaching content of ASICDP is entirely replicated from Hardware Acceleration Design Methodology (HADM) released by the present authors on "China University MOOC," the largest MOOC platform in China. Meanwhile, with the support of the "Xilinx & Ministry of Education University-Industry Collaborative Education Program," an FPGA development board called the "Spartan Edge Accelerator Board" (SEA Board) designed by the authors was used in the experimental teaching of the ASICDP. This method can be used to establish the linkage between online courses and offline experiments, and cultivate students' practical VLSI design and FPGA prototype verification skills. It is believed that for educators that want to improve courses related to ASIC design and FPGA prototype verification, Online-MOOC + Offline-Inexpensive FPGA Board is an effective method with lower cost that is easily promotable and replicated. Zhixiong Di, Yongming Tang, Jiahua Lu, Zhaoyang Lv |
ACM Great Lakes Symposium on VLSI | 1 |