EDBT 2026 Demo / reviewers in the wild / expert
Gushu Li
dblp:163/3591
· DBLP profile ↗
23ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0002-6233-0334ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 6 first-author · 15 since 2021Software engineering, systems software and programming languages · 15 · 5 first-author · 12 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QTurbo: A Robust and Efficient Compiler for Analog Quantum SimulationabstractAnalog quantum simulation leverages native hardware dynamics to emulate complex quantum systems with great efficiency by bypassing the quantum circuit abstraction. However, conventional compilation methods for analog simulators are typically labor-intensive, prone to errors, and computationally demanding. This paper introduces QTurbo, a powerful analog quantum simulation compiler designed to significantly enhance compilation efficiency and optimize hardware execution time. By generating precise and noise-resilient pulse schedules, our approach ensures greater accuracy and reliability, outperforming the existing state-of-the-art approach. Junyu Zhou 0005, Yuhao Liu 0017, Shize Che, Anupam Mitra, Efekan Kökcü, Ermal Rrapaj, Costin Iancu, Gushu Li |
ASPLOS (1) | 8 |
| 2026 | AlphaSyndrome: Tackling the Syndrome Measurement Circuit Scheduling Problem for QEC CodesabstractQuantum error correction (QEC) is essential for scalable quantum computing, yet repeated syndrome-measurement cycles dominate its spacetime and hardware cost. Although stabilizers commute and admit many valid execution orders, different schedules induce distinct error-propagation paths under realistic noise, leading to large variations in logical error rate. Outside of surface codes, effective syndrome-measurement scheduling remains largely unexplored. We present AlphaSyndrome, an automated synthesis framework for scheduling syndrome-measurement circuits in general commuting-stabilizer codes under minimal assumptions: mutually commuting stabilizers and a heuristic decoder. AlphaSyndrome formulates scheduling as an optimization problem that shapes error propagation to (i) avoid patterns close to logical operators and (ii) remain within the decoder's correctable region. The framework uses Monte Carlo Tree Search (MCTS) to explore ordering and parallelism, guided by code structure and decoder feedback. Across diverse code families, sizes, and decoders, AlphaSyndrome reduces logical error rates by 80.6% on average (up to 96.2%) relative to depth-optimal baselines, matches Google's hand-crafted surface-code schedules, and outperforms IBM's schedule for the Bivariate Bicycle code. Yuhao Liu 0017, Shuohao Ping, Junyu Zhou 0005, Ethan Decker, Justin Kalloor, Mathias Weiden, Kean Chen, Yunong Shi, Ali Javadi-Abhari, Costin Iancu, Gushu Li |
ASPLOS (2) | 11 |
| 2026 | Kernpiler: Compiler Optimization for Quantum Hamiltonian Simulation with Partial TrotterizationabstractDescription This artifact contains the core implementation of Kernpiler, a compiler framework for optimizing quantum circuits, and supports full reproducibility of all experimental results presented in the associated paper. The artifact includes all code, data pipelines, and scripts required to regenerate Figures 5–11. System Used for Data Collection NVIDIA A100 GPU with 80GB memory AMD EPYC 9654P 96-core processor x86_64 Linux system Python 3.13.5 Experiments may be computationally intensive, but they are fully parallelizable across multiple devices. Installation Clone or download the repository and navigate to the project directory. Create a virtual environment: python3 -m venv validate source validate/bin/activate Install dependencies: python -m pip install -r requirements.txt Install torch-scatter: python -m pip install --no-cache-dir torch-scatter -f https://data.pyg.org/whl/torch-2.10.0+cu128.html Experiment Workflow All experiment scripts are located in: src/compiler/optimization_passes/experiments Each figure can be reproduced by running its corresponding data collection and graphing scripts: Figure 5exp_gatecount_datacollection.py→ graph using:graph_data_scripts/graph_absolute.py Figure 6exp_partition_scaling_datacollection_o1exp_partition_scaling_datacollection→ graph using:graph_data_scripts/graph_o1_o2_side_by_side.py Figure 7exp_runtime_per_pass.py→ graph using:graph_data_scripts/graph_runtime_per_pass.py Figure 8exp_partition_scaling_datacollectiono1_phoenixFT→ graph using:graph_data_scripts/graph_firstorder_scalingFT.py Figure 9exp_partitionalgvsrandom.py→ graph using:graph_data_scripts/graph_partition_vs_random.py Figure 10exp_scaling_data_rewriteradius.py→ graph using:graph_data_scripts/graph_scalingdata.py Figure 11exp_error_scaling_systemsize.py→ output generated directly (no additional graph script required) Execution Notes All experiments are independent Parallel execution is supported Runtime varies depending on system size and hardware Ethan Decker, Lucas Goetz, Evan McKinney, Erik Gustafson, Junyu Zhou 0005, Alex K. Jones, Ang Li 0006, Alexander Schuckert, Samuel A. Stein, Eleanor Crane, Gushu Li |
ISCA | 12 |
| 2025 | Verifying Fault-Tolerance of Quantum Error Correction CodesabstractAbstract Quantum computers have advanced rapidly in qubit count and gate fidelity. However, large-scale fault-tolerant quantum computing still relies on quantum error correction code (QECC) to suppress noise. Manually or experimentally verifying the fault-tolerance property of complex QECC implementation is impractical due to the vast error combinations. This paper formalizes the fault-tolerance of QECC implementations within the language of quantum programs. By incorporating the techniques of quantum symbolic execution, we provide an automatic verification tool for quantum fault-tolerance. We evaluate and demonstrate the effectiveness of our tool on a universal set of logical operations across different QECCs. Kean Chen, Yuhao Liu 0017, Wang Fang 0001, Jennifer Paykin, Xin-Chuan Wu, Albert T. Schmitz, Steve Zdancewic, Gushu Li |
CAV (4) | 8 |
| 2025 | HATT: Hamiltonian Adaptive Ternary Tree for Optimizing Fermion-to-Qubit MappingabstractThis paper introduces the Hamiltonian-Adaptive Ternary Tree (HATT) framework to compile optimized Fermion-to-qubit mapping for specific Fermionic Hamiltonians. In the simulation of Fermionic quantum systems, efficient Fermion-toqubit mapping plays a critical role in transforming the Fermionic system into a qubit system. HATT utilizes ternary tree mapping and a bottom-up construction procedure to generate Hamiltonian aware Fermion-to-qubit mapping to reduce the Pauli weight of the qubit Hamiltonian, resulting in lower quantum simulation circuit overhead. Additionally, our optimizations retain the important vacuum state preservation property in our Fermion-toqubit mapping and reduce the complexity of our algorithm from $O\left(N^{4}\right)$ to $O\left(N^{3}\right)$. Evaluations on various Fermionic systems demonstrate $5 \sim 25 \%$ reduction in Pauli weight, gate count, and circuit depth, alongside excellent scalability to larger systems. Experiments on the Ionq device also show the advantages of HATT in noise resistance in quantum simulations. Yuhao Liu 0017, Kevin Yao, Jonathan Hong, Julien Froustey, Ermal Rrapaj, Costin Iancu, Gushu Li, Yunong Shi |
HPCA | 7 |
| 2025 | MarQSim: Reconciling Determinism and Randomness in Compiler Optimization for Quantum SimulationabstractQuantum Hamiltonian simulation, fundamental in quantum algorithm design, extends far beyond its foundational roots, powering diverse quantum computing applications. However, optimizing the compilation of quantum Hamiltonian simulation poses significant challenges. Existing approaches fall short in reconciling deterministic and randomized compilation, lack appropriate intermediate representations, and struggle to guarantee correctness. Addressing these challenges, we present MarQSim, a novel compilation framework. MarQSim leverages a Markov chain-based approach, encapsulated in the Hamiltonian Term Transition Graph, adeptly reconciling deterministic and randomized compilation benefits. Furthermore, we formulate a Minimum-Cost Flow model that can tune transition matrices to enforce correctness while accommodating various optimization objectives. Experimental results demonstrate MarQSim’s superiority in generating more efficient quantum circuits for simulating various quantum Hamiltonians while maintaining precision. Xiuqi Cao, Junyu Zhou 0005, Yuhao Liu 0017, Yunong Shi, Gushu Li |
Proc. ACM Program. Lang. | 5 |
| 2024 | Fermihedral: On the Optimal Compilation for Fermion-to-Qubit EncodingabstractThis paper introduces Fermihedral, a compiler framework focusing on discovering the optimal Fermion-to-qubit encoding for targeted Fermionic Hamiltonians. Fermion-to-qubit encoding is a crucial step in harnessing quantum computing for efficient simulation of Fermionic quantum systems. Utilizing Pauli algebra, Fermihedral redefines complex constraints and objectives of Fermion-to-qubit encoding into a Boolean Satisfiability problem which can then be solved with high-performance solvers. To accommodate larger-scale scenarios, this paper proposed two new strategies that yield approximate optimal solutions mitigating the overhead from the exponentially large number of clauses. Evaluation across diverse Fermionic systems highlights the superiority of Fermihedral, showcasing substantial reductions in implementation costs, gate counts, and circuit depth in the compiled circuits. Real-system experiments on IonQ's device affirm its effectiveness, notably enhancing simulation accuracy. Yuhao Liu 0017, Shize Che, Junyu Zhou 0005, Yunong Shi, Gushu Li |
ASPLOS (3) | 5 |
| 2024 | Fast Virtual Gate Extraction For Silicon Quantum Dot DevicesabstractSilicon quantum dot devices stand as promising candidates for large scale quantum computing due to their extended coherence times compact size, and recent experimental demonstrations of sizable qubit arrays. Despite the great potential, controlling these arrays remains a significant challenge. This paper introduces a new virtual gate extraction method to quickly establish orthogonal control on the potentials for individual quantum dots. Leveraging insights from the device physics, the proposed approach significantly re duces the experimental overhead by focusing on crucial regions around charge state transition. Furthermore, by employing an efficient voltage sweeping method, we can efficiently pinpoint these charge state transition lines and filter out erroneous points. Exper imental evaluation using real quantum dot chip datasets demon strates a substantial 5.84× to 19.34× speedup over conventional methods, thereby showcasing promising prospects for accelerating the scaling of silicon spin qubit devices. Shize Che, Seongwoo Oh, Haoyun Qin, Yuhao Liu 0017, Anthony Sigillito, Gushu Li |
DAC | 6 |
| 2024 | Bosehedral: Compiler Optimization for Bosonic Quantum ComputingabstractBosonic quantum computing, based on the infinite-dimensional qumodes, has shown promise for various practical applications that are classically hard. However, the lack of compiler optimizations has hindered its full potential. This paper introduces Bosehedral, an efficient compiler optimization framework for (Gaussian) Boson sampling on Bosonic quantum hardware. Bosehedral overcomes the challenge of handling infinite-dimensional qumode gate matrices by performing all its program analysis and optimizations at a higher algorithmic level, using a compact unitary matrix representation. It optimizes qumode gate decomposition and logical-to-physical qumode mapping, and introduces a tunable probabilistic gate dropout method. Overall, Bosehedral significantly improves the performance by accurately approximating the original program with much fewer gates. Our evaluation shows that Bosehedral can largely reduce the program size but still maintain a high approximation fidelity, which can translate to significant end-to-end application performance improvement. Junyu Zhou 0005, Yuhao Liu 0017, Yunong Shi, Ali Javadi-Abhari, Gushu Li |
ISCA | 5 |
| 2023 | OneQ: A Compilation Framework for Photonic One-Way Quantum ComputationabstractIn this paper, we propose OneQ, the first optimizing compilation framework for one-way quantum computation towards realistic photonic quantum architectures. Unlike previous compilation efforts for solid-state qubit technologies, our innovative framework addresses a unique set of challenges in photonic quantum computing. Specifically, this includes the dynamic generation of qubits over time, the need to perform all computation through measurements instead of relying on 1-qubit and 2-qubit gates, and the fact that photons are instantaneously destroyed after measurements. As pioneers in this field, we demonstrate the vast optimization potential of photonic one-way quantum computing, showcasing the remarkable ability of OneQ to reduce computing resource requirements by orders of magnitude. Hezi Zhang, Anbang Wu, Gushu Li, Hassan Shapourian, Alireza Shabani, Yufei Ding 0001 |
ISCA | 4 |
| 2022 | Paulihedral: a generalized block-wise compiler optimization framework for Quantum simulation kernelsabstractThe quantum simulation kernel is an important subroutine appearing as a very long gate sequence in many quantum programs. In this paper, we propose Paulihedral, a block-wise compiler framework that can deeply optimize this subroutine by exploiting high-level program structure and optimization opportunities. Paulihedral first employs a new Pauli intermediate representation that can maintain the high-level semantics and constraints in quantum simulation kernels. This naturally enables new large-scale optimizations that are hard to implement at the low gate-level. In particular, we propose two technology-independent instruction scheduling passes, and two technology-dependent code optimization passes which reconcile the circuit synthesis, gate cancellation, and qubit mapping stages of the compiler. Experimental results show that Paulihedral can outperform state-of-the-art compiler infrastructures in a wide-range of applications on both near-term superconducting quantum processors and future fault-tolerant quantum computers. Gushu Li, Anbang Wu, Yunong Shi, Ali Javadi-Abhari, Yufei Ding 0001, Yuan Xie 0001 |
ASPLOS | 1 |
| 2022 | A synthesis framework for stitching surface code with superconducting quantum devicesabstractQuantum error correction (QEC) is the central building block of fault-tolerant quantum computation but the design of QEC codes may not always match the underlying hardware. To tackle the discrepancy between the quantum hardware and QEC codes, we propose a synthesis framework that can implement and optimize the surface code onto superconducting quantum architectures. In particular, we divide the surface code synthesis into three key subroutines. The first two optimize the mapping of data qubits and ancillary qubits including syndrome qubits on the connectivity-constrained superconducting architecture, while the last subroutine optimizes the surface code execution by rescheduling syndrome measurements. Our experiments on mainstream superconducting architectures demonstrate the effectiveness of the proposed synthesis framework. Especially, the surface codes synthesized by the proposed automatic synthesis framework can achieve comparable or even better error correction capability than manually designed QEC codes. Anbang Wu, Gushu Li, Hezi Zhang, Gian Giacomo Guerreschi, Yufei Ding 0001, Yuan Xie 0001 |
ISCA | 2 |
| 2022 | AutoComm: A Framework for Enabling Efficient Communication in Distributed Quantum ProgramsabstractDistributed quantum computing (DQC) is a promising approach to extending the computational power of near-term quantum hardware. However, the non-local quantum communication between quantum nodes is much more expensive and error-prone than the local quantum operation within each quantum device. Previous DQC compilers focus on optimizing the implementation of each non-local gate and adopt similar compilation designs to single-node quantum compilers. The communication patterns in distributed quantum programs remain unexplored, leading to a far-from-optimal communication cost. In this paper, we identify burst communication, a specific qubit-node communication pattern that widely exists in various distributed quantum programs and can be leveraged to guide communication overhead optimization. We then propose AutoComm, an automatic compiler framework to extract burst communication patterns from input programs and then optimize the communication steps of burst communication discovered. Compared to state-of-the-art DQC compilers, experimental results show that our proposed AutoComm can reduce the communication resource consumption and the program latency by 72.9% and 69.2% on average, respectively. Anbang Wu, Hezi Zhang, Gushu Li, Alireza Shabani, Yuan Xie 0001, Yufei Ding 0001 |
MICRO | 3 |
| 2022 | STPAcc: Structural TI-Based Pruning for Accelerating Distance-Related Algorithms on CPU-FPGA PlatformsabstractAs a promising solution to boost the performance of distance-related algorithms (e.g.,$K$-means and KNN), FPGA-based acceleration attracts lots of attention, but also comes with numerous challenges. In this work, we propose,STPAcc, an optimization framework based on structural triangle-inequality (TI)-based pruning (STP) for accelerating distance-related algorithms on CPU-FPGA platforms. STPAcc provides a domain-specific language to unify distance-related algorithms effectively, a structural TI-based pruning strategy to remove unnecessary distance computations, a coarse-grained workload partitioning and mapping strategy to fully exploit the potentials of the CPU-FPGA platform, and fine-grained hardware optimizations to further improve performance on the FPGA. Intensive experiments show that STPAcc designs achieve$31.42\times $speedup and$99.63\times $better energy efficiency on average over standard CPU-based implementations. Boyuan Feng, Gushu Li, Lei Deng 0003, Yuan Xie 0001, Yufei Ding 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | TiAcc: Triangle-inequality based Hardware Accelerator for K-means on FPGAsabstractK-means is one of the most important unsuper-vised learning algorithms. In this paper, we present TiAcc, a triangle-inequality based K-means hardware accelerator on FPGAs. TiAcc highlights itself with an algorithm-hardware co-design strategy tailored for K-means clustering. Specifically, TiAcc leverages a novel triangle-inequality based filtering to eliminate unnecessary distance computations without changing the final clustering results. Meanwhile, it employs a pipeline decoupling approach to mitigate the irregularity of the remaining computations, and an efficient hardware architecture design to fully exploit the pipeline and parallel processing capability of FPGAs. Moreover, TiAcc provides parameterized configuration knobs that can minimize the manual efforts in the arduous hardware design process and provides flexibility to optimize hardware designs for a variety of datasets with different sizes and dimensionalities. Intensive experiments show that TiAcc achieves an average 4.94× speedup and significant energy efficiency (average 74.22 ×) compared with an optimized K-means running on a server-grade Xeon CPU. Boyuan Feng, Gushu Li, Georgios Tzimpragos, Lei Deng 0003, Yuan Xie 0001, Yufei Ding 0001 |
CCGRID | 3 |
| 2021 | Software-Hardware Co-Optimization for Computational Chemistry on Superconducting Quantum ProcessorsabstractComputational chemistry is the leading application to demonstrate the advantage of quantum computing in the near term. However, large-scale simulation of chemical systems on quantum computers is currently hindered due to a mismatch between the computational resource needs of the program and those available in today’s technology. In this paper we argue that significant new optimizations can be discovered by co-designing the application, compiler, and hardware. We show that multiple optimization objectives can be coordinated through the key abstraction layer of Pauli strings, which are the basic building blocks of computational chemistry programs. In particular, we leverage Pauli strings to identify critical program components that can be used to compress program size with minimal loss of accuracy. We also leverage the structure of Pauli string simulation circuits to tailor a novel hardware architecture and compiler, leading to significant execution overhead reduction by up to 99%. While exploiting the high-level domain knowledge reveals significant optimization opportunities, our hardware/software framework is not tied to a particular program instance and can accommodate the full family of computational chemistry problems with such structure. We believe the co-design lessons of this study can be extended to other domains and hardware technologies to hasten the onset of quantum advantage. Gushu Li, Yunong Shi, Ali Javadi-Abhari |
ISCA | 1 |
| 2021 | GNNAdvisor: An Adaptive and Efficient Runtime System for GNN Acceleration on GPUs
Boyuan Feng, Gushu Li, Shuangchen Li, Lei Deng 0003, Yuan Xie 0001, Yufei Ding 0001 |
OSDI | 3 |
| 2021 | Palleon: A Runtime System for Efficient Video Processing toward Dynamic Class Skew
Boyuan Feng, Gushu Li, Yuan Xie 0001, Yufei Ding 0001 |
USENIX ATC | 3 |
| 2020 | Towards Efficient Superconducting Quantum Processor Architecture DesignabstractMore computational resources (i.e., more physical qubits and qubit connections) on a superconducting quantum processor not only improve the performance but also result in more complex chip architecture with lower yield rate. Optimizing both of them simultaneously is a difficult problem due to their intrinsic trade-off. Inspired by the application-specific design principle, this paper proposes an automatic design flow to generate simplified superconducting quantum processor architecture with negligible performance loss for different quantum programs. Our architecture-design-oriented profiling method identifies program components and patterns critical to both the performance and the yield rate. A follow-up hardware design flow decomposes the complicated design procedure into three subroutines, each of which focuses on different hardware components and cooperates with corresponding profiling results and physical constraints. Experimental results show that our design methodology could outperform IBM's general-purpose design schemes with better Pareto-optimal results.,0 Gushu Li, Yufei Ding 0001, Yuan Xie 0001 |
ASPLOS | 1 |
| 2020 | Eliminating Redundant Computation in Noisy Quantum Computing SimulationabstractNoisy Quantum Computing (QC) simulation on a classical machine is very time consuming since it requires Monte Carlo simulation with a large number of error-injection trials to model the effect of random noises. Orthogonal to existing QC simulation optimizations, we aim to accelerate the simulation by eliminating the redundant computation among those Monte Carlo simulation trials. We observe that the intermediate states of many trials can often be the same. Once these states are computed in one trial, they can be temporarily stored and reused in other trials. However, storing such states will consume significant memory space. To leverage the shared intermediate states without introducing too much storage overhead, we propose to statically generate and analyze the Monte Carlo simulation simulation trials before the actual simulation. Those trials are reordered to maximize the overlapped computation between two consecutive trials. The states that cannot be reused in follow-up simulation are dropped, so that we only need to store a few states. Experiment results show that the proposed optimization scheme can save on average 80% computation with only a small number of state vectors stored. In addition, the proposed simulation scheme demonstrates great scalability as more computation can be saved with more simulation trials or on future QC devices with reduced error rates. Gushu Li, Yufei Ding 0001, Yuan Xie 0001 |
DAC | 1 |
| 2020 | Projection-based runtime assertions for testing and debugging Quantum programsabstractIn this paper, we propose Proq, a runtime assertion scheme for testing and debugging quantum programs on a quantum computer. The predicates in Proq are represented by projections (or equivalently, closed subspaces of the state space), following Birkhoff-von Neumann quantum logic. The satisfaction of a projection by a quantum state can be directly checked upon a small number of projective measurements rather than a large number of repeated executions. On the theory side, we rigorously prove that checking projection-based assertions can help locate bugs or statistically assure that the semantic function of the tested program is close to what we expect, for both exact and approximate quantum programs. On the practice side, we consider hardware constraints and introduce several techniques to transform the assertions, making them directly executable on the measurement-restricted quantum computers. We also propose to achieve simplified assertion implementation using local projection technique with soundness guaranteed. We compare Proq with existing quantum program assertions and demonstrate the effectiveness and efficiency of Proq by its applications to assert two sophisticated quantum algorithms, the Harrow-Hassidim-Lloyd algorithm and Shor’s algorithm. Gushu Li, Li Zhou 0013, Nengkun Yu, Yufei Ding 0001, Mingsheng Ying, Yuan Xie 0001 |
Proc. ACM Program. Lang. | 1 |
| 2019 | Tackling the Qubit Mapping Problem for NISQ-Era Quantum DevicesabstractDue to little considerations in the hardware constraints, e.g., limited connections between physical qubits to enable two-qubit gates, most quantum algorithms cannot be directly executed on the Noisy Intermediate-Scale Quantum (NISQ) devices. Dynamically remapping logical qubits to physical qubits in the compiler is needed to enable the two-qubit gates in the algorithm, which introduces additional operations and inevitably reduces the fidelity of the algorithm. Previous solutions in finding such remapping suffer from high complexity, poor initial mapping quality, and limited flexibility and control. To address these drawbacks mentioned above, this paper proposes a SWAP-based Bidirectional heuristic search algorithm (SABRE), which is applicable to NISQ devices with arbitrary connections between qubits. By optimizing every search attempt, globally optimizing the initial mapping using a novel reverse traversal technique, introducing the decay effect to enable the trade-off between the depth and the number of gates of the entire algorithm, SABRE outperforms the best known algorithm with exponential speedup and comparable or better results on various benchmarks. Gushu Li, Yufei Ding 0001, Yuan Xie 0001 |
ASPLOS | 1 |
| 2015 | A STT-RAM-based low-power hybrid register file for GPGPUsabstractRecently, general-purpose graphics processing units (GPGPUs) have been widely used to accelerate computing in various applications. To store the contexts of thousands of concurrent threads on a GPU, a large static random-access memory (SRAM)-based register file is employed. Due to high leakage power of SRAM, the register file consumes 20% to 40% of the total GPU power consumption. Thus, hybrid memory system, which combines SRAM and the emerging non-volatile memory (NVM), has been employed for register file design on GPUs. Although it has shown strong potential to alleviate the power issue of GPUs, existing hybrid memory solutions might not exploit the intrinsic feature of GPU register file. By leveraging the warp schedule on GPU, this paper proposes a hybrid register architecture which consists of a NVM-based register file and mixed SRAM-based write buffers with a warp-aware write back strategy. Simulation results show that our design can eliminate 64% of write accesses to NVM and reduce power of register file by 66% on average, with only 4.2% performance degradation. After we apply the power gating technique, the register power is further reduced to 25% of SRAM counterpart on average. Gushu Li, Xiaoming Chen 0003, Guangyu Sun 0003, Henry Hoffmann, Yongpan Liu, Yu Wang 0002, Huazhong Yang |
DAC | 1 |