EDBT 2026 Demo / reviewers in the wild / expert
Hui Wu 0007
dblp:17/995-7
· DBLP profile ↗
16ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0008-9150-5608ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PVT-Insensitive Matrix-Vector Multiplier and ADC-less Tiling for Analog Computing and AI Accelerators
Lianlong Sun, Hui Wu 0007 |
ISLPED | 5 |
| 2026 | SLAM: Extending the Reach of Ising Machine AdvantageabstractDynamics-based Ising machines (IMs) are a promising substrate for high-speed combinatorial optimization and sampling. For problems that fit within their fixed capacity, they provide orders-of-magnitude speedups over conventional software algorithms. Once a problem exceeds hardware capacity, somewhat surprisingly, they become practically useless: Previous works generally relegate them to isolated sub-solvers. As our analysis will show, this approach produces no clear time or energy benefits, losing any inherent advantage of hardware IMs.After analyzing the shortcomings of previous hybrid proposals, we introduce a better method to extend IM advantage beyond hardware capacity. We call it the Stepped Large-neighborhood Annealing Method (SLAM), which drastically improves performance on combinatorial benchmarks with minimal host-side computation.We also propose novel architectural support to implement SLAM with low data movement overheads. The resulting augmented IM continues to exploit dynamical systems-enabled parallelism and thus maintains order of magnitude time-to-solution and energy-to-solution advantages over CPU-based algorithms. As a side benefit, our approach also significantly reduces the impact of device variation on solution quality, another perennial issue for analog optimizers. Matthew X. Burns, Zahra Azad, Yongchao Liu 0003, Tong Geng, Hui Wu 0007, Michael C. Huang 0001 |
IEEE Trans. Computers | 5 |
| 2025 | Integrated Hardware Annealing Based on Langevin Dynamics for Ising MachinesabstractIsing machines are non-von Neumann machines designed to solve combinatorial optimization problems (COP) by searching for the ground state, or the lowest energy configuration, within the Ising model. However, Ising machines often face the challenges of getting trapped in local minima due to the complex energy landscapes. Hardware annealing algorithms help mitigate this issue by using a probabilistic approach to steer the system toward the ground state. In this paper, we present a hardware annealing algorithm for Ising machines based on Langevin dynamics, a stochastic perturbation by random noise. Theoretical analysis, system-level design, and detailed circuit design are carried out. We evaluate the performance of the algorithm through chip-level simulation using a standard 65-nm CMOS technology to demonstrate the algorithm's efficacy. The results show that the proposed hardware annealing algorithm effectively guides the system to reach the ground state with a probability of 86.5%, significantly improving the solution quality by 97.5%. Further, we compare the algorithm with state-of-the-art hardware annealing methods through behavioral-level simulations, highlighting its improved solution quality alongside a 50% reduction in time- to-solution. Yongchao Liu 0003, Lianlong Sun 0001, Michael C. Huang 0001, Hui Wu 0007 |
DATE | 4 |
| 2025 | Ising machine based on charge re-distributionabstractIsing machines have attracted significant attention for solving quadratic unconstrained binary optimization (QUBO) with high performance. Among them, CMOS-compatible dynamics-based designs are promising approaches with high efficiency, scalability, and potential extension to more general combinatorial optimization problems (COPs) beyond QUBO problems. However, these machines can be limited by poor solution quality due to variations and leakage, or by long solution times resulting from inefficient annealing. In this paper, we propose a novel all-to-all connected Ising machine based on charge redistribution, capable of solving generic QUBO problems. To enhance solution quality, we apply a stochastic cycling annealing technique. A 50-spin Ising machine was developed using commercial 65nm CMOS technology, with both system-level and circuit-level designs implemented. Additionally, behavior-level simulations were conducted to evaluate the performance of a larger system. Simulation results show that the proposed Ising machine delivers competitive solution quality with reduced time-to-solution, making it a strong candidate for solving COPs. Yongchao Liu 0003, Lianlong Sun 0001, Matthew X. Burns, Michael C. Huang 0001, Hui Wu 0007 |
ISCAS | 5 |
| 2023 | Supporting Energy-based Learning with an Ising Machine substrate: a Case Study on RBMabstractNature apparently does a lot of computation constantly. If we can harness some of that computation at an appropriate level, we can potentially perform certain type of computation (much) faster and more efficiently than we can do with a von Neumann computer. Indeed, many powerful algorithms are inspired by nature and are thus prime candidates for nature-based computation. One particular branch of this effort that has seen some recent rapid advances is Ising machines. Some Ising machines are already showing better performance and energy efficiency for optimization problems. Through design iterations and co-evolution between hardware and algorithm, we expect more benefits from nature-based computing systems in the future. In this paper, we make a case for an augmented Ising machine suitable for both training and inference using an energy-based machine learning algorithm. We show that with a small change, the Ising substrate accelerates key parts of the algorithm and achieves non-trivial speedup and efficiency gain. With a more substantial change, we can turn the machine into a self-sufficient gradient follower to virtually complete training entirely in hardware. This can bring about 29x speedup and about 1000x reduction in energy compared to a Tensor Processing Unit (TPU) host. Uday Kumar Reddy Vengalam, Yongchao Liu 0003, Tong Geng, Hui Wu 0007, Michael C. Huang 0001 |
MICRO | 4 |
| 2019 | To Stack or Not To Stackabstract3D memory technology, such as Micron's hybrid memory cube (HMC), has re-energized the architectural pursuit of computation very close to, or inside the memory chip. Such a design falls into the broader category of near-data processing (NDP). The motivation for such design is because the current Von Neumann architecture of chip-multiprocessors is thought to make data movement expensive. Current NDP work focuses on the possibility of architecting computation engines, such as accelerators, cores, or graphic processing units right below the memory layers and inside the logic layer of the HMC sub-system. However, such a stacking design does present a number of technical challenges such as heat dissipation, power supply, etc. While these challenges can certainly be overcome, and needs to be addressed, in this work, we seek to answer a related question of whether it is necessary to stack general-purpose computation engines, directly inside the memory unit, in order to achieve the performance potential of NDP system; thus, to stack or not to stack. We show that, with computing models used in current NDP designs, placing the computation engines very close to, but outside the memory system (not stacking) can provide comparable performance without significant energy costs. This can be achieved without inventing any new technology, but utilizing current state-of-the-art high-speed link design practices. Richard Afoakwa, Lejie Lu, Hui Wu 0007, Michael C. Huang 0001 |
PACT | 3 |
| 2019 | Concurrent Multipoint-to-Multipoint Communication on Interposer ChannelsabstractChip-to-chip communication for next generation computing will require larger bandwidth density to support ever increasing data traffic between processors, memories and I/O. 3-D integration enables a large number of processor and memory chips to be densely packed on an interposer with fine-pitch interconnect lanes. Advanced signaling techniques such as pulse amplitude modulation (PAM) can be employed to improve bandwidth per lane. Most recent work on interposer-based chip-to-chip interconnects focus primarily on point-to-point serial links. Without adding costly routers, these designs will severely limit the overall system level concurrency. In this paper, we propose an ultrahigh-speed multipoint-to-multipoint link design for interposer channels, which supports PAM signaling. Each node on the link can send, receive, drop, or relay data at line rate without complex routing. This design enables splitting the physical link into segments, and allows multicast/broadcast. A proof-of-concept system prototype with up to 16 nodes integrated on a silicon interposer with up to 22-mm node spacing is designed and evaluated using circuit and system simulations. The PAM-4 transceiver and link interface circuits at each node are implemented using a standard 130-nm SiGe BiCMOS technology. The transceiver can achieve a data rate of 40-Gb/s/lane, with channel loss of -3.5 dB per segment at Nyquist frequency, and energy efficiency between 1.29-pJ/b between two neighboring nodes or 0.21-pJ/b more per additional nodes. Using a cycle-level system simulation, such a high-concurrency communication fabric can improve overall performance between 2% to 18% over baseline. Lejie Lu, Richard Afoakwa, Michael C. Huang 0001, Hui Wu 0007 |
ISLPED | 4 |
| 2018 | High Swing Pulse-Amplitude Modulation of Transmission Line Links for On-Chip CommunicationabstractWith ever increasing core count of chip-multiprocessors (CMPs), the network-on-chip (NoC) fabric continues to be an important component for performance and energy. We propose the use of high voltage swing serial links as the backbone NoC. We designed transmitter drivers to deliver a high output swing and enable high speed Pulse-Amplitude Modulation (PAM-4 and PAM-8) transmissions. We show that with careful circuit-level transceiver design, coupled with system level architectural utilization of such links, it is possible to drive up to 8 cm of on-chip transmission line at diverse adaptive modulations. Using such a design, experimental analysis shows an average of 1.4× performance improvement over baseline. The overall energy-delay product improvement is 1.75×. Richard Afoakwa, Lejie Lu, Yong Wang 0026, Hui Wu 0007, Michael C. Huang 0001 |
ISCAS | 4 |
| 2018 | An Energy-Efficient High-Swing PAM-4 Voltage-Mode TransmitterabstractAs the data rate of high-speed I/Os continues to increase, four-level pulse amplitude modulation (PAM-4) is adopted to improve the bandwidth density and link margin at 50 Gb/s and beyond. Compared to non-return-to-zero (NRZ) signaling, however, the PAM-4 eye height is reduced, which calls for larger transmitter swing to maintain signal-to-noise-ratio. A new energy-efficient transmitter is proposed to generate large swing PAM-4 signals with a cascode voltage-mode driver and supporting pre-drivers and logic circuits. By reconfiguring the pull-up and pull-down branches based on the transmit data and steering the bypass currents, the proposed voltage-mode driver significantly reduces power consumption compared to conventional implementation while maintaining impedance matching. Voltage stacking technique is adopted for pre-drivers to further improve energy efficiency. To demonstrate the new transmitter design, a prototype 56 Gb/s PAM-4 transmitter is designed using a generic 28-nm CMOS technology with a 2-V power supply voltage. It achieves a overall output swing of 2 V and a minimum eye height of 490 mV with good linearity (98.7% level separation mismatch ratio). Compared to a conventional voltage-mode transmitter design with the same swing, the static power consumption of the new transmitter is reduced almost by half (from 30 mW to 16 mW), and its overall energy efficiency improves from 0.7 pJ/b to 0.5 pJ/b. Lejie Lu, Yong Wang 0026, Hui Wu 0007 |
ISLPED | 3 |
| 2017 | Design high bandwidth-density, low latency and energy efficient on-chip interconnectabstractFor future high-performance computing chips, on-chip interconnect requires large bandwidth-density, low latency, and high energy-efficiency, which pose significant design challenges. This paper presents a design space exploration of transmission line based on-chip interconnect. First, we conduct an optimization of on-chip transmission lines to minimize the size, channel loss and inter-symbol-interference (ISI), hence to maximize the bandwidth-density. Based on the result, differential coplanar waveguide (CPW) with 55-μm pitch size is chosen as the transmission line topology. Next, channel capacities of channel lengths from 2 to 8 cm are characterized based on time-domain pulse responses. Various equalizers are studied, which are used to increase the ISI-limited channel capacity. To make better use of the large equalized channel capacity, pulse amplitude modulation (PAM) is employed instead of traditional non-return-to-zero (NRZ) signaling. A link budget analysis is then conducted to find the optimal modulation format for each channel. To verify our analyses, several transceivers are designed in 28-nm CMOS technology. An 84-Gb/s PAM-8 transceiver achieves 1.5-Gb/s/μm bandwidth-density over a 4-cm channel. The unrepeated bandwidth-density is 6.1 Gb/s/μm·cm, which is almost 10 times larger compared to prior work. Yong Wang 0026, Hui Wu 0007 |
ISLPED | 2 |
| 2012 | Enhancing effective throughput for transmission line-based busabstractMain-stream general-purpose microprocessors require a collection of high-performance interconnects to supply the necessary data movement. The trend of continued increase in core count has prompted designs of packet-switched network as a scalable solution for future-generation chips. However, the cost of scalability can be significant and especially hard to justify for smaller-scale chips. In contrast, a circuit-switched bus using transmission lines and corresponding circuits offers lower latencies and much lower energy costs for smaller-scale chips, making it a better choice than a full-blown network-on-chip (NoC) architecture. However, shared-medium designs are perceived as only a niche solution for small- to medium-scale chips. In this paper, we show that there are many low-cost mechanisms to enhance the effective throughput of a bus architecture. When a handful of highly cost-effective techniques are applied, the performance advantage of even the most idealistically configured NoCs becomes vanishingly small. We find transmission line-based buses to be a more compelling interconnect even for large-scale chip-multiprocessors, and thus bring into doubt the centrality of packet switching in future on-chip interconnect. Aaron Carpenter, Jianyun Hu, Övünç Kocabas, Michael C. Huang 0001, Hui Wu 0007 |
ISCA | 5 |
| 2011 | A case for globally shared-medium on-chip interconnectabstractAs microprocessor chips integrate a growing number of cores, the issue of interconnection becomes more important for overall system performance and efficiency. Compared to traditional distributed shared-memory architecture, chip-multiprocessors offer a different set of design constraints and opportunities. As a result, a conventional packet-relay multiprocessor interconnect architecture is a valid, but not necessarily optimal, design point. For example, the advantage of off-the-shelf interconnect and the in-field scalability of the interconnect are less important in a chip-multiprocessor. On the other hand, even with worsening wire delays,packet switching represents a non-trivial component of overall latency. Aaron Carpenter, Jianyun Hu, Michael C. Huang 0001, Hui Wu 0007 |
ISCA | 5 |
| 2011 | A design space exploration of transmission-line links for on-chip interconnect
Aaron Carpenter, Jianyun Hu, Michael C. Huang 0001, Hui Wu 0007, Peng Liu 0016 |
ISLPED | 4 |
| 2010 | An intra-chip free-space optical interconnectabstractContinued device scaling enables microprocessors and other systems-on-chip (SoCs) to increase their performance, functionality, and hence, complexity. Simultaneously, relentless scaling, if uncompensated, degrades the performance and signal integrity of on-chip metal interconnects. These systems have therefore become increasingly communications-limited. The communications-centric nature of future high performance computing devices demands a fundamental change in intra- and inter-chip interconnect technologies. Alok Garg, Berkehan Ciftcioglu, Jianyun Hu, Ioannis Savidis, Rebecca Berman, Peng Liu 0016, Michael C. Huang 0001, Hui Wu 0007, Eby G. Friedman, Gary Wicks, Duncan Moore |
ISCA | 11 |
| 2008 | Injection-Locked Clocking: A Low-Power Clock Distribution Scheme for High-Performance MicroprocessorsabstractWe propose injection-locked clocking (ILC) to combat deteriorating clock skew and jitter, and reduce power consumption in high-performance microprocessors. In the new clocking scheme, injection-locked oscillators are used as local clock receivers. Compared to conventional clocking with buffered trees or grids, ILC can achieve better power efficiency, lower jitter, and much simpler skew compensation thanks to its built-in deskewing capability. Unlike other alternatives, ILC is fully compatible with conventional clock distribution networks. In this paper, a quantitative study based on circuit and microarchitectural-level simulations is performed. Alpha21264 is used as the baseline processor, and is scaled to 0.13 m and 3 GHz. Simulations show 20- and 23-ps jitter reduction, 10.1% and 17% power savings in two ILC configurations. A test chip distributing 5-GHz clock is implemented in a standard 0.18- m CMOS technology and achieved excellent jitter performance and a deskew range up to 80 ps. Aaron Carpenter, Berkehan Ciftcioglu, Alok Garg, Michael C. Huang 0001, Hui Wu 0007 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2006 | A reconfigurable, multi-gigahertz pulse shaping circuit based on distributed transversal filtersabstractA distributed transversal filter (DTF) is a good candidate for ultrafast pulse shaping in wideband systems like UWB. This paper presents the circuit analysis of DTFs based on transmission line theory. We also show the detailed design procedure of DTFs. A 5-tap prototype DTF is designed and fabricated using microwave PCB substrate and pHEMT discrete transistors. Simulation and initial measurement results demonstrate its pulse shaping capability. Yunliang Zhu, Jonathan D. Zuegel, John R. Marciante, Hui Wu 0007 |
ISCAS | 4 |