VLDB 2026 Research / reviewers in the wild / expert
Hui Wang 0023
dblp:39/721-23
· DBLP profile ↗
19ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-4688-1451ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 13 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SA-ANT: Efficient Low-Bit Group-Wise Quantization for Large Language Models via Sign-Asymmetric Adaptive Numeric TypeabstractLarge language models (LLMs) have demonstrated remarkable potential across diverse domains; meanwhile, their large parameter sizes pose substantial inference costs, motivating the need for efficient low-bit quantization. Group-wise quantization, which adopts finer granularity, has been widely used to improve low-bit quantization performance. Several adaptive numeric types have been proposed to further enhance low-bit group-wise quantization; however, they construct quantization grids based on symmetric numeric types, which limits their ability to model asymmetric distributions. To address this limitation, we propose SA-ANT, a sign-asymmetric adaptive numeric type for efficient low-bit group-wise quantization. SA-ANT constructs quantization grids separately on the positive and negative sides, enabling adaptive support for asymmetric and non-uniform distributions. Furthermore, the carefully designed SA-ANT not only reduces quantization errors but also ensures a unified computing across different sub numeric types, thereby facilitating hardware efficiency. To accelerate LLM inference, we develop (1) a quantization framework that transforms LLM weights into the SA-ANT and adaptively selects the sub numeric type for each group, and (2) an accelerator that maps SA-ANT inference to low-bit INT operations. Experimental results show that SA-ANT delivers 3.92%– 5.57% higher accuracy than state-of-the-art adaptive numeric types under 3-bit weight quantization, while also enabling 7.84%– 44.65% area savings and 7.80%–43.88% power reductions. Xinkuang Geng, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 3 |
| 2026 | Bridging the Power Estimation Gap: A GNN-Based Prediction Model for Approximate Logic SynthesisabstractApproximate computing is an application-related paradigm that trades limited accuracy for improvements in hardware cost. As a key technique of approximate computing, approximate logic synthesis (ALS) automatically generates approximate circuits with reduced area, power, and delay while satisfying predefined quality-of-result (QoR) constraints. However, in typical gate-level ALS workflows, synthesis tools are invoked at the final stage for optimization, leading to a discrepancy between the circuit in design space exploration (DSE) and the final obtained circuit. Thus, the power estimation for a candidate circuit during DSE may exhibit a significant gap from the actual power consumed by its post-synthesis circuit. This gap may mislead the DSE to a sub-optimal design. To address this issue, we propose a graph neural network (GNN)-based power prediction model that operates on gate-level circuits. The model incorporates multi-head channel attention, which extracts high-level topological and functional features that correlate with power dissipation and implicitly captures the optimization behavior of synthesis tools. Thus, it enables a direct prediction of post-synthesis power from pre-synthesis gate-level circuits. Experimental results show that the proposed model improves the concordance index (C-index) for power ranking by up to 14.0% over traditional methods. Furthermore, we construct an ALS framework by integrating the proposed model with Cartesian genetic programming (CGP). Compared to state-of-the-art ALS approaches, our GNN-CGP framework generates circuits with up to 26.8% power savings under the same error constraints. Fuxuan Li, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 4 |
| 2026 | A 28-Gbps 28-nm CMOS Wireline Receiver Analog Front End With An Adaptive CTLE Enabled By A Hexagonal Mask Eye-Opening Monitor
Yixing Jiang, Runtao Huo, Weihao Jie, Minmin You, Hui Wang 0023 |
ISCAS | 7 |
| 2026 | A 10-Bit Successive Reference Approximation ADC Enabled By Switched-Capacitor Dynamic Reference Generation With Pipelined Pre-charging And Voltage Ripple Cancellation
Hanlin Xu, Runtao Huo, Honglan Jiang, Minmin You, Hui Wang 0023 |
ISCAS | 7 |
| 2026 | LUT-ALMs: Trading Off Accuracy and Power for Approximate Logarithmic Multipliers via LUT OptimizationabstractLogarithmic multiplier (LM) converts fixed-point (FxP) input operands to logarithmic numbers and performs multiplication with simple shift and addition operations, which achieves distinct power reduction, yet with significant single-sided errors. This paper proposes to fuse error compensation with logarithmic conversion by using customized look-up tables (LUTs). To avoid the use of large LUTs, partition strategies are designed for the optimization of LUTs. In addition, to effectively balance the accuracy and hardware costs, two iterative algorithms are proposed for generating precision-configurable LUTs. Based on the optimized LUTs, high-accuracy and low-power approximate LMs (LUT-ALMs) are constructed for 8-bit and 16-bit multiplications. Furthermore, to enhance the flexibility of data type and bit width, a mixed-mode LM (MM-ALM) supporting eight multiplication modes is devised. Compared with the exact 8-bit signed multiplier from Synopsys DesignWare (DW) library (DW-Exact), LUT-ALMs show up to 30.96% and 23.99% reductions in the power-delay product (PDP) and area-delay product (ADP), respectively, with a mean relative error distance (MRED) of 3.45%. Compared with state-of-the-art approximate 8- bit signed multipliers, LUT-ALMs form the Pareto front in terms of PDP and MRED. For 16-bit multiplication, LUT-ALMs obtain up to 70.90% and 62.13% savings in PDP and ADP with a MRED of 2.75%, compared with the corresponding DW-Exact. Compared with the corresponding mixed-mode exact multiplier constructed of DW multipliers, MM-ALM performing 8-bit multiplication achieves up to 57.44% and 47.99% reductions in PDP and ADP, respectively. When performing 4-bit multiplications, MM-ALM can save up to 37.16% savings in PDP. With lower hardware over-heads, LUT-ALMs and MM-ALM present comparable accuracy to the corresponding exact designs in the considered convolutional neural networks (CNNs) and image processing applications. The hardware description of the devised LMs is open-sourced athttps://anonymous.4open.science/r/LUT-ALM-5FD1. Xinkuang Geng, Xiaolu Hu, Hui Wang 0023, Jianfei Jiang 0001, Qin Wang 0009, Siting Liu 0001, Jie Han 0001, Honglan Jiang |
IEEE Trans. Computers | 4 |
| 2026 | Guest Editorial: Special Section on the Asia Pacific Conference on Circuits and Systems - APCCAS 2025
Junghwan Han, Won-Young Lee, Hui Wang 0023 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | EPIC: Error PredIction and Correction for Power-Efficient Voltage Underscaling Multiply-Accumulate UnitabstractMatrix multiplication dominates the power consumption in compute-intensive applications such as deep neural networks (DNNs), spurring intensive investigations into power-efficient multiply-accumulate (MAC) units. Among the mainstream low-power design methodologies, voltage underscaling can achieve effective power savings yet induce timing errors that may lead to catastrophic accuracy loss. In this paper, we propose an error prediction and correction framework (denoted as EPIC) for arbitrary MAC unit under voltage underscaling, which predicts the timing errors and samples the correct output by using a delay-tunable clock. A prediction bits searching algorithm is proposed to enhance the prediction accuracy with low hardware cost, resulting in up to 100% accuracy. While preserving the accuracy, EPIC achieves up to 52% power savings over the corresponding MAC operating at nominal voltage. With transistor-level optimizations, EPIC incurs only 8% area and 1% power overheads, achieving 100% error correction under a voltage underscaling ratio of $\mathbf{0. 7 4}$. Compared to state-of-the-art error resilient circuit designs, EPIC consumes 60%-88% less area. Additionally, to achieve the accuracy performance of EPIC in error-resilient applications, we propose a simulation workflow involving precise timing features, enabling an accurate simulation of voltage underscaling MAC in large-scale applications. The experimental results show that, under voltage-underscaling, the MAC with EPIC consumes 11% less power than the one without EPIC, when a same accuracy as exact implementation is required in multi-layer perceptron (MLP). Tongjing Wu, Xiaolu Hu, Siting Liu 0001, Hui Wang 0023, Weifeng He, Zhigang Mao, Honglan Jiang |
DAC | 5 |
| 2025 | Lookup Table Refactoring: Towards Efficient Logarithmic Number System Addition for Large Language ModelsabstractCompared to integer quantization, logarithmic quantization aligns more effectively with the long-tailed distribution of data in large language models (LLMs), resulting in lower quantization errors. Moreover, the logarithmic number system (LNS) employs a fixed-point adder to perform multiplication, indicating a potential reduction in computational complexity for LLM accelerators that require extensive multiply-accumulate (MAC) operations. However, a key bottleneck is that LNS addition requires complex nonlinear functions, which are typically approximated using lookup tables (LUTs). This study aims to reduce the hardware resources needed for LUTs in LNS addition while maintaining high precision. Specifically, we investigate the specific nature of addition operations within LLMs; the relationship between the hardware parameters of the LUT and the computing errors is then mathematically derived. Based on these insights, we propose LUT refactoring to optimize the LUT for enhanced efficiency in LNS addition. With 10.93% and 19.78% reductions in area-delay product (ADP) and power-delay product (PDP), respectively, LUT refactoring results in an accuracy improvement of up to 33.5% in LLM benchmarks compared to the naive design. When compared to integer quantization, our method achieves higher accuracy while reducing area by 18.27% and power by 42.61%. Xinkuang Geng, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 3 |
| 2025 | Segment-Wise Accumulation: Low-Error Logarithmic Domain Computing for Efficient Large Language Model InferenceabstractLogarithmic domain computing (LDC) has great potential for reducing quantization errors and computational complexity in Large Language Models (LLMs). While logarithmic multiplication can be efficiently implemented using fixed-point addition, the primary challenge in multiply-accumulate (MAC) operations is balancing the precision of logarithmic adders with their hardware overhead. Through a detailed analysis of the errors inherent in LDC-based LLMs, we propose segment-wise accumulation (SWA) to mitigate these errors. In addition, a processing element (PE) is introduced to enable SWA in the systolic array architecture. Compared with the accumulation scheme devised for enhancing floating-point computing, the proposed SWA facilitates the integration into existing accelerator architectures, resulting in lower hardware overhead. The experimental results show that SWA allows LDC under low-precision configurations to achieve remarkable accuracy in LLMs, demonstrating higher hardware efficiency than merely increasing the precision of individual computations. Our method, while maintaining a lower hardware overhead than traditional LDC, achieves more than 13.9% improvement in average accuracy across multiple zero-shot benchmarks in LLAMA-2-7B. Furthermore, compared to integer domain computing, a logarithmic processing element array based on the proposed SWA yields reductions of 24.6% in area and 42.3% in power, while achieving higher accuracy. Xinkuang Geng, Yunjie Lu, Hui Wang 0023, Honglan Jiang |
DATE | 3 |
| 2025 | Low-Power Multiplier Designs by Leveraging Correlations of 2$\times$×2 Encoded Partial ProductsabstractMultipliers, particularly those with small bit widths, are essential for modern neural network (NN) applications. In addition, multiple-precision multipliers are in high demand for efficient NN accelerators; therefore, recursive multipliers used in low-precision fusion schemes are gaining increasing attention. In this work, we design exact recursive multipliers based on customized approximate full adders (AFAs) for low-power purposes. Initially, the partial products (PPs) encoded by 2×2 multiplications are analyzed, which reveals the correlations among adjacent PPs. Based on these correlations, we propose 4×4 recursive multiplier architectures where certain full adders (FAs) can be simplified without affecting the correctness of the multiplication. Manually and synthesis tool-based FA simplifications are performed separately. The obtained 4×4 multipliers are then used to construct 8×8 multipliers based on a low-power recursive architecture. Finally, the proposed signed and unsigned 4×4 and 8×8 multipliers are evaluated using a 28nm CMOS technology. Compared with DesignWare (DW) multipliers, the proposed signed and unsigned 4×4 multipliers achieve power reductions of 16.5% and 11.6%, respectively, without compromising area or delay; alternatively, the delay can be reduced by 20.9% and 39.4%, respectively, without compromising power or area. For signed and unsigned 8×8 multipliers, the maximum power reductions are 9.7% and 13.7%, respectively, albeit with a trade-off in area. Siting Liu 0001, Hui Wang 0023, Qin Wang 0009, Fabrizio Lombardi, Zhigang Mao, Honglan Jiang |
IEEE Trans. Computers | 3 |
| 2025 | A Reference Oversampling PLL With a FoMREF of -240.1 dB Enabled By a Capacitive Parasitic-Proof Ring Oscillator and a Time-Multiplexed Gm StageabstractThis paper presents a compact ring-oscillator (RO)-based phase-locked loop (PLL) implemented upon the principle that reference oversampling essentially boosts the reference frequency and thus extends the achievable bandwidth to such an extent that an area-efficient RO can be used without significantly sacrificing the phase noise (PN) and jitter performance when compared to conventional LC-based PLLs. By employing an analog reference oversampling PLL structure, RO noise is greatly suppressed by taking advantage of such an extended maximum PLL bandwidth. In conjunction with power- and spur-reduction techniques including a low-power time-multiplexed Gm stage and a capacitive parasitic-proof RO, this work implements a compact and low PN PLL without requiring complicated calibration or additional power/area penalties. Fabricated in a standard$0.18~\mu $m CMOS technology, the proposed PLL occupies an active area of 0.41 mm2. When operating at 1.6 GHz, the proposed PLL achieves an rms jitter of 585 fs with 5.7 mW power consumption, yielding a FoMREFof -240.1 dB. Xueke Cai, Tong Zhang 0030, Jianjun Zhou 0002, Howard Yang, Honglan Jiang, Yongfu Li 0002, Hui Wang 0023 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2024 | Design and Analysis of a Family of pW-Level Sub-1V CMOS VRGs by Stacking a Current-Source Transistor and a Resistive-Load TransistorabstractThis paper presents the design and analysis of a family of voltage reference generators (VRGs) based on the stacking of a current-source transistor MIand a resistive-load transistor MR, i.e., stacking of MI,R(SMIR) in standard CMOS technology for sub-1V and sub-nW operation. Design guidelines are provided to obtain the reference voltage for various temperature characteristics, namely proportional to absolute temperature (PTAT), complementary to absolute temperature (CTAT), and constant with temperature (CWT), by appropriately sizing the two transistors as current source and load respectively. The proposed 6 such SMIR VRGs in 65nm consume an average power less than 4pW and occupy an area less than 30µm × 100µm when measured from 20 different samples. All VRGs operate at a minimum supply voltage of 0.4V and achieve an average line regulation better than 0.3%/V. For the CWT VRGs, an average temperature coefficient better than 260.9ppm/°C is achieved from -20°C to 80°C. Tong Zhang 0030, Dingguo Zhang, Jing Jin 0005, Patrick P. Mercier, Hui Wang 0023 |
ISCAS | 5 |
| 2024 | A 0.83-pJ/b 20-Gb/s/Pin Single-Ended Transceiver With AC/DC-Coupled Pre-Emphasis FFE and Edge-Dependent Phase-Modulation DFE for Low-Power Memory ControllersabstractThis article presents an energy-efficient single-ended transceiver featuring the proposed AC/DC-coupled pre-emphasis feed-forward equalizer (PE-FFE) and edge-dependent phase-modulation decision feedback equalizer (PM-DFE) for low-power memory controllers. Specifically: 1) on the transmitter (Tx) side, an AC/DC-coupled PE-FFE is implemented in a ground-terminated Tx to minimize the equalization (EQ) power and maximize the output swing; 2) on the receiver (Rx) side, an edge-dependent PM-DFE operating on the full-swing signal edges is proposed for time-domain EQ, which improves sampling margin and reduces the linear EQ requirement as well as the power consumption. The design, fabricated in 22-nm CMOS, achieves a data rate of 20 Gb/s/pin with a 41 mV increase in the Tx output eye height and a 0.12 UI increase in the Rx sampling margin over a channel with a 10.3 dB loss. Measurement results reveal an energy efficiency of 0.45 pJ/b and 0.38 pJ/b for the Tx and the Rx, respectively, and a figure-of-merit of 0.081 pJ/b/dB (Tx+Rx). Jing Jin 0005, Xiaoming Liu 0008, Hui Wang 0023, Huzhi Tang, Yuekang Guo, Tingting Mo, Jianjun Zhou 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2017 | Silicon-Integrated High-Density Electrocortical InterfacesabstractRecent demand and initiatives in brain research have driven significant interest toward developing chronically implantable neural interface systems with high spatiotemporal resolution and spatial coverage extending to the whole brain. Electroencephalography-based systems are noninvasive and cost efficient in monitoring neural activity across the brain, but suffer from fundamental limitations in spatiotemporal resolution. On the other hand, neural spike and local field potential (LFP) monitoring with penetrating electrodes offer higher resolution, but are highly invasive and inadequate for long-term use in humans due to unreliability in long-term data recording and risk for infection and inflammation. Alternatively, electrocorticography (ECoG) promises a minimally invasive, chronically implantable neural interface with resolution and spatial coverage capabilities that, with future technology scaling, may meet the needs of recently proposed brain initiatives. In this paper, we discuss the challenges and state-of-the-art technologies that are enabling next-generation fully implantable high-density ECoG interfaces, including details on electrodes, data acquisition front-ends, stimulation drivers, and circuits and antennas for wireless communications and power delivery. Along with state-of-the-art implantable ECoG interface systems, we introduce a modular ECoG system concept based on a fully encapsulated neural interfacing acquisition chip (ENIAC). Multiple ENIACs can be placed across the cortical surface, enabling dense coverage over wide area with high spatiotemporal resolution. The circuit and system level details of ENIAC are presented, along with measurement results. Sohmyung Ha, Abraham Akinin, Jiwoong Park, Chul Kim, Hui Wang 0023, Christoph Maier, Patrick P. Mercier, Gert Cauwenberghs |
Proc. IEEE | 5 |
| 2016 | A 14.5 pW, 31 ppm/°C resistor-less 5 pA current reference employing a self-regulated push-pull voltage reference generatorabstractThis paper presents a gate-leakage-based supply- and temperature-stabilized current reference generator that can output currents as low as 5 pA with minimal power overhead. The output reference current is generated by driving a set of gate-leakage transistors designed to have opposing temperature coefficients with a stabilized voltage reference. Low-power operation is achieved by generating the voltage reference via a novel two-stage, 4T push-pull structure that can operate at a low supply voltage, and driving this reference to the gate-leakage transistors via a low-voltage self-biased amplifier. Designed in a 65 nm CMOS process, the proposed current reference generator is simulated to consume 14.5 pW at a 0.5 V supply voltage. Due to the push-pull structure and complementary gate-leakage transistors, the design achieves a temperature stability of 31 ppm/°C from 0 °C to 100 ° C, and a line sensitivity of 0.94%/V averaged across 500 Monte Carlo samples, thereby enabling an ultra-low-power, area-efficient, and temperature- and supply-stabilized current reference solution at pA-levels. Hui Wang 0023, Patrick P. Mercier |
ISCAS | 1 |
| 2012 | A novel overlapping coil structure for dual band telemetry systemabstractIn this paper, a novel overlapping coil structure for dual band power and data telemetry systems is proposed. Through carefully choosing the coil structural parameters, overlapping structure can achieve both low power interference and ease of manufacturing at the same time. Two design examples are presented and analyzed. Compared with conventional structures which set coils on the same plane, the overlapping structure shows higher power interference rejection and better overall signal quality. Peijun Wang, Yina Tang, Hui Wang 0023, Guoxing Wang |
ISCAS | 3 |
| 2012 | Anti-interference pseudo-differential wideband LNA for DVB-S.2 RF tunersabstractA novel pseudo-differential wideband Low Noise Amplifier (LNA) for DVB-S.2 Radio Frequency (RF) tuners is proposed. Based on narrow-band source degenerated structure, the proposed wideband LNA, covering the Digital Video Broadcast-Satellite.2 (DVB-S.2) band, demonstrates a higher gain and a lower Noise Figure (NF) than traditional wideband LNAs. Furthermore, a pseudo-differential topology is proposed to separate and cancel the in-band interferences. Designed and simulated in a 0.18 um CMOS, the proposed wideband LNA demonstrates an NF better than 2.5 dB and an input matching better than -10 dB with a gain higher than 18 dB over the operating band from 950 MHz to 2150 MHz. The LNA also achieves a typical 30 dB suppression of the interferences from the GSM signals. The simulated input-referred 3rd-order intercept point (IIP3) is 9.45 dBm and the total current consumption is 12 mA from a 1.8 V power supply. Hui Wang 0023, Wufeng Wang, Jing Jin 0005, Dongpo Chen, Jianjun Zhou 0002 |
ISCAS | 1 |
| 2004 | An adaptive motion estimation algorithm based on evolution strategiesabstractBased on evolution strategies (ESs), a novel adaptive motion estimation search algorithm (AESME) is presented. ESs consider evolutionary progress on the phenotype level. In contrast, genetic algorithms focus on heredity genetic mechanisms on the chromosome level. In ESs, the mutation operation accords with the normal distribution law. In the AESME algorithm, the (/spl mu/, /spl lambda/)-ES algorithm is adopted to block motion estimation, and the adaptive scheme is advanced to improve the convergence rate on the basis of the 1/5 success rule. Experimental results demonstrate that this algorithm has similar performance to that of the full-search (FS) algorithm, and owing to the inherent parallelism and low complexity of ESs, AESME is suitable for VLSI implementation. Hui Wang 0023, Zhigang Mao |
ICASSP (3) | 1 |
| 2004 | An adaptive motion estimation algorithm based on evolution strategies with correlated mutations
Hui Wang 0023, Zhigang Mao |
ICIP | 1 |