VLDB 2026 Research / reviewers in the wild / expert
Siqin Liu
dblp:97/10372
· DBLP profile ↗
10ranked-venue papers
7as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MONET: A Mixture-of-Experts Accelerator with a Multicast-Optimized Two-Tier Network-on-ChipabstractThe growing complexity of Mixture-of-Experts (MoE) models in machine learning applications demands innovative hardware solutions to address their unique computational and data movement challenges. Some of the critical challenges facing MoE models include sparse activation, dynamic token routing and irregular computation patterns that lead to low utilization and higher communication latency. In this paper, we introduce MONET, a novel two-tier Network-on-Chip (NoC) architecture designed to efficiently execute MoE workloads by co-optimizing compute, memory, and interconnect subsystems. The first tier consists of a reconfigurable systolic processing element (PE) island, executing both gating and expert computations, with runtime-configurable support for sparse/dense operations, expert reordering, and activation functions. The second tier incorporates a dual mesh network connecting a grid of PE islands; one network manages input token delivery with a broadcast scheme optimized for the gating phase of MoE, while the other is tailored for efficient inter-expert communication necessary for result aggregation. Evaluated on MoE benchmarks, MONET demonstrates up to 8.5× lower latency and over 6× better energy efficiency compared to state-of-the-art MoE accelerators. Siqin Liu, Maya Roediger, Avinash Karanth |
DATE | 1 |
| 2025 | Language semantics to support secure computation and communication in embedded systems via hardware monitorsabstractAs embedded systems with manycores and Network-on-Chips (NoCs) become ubiquitous, emerging hardware and software vulnerabilities have made it challenging to ensure system integrity especially when third-party intellectual property (IP) is used for rapid prototyping. Prior works have evaluated hardware monitors for ensuring correctness of the system by threat assessment and effective mitigation. However, none have evaluated models that combine both computation (processor pipeline) and communication (NoC) vulnerabilities simultaneously. In this paper, we propose a high-level policy language called d-GUARD that is used to define runtime security policies that can be compiled into hardware monitors. The advantage of this new language is the ability to dynamically change policies based on program’s runtime behavior. To translate high-level policies into low-level hardware monitors, we describe a compiler for d-GUARD that synthesizes policies into Verilog modules. Instead of simply evaluating the design of secure policies for processor pipelines, we extend to secure NoC microarchitectures, including policies for links and routers, as well as policies to prevent Denial-of-Service (DoS) attacks. To mitigate attacks against secure microarchitectures, we also propose fault-tolerant routing approaches to avoid rogue routers when the number of policy violations exceeds a certain threshold. Our secure policies for processor pipelines and NoC microarchitectures consume marginal area and power overhead when compared to baseline making it well suited for low-cost embedded systems. • A high-level security policy language called D-GUARD, which is expressive, modular and compatible with power-efficient hardware is proposed. • The proposed dynamic policies implemented at various levels of the embedded enable low-cost monitors that can be modified based on application demands. • The runtime costs of the processor pipeline using OptimSoC, an open-source implementation of OpenRISC1000 architecture where processor policies are evaluated, which showed less than 1% overhead in power and area when implemented in 14 nm and 45 nm technologies and no performance overhead when implemented on BEEBS benchmark suite. Garett Cunningham, Siqin Liu, Harsha Chenji, David W. Juedes, Avinash Karanth |
Integr. | 2 |
| 2025 | GreeNX: An Energy-Efficient and Sustainable Approach to Sparse Graph Convolution Networks Accelerators Using DVFSabstractGraph convolutional networks (GCNs) have emerged as an effective approach to extend deep learning algorithms for graph-based data analytics. However, GCNs implementation over large, sparse datasets presents challenges due to irregular computation and dataflow patterns. Specialized GCN accelerators have emerged to deliver superior performance over generic processors. However, prior techniques that include specialized datapaths, optimized sparse computation, and memory access patterns, handle different phases of GCNs differently which results in excess energy consumption and reduced throughput due to sub-optimal dataflows. In this paper, we propose GreeNX, a computation and communication-aware GCN accelerator that uniformly applies three complementary techniques to all phases of GCN. First, we abstract two cascaded sparse-dense matrix multiplications that uniformly process the computation in both aggregation and combination phases of GCNs to improve throughput. Second, to mitigate the overheads of processing irregular sparse data, we develop a dynamic-voltage-and-frequency-scaling (DVFS) scheme by grouping a row of processing elements (PEs) that dynamically changes the applied voltage/frequency (V/F) to improve energyefficiency. Third, we conduct a comprehensive carbon footprint evaluation, analyzing both embodied and operational emissions for GCNs. Extensive simulation and experiments validate that our GreeNX consistently reduces memory accesses and energy consumption leading to an average 7.3× speedup and 5.6× energy savings on six real-world graph datasets over several state-of-the-art GCN accelerators including HyGCN, AWB-GCN, GCoD, GRIP, IGCN, and LW-GCN Siqin Liu, Prakash Chand Kuve, Avinash Karanth |
IEEE Trans. Sustain. Comput. | 1 |
| 2024 | SNAC: Mitigation of Snoop-Based Attacks with Multi-Tier Security in NoC ArchitecturesabstractNetwork-on-chips (NoCs) are crucial for multicore and manycore System-on-Chip (SoC) architectures. However, the integration of third-party Intellectual Property (IP) cores in SoCs has introduced hardware vulnerabilities. Snoop-based attacks exploit these vulnerabilities by inserting malicious Hardware Trojans into routers, allowing them to extract sensitive information as packets traverse the NoC. To address these security concerns, we propose SNAC: Mitigation of Snoop-based Attacks in NoCs. SNAC employs a three-tier architecture with increasing security levels, each with proportional power and latency overheads. The first tier introduces path randomization to prevent attackers from predicting packet routes. In the second tier, we encrypt source and destination information using lightweight backward XoR encryption. The third tier combines techniques from tiers one and two, extending obfuscation along with path randomization. SNAC was evaluated using synthetic and real-world benchmarks. Our results show that SNAC incurs dynamic power overheads of 4.2%, 3.9%, and 6.1% for Tiers 1, 2, and 3 respectively, with area overheads of 6.2%, 4.2%, and 9.2%. Siqin Liu, Saumya Chauhan, Avinash Karanth |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | HSCONN: Hardware-Software Co-Optimization of Self-Attention Neural Networks for Large Language ModelsabstractSelf-attention models excel in natural language processing and computer vision by capturing contextual information but encounter several challenges such as efficient data movement, quadratic computational complexity, and excessive memory accesses. Sparse attention techniques emerge as a solution, however, their irregular or regular patterns, coupled with costly data pre-processing, diminish their hardware efficiency. This paper introduces HSCONN, an energy-efficient hardware accelerator for self-attention, mitigating computational and memory overheads. HSCONN employs dynamic voltage and frequency scaling (DVFS) along with exploiting dynamic sparsity in matrix multiplication, thereby optimizing energy efficiency. The approach includes a row-wise pruning algorithm and independent voltage/frequency islands for processing elements, exploiting additional sparsity to reduce memory access and overall energy consumption. Experiments in natural language processing showcase HSCONN’s remarkable speedups (1952 ×, 615 ×) and energy reductions (up to 820 ×, 113 ×) over CPU and GPU architectures. Compared to A3, SpAtten, and Sanger, HSCONN demonstrates superior speedup (1.71 ×, 1.25 ×, 1.47 ×) and higher energy efficiency (1.5 ×, 1.7 ×, 1.4 ×). Siqin Liu, Prakash Chand Kuve, Avinash Karanth |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | The Modeling of 3-D DC Resistivity Based on Integral Equation of PotentialabstractWe perform a 2D Fourier transform along the horizontal direction on the 3D integration problem for the DC anomalous potential in the spatial domain, transforming it into a 1D integration problem with independent solutions at different wavenumbers. And the resulting 1D integral equation in the wavenumber domain is expressed as a sum of element integrals. We then employ the shape function method, using a quadratic shape function to characterize the scattering current density for each element and compute analytical expressions for the element integrals. Finally, we performed a 2D inverse Fourier transform of the wavenumber-domain anomalous potential and electric field to obtain their values in the spatial domain, which are modified using iterative operators. This approach fully leverages the efficiency of the Fourier transform method and the high accuracy of the shape function integration method to achieve faster and more accurate solutions to the 3D DC resistivity numerical simulation problem. Two examples demonstrate the accuracy and efficiency of our proposed algorithm. Numerical examples were used to analyze the relationship between the number of iterations and the anomaly conductivity difference during the convergence of the algorithm. It is shown that the algorithm converges faster for high resistance anomalies and can be applied to simulate models with significantly different resistivities between the background medium and the anomaly. Jiaxuan Ling, Shiwei Wei, Siqin Liu, Shuliu Wei, Lihua He |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Exploiting Wireless Technology for Energy-Efficient Accelerators With Multiple Dataflows and PrecisionabstractAs model size and the number of layers increase, Deep Neural Networks (DNNs) demand enormous computational power and throughput to meet exceedingly high prediction accuracy’s of today’s machine learning (ML) applications. Spatial hardware accelerators have been proposed that optimize the dataflow and exploit sparsity to provide a significant decrease in power consumption. As spatial architectures are traditionally designed with metallic interconnects, significant power is expended for data movement for different dataflows. In this paper, we exploit extended wireless technology to design a power-efficient and high-throughput DNN accelerator, e-WiNN, that can be configured for all representative dataflows and arithmetic precisions. We leverage novel circuit design by utilizing Dadda-algorithm based Multiply-and-Accumulate (MAC) circuits for 4-bit, 8-bit and 16-bit inputs to reduce area, power and delay constraints in 14 nm predictive technology. Our novel wireless transmitter integrates on- off keying (OOK) modulator with power amplifier that results in significant energy savings. To reduce the area overhead, we cluster wireless transceivers into groups of four such that both weights and input features can be effectively multicast to reduce the data movement. The energy efficient transceiver circuit is implemented in state-of-the-art BSIM 32 nm FinFET technology model and our link budget considers required RF power for different frequencies and inter-PE distance at three different antenna directivities including isotropic. Our detailed RTL modeling and cycle-accurate simulation results show that e-WiNN achieves 36.3% latency reduction and 76.1% energy saving when compared to state-of-art wire interconnected accelerators; 70.3% area reduction and 41.6% energy saving at the cost of 11% latency increase when compared to prior wireless accelerators on various neural networks (AlexNet, VGG16, and ResNet-9/50). Siqin Liu, Talha Furkan Canan, Harsha Chenji, Soumyasanta Laha, Savas Kaya, Avinash Karanth |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Dynamic Voltage and Frequency Scaling to Improve Energy-Efficiency of Hardware AcceleratorsabstractNeural networks (NNs) have been used in a wide variety of artificial intelligence (AI) applications, including speech recognition, image recognition, automatic robotics, and games. State-of-the-art NNs provide high prediction accuracy at the expense of massive computation that involves large model parameters which consume substantial energy. Though sparse NN s have emerged to reduce the computation and storage overhead, existing specialized DNN accelerators cannot maximize the energy savings when exploiting both dynamic and static sparsity, especially for irregular NNs. In this paper, we propose a dynamic voltage and frequency scaling (DVFS) based hardware accelerator that effectively exploits the dynamic and static sparsity of NNs with dynamic voltage/frequency (V/F) scaling and power gating techniques to reduce both static and dynamic power. To explore the efficiency of DVFS implementation at different granularities, we evaluate both coarse-grained and fine-grained DVFS implementation with different design trade-offs. Further, our proposed DVFS model predicts the dynamic computation workloads as well as V/ F pairs to be supplied to processing elements (PEs) in the hardware intelligently through pre-trained weight vectors. The machine learning based prediction algorithm is deployed to improve the DVFS mode selection accuracy. Our simulation results on AlextNet, VGG16, and ResNet50 show that we can achieve an average dynamic energy savings of 59–66 % and an average static power reduction of 69–80 % compared to the baseline. Siqin Liu, Avinash Karanth |
HiPC | 1 |
| 2021 | WiNN: Wireless Interconnect based Neural Network AcceleratorabstractDeep Neural Networks (DNNs) have demonstrated promising performance in accuracy for several applications such as image processing, speech recognition, and autonomous systems and vehicles. Spatial accelerators have been proposed to achieve high parallelism with arrays of processing elements (PE) and energy efficient data movement using traditional Network-on-Chip (NoC) architectures. However, larger DNN models impose high bandwidth and low latency communication demands between PEs, which is a fundamental challenge for metallic NoC architectures. In this paper, we propose WiNN, a wireless and wired interconnected neural network accelerator that employs on-chip wireless links to provide high network bandwidth and single cycle multicast communication. We design separate wireless networks modulated with two different frequency bands one each for the weights and input Highly directional antennas are implemented to avoid noise and interference. We propose multicast-for-wireless (MW) dataflow for our proposed accelerator that efficiently exploits the wireless channels’ multicast capabilities to reduce the communication overheads. Our novel wireless transmitter integrates on-off keying (OOK) modulator with power amplifier that results in significant energy savings. Our simulation results show that WiNN achieves 74% latency reduction and 37.5% energy saving when compared to state-of-art metallic link-based accelerators, 38.1% latency reduction and 19.4% energy saving when compared to prior wireless accelerators for various neural networks (AlexNet, VGG16, and ResNet-50). Siqin Liu, Sushanth Karmunchi, Avinash Karanth, Soumyasanta Laha, Savas Kaya |
ICCD | 1 |
| 2011 | Equivalent Sampling Oscilloscope with External Delay Embedded SystemabstractThe internal delay methods for sequential equivalent sampling are widely used in the digital storage oscilloscope and limited completely by maximum operating frequency of embedded system. A new external delay technology for equivalent sampling oscilloscope is presented in this paper. The technology is based on the external programmable delay chips, which provide much shorter delay time for equivalent sampling rate as well as much higher operating frequency. With real time sampling and equivalent time sampling, the embedded system could perform rapid and effective measurement for the periodic signal which is either fast-varying or slow-varying and also be able to take real time sampling to the single-shot signals. Jingzhu Yang, Siqin Liu, Chunsheng Zhu, Fei Hao 0001 |
HPCC | 2 |