EDBT 2026 Demo / reviewers in the wild / expert
Kaiyuan Yang 0001
dblp:81/11242
· DBLP profile ↗
19ranked-venue papers
1as first author
13since 2021 · last 2025
0000-0001-7220-9389ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 7 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Computer networks · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RAP: Reconfigurable Automata ProcessorabstractRegular pattern matching is essential for applications such as text processing, malware detection, network security, and bioinformatics.Recent in-memory automata processors have significantly advanced the energy and memory efficiency over conventional computing platforms.However, these processors are typically optimized only for one type of automata, limiting their capability to efficiently support regex processing under diverse real-world workloads.This paper presents RAP, the first reconfigurable in-memory automata processor for efficient regular pattern matching across diverse workloads.It supports Nondeterministic Finite Automata (NFA), Nondeterministic Bit Vector Automata (NBVA), and Linear NFA (LNFA) through reconfigurable architecture and circuit designs, and a compiler for translation.RAP is evaluated in 28nm CMOS PDK, achieving 1.2-1.5×higher energy efficiency and 1.3-2.5×higher compute density compared to SotA automata processors for NFA (CA and CAMA) over diverse real-world benchmarks.It also achieves 1.6× higher compute density and similar energy efficiency as BVAP, a SotA optimized for bounded repetitions.Finally, RAP is >100× and >1000× more energy efficient than SotA GPU and CPU solutions. Ziyuan Wen, Alexis Le Glaunec, Konstantinos Mamouras, Kaiyuan Yang 0001 |
ISCA | 4 |
| 2025 | Dual-Resonance Magnetoelectric Power and Data Links for Miniaturized Wireless Bio-ImplantsabstractMiniature, battery-free implants promise transformative bio-electronic therapies by enabling minimally invasive implantation procedures, reducing risk, and extending device lifetime. Among all wireless power and data transfer (WPDT) modalities, magnetoelectrics (ME) has emerged as a particularly promising solution for millimeter-scale implants, boasting lower tissue attenuation and higher power transfer efficiency over conventional inductive and ultrasonic methods. However, as an acoustic resonator, ME devices face an inherent tradeoff between Q-factor and bandwidth, limiting their ability to simultaneously achieve high-speed communication and efficient wireless power transfer (WPT). To fundamentally circumvent the challenge, this paper presents dual-resonance ME WPDT that exploits the unique multimode resonances of ME transducers to realize WPT and communication at distinct frequencies. Based on this dual-resonance principle, we demonstrate reconfigurable active and passive schemes for different biomedical applications, with a proof-of-concept system including a miniature implant and an external transceiver. The active scheme achieves 60 kbps at operational distances of 6 cm with 2.5 mW implant power, while the passive backscatter offers 20 kbps continuous streaming at 4 cm with negligible power, demonstrating the first reported non-interrupted ME WPDT system and more than twice the data rate of previous ME backscatter methods. Both schemes further support on-off keying (OOK) and binary phase shift keying (BPSK) modulations, providing additional flexibility to tailor communication needs between power efficiency and robustness. The complete prototype system was validated through comprehensive in-vitro experiments and in-vivo EMG streaming in a rodent model. Wei Wang 0433, Ellie C. Chen, Naveed H. Ahmed, Wonjune Kim, Yiwei Zou, Joshua E. Woods, Yumin Su, Huan-Cheng Liao, Jacob T. Robinson, Kaiyuan Yang 0001 |
MobiCom | 10 |
| 2025 | MBSNTT: A Highly Parallel Digital In-Memory Bit-Serial Number Theoretic Transform AcceleratorabstractConventional cryptographic systems protect the data security during communication but give third-party cloud operators complete access to compute decrypted user data. Homomorphic encryption (HE) promises to rectify this and allow computations on encrypted data to be done without actually decrypting it. However, HE encryption requires several orders of magnitude higher latency than conventional encryption schemes. Number theoretic transform (NTT), a polynomial multiplication algorithm, is the bottleneck function in HE. In traditional architectures, memory accesses and support for parallel operations limit NTT’s throughput and energy efficiency. Processing in memory (PIM) is an interesting approach that can maximize parallelism with high-energy efficiency. To enable HE on resource-constrained edge devices, this article presents MBSNTT, a digital in-memory Multi-Bit-Serial NTT accelerator, achieving high parallelism and energy efficiency for NTT with minimized area. MBSNTT features a novel multi-bit-serial modular multiplication algorithm and PIM implementation that computes all modular multiplications in an NTT in parallel. It further adopts a constant geometry NTT data flow for efficient transition between NTT stages and different cores. Our evaluation shows that MBSNTT achieves$1.62\times $($19.08\times $) higher throughput and$64.9\times $($2.06\times $) lower energy than state-of-the-art PIM NTT accelerators Crypto-PIM (MeNTT), at a polynomial order of 8 K and bit width of 128. Akhil Reddy Pakala, Zhiyu Chen 0003, Kaiyuan Yang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | BVAP: Energy and Memory Efficient Automata Processing for Regular Expressions with Bounded RepetitionsabstractRegular pattern matching is pervasive in applications such as text processing, malware detection, network security, and bioinformatics. Recent studies have demonstrated specialized in-memory automata processors with superior energy and memory efficiencies than existing computing platforms. Yet, they lack efficient support for the construct of bounded repetition that is widely used in regular expressions (regexes). This paper presents BVAP, a software-hardware co-designed in-memory Bit Vector Automata Processor. It is enabled by a novel theoretical model called Action-Homogeneous Non-deterministic Bit Vector Automata (AH-NBVA), its efficient hardware implementation, and a compiler that translates regexes into hardware configurations. BVAP is evaluated with a cycle-accurate simulator in a 28nm CMOS process, achieving 67-95% higher energy efficiency and 42-68% lower area, compared to state-of-the-art automata processors (CA, eAP, and CAMA), across a set of real-world benchmarks. Ziyuan Wen, Lingkun Kong, Alexis Le Glaunec, Konstantinos Mamouras, Kaiyuan Yang 0001 |
ASPLOS (2) | 5 |
| 2024 | Lightweight Machine Learning and Embedded Security Engine for Physical-Layer Identification of Wireless IoT NodesabstractSecuring low-power Internet-of- Things (IoT) sensor nodes presents a critical challenge for the widespread adoption of IoT technology, given their inherent limitations in energy, computation, and storage resources. As a promising alternative to conventional wireless security approaches based on cryptography, there has been a growing interest in RF physical-layer security, especially RF fingerprinting, which offers the promise of reduced overhead and energy consumption. In this work, we present an artificial neural network (ANN) model tailored to identify IoT transmitters by harnessing their unique power spectral density (PSD). The network is designed to be lightweight and can be readily implemented on resource-constrained IoT nodes. Combined with our customized radio frontend, we achieve superior identification performance. In the measurements, we can reliably identify 240 devices with a 99 % accuracy on trained distances and 40 devices with an above 95 % accuracy at an unknown distance that is excluded from the training data. These results demonstrate significant improvement in robustness, reliability, and identification accuracy over prior art while ensuring compatibility with resource-constrained IoT nodes. Qiufeng Rui, Noah Elzner, Qiang Zhou 0012, Ziyuan Wen, Yan He 0002, Kaiyuan Yang 0001, Taiyun Chi |
ICC | 6 |
| 2023 | CASA: An Energy-Efficient and High-Speed CAM-based SMEM Seeding Accelerator for Genome AlignmentabstractGenome analysis is a critical tool in medical and bioscience research, clinical diagnostics and treatment, and disease control and prevention. Seed and extension-based alignment is the main approach in the genome analysis pipeline, and BWA-MEM2, a widely acknowledged tool for genome alignment, performs seeding by searching for super maximal exact match (SMEM). The computation of SMEM searching requires high memory bandwidth and energy consumption, which becomes the main performance bottleneck in BWA-MEM2. State-of-the-Art designs like ERT and GenAx have achieved impressive speed-ups of SMEM-based genome alignment. However, they are constrained by frequent DRAM fetches or computationally intensive intersection calculations for all possible k-mers at every read position. Yi Huang 0036, Lingkun Kong, Dibei Chen, Zhiyu Chen 0003, Jianfeng Zhu 0001, Konstantinos Mamouras, Shaojun Wei, Kaiyuan Yang 0001, Leibo Liu |
MICRO | 9 |
| 2022 | CAMA: Energy and Memory Efficient Automata Processing in Content-Addressable MemoriesabstractAccelerating finite automata processing is critical for advancing real-time analytic in pattern matching, data mining, bioinformatics, intrusion detection, and machine learning. Recent in-memory automata accelerators leveraging SRAMs and DRAMs have shown exciting improvements over conventional digital designs. However, the bit-vector representation of state transitions used by all state-of-the-art (SOTA) designs is only optimal in processing worst-case completely random patterns, while a significant amount of memory and energy is wasted in running most real-world benchmarks.We present CAMA, a Content-Addressable Memory (CAM) enabled Automata accelerator for processing homogeneous non-deterministic finite automata (NFA). A radically different state representation scheme, along with co-designed novel circuits and data encoding schemes, greatly reduces energy, memory, and chip area for most realistic NFAs. CAMA is holistically optimized with the following major contributions: (1) a 16 × 256 8-transistor (8T) CAM array for state matching, replacing the 256 × 256 6T SRAM array or two 16×256 6T SRAM banks in state-of-the-art (SOTA) designs; (2) a novel encoding scheme that enables content searching within 8T SRAMs and adapts to different applications; (3) a reconfigurable and scalable architecture that improves efficiency on all tested benchmarks, without losing support for any NFA that’s compatible with SOTA designs; (4) an optimization framework that automates the choice of encoding schemes and maps a given NFA to the proposed hardware.Two versions of CAMA, one optimized for energy (CAMA-E) and the other for throughput (CAMA-T), are comprehensively evaluated in a 28nm CMOS process, and across 21 real-world and synthetic benchmarks. CAMA-E achieves 2.1×, 2.8 ×, and 2.04× lower energy than CA, 2-stride Impala, and eAP. CAMA-T shows 2.68×, 3.87× and 2.62 × higher average compute density than 2-stride Impala, CA, and eAP. Both versions reduce the chip area required for the largest tested benchmark by 2.48× over CA, 1.91× over 2-stride Impala, and 1.78× over eAP. Yi Huang 0036, Zhiyu Chen 0003, Dai Li, Kaiyuan Yang 0001 |
HPCA | 4 |
| 2022 | F8Net: Fixed-Point 8-bit Only Multiplication for Network Quantization
Qing Jin, Jian Ren 0005, Richard Zhuang, Sumant Hanumante, Zhengang Li 0001, Zhiyu Chen 0003, Yanzhi Wang 0001, Kaiyuan Yang 0001, Sergey Tulyakov |
ICLR | 8 |
| 2022 | Magnetoelectric backscatter communication for millimeter-sized wireless biomedical implantsabstractThis paper presents the design, implementation, and experimental evaluation of a wireless biomedical implant platform exploiting the magnetoelectric effect for wireless power and bi-directional communication. As an emerging wireless power transfer method, magnetoelectric is promising for mm-scaled bio-implants because of its superior misalignment sensitivity, high efficiency, and low tissue absorption compared to other modalities [46, 59, 60]. Utilizing the same physical mechanism for power and communication is critical for implant miniaturization, but low-power magnetoelectric uplink communication has not been achieved yet. For the first time, we design and demonstrate near-zero power magnetoelectric backscatter from the mm-sized implants by exploiting the converse magnetostriction effects. Zhanghao Yu, Fatima T. Alrashdan, Wei Wang 0433, Matthew Parker, Frank Y. Chen, Joshua E. Woods, Zhiyu Chen 0003, Jacob T. Robinson, Kaiyuan Yang 0001 |
MobiCom | 10 |
| 2022 | Software-hardware codesign for efficient in-memory regular pattern matchingabstractRegular pattern matching is used in numerous application domains, including text processing, bioinformatics, and network security. Patterns are typically expressed with an extended syntax of regular expressions. This syntax includes the computationally challenging construct of bounded repetition or counting, which describes the repetition of a pattern a fixed number of times. We develop a specialized in-memory hardware architecture that integrates counter and bit vector modules into a state-of-the-art in-memory NFA accelerator. The design is inspired by the theoretical model of nondeterministic counter automata (NCA). A key feature of our approach is that we statically analyze regular expressions to determine bounds on the amount of memory needed for the occurrences of bounded repetition. The results of this analysis are used by a regex-to-hardware compiler in order to make an appropriate selection of counter or bit vector modules. We evaluate our hardware implementation using a simulator based on circuit parameters collected by SPICE simulation in TSMC 28nm CMOS process. We find that the use of counter and bit vector modules outperforms unfolding solutions by orders of magnitude. Experiments concerning realistic workloads show up to 76% energy reduction and 58% area reduction in comparison to CAMA, a recently proposed in-memory NFA accelerator. Lingkun Kong, Qixuan Yu 0001, Agnishom Chattopadhyay, Alexis Le Glaunec, Yi Huang 0036, Konstantinos Mamouras, Kaiyuan Yang 0001 |
PLDI | 7 |
| 2022 | MeNTT: A Compact and Efficient Processing-in-Memory Number Theoretic Transform (NTT) AcceleratorabstractLattice-based cryptography (LBC), exploiting learning with error (LWE) problems, is a promising candidate for postquantum cryptography. The number theoretic transform (NTT) is the latency- and energy-dominant process in the computation of LWE problems. This article presents a compact and efficient in-MEmory NTT accelerator, named MeNTT, which explores an optimized computation in and near a 6T SRAM array. Specifically designed peripherals enable fast and efficient modular operations. Moreover, a novel mapping strategy reduces the data flow between the NTT stages into a unique pattern, which greatly simplifies the routing among processing units (i.e., the SRAM column in this work), reducing the energy and area overheads. The accelerator achieves significant latency and energy reductions over prior arts. Dai Li, Akhil Reddy Pakala, Kaiyuan Yang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2021 | NPAS: A Compiler-Aware Framework of Unified Network Pruning and Architecture Search for Beyond Real-Time Mobile AccelerationabstractWith the increasing demand to efficiently deploy DNNs on mobile edge devices, it becomes much more important to reduce unnecessary computation and increase the execution speed. Prior methods towards this goal, including model compression and network architecture search (NAS), are largely performed independently, and do not fully consider compiler-level optimizations which is a must-do for mobile acceleration. In this work, we first propose (i) a general category of fine-grained structured pruning applicable to various DNN layers, and (ii) a comprehensive, compiler automatic code generation framework supporting different DNNs and different pruning schemes, which bridge the gap of model compression and NAS. We further propose NPAS, a compiler-aware unified network pruning and architecture search. To deal with large search space, we propose a meta-modeling procedure based on reinforcement learning with fast evaluation and Bayesian optimization, ensuring the total number of training epochs comparable with representative NAS frameworks. Our framework achieves 6.7ms, 5.9ms, and 3.9ms ImageNet inference times with 78.2%, 75% (MobileNet-V3 level), and 71% (MobileNet-V2 level) Top-1 accuracy respectively on an off-the-shelf mobile phone, consistently outperforming prior work. Zhengang Li 0001, Geng Yuan, Wei Niu 0002, Pu Zhao 0001, Yanyu Li, Yuxuan Cai 0001, Xuan Shen, Zheng Zhan 0001, Zhenglun Kong, Qing Jin, Zhiyu Chen 0003, Sijia Liu 0001, Kaiyuan Yang 0001, Bin Ren 0002, Yanzhi Wang 0001, Xue Lin 0001 |
CVPR | 13 |
| 2021 | MC2-RAM: An In-8T-SRAM Computing Macro Featuring Multi-Bit Charge-Domain Computing and ADC-Reduction Weight EncodingabstractIn-memory computing (IMC) is a promising hardware architecture to circumvent the memory walls in data-intensive applications, like deep learning. Among various memory technologies, static random-access memory (SRAM) is promising thanks to its high computing accuracy, reliability, and scalability to advanced technology nodes. This paper presents a novel multi-bit capacitive convolution in-SRAM computing macro for high accuracy, high throughput and high efficiency deep learning inference. It realizes fully parallel charge-domain multiply-and-accumulate (MAC) within compact 8-transistor 1-capacitor (8T1C) SRAM arrays that is only 41% larger than the standard 6T cells. It performs MAC with multi-bit activations without conventional digital bit-serial shift-and-add schemes, leading to drastically improved throughput for high-precision CNN models. An ADC-reduction encoding scheme complements the compact sram design, by reducing the number of needed ADCs by half for energy and area savings. A $576 \times 130$ macro with 64 ADCs is evaluated in 65nm with post-layout simulations, showing 4.60 TOPS/mm2compute density and 59.7 TOPS/W energy efficiency with 4/4-bit activations/weights. The MC2- RAM also achieves excellent linearity with only 0.14 mV (4.5% of the LSB) standard deviation of the output voltage in Monte Carlo simulations. Zhiyu Chen 0003, Qing Jin, Jingyu Wang 0003, Yanzhi Wang 0001, Kaiyuan Yang 0001 |
ISLPED | 5 |
| 2019 | IoT2 - the Internet of Tiny Things: Realizing mm-Scale Sensors through 3D Die StackingabstractThe Internet of Things (IoT) is a rapidly evolving application space. One of the fascinating new fields in IoT research is mm-scale sensors, which make up the Internet of Tiny Things (IoT2). With their miniature size, these systems are poised to open up a myriad of new application domains. Enabled by the unique characteristics of cyber-physical systems and recent advances in low-power design and bare-die 3D chip stacking, mm-scale sensors are rapidly becoming a reality. In this paper, we will survey the challenges and solutions to 3D-stacked mm-scale design, highlighting low-power circuit issues ranging from low-power SRAM and miniature neural network accelerators to radio communication protocols and analog interfaces. We will discuss system-level challenges and illustrate several complete systems and their merging application spaces. Sechang Oh 0001, Minchang Cho, Xiao Wu 0002, Yejoong Kim, Li-Xuan Chuo, Wootaek Lim, Pat Pannuto, Suyoung Bang, Kaiyuan Yang 0001, Hun-Seok Kim, Dennis Sylvester, David T. Blaauw |
DATE | 9 |
| 2017 | Rectified-linear and recurrent neural networks built with spin devicesabstractA spin synapse with analog programmability using all charge current is proposed. Compared with using spin current, the proposed all-charge-current synapses can be placed in a larger cross-bar array to form a denser and larger neural network. Using the current summation, DOT product can be realized. We further employ a compact racetrack converter as the neuron to implement a rectified-linear neural network, saving area by 67% and energy by 69% compared with a spin binary-threshold neural network while achieving similar accuracy with MNIST digit recognition benchmark. Storing the domain wall motion in a time-based fashion, a recurrent neural network can be realized for time-involved inference tasks. Qing Dong 0001, Kaiyuan Yang 0001, Laura Fick, David T. Blaauw, Dennis Sylvester |
ISCAS | 2 |
| 2017 | Low-Power and Compact Analog-to-Digital Converter Using Spintronic Racetrack Memory DevicesabstractCurrent-induced domain wall (DW) motion in spintronic racetrack memory promises energy-efficient analog computation using compact magnetic nanowires. This paper explores the feasibility of analog-to-digital converter (ADC) based on current-induced DW motion and introduces an n-bit ADC using n racetrack magnetic nanowires. With each magnetic nanowire having a different configuration granularity, an n-bit binary or gray code is generated simultaneously. The proposed ADC structure achieves 21 fJ/conversion-step at 20 MHz with an area of about 10 μm2. The racetrack ADC is suitable for applications requiring dense ADC arrays, such as image sensors. This paper describes one ultrahigh speed digital pixel sensor imaging system benefiting from the racetrack ADC. Qing Dong 0001, Kaiyuan Yang 0001, Laura Fick, David Fick, David T. Blaauw, Dennis Sylvester |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | A2: Analog Malicious HardwareabstractWhile the move to smaller transistors has been a boon for performance it has dramatically increased the cost to fabricate chips using those smaller transistors. This forces the vast majority of chip design companies to trust a third party -- often overseas -- to fabricate their design. To guard against shipping chips with errors (intentional or otherwise) chip design companies rely on post-fabrication testing. Unfortunately, this type of testing leaves the door open to malicious modifications since attackers can craft attack triggers requiring a sequence of unlikely events, which will never be encountered by even the most diligent tester. In this paper, we show how a fabrication-time attacker can leverage analog circuits to create a hardware attack that is small (i.e., requires as little as one gate) and stealthy (i.e., requires an unlikely trigger sequence before effecting a chip's functionality). In the open spaces of an already placed and routed design, we construct a circuit that uses capacitors to siphon charge from nearby wires as they transition between digital values. When the capacitors fully charge, they deploy an attack that forces a victim flip-flop to a desired value. We weaponize this attack into a remotely-controllable privilege escalation by attaching the capacitor to a wire controllable and by selecting a victim flip-flop that holds the privilege bit for our processor. We implement this attack in an OR1200 processor and fabricate a chip. Experimental results show that our attacks work, show that our attacks elude activation by a diverse set of benchmarks, and suggest that our attacks evade known defenses. Kaiyuan Yang 0001, Matthew Hicks, Qing Dong 0001, Todd M. Austin, Dennis Sylvester |
IEEE Symposium on Security and Privacy | 1 |
| 2015 | Racetrack converter: A low power and compact data converter using racetrack spintronic devicesabstractCurrent-induced domain wall motion in racetrack memory promises energy-efficient analog computation using compact magnetic nanowires. This paper explores the feasibility of data converters based on current-induced domain wall motion and introduces an n-bit ADC using n racetrack magnetic nanowires. With each magnetic nanowire having a different configuration granularity, an n-bit binary or gray code is generated simultaneously. The proposed ADC structure achieves 21fJ/conversion-step at 20MHz with area under 10μm2. The racetrack ADC is suitable for applications requiring dense ADC arrays, such as image sensors. This paper describes one ultra-high speed DPS imaging system benefiting from the racetrack ADC. Qing Dong 0001, Kaiyuan Yang 0001, Laura Fick, David Fick, David T. Blaauw, Dennis Sylvester |
ISCAS | 2 |
| 2012 | A transformer-based filtering technique to lower LC-oscillator phase noiseabstractPhase noise of oscillators can be improved through noise filtering. This paper presents a transformer-based harmonic filtering technique to improve the phase noise performance of the fully differential CMOS LC oscillators. The primary and secondary of the transformer are placed at the sources of NMOS and PMOS differential pairs, respectively. Due to the extra degree of freedom introduced by the transformer, the filter can resonate at two different frequencies, which are designed to be the second and fourth harmonics. The analysis of the filtering method is presented and principles of operation are discussed in details. The performance is compared with other filtering techniques, and the simulation results show improved FoM compared to published literature. Qing Jin, Kaiyuan Yang 0001, Chunyuan Zhou, Lei Zhang 0033, Yan Wang 0023, Zhiping Yu, Weidong Geng |
ISCAS | 2 |