EDBT 2026 Demo / reviewers in the wild / expert
Yuyao Kong
dblp:237/1003
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Event-Centric Structural Modeling for Zero-Shot Video Moment RetrievalabstractZero-Shot Video Moment Retrieval (VMR) aims to localize a specific temporal segment in an untrimmed video that corresponds to a natural language query by leveraging frozen vision-language models, eliminating the need for costly temporal annotations. Fundamentally, an untrimmed video consists of a sequence of atomic events with varying lengths. A critical challenge in VMR is thus to disentangle these events into coherent candidates and accurately identify the one that best aligns with the query. However, conventional methods often overlook this inherent event structure: rigid sliding windows tend to fragment semantically coherent events, while uniform pooling across frames allows ambiguous boundary noise to dilute the relevance of the discriminative core. To mitigate these issues, we propose an Event-Centric Structural Modeling (ECSM) framework. Specifically, our approach replaces rigid windowing with adaptive global segmentation to preserve event integrity. Furthermore, we introduce Gaussian-based weighting to highlight the discriminative core of events while suppressing boundary interference. Extensive experiments on Charades-STA and ActivityNet Captions demonstrate that our method achieves state-of-the-art performance in the training-free setting and exhibits robust generalization across diverse out-of-distribution scenarios. Code is available at https://github.com/youziizii/ECSM. Yongxiu Xu, Yuyao Kong, Gaopeng Gou |
ICMR | 3 |
| 2026 | 3DGauCIM: Accelerating Static/Dynamic 3D Gaussian Splatting via Digital CIM for High Frame Rate Real-Time Edge RenderingabstractDynamic 3D Gaussian splatting (3DGS) extends static 3DGS to render dynamic scenes, enabling AR/VR applications with moving objects. However, implementing dynamic 3DGS on edge devices faces challenges: (1) Loading all Gaussian parameters from DRAM for frustum culling incurs high energy costs. (2) Increased parameters for dynamic scenes elevate sorting latency and energy consumption. (3) Limited on-chip buffer capacity with higher parameters reduces buffer reuse, causing frequent DRAM access. (4) Dynamic 3DGS operations are not readily compatible with digital compute-in-memory (DCIM). These challenges hinder real-time performance and power efficiency on edge devices, leading to reduced battery life or requiring bulky batteries. To tackle these challenges, we propose algorithm-hardware co-design techniques. At the algorithmic level, we introduce three optimizations: (1) DRAM-access reduction frustum culling to lower DRAM access overhead, (2) Adaptive tile grouping to enhance on-chip buffer reuse, and (3) Adaptive interval initialization Bucket-Bitonic sort to reduce sorting latency. At the hardware level, we present a DCIM-friendly computation flow that is evaluated using the measured data from a 16 nm DCIM prototype chip. Our experimental results on Large-Scale Real-World Static/Dynamic Datasets demonstrate the ability to achieve high frame rate real-time rendering exceeding 200 frames per second (FPS) with minimal power consumption—merely 0.28 W for static Large-Scale Real-World scenes and 0.63 W for dynamic Large-Scale Real-World scenes. This work successfully addresses the significant challenges of implementing static/dynamic 3DGS technology on resource-constrained edge devices. Wei-Hsing Huang, Cheng-Jhih Shih, Jian-Wei Su, Samuel Wade Wang, Vaidehi Garg, Yuyao Kong, Jen-Chun Tien, Nealson Li, Arijit Raychowdhury, Meng-Fan Chang, Yingyan (Celine) Lin, Shimeng Yu |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2026 | Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale SystemsabstractRecent developments have introduced Kolmogorov– Arnold networks (KANs), an innovative architectural paradigm capable of replicating conventional deep neural network (DNN) capabilities while utilizing significantly reduced parameter counts through the employment of parameterized B-spline functions incorporating trainable coefficients. Nevertheless, the B-spline functional components inherent to KAN architectures introduce distinct hardware acceleration complexities. While B-spline function evaluation can be accomplished through lookup table (LUT) implementations that directly encode functional mappings, thus minimizing computational overhead, such approaches continue to demand considerable circuit infrastructure, including LUTs, multiplexers, decoders, and associated components. This work presents an algorithm-hardware co-design approach for KAN acceleration. At the algorithmic level, techniques include alignment–symmetry and PowerGap KAN hardware-aware quantization, KAN sparsity-aware mapping strategy, and circuit-level techniques include N:1 time modulation dynamic voltage input generator with analog-compute-in-memory (ACIM) circuits. Furthermore, this work conducts comprehensive evaluations on large-scale KAN networks to validate the proposed methodologies. Nonideality factors, including partial sum deviations arising from process variations, have been evaluated with the statistics measured from the TSMC 22-nm RRAM-ACIM prototype chips. Utilizing optimally determined KAN hyperparameters in conjunction with circuit optimizations implemented and evaluated at the 22-nm technology node, despite the model sizes for large-scale tasks in this work increasing by 435 K$\times $to 756 K$\times $compared to tiny-scale tasks in previous work, the area overhead increases by only 26 K$\times $to 40 K$\times $, with power consumption rising by merely$48\times $to$93\times $, while accuracy degradation remains minimal at 0.11%–0.22%, thereby demonstrating the scaling potential of our proposed architecture. Wei-Hsing Huang, Jianwei Jia, Yuyao Kong, Faaiq G. Waqar, Tai-Hao Wen, Meng-Fan Chang, Shimeng Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge InferenceabstractRecently, a novel model named Kolmogorov-Arnold Networks (KAN) has been proposed with the potential to achieve the functionality of traditional deep neural networks (DNNs) using orders of magnitude fewer parameters by parameterized B-spline functions with trainable coefficients. However, the B-spline functions in KAN present new challenges for hardware acceleration. Evaluating the B-spline functions can be performed by using lookup tables (LUTs) to directly map the B-spline functions, thereby reducing computational resource requirements. However, this method still requires substantial circuit resources (LUTs, MUXs, decoders, etc.). For the first time, this paper employs an algorithm-hardware co-design methodology to accelerate KAN. The proposed algorithm-level techniques include Alignment-Symmetry and PowerGap KAN hardware aware quantization, KAN sparsity aware mapping strategy, and circuit-level techniques include N:1 Time Modulation Dynamic Voltage input generator with analog-CIM (ACIM) circuits. The impact of non-ideal effects, such as partial sum errors caused by the process variations, has been evaluated with the statistics measured from the TSMC 22nm RRAM-ACIM prototype chips. With the best searched hyperparameters of KAN and the optimized circuits implemented in 22 nm node, we can reduce hardware area by 41.78x, energy by 77.97x with 3.03% accuracy boost compared to the traditional DNN hardware. Wei-Hsing Huang, Jianwei Jia, Yuyao Kong, Faaiq G. Waqar, Tai-Hao Wen, Meng-Fan Chang, Shimeng Yu |
ASP-DAC | 3 |
| 2025 | Digital Compute-in-Memory Ising Annealer with Ferroelectric Capacitor-Based nvSRAM for Combinatorial Optimization ProblemsabstractCombinatorial optimization problems (COPs) have a wide range of applications. The Ising model-based annealer is gaining attention for its efficiency and speed in finding approximate solutions. However, building an Ising machine that is area- and energy-efficient, scalable, and with low compute latency in CMOS is challenging. In this paper, we present a digital compute-in-memory (DCIM) Ising annealer that uses ferroelectric capacitor (FeCap)-based nvSRAM to solve COPs like the Traveling Salesman Problem (TSP). By using weak recall operations, our design eliminates the need to reload weights, significantly reducing energy consumption and speeding up processing compared to other approaches. Simulations using a 16nm PDK demonstrate that our nvSRAM-based DCIM array maintains accuracy while reducing latency by up to 55.0% and energy by 49.6% compared to prior work implemented with conventional SRAM DCIM array. Algorithm validation further shows that the random noise introduced by weak recall can be effectively utilized in the annealing process. Yuyao Kong, Jianwei Jia, Anni Lu, Faaiq G. Waqar, Yuan-Chun Luo, Hai Li 0001, Ian A. Young, Shimeng Yu |
ISCAS | 1 |
| 2023 | From macro to microarchitecture: reviews and trends of SRAM-based compute-in-memory circuits
Zhaoyang Zhang 0008, Jinwu Chen, Xi Chen 0107, An Guo 0001, Bo Wang 0023, Tianzhu Xiong, Yuyao Kong, Xingyu Pu, Shengnan He, Xin Si, Jun Yang 0006 |
Sci. China Inf. Sci. | 7 |
| 2022 | SNNIM: A 10T-SRAM based Spiking-Neural-Network-In-Memory architecture with capacitance computationabstractSpiking-Neural-Networks (SNN) have natural advantages in high-speed signal processing and big data operation. However, due to the complex implementation of synaptic arrays, SNN based accelerators may face low area utilization and high energy consumption. Computing-In-Memory (CIM) shows great potential in performing intensive and high energy efficient computations. In this work, we proposed a JOT-SRAM based Spiking-Neural-Network-In-Memory architecture (SNNIM) with 28nm CMOS technology node. A compact JOT-SRAM bit-cell was developed to realize signed 5bit synapses arrays and configurable bias arrays (SYBIA). The soma array based standard 8T-SRAM (SMTA) stores the soma membrane voltage and the threshold value. A capacitance computation scheme (CCA) between them was proposed to support various SNN operations. The proposed SNNIM achieved energy efficiency of 25.18 TSyOPSI. And the proposed SNNIM achieved 1.79+× better array efficiency compared with previous works. Bo Wang 0023, Xiang Li 0147, Anran Yin, Zhongyuan Feng, Yuyao Kong, Tianzhu Xiong, Haiming Hsu, Yongliang Zhou, An Guo 0001, Jun Yang 0006, Xin Si |
ISCAS | 7 |
| 2019 | In-memory Processing based on Time-domain CircuitabstractDeep Neural Networks (DNN) have emerged as a dominant algorithm for machine learning (ML). High performance and extreme energy efficiency are critical for deployments of DNN, especially in mobile platforms such as autonomous vehicles, cameras, and other devices of internet of things. However, DNNs lead to massive data movement and memory accesses, which prevents it from being integrated into always-on Internet-of-Things (IoT) devices. Recently, computing in-memory (CIM) architectures embeds analog computation circuits in/near the memory arrays. It significantly reduces data movement energy. This paper summarizes the most recent novel methods on the CIM architectures based on the time-domain computation. Compared with voltage-domain and frequency-domain analog computing method, time-domain computation provides more flexibility, higher accuracy and greater scalability for larger neural networks. Thereafter, the first in-memory binary weight network (BWN) processor based on pulse-width modulation in which the feature is stored in memory is also presented. This work significantly reduces memory accesses (4x), and achieves state-of-the-art peak energy efficiency of 119.7TOPS/W. Yuyao Kong, Jun Yang 0006 |
ACM Great Lakes Symposium on VLSI | 1 |