VLDB 2026 Research / reviewers in the wild / expert
Yirong Kan
dblp:272/2608
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-4070-0672ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mutual Information Loss for Enhancing Class-Wise Representation in Spiking Neural Networks
Yiling Yang, Yirong Kan, Yasuhiko Nakashima |
ISCAS | 2 |
| 2026 | FPSpike: A Fully Parallel and Reconfigurable Architecture for Accelerating Spiking Neural Networks With Structured SparsityabstractThis paper introduces FPSpike, a fully-parallel and reconfigurable architecture for accelerating spiking neural networks (SNNs) with structured sparsity. By introducing structured sparse synaptic connections, the neuron computation and weight storage costs are significantly reduced, while the wiring constraints of hardware implementation are alleviated, thus realizing a fully parallel SNN architecture on a single chip. Furthermore, FPSpike achieves high reconfigurability through local connections, allowing the network topology to be flexibly partitioned into multiple distinct hardware cores without redundancy, thereby enabling multi-task spatial parallel processing. In particular, a series of software-hardware co-designs are performed to improve the performance of FPSpike. Experimental results with FPGA-based implementation demonstrate that FPSpike achieves inference throughput improvements of$2.5\times \sim 43.9\times $and$3.2\times \sim 93.8\times $compared to Intel i7-12700K CPU and NVIDIA RTX3060Ti GPU, respectively. FPSpike also achieves performance speedups of$2.0\times \sim 20.2\times $and energy efficiency improvements of$1.3\times \sim 7.7\times $compared to the state-of-the-art FPGA-based SNN accelerators. Yirong Kan, Man Wu, Yasuhiko Nakashima |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2026 | DystoHD: an area-efficient hyperdimensional computing system with dynamic hypervector generation for memory-constrained devices
Yirong Kan, Man Wu, Yasuhiko Nakashima |
J. Supercomput. | 2 |
| 2025 | Buffer and Splitter Insertion for Adiabatic Quantum-Flux-Parametron CircuitsabstractThe extremely low-bit energy characteristic of the adiabatic quantum-flux-parametron (AQFP) circuit makes it a promising candidate for highly energy-efficient computing systems. However, in contrast with conventional circuit design, general logic synthesis tools can not make sure that the circuit functionality of generated AQFP circuits is correct. AQFP circuits require buffer and splitter insertion for dataflow synchronization at all clock phases of the circuit and multifan-out driving. Notably, buffers and splitters inserted take up much area and delay in AQFP circuits, also causing a significant increase in energy dissipation. To address this problem, this article analyses in detail why buffer and splitter insertion is necessary for AQFP circuits and proposes a global optimization framework for this purpose. This framework consists of three parts: 1) logic level assignment; 2) splitter tree generation; and 3) buffer insertion. An integer linear programming algorithm is proposed for the logic level assignment to estimate the globally optimal number of inserted buffers and splitters. Subsequently, a dynamic programming-based multiway search tree generation algorithm is proposed to construct an optimal splitter tree for each net of the input circuit. Moreover, three optimization strategies are proposed to further enhance the effectiveness and efficiency of our framework. Experimental results on ISCAS’85 and EPFL benchmarks demonstrate the effectiveness and efficiency of our proposed framework compared with the state-of-the-art, particularly with significant advantages on large circuits. Rongliang Fu, Mengmeng Wang 0006, Yirong Kan, Olivia Chen, Nobuyuki Yoshikawa, Bei Yu 0001, Tsung-Yi Ho |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | A Fully-Parallel Reconfigurable Spiking Neural Network Accelerator with Structured Sparse ConnectionsabstractIn this work, we present a fully parallel reconfigurable spiking neural network (SNN) accelerator for various applications of edge computing. In contrast to conventional fully connected and irregular sparse network topologies, structured sparse synaptic connections are introduced to implement SNN in fully parallel on hardware by significantly reducing neuron computation and weight memory costs. Benefiting from the hardware-friendly symmetric SNN topology, the proposed accelerator is flexibly configured into multiple classifiers without hardware redundancy to support various tasks. Furthermore, to compress the layer depth of individual classifiers to reduce latency, layer-skip connections are introduced to alleviate the vanishing gradient problem in training. Various experiments are conducted to explore the optimal settings of the proposed accelerator. The proposed SNN accelerator is verified on the Xilinx ZCU102 FPGA. The results show that the proposed SNN accelerator achieves accuracy of 97.1% and energy efficiency of 40.1 GSOP/s/W on the MNIST dataset. Yirong Kan, Yasuhiko Nakashima |
ISCAS | 2 |
| 2024 | Bisection Neural Network Toward Reconfigurable Hardware ImplementationabstractA hardware-friendly bisection neural network (BNN) topology is proposed in this work for approximately implementing massive pieces of complex functions in arbitrary on-chip configurations. Instead of the conventional reconfigurable fully connected neural network (FC-NN) circuit topology, the proposed hardware-friendly topology performs NN behaviors in a bisection structure, in which each neuron includes two constant synapse connections for both inputs and outputs. Compared with the FC-NN one, the reconfiguration of the BNN circuit topology eliminates the remarkable amount of dummy synapse connections in hardware. As the main target application, this work aims at building a general-purpose BNN circuit topology that offers a great amount of NN regressions. To achieve this target, we prove that the NN behaviors of the FC-NN circuit topologies can be migrated to the BNN circuit topologies equivalently. We introduce two approaches including the refining training algorithm and the inverted-pyramidal strategy to further reduce the number of neurons and synapses. Finally, we conduct the inaccuracy tolerance analysis to suggest the guideline for ultra-efficient hardware implementations. Compared with the state-of-the-art FC-NN circuit topology-based TrueNorth baseline, the proposed design can achieve 17.8- 22.2× hardware reduction and less than 1% inaccuracy. Yirong Kan, Sa Yang, Yasuhiko Nakashima |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | A Global Optimization Algorithm for Buffer and Splitter Insertion in Adiabatic Quantum-Flux-Parametron CircuitsabstractAs a highly energy-efficient application of low-temperature superconductivity, the adiabatic quantum-flux-parametron (AQFP) logic circuit has characteristics of extremely low-power consumption, making it an attractive candidate for extremely energy-efficient computing systems. Since logic gates are driven by the alternating current (AC) serving as the clock signal in AQFP circuits, plenty of AQFP buffers are required to ensure that the dataflow is synchronized at all logic levels of the circuit. Meanwhile, since the currently developed AQFP logic gates can only drive a single output, splitters are required by logic gates to drive multiple fan-outs. These gates take up a significant amount of the circuit's area and delay. This paper proposes a global optimization algorithm for buffer and splitter (B/S) insertion to address the issues above. The B/S insertion is first identified as a combinational optimization problem, and a dynamic programming formulation is presented to find the global optimal solution. Due to the limitation of its impractical search space, an integer linear programming formulation is proposed to explore the global optimization of B/S insertion approximately. Experimental results on the ISCAS'85 and simple arithmetic benchmark circuits show the effectiveness of the proposed method, with an average reduction of 8.22% and 7.37% in the number of buffers and splitters inserted compared to the state-of-the-art methods from ICCAD'21 and DAC'22, respectively. Rongliang Fu, Mengmeng Wang 0006, Yirong Kan, Nobuyuki Yoshikawa, Tsung-Yi Ho, Olivia Chen |
ASP-DAC | 3 |
| 2022 | Automatic Sleep Staging via Frequency-Wise Spiking Neural NetworksabstractIdentifying sleep stages is a fundamental step for both early detection of disease and neuroscientific exploration. Automatic sleep staging is classic research for replacing the time-consuming gold-standard manual staging procedure. Recently, promising results have been achieved on automatic staging by extracting spatio-temporal features via deep neural networks from electroencephalogram (EEG). However, such methods fail to consistently yield good performance due to a missing piece in data representation: the dynamic fluctuations of neurons on top of EEG features that is non-trivial for automatic sleep staging task. This paper introduces a biomimicry spiking neural network (SNN) to map the aforementioned features serving for automatic sleep staging. Such SNNs are designed as an array of encoders that converts frequency-specific features into long-term spiking coding and then a popular ANN model is used as the staging machine by absorbing the stage-dependent spiking representation. For proof-of-concept, the performance of the proposed framework is demonstrated by introducing multiple sleep datasets. The experimental results showed that the proposed method achieved a competitive stage scoring performance, especially for Wake, N2, and N3, with higher Precision of 0.94, 0.87, and 0.86. Moreover, the ablation studies prove the SNN has the potential for extracting the neuron’s variation features. Haohui Jia, Ziwei Yang 0002, Pei Gao, Man Wu, Chen Li 0027, Yirong Kan |
BIBM | 6 |
| 2022 | Online Learning of Parameters for Modeling User Preference Based on Bayesian NetworkabstractBy analyzing users’ behavior data for personalized services, most state-of-the-art methods for user preference modeling are often based on batch-mode machine learning algorithms, where all rating data are assumed to be available throughout the training process. However, data in the real world often arrives sequentially and user preference may change dynamically. The real-time characteristics of rating data make the algorithms for preference modeling challenging to suit real-world online applications. By the user preference model (UPM) based on Bayesian network with a latent variable (BNLV), uncertain relationships among relevant attributes of users, objects and ratings could be represented, in which user preference is represented by the latent variable. In this paper, we propose an online approach for parameter learning of UPM. Specifically, we first extend the classic Voting EM algorithm by using Bayesian estimation in terms of the situation with latent variables. Consequently, we propose the algorithm for learning parameters of UPM from few and sequentially-changing rating data to reflect the gradually changing preferences. Finally, we test the effectiveness of our proposed algorithm by conducting experiments on various datasets. Experimental results demonstrate the superiority of our method in various measurements. Yirong Kan, Kun Yue, Hao Wu 0010, Xiaodong Fu, Zhengbao Sun |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |
| 2022 | MuGRA: A Scalable Multi-Grained Reconfigurable Accelerator Powered by Elastic Neural NetworkabstractA massive core computing architecture is developed for accelerating arbitrary calculations in fully parallel with high speed and low cost. The proposed architecture is reconfigurable in fine-grained (arbitrary functions), mid-grained (flexible function feature, accuracy, and number of operands), and coarse-grained (organization of cores). By implementing a large scale of novel bisection neural network (BNN) on hardware, the re-configuration is conducted by partitioning entire BNN into any specific pieces without redundancy. Each piece of BNN retrieves the arbitrary function approximately. By reconfiguring the BNN topology in software, we can easily adjust dimensions of the computing kernel without rewiring, and achieve a wide range of trade-offs between accuracy and efficiency in hardware. In this manner, the multi-grained reconfigurable accelerator (MuGRA) is achieved. Since MuGRA is flexible in all grained levels, various configurations for each validation are demonstrated with rich options of performance-cost matrix. From the FPGA implementation results, compared with other traditional function approximation methods, our method provides fewer parameter storage requirements. The comparison against related works proves that our accelerator effectively reduces the calculation latency with slight accuracy loss. Yirong Kan, Man Wu, Yasuhiko Nakashima |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | DiaNet: An elastic neural network for effectively re-configurable implementationabstractAn elastic neural network is developed and evolved towards effectively re-configurable hardware in fully parallel on chip. The original prototype of DiaNet is organized as a symmetrical bisection neural network, which is feasible to be partitioned into arbitrary pieces of neural networks (NNs) without redundancy. To prevent the depth explosion in implementing complex tasks (complicated pattern recognition for instance), the evolution of DiaNets is investigated in this work. By using the I/O layer integration technology which enables all neurons in the hidden layer of DiaNet to receive inputs, the number of layers is reduced to 8.8% of DiaNet prototype. In this manner, the DiaNet topology is feasible to implement complex NNs without the risk of depth explosion. Moreover, the skip connection technology is proposed to avoid the gradient vanishing due to deep learning, which is significant to DiaNets especially. Compared with the LeNet5 model as state-of-the-art, the evolved DiaNet topology achieves the parameter reduction of 90.86% for MNIST recognition with the negligible loss of accuracy. To reduce hardware utilization, the sensitivity to the decline of computational precision and bit-width is investigated to suggest the guideline for efficient hardware implementations. Finally, the effectiveness of DiaNet is verified by the proposed re-configurable architecture on FPGA with the power reduction of 10.8% compared to state-of-the-art implementations. Man Wu, Yirong Kan, Tati Erlina, Yasuhiko Nakashima |
Neurocomputing | 2 |