EDBT 2026 Demo / reviewers in the wild / expert
Song Jia
dblp:29/6478
· DBLP profile ↗
17ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-4704-2510ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CHiP-NoC: A Congestion-Adaptive Dual-Mode Neuromorphic NoC with Hybrid Spike Compression
Yipeng Gao, Yi Zhong 0002, Yingying Cui, Song Jia, Yuan Wang 0001 |
ISCAS | 4 |
| 2025 | CROSSCUT: A Multi-Core Neuromorphic Accelerator Improving Resource-UtilizationabstractNeuromorphic computing is attracting significant attention due to its bio-mimetic characteristics. Consequently, neuromorphic hardware platforms have emerged as innovative computing architectures for acceleration. However, the fixed nature of data flow and resources leads to considerable inefficiencies in storage and computation, thereby limiting both utilization efficiency and overall performance. This severely hinders the deployment of edge artificial intelligence (AI) models. To address these issues, we present a multi-core neuromorphic accelerator named CROSSCUT. This crossbar-based system supports both spiking neural network (SNN) and artificial neural network (ANN) paradigms and has a capacity of 256K neurons and 288M synapses. By leveraging the Neuron Package Mechanism (NPM) and Synapse Compress Mechanism (SCM), CROSSCUT can increase input data scale by 64 times and reduce wasted resources and computations by 46.7%, ensuring high compatibility with diverse network structures in machine learning models. Additionally, a Tree-Mesh hybrid network on chip (NoC) is constructed for inter-core communication. Implemented on Xilinx XCVU9P FPGA, CROSSCUT can achieve a peak performance of 431.9 GSOPS/s and 121.13 GSOPS/W energy efficiency. The inference accuracy on MNIST is 98.2%. Youming Yang 0002, Yi Zhong 0002, Zilin Wang 0001, Tao Zhang 0140, Li Lun, Yingying Cui, Xiaoxin Cui, Song Jia, Yuan Wang 0001 |
ISCAS | 8 |
| 2025 | Enabling Fault Isolation in Fault Injection for Automotive-grade Circuits SimulationabstractThis paper proposes a method to enable fault isolation in non-intrusive fault injection at RTL and GL with a high-speed feature. The key scheme is to minimize the number of isolated modules inserted while maintaining the purity of the circuit source code called half-intrusive fault inject technique (HIFI). The proposed scheme integrates the advantages of intrusive and non-intrusive fault injection technique and enables fault isolation for the former. Simulation can be achieved by proposed HIFI with significantly 5× accelerated run-time than conventional intrusive technique. Besides, this work used this technique to build an open-source fault injection tool to help learners and researchers in the field of functional safety of integrate circuits. Pingsheng Zhang, Zenan Yan, Yuan Wang 0001, Song Jia |
ISCAS | 7 |
| 2024 | NeuroREC: A 28-nm Efficient Neuromorphic Processor for Radar Emitter ClassificationabstractRadar emitter classification (REC) plays an important role in modern warfare. Traditional REC methods have difficulty identifying complex radar signals in the present day. Inspired by biology, spiking neural networks (SNNs) have gradually gained widespread attention due to their low power characteristics. Compared with convolutional neural networks (CNNs), SNNs are more suitable for application in the field of REC. The reason is that SNN can not only maintain higher accuracy in the presence of noise interference, but also reduce the power consumption of mobile devices. However, it is challenging to make full use of the input sparsity of radar emitter signals and the weight sparsity of pruned SNN models. In this paper, a 28-nm neuromorphic processor for REC named NeuroREC is proposed. It uses matrix compression algorithms to store sparse weights on chip, and designs corresponding spike detection circuits for this purpose. As a single-core design, we propose a ping-pong running mechanism to alleviate the imbalance between IO throughput and peak performance. Two SNN models for classifying RadioML2016.b and RadioML2018.a datasets are deployed on the chip, achieving competitive accuracy with only 8 timesteps, and demonstrating better robustness than CNN. Fabricated in 28-nm CMOS process, NeuroREC runs at frequencies ranging from 22.5MHz to 744MHz. Under specific sparsity conditions, it can reach an energy efficiency of 7.22TSOP/W for 8-bit weight. Zilin Wang 0001, Zehong Ou, Yi Zhong 0002, Youming Yang 0002, Li Lun, Hufei Li, Jian Cao 0002, Xiaoxin Cui, Song Jia, Yuan Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2024 | Marmotini: A Weight Density Adaptation Architecture With Hybrid Compression Method for Spiking Neural NetworkabstractBrain-inspired spiking neural network (SNN) has recently attracted widespread interest owing to its event-driven nature and relatively low-power hardware for transmitting highly sparse binary spikes. To further improve energy efficiency, some matrix compression algorithms are used for weight storage. However, the weight sparsity of different layers varies greatly. For a multicore neuromorphic system, it is difficult for the same compression algorithm to adapt to all the layers of SNN model. In this work, we propose a weight density adaptation architecture with hybrid compression method for SNN, named Marmotini. It is a multicore heterogeneous design, including three types of cores to complete computation of different weight sparsity. Benefiting from the hybrid compression method, Marmotini minimizes the waste of neurons and weights as much as possible. Besides, for better flexibility, a reconfigurable core that can be configured to compute convolutional layer or fully connected layer is proposed. Implemented on Xilinx Kintex UltraScale XCKU115 field-programmable gate array (FPGA) board, Marmotini can operate at 150-MHz frequency, achieving 244.6-GSOP/s peak performance and 54.1-GSOP/W energy efficiency at 0% spike sparsity. Zilin Wang 0001, Yi Zhong 0002, Zehong Ou, Youming Yang 0002, Xiaoxin Cui, Song Jia, Yuan Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2019 | An algorithm for calculating coverage rate of WSNs based on geometry decomposition approach
Bailing Wang, Song Jia, Haohan Hong |
Peer-to-Peer Netw. Appl. | 3 |
| 2017 | A reliable true random number generator based on novel chaotic ring oscillatorabstractA novel true random number generator is proposed and implemented on XC6SLX16. It consumes 44 LUTs and generates output bitrate at 125 Mbps without post-processing, or 2000 Mbps with post-processing. The underlying mechanism of chaotic dynamics in Boolean chaotic oscillator is also researched. With the utilization of the proposed entropy source, the new scheme can precede referenced designs in reliability, resource consumption, output bitrate, design simplicity and requirement of post-processing. Yunfan Yang, Song Jia, Yuan Wang 0001, Shaonan Zhang |
ISCAS | 2 |
| 2016 | A novel low-power and high-speed dual-modulus prescaler based on extended true single-phase clock logicabstractA novel low-power and high-speed dual-modulus prescaler based on extended true single-phase clock (E-TSPC) scheme is presented. By restricting the short-circuit current in noncritical branchs, the design reduces the major source of power dissipation in E-TSPC scheme. The presented design enhances the maximum working frequency with shorter critical path and lower load capacitances. Simulation results in SMIC 40nm process show that compared with referenced E-TSPC based designs at least 61.2% (divide-by-2) and 41.1% (divide-by-3) reduction in power delay product (PDP) can be achieved by the proposed design. Song Jia, Yuan Wang 0001 |
ISCAS | 1 |
| 2016 | A novel low-leakage power-rail ESD clamp circuit with adjustable triggering voltage and superior false-triggering immunity for nanoscale applicationsabstractThis work presents a novel power-rail electrostatic discharge (ESD) clamp circuit for nanoscale applications. By skillfully incorporating transient and static ESD detection mechanisms into its detection circuit, the proposed circuit achieves a wide range of adjustable triggering voltage (Ft1) while maintaining low standby leakage current (Ileak). Besides, the proposed circuit achieves significantly-improved false-triggering immunity compared with the transient-triggered circuit. All investigated circuits are fabricated in a 65-nm CMOS process. Simulation and test results have both confirmed the superiority of the proposed circuit. In addition, the proposed circuit achieves similar triggering behaviors in both transmission line pulsing (TLP) and very fast TLP (VF-TLP) tests. Guangyi Lu, Yuan Wang 0001, Jian Cao 0002, Song Jia, Xing Zhang 0002 |
ISCAS | 4 |
| 2016 | Delay-locked loop based frequency quadrupler with wide operating range and fast locking characteristicsabstractA wide operating range and fast locking delay-locked loop (DLL) based frequency quadrupler that includes an eight-phase-clock generator and an edge combiner is proposed. The eight-phase-clock generator is composed of a coarse-code generator, a fine-code generator and a digital controlled delay line, which uses four differential delay units to generate equally spaced eight-phase clocks. The coarse-code generator adopts a time-to-digital scheme to achieve short locking time and wide operating range. A fine-code digital-to-analog converter in the fine-code generator converts the fine codes to analog voltage for high precision. Moreover, the novel edge-combiner circuit combines the eight-phase clocks to x4 frequency output with 50% duty cycle ratio. Experimental results in a 65-nm CMOS process show this frequency multiplier can cover a frequency range from 320 MHz to 2.4 GHz and cost 5~40 cycles to finish locking. Yuan Wang 0001, Yuequan Liu, Mengyin Jiang, Song Jia, Xing Zhang 0002 |
ISCAS | 4 |
| 2016 | Area-efficient transient power-rail electrostatic discharge clamp circuit with mis-triggering immunity in a 65-nm CMOS process
Yuan Wang 0001, Guangyi Lu, Haibing Guo, Jian Cao 0002, Song Jia, Xing Zhang 0002 |
Sci. China Inf. Sci. | 5 |
| 2015 | A low-power high-speed 32/33 prescaler based on novel divide-by-4/5 unit with improved true single-phase clock logicabstractIn this work, new design techniques that aim to reduce power consumption of true single-phase clock-based (TSPC) prescalers is presented. The structure of divide-by-4/5 frequency divider is simplified, and its performance is compared with previous work to demonstrate the improvement. Simulation results show at least a 25% reduction of power consumption is achieved by the proposed unit. In the 32/33 dual modulus prescaler, a critical path cutting scheme is introduced to improve speed to the limit decided by the divide-by-4/5 unit. Song Jia, Shilin Yan, Yuan Wang 0001, Ganggang Zhang |
ISCAS | 1 |
| 2015 | 180.5Mbps-8Gbps DLL-based clock and data recovery circuit with low jitter performanceabstractA wide range delay-locked loop (DLL) based clock and data recovery (CDR) circuit including coarse and fine tune blocks is proposed in this paper. The coarse tune block adopts a time to digital converter and digital control delay line to widen the frequency capture range, reduce locking time and prevent the false locking problem. In the fine tune block, a novel phase detector combines the tasks of sampling and charge-pump using half rate clock. Starting-control circuit can ensure CDR takes full use of the delay range provided by voltage control delay line. Moreover, a fully analog DLL technique is applied to exploit the benefits of low skew and jitter performance. The simulation result shows the proposed CDR can cover a wide frequency range from 180.5Mbps to 8Gbps, while the peak-to-peak jitter of recovery clock is 2.7ps at 200Mbps and 1.06ps at 8Gbps. Fabricated in a 65nm CMOS process, this design dissipates 9.9mW and 22.9mW respectively at 200 Mbps and 8Gbps from a 1.2 V supply. Yuequan Liu, Yuan Wang 0001, Song Jia, Xing Zhang 0002 |
ISCAS | 3 |
| 2015 | Investigation on the layout strategy of ggNMOS ESD protection devices for uniform conduction behavior and optimal width scaling
Guangyi Lu, Yuan Wang 0001, Lizhong Zhang, Jian Cao 0002, Song Jia, Xing Zhang 0002 |
Sci. China Inf. Sci. | 5 |
| 2012 | Design of novel, semi-transparent flip-flops (STFF) for high speed and low power application
Xiayu Li, Song Jia, Yuan Wang 0001, Ganggang Zhang |
Sci. China Inf. Sci. | 2 |
| 2010 | Low swing drivers based on charge redistribution
Fengfeng Wu, Song Jia, Yuan Wang 0001, Ganggang Zhang |
Sci. China Inf. Sci. | 2 |
| 2003 | Design Retargetable Platform System for Microprocessor Functional TestabstractMicroprocessors are extremely versatile and complexity that present significantly test challenges. This paper describes a retargetable functional test platform system design for various microprocessors. Characterized by configurable test environment generator, retargetable assembler and strong ATPG the platform system could automatically produce different test environment and assemble out relative test codes to adapt to the microprocessor under test. Experiments show that the platform system works correctly, flexibly and efficiently. Wennan Feng, Song Jia, Anping Jiang, Lijiu Ji |
Asian Test Symposium | 3 |