EDBT 2026 Demo / reviewers in the wild / expert
Delong Shang
dblp:68/6905
· DBLP profile ↗
22ranked-venue papers
8as first author
5since 2021 · last 2024
0009-0000-6674-2347ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 8 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Design of High-Performance while Energy-Efficient Microprocessor with Novel Asynchronous Techniques: (PhD Forum Paper)abstractDesign of high-performance processors with low power is the primary goal of many contemporary and futuristic applications, while asynchronous circuit is one of the solution. This brief presents a novel RISC-V microprocessor architecture which is capable of achieving these requirements by using two proposed asynchronous techniques. The first contribution is a relative-timing based asynchronous controller (RTAC), which is very small (only 7 gates) and low-power when compared to other existing controllers. The second contribution is a novel quasi-2phase handshake protocol which combines the benefits of the traditional 2-phase and 4-phase protocols, it allows a simple 4-phase asynchronous controller to achieve the same throughput as a 2-phase circuit. The simulation results proved that an asynchronous multiplier adopted the quasi-2phase protocol has approximately 1.95 times throughput higher than the 4-phase one. Finally, the key technologies of this work will be thoroughly validated in a 3-stage 32bits RISC-V microprocessor to further demonstrate the feasibility of the proposed techniques. Xiqin Tang, Delong Shang |
ASAP | 2 |
| 2024 | Multiscale Residual Network with Dynamic Depthwise Convolution for Multi-label 12-Lead ECG ClassificationabstractAccurate and efficient classification of electrocardiogram (ECG) signals is an important task in clinical diagnosis and patient care. Existing approaches of ECG signal classification suffer from problems such as feature redundancy, overfitting, increased computational complexity, and decreased real-time performance. The objective of this study is to propose an effective and efficient solution for ECG signal classification. We introduce a novel end-to-end multi-label classification model named MSR-D2CNN that combines a multi-scale dynamic depthwise convolution module (MSD2C-Blocks) with residual skip connections. MSD2C-Blocks integrate the channel attention mechanism Squeeze-and-Excitation(SE) with the ECG dynamic convolution (EDC) module. The model can capture features of different sizes and shapes of ECG signals while maintaining a lightweight structure. We evaluate the proposed model on the PTB-XL dataset, which consists of 12-lead ECG signals and covers 6 different label settings. Our proposed model achieves excellent classification performance while maintaining a low parameter count and computational complexity. Specifically, the parameter count is 0.827M, FLOPs is 134.563M, and Macro-AUC scores are 0.9350, 0.9360, 0.9339, 0.9341, 0.8765, and 0.9668 across the six different label settings. Our proposed solution provides an effective and efficient approach to ECG signal classification, overcoming the limitations of existing methods. The MSR-D2CNN model can accurately classify ECG signals while maintaining a low parameter count and computational complexity, making it practical for real-world scenarios. Linhai Xie, Yilei Man, Delong Shang |
IJCNN | 3 |
| 2024 | A 409mV, Sub-10nW Power-on Reset Circuit Using Adaptive Accuracy Adjustment for Low Voltage ApplicationsabstractA low power power-on reset (POR) circuit with low temperature coefficient, low quiescent current and small area is proposed in this paper. The POR circuit samples the power supply voltage through a high threshold transistor and then converts the voltage to a current, which will be compared with a native NMOS based current reference to obtain the reset signal. In order to reduce the power consumption in steady state, an adaptive accuracy adjustment mechanism is employed in the proposed POR circuit. The POR circuit uses a low accuracy but energy efficient structure to monitor the supply voltage in steady state, and when a voltage drop is observed, the POR circuit quickly switches to a high accuracy mode to get an accurate brown-out detection trip voltage. The POR circuit is implemented in 55nm CMOS process and the active area is just 107μm2. Post-layout simulation results show that the POR circuit has a POR trip voltage of 409mV, a static power of 9.11nW at a supply voltage of 0.45V. Besides, the temperature coefficient of the proposed POR circuit is only 31.76μV/°C over a temperature range of -40°C to 125°C. Heng You, Dashan Shi, Delong Shang, Shushan Qiao |
ISCAS | 3 |
| 2024 | Differentiable architecture search with multi-dimensional attention for spiking neural networks
Yilei Man, Linhai Xie, Shushan Qiao, Delong Shang |
Neurocomputing | 5 |
| 2024 | A 1-8b Reconfigurable Digital SRAM Compute-in-Memory Macro for Processing Neural NetworksabstractThis work presents a 1-8b reconfigurable digital SRAM compute-in-memory (CIM) macro, which significantly improves array utilization and energy efficiency under different input and weight configurations compared to previous works. To ensure the array utilization under different configurations, a row-based bitwise-summation-first digital CIM architecture is proposed. In addition, to realize flexible switching between signed and unsigned operations, a complete 2’s complement encoding method is adopted, which makes the computation of the sign bits consistent with that of the magnitude bits when performing signed operations, thus ensuring that each row of the CIM array can store the sign of the weight. Due to the support of reconfigurable bit width, the proposed CIM macro can be widely used in various neural networks for optimal efficiency. In order to better apply the CIM macro to binarized neural networks, a configurable bitwise multiplier is presented, which supports both AND and XNOR operations. Moreover, since the power consumption of the adder tree occupies a major part of the digital CIM macro, a 4–2 compressor based adder tree is presented to further improve the energy efficiency. Measurement results based on 55nm CMOS process show that the proposed CIM macro achieves an energy efficiency of up to 2238TOPS/W at 1b/1b and 44.82TOPS/W at 4b/4b MAC operations. Heng You, Delong Shang, Shushan Qiao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2018 | Approximate Fixed-Point Elementary Function Accelerator for the SpiNNaker-2 Neuromorphic ChipabstractNeuromorphic chips are used to model biologically inspired Spiking-Neural-Networks (SNNs) where most models are based on differential equations. Equations for most SNN algorithms usually contain variables with one or more excomponents. SpiNNaker is a digital neuromorphic chip that has so far been using pre-calculated look-up tables for exponential function. However this approach is limited because the memory requirements grow as more complex neural models are developed. To save already limited memory resources in the next generation SpiNNaker chip, we are including a fast exponential function in the silicon. In this paper we analyse iterative algorithms for elementary functions and show how to build a single hardware accelerator for exp and natural log, for a neuromorphic chip prototype, to be manufactured in a 22 nm FDSOI process. We present the accelerator that has algorithmic level approximation control, allowing it to trade precision for latency and energy efficiency. As an addition to neuromorphic chip application, we provide analysis of a parameterized elementary function unit that can be tailored for other systems with different power, area, accuracy and latency constraints. Mantas Mikaitis, David R. Lester, Delong Shang, Steve Furber, Gengting Liu, Jim D. Garside, Stefan Scholze, Sebastian Höppner, Andreas Dixius |
ARITH | 3 |
| 2016 | Low power voltage sensing through capacitance to digital conversionabstractCapacitance sensors are widely used for sensing physical parameters. Conventional capacitance to digital methods use complex analog ADC techniques which are power hungry. Recently a fully digital solution was proposed with improved power consumption. This paper describes a number of problems in that solution, analyzes these problems, and proposes a new design free of these problems. A voltage senor as an example was designed based on the proposed capacitance to digital conversion in this paper. The new method achieves the same accuracy with less than half the circuit size, and 25% and 33% savings on power and energy consumption. Delong Shang, Yuqing Xu, Kaiyuan Gao, Fei Xia 0001, Alexandre Yakovlev |
DDECS | 1 |
| 2014 | Asynchronous design for new on-chip wide dynamic range power electronicsabstractAsynchronous circuits will play an important role in microelectronic systems in the future, especially in energy harvesting and autonomous (EHA) systems where such circuits will be able to offer robustness and deliver high efficiency in a wide range of power-energy conditions. The concept of Capacitor Bank Block (CBB) mechanisms was proposed to form the basis of electronics for powering asynchronous loads. These mechanisms will benefit EHA systems by enabling effective co-scheduling of computational tasks and energy supply. This paper demonstrates how the CBB mechanisms can themselves be controlled by asynchronous circuits, thereby forming a new type of power delivery units (PDU) that will be able to deliver power to intelligent digital logic in future EHA systems. These PDUs are superior to traditional power converters largely because the latter can only regulate sufficiently high power and energy levels (regular and periodic) as well as their controllers require stable power levels themselves. This makes them unsuitable for intermittent and sporadic conditions inherent to EHA systems. In this paper, a novel asynchronous control for the CBB is described. Experiments and analysis of the new PDUs, comprising CBBs and asynchronous control, are presented and discussed in detail. Delong Shang, Xuefu Zhang, Fei Xia 0001, Alexandre Yakovlev |
DATE | 1 |
| 2014 | Asynchronously assisted FPGA for variabilityabstractThe effect of variability has become increasingly significant as a result of technology geometry scaling. This paper describes Asynchronous Assisting Logic (AAL) blocks and the method of introducing them into modern FPGA architecture, in order to increase tolerance of the wide range latency variations caused by parametric variation, and temperature and supply voltage fluctuations. The proposed method leverages the availability of variation maps and suggests deploying configurable AAL blocks only into the variation critical paths - reinforcing rather rerouting/remapping. This method reduces the size overhead significantly which normally will be incurred by fully asynchronous designs. The proposed technique maintains the existing FPGA architecture allowing potential reuse of design flow. Simulations show correct functionality given regularly variable, randomly variable and capacitor switching energy harvester voltage supplies. Hock Soon Low, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
FPL | 2 |
| 2013 | High-order reconfigurable FIR filter design based on statistical analysis of CSD coefficientsabstractPerformance and power consumption are two important aspects of reconfigurable finite impulse response (FIR) filters. In this paper, a new reconfigurable FIR architecture named SCSD FIR is proposed. This proposed architecture has been synthesized in 0.13 μm technology with a precision of 16 bits. Design example shows that FIR with 551 taps achieves 21.3% area reduction and 19.0% power reduction compared with the most efficient existing architecture of FIRs. It also achieves an improvement of 13.1% in speed. Furthermore, these advantages are expanded as the taps of filters are increasing. Rui Jia, Rui Chen 0014, Delong Shang, Haigang Yang |
FPT | 5 |
| 2013 | Wide-range, reference free, on-chip voltage sensor for variable Vdd operationsabstractIn future systems with relatively unreliable and unpredictable energy sources such as harvesters, the system Vdd may become non-deterministic. Reliable and accurate on-chip voltage sensors are therefore indispensible for the power and computation management of such systems. Stable and known references are also difficult to obtain in this environment. This paper describes a reference-free voltage sensor implemented using a speed independent (SI) SRAM cell and an inverter chain. It can work under a wide range of Vdd, and provides accurate measurements of Vdd over this operating range with a precision range from 50mV to 10mV. Unlike existing methods, the voltage information is directly generated as a digital code without any analog circuits. This is realized by exploiting the inherently different latency behaviors of different types of circuits under different Vdd. Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 1 |
| 2013 | Concurrent Multiresource Arbiter: Design and ApplicationsabstractThis paper presents a novel type of asynchronous arbiter that allocates M interchangeable resources among N clients. This arbiter enables the concurrent utilization of multiple resources and is a useful device in various load-balancing circuits. Dedicated request signals from the resources and the clients are used in pairs to form each new grant. The 2 × 2 arbiter is examined as an accessible special case of the N × M arbiter. A concurrent implementation is compared to fully sequential design. It is shown that the sequential design can be more practical when the time between a grant and the withdrawal of the initial request is small. The concurrent design provides higher performance in a system with a longer resource utilization time. A scalable tiled structure is developed to extend the arbiter structure beyond 2 × 2 to support N clients and M resources. Models and subsequent implementations of the tiles are presented. The tiles can be repeated without the use of additional connecting logic, enabling the construction of arbiters of larger sizes. Several examples demonstrate the usage of the arbiter. Stanislavs Golubcovs, Delong Shang, Fei Xia 0001, Andrey Mokhov, Alexandre Yakovlev |
IEEE Trans. Computers | 2 |
| 2012 | Ultra-low power transmitterabstractThis paper presents a design of an ultra-low power UWB transmitter based on 4thand 5thderivative Gaussian pulse shapes implemented in UMC 90nm CMOS technology. The simulations show 119mV peak to peak pulse amplitude and the pulse width of 240 ps for the 5thderivative Gaussian pulse and 99.71mV pulse amplitude and 190 ps pulse width for the 4thderivative Gaussian pulse. Power consumption of the pulse generators are calculated 30.11 uW and 21.5 uW for the 5thand 4thderivative Gaussian pulse respectively at a 100MHz pulse repeating frequency (PRF). Ultra-low power radio transmission is important in such application contexts as wireless network nodes and sensors powered by energy harvesters. Mohsen Ghasempour, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 2 |
| 2011 | Variation tolerant asynchronous FPGA (abstract only)abstractThis paper describes the realization of an interconnect Delay Insensitive (DI) FPGA architecture with distributed asynchronous control. This architecture maintains the basic block structure of traditional FPGAs allowing the potential use of existing FPGA design tools in block design. This asynchronous FPGA architecture is mainly aimed at tolerating the unpredictable delay variations caused by process and environment variations in current and future VLSI technology nodes and also targets low power operations, including modes such as dynamic voltage scaling and variable Vdd, as in applications featuring energy harvesting. This is achieved by making the longer inter-block interconnects DI, keeping the computational logic single-rail, and removing global clocks. Hock Soon Low, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
FPGA | 2 |
| 2011 | A Novel Power Delivery Method for Asynchronous Loads in Energy Harvesting SystemsabstractFor systems depending on power harvesting, a fundamental contradiction in the power delivery chain has existed between conventional synchronous computational loads requiring relatively stable Vdd and power harvesters unable to supply it. DC/DC conversion has therefore been an integral part of such systems to resolve this contradiction. On the other hand, asynchronous computational loads, in addition to their potential power-saving capabilities, can be made tolerant to a much wider range of Vdd variance. This may open up opportunities for much more energy efficient methods of power delivery. This article presents in-depth investigations into the behavior and performance of different on-chip power delivery methods driving both asynchronous and synchronous loads directly from a harvester source. A novel power delivery method, which employs a capacitor bank for adaptively storing the energy from power harvesters depending on load and source conditions, is developed. Its advantages, especially when driving asynchronous loads, are demonstrated through comprehensive comparative analysis. Xuefu Zhang, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2010 | Stochastic analysis of power, latency and the degree of concurrencyabstractConcurrent processing has become the default mode of operation in on-chip systems. Silicon has become cheap enough for having hardware facilities to support very large scale concurrent processing on chip. As a result the availability and applicability of power is becoming more of a limiting factor than logic. However, the advantage of parallelism in reducing power consumption will soon become unrealistic because of the limited scope of reducing Vdd beyond threshold voltage, leaving the reduction of concurrency (through the partial shut-down of system blocks) as a realistic means of reducing power consumption when needed. A stochastic modelling approach is presented in this paper which can integrate the degree of concurrency as a parameter into power and latency analysis. This will facilitate a system design and management regime where the degree of concurrency is used as a means of control to achieve power and performance goals. Yuan Chen 0002, Isi Mitrani, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 3 |
| 2010 | Asynchronous FPGA architecture with distributed controlabstractAsynchronous techniques have become more significant with continued scaling of VLSI technologies. This paper proposes an asynchronous FPGA architecture. Different from previous methods of introducing asynchrony into FPGAs, our method seeks to preserve the current FPGA cell structure as much as possible, whilst achieving delay insensitivity in the inter-cell interconnects. By using David Cells as the central technique in the delay insensitive clock replacement, this method is conducive to the establishment of an automatic design and synthesis flow. It also particularly caters for low power designs, where current FPGA solutions are not effective yet. Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 1 |
| 2010 | Highly parallel multi-resource arbitersabstractMulti-resource multi-client arbiters are becoming more important in on-chip systems because of the increasing significance of dynamic, run-time, allocation of various system performance resources such as power and computation and communication facilities. Arbiters, for example, can be used to limit the amount of concurrency for regulating voltage droops, and for balancing load and traffic. This paper describes the design of multi-resource arbiters with high degrees of concurrency. By using freezing logic, this design method guarantees correct computation whilst simplifies the implementation. Quick release mechanisms and the implementation of the multi-token concept through the duplication of the client requests help improve the efficiency. Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 1 |
| 2007 | Registers for Phase Difference Based LogicabstractA logic design style known as phase difference-based logic (PDBL) has several benefits with respect to security and testing. An existing design method for PDBL circuits has so far been lacking an important component, a register. In this paper, we present the design of a speed independent PDBL register and a timed PDBL register, which can be used in asynchronous or synchronous circuits. Comparisons are presented in terms of speed, size, and power consumption. Delong Shang, Alexandre Yakovlev, Albert Koelmans, Danil Sokolov, Alexandre V. Bystrov |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2006 | Low-Cost Online Testing of Asynchronous HandshakesabstractA new low-cost low-complexity checker for online testing of asynchronous interfaces in globally-asynchronous locally-synchronous circuits is proposed. The solution is fully based upon the standard gate libraries. The checker itself is fully offline testable. It also provides a fault-locating functionality, which is achieved by combining the online mode with scan techniques Delong Shang, Alexandre Yakovlev, Frank P. Burns, Fei Xia 0001, Alexandre V. Bystrov |
ETS | 1 |
| 2005 | On-Line Testing of Globally Asynchronous CircuitsabstractThe problem of on-line testing of asynchronous circuits is analyzed, several infrastructures are proposed and a self-checking tree checker is designed. The checker uses on-demand self-test, which reduces power consumption and guarantees bounded self-test period. Objects under test are tested by observing protocols at their primary inputs and outputs. The checker does not slow down the functional system, as it only samples the signals. The protocols are checked by identifying enabled and refused signal transitions in each state of the system. The fault coverage of internal faults of the checker is calculated. Simulation results are included. Delong Shang, Alexandre V. Bystrov, Alexandre Yakovlev, Deepali Koppad |
IOLTS | 1 |
| 2004 | An Asynchronous Synthesis Toolset Using VerilogabstractWe present a new CAD tool set for generating asynchronous circuits from high-level Verilog level-sensitive specifications. Initially, high-level Verilog descriptions are compiled and converted into a novel intermediate Petri net format. The intermediate format is subsequently passed to optimization tools and mapping tools where it is directly mapped into asynchronous datapath and control circuits using David cells (DCs). Finally, logic optimization tools are applied to generate speed-independent (SI) circuits. The speed independent circuits generated perform well compared to circuits generated by existing asynchronous tools. Frank P. Burns, Delong Shang, Albert Koelmans, Alexandre Yakovlev |
DATE | 2 |