Shuyuan Yu

dblp:136/0494 · DBLP profile ↗
← Back
13ranked-venue papers
10as first author
8since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2024 Fast and Scaled Counting-Based Stochastic Computing Divider Design
abstract
This article presents novel designs for stochastic computing (SC)-based dividers, which promise low latency, high energy efficiency as well as high accuracy for error-tolerant arithmetic operations. We first introduce CBDIV, which is based on the recently proposed counter-based SC concept and correlation based SC to perform division. Then we introduce FSCDIV, which further improves the accuracy of CBDIV by applying a scaling strategy and mitigating the latency by optimizing the counting scheme. The FSCDIV will equally scale up the divider and dividend before the division process, and thereby avoid large relative error when both input values of the divider and dividend are small. The proposed fast counting method, accelerates FSCDIV by counting new bit pair (0-1 pair) among only half of the stochastic number bitstream instead of the entire bitstream, resulting in almost half of the counting latency and one-fourth of the overall division operation latency. The experimental results demonstrate that the proposed CBDIV, implemented in a 32nm technology node, outperforms state-of-the-art works by 77.8% in accuracy, 37.1% in delay, 21.5% in area, 50.6% in area delay product (ADP), and 25.9% in power consumption. Compared to the fixed-point division baseline, CBDIV also achieves a 31.9% reduction in energy consumption and is more energy-efficient than existing SC-based dividers for binary inputs and outputs required in efficient image processing implementations. Moreover, we demonstrate that FSCDIV improves delay by 56.4%, ADP by 16.0%, energy consumption by 45.0%, and accuracy by 61.2%. We also evaluate CBDIV and FSCDIV designs in a contrast stretch image processing workload, and the results show that the proposed designs can improve the image quality by up to 18.3 dB on average when compared to state-of-the-art works.
Shuyuan Yu, Maliha Tasnim, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 PAALM: Power Density Aware Approximate Logarithmic Multiplier Design
abstract
Approximate hardware designs can lead to significant power or energy reduction. However, a recent study showed that approximated designs might lead to unwanted higher temperature and related reliability issues due to the increased power density. In this work, we try to mitigate this important problem by proposing a novel power density aware approximate logarithmic multiplier (called PAALM) design for the first time. The new multiplier design is based on the approximate logarithmic multiplier (ALM) framework due to its rigorous mathematics based foundation. The idea is to re-design the high computing switch activities of existing ALM designs based on equivalent mathematical formula so that the power density can be reduced at no accuracy loss while at costs of some area overheads. Our results show that the proposed PAALM design can improve 11.5%/5.7% of power density and 31.6%/70.8% of area with 8/16-bit precision when compared with the fixed-point multiplier baseline, respectively. And also achieves extremely low error bias: -0.17/0.08 for 8/16-bit precision, respectively. On top of this, we further implement the PAALM design in a Convolutional Neural Network (CNN) and test it on CIFAR10 dataset. The results show that with error compensation, PAALM can achieve the same inference accuracy as the fixed-point multiplier baseline. We also evaluate the PAALM in a discrete cosine transformation (DCT) application. The results show that with error compensation, PAALM can improve the image quality of 8.6dB in average when compared to the ALM design.
Shuyuan Yu, Sheldon X.-D. Tan
ASP-DAC1
2023 MAGIC-DHT: Fast in-memory computing for Discrete Hadamard Transform
abstract
Discrete Hadamard transform (DHT) is a signal processing tool that decomposes an arbitrary input vector into a superposition of Walsh functions. Due to its wide range of applications in processing big data, a fast and energy-efficient hardware design for DHT with high throughput capability is essential. Processing in memory (PIM) allows the in-place computation to reduce the data traffic, which is a major speed bottleneck in the existing computing. In this work, we propose an efficient hybrid parallel PIM-based computation for DHT. Our proposed method explores the recursive computation of DHT and is based on the memristor-aided logic (MAGIC) gates in which the arithmetic operations are carried out via simple logic NOR operation. We propose two in-memory computing methods for the DHT encoding process. At the arithmetic level, to improve efficiency, we propose to share the intermediate results between addition and subtraction in DHT in the first method called MAGIC-DHT-1D which provides an average speedup of 1.12× over the recently proposed DigitalPIM for 1D DHT. Furthermore,MAGIC-DHT-1D also outperforms SIMPLER in terms of energy and energy density in average. We also propose a second method, called MAGIC-DHT-2D, to share the carrier independent computation cycles among multi-bit parallel addition and subtraction. At the algorithm level, we also explore both row and column-based PIM NOR computing in the same crossbar to avoid the transposition operation required in the 2D DHT process. MAGIC-DHT-2D provides an average speedup of 4.84× and 7.25× over two state-of-the-art methods DigitalPIM and SIMPLER, respectively for each each complete set of 2D DHT computing cycles. Our numerical results further show that our proposed optimized methods can lead up to 56.19× and 6.90× speed-up, as well as 57.84× and 5.96× higher throughput over NVIDIA RTX Titan GPU to compute 1D DHT and 2D DHT, respectively.
Maliha Tasnim, Chinmay Raje, Shuyuan Yu, Elaheh Sadredini, Sheldon X.-D. Tan
Integr.3
2022 HEALM: Hardware-Efficient Approximate Logarithmic Multiplier with Reduced Error
abstract
In this work, we propose a new approximate logarithm multipliers (ALM) based on a novel error compensation scheme. The proposed hardware-efficient ALM, named HEALM, first determines the truncation width for mantissa summation in ALM. Then the error compensation or reduction is performed via a lookup table, which stores reduction factors for different regions of input operands. This is in contrast to an existing approach, in which error reduction is performed independently of the width truncation of mantissa summation. As a result, the new design will lead to more accurate result with both reduced area and power. Furthermore, different from existing approaches which will either introduce resource overheads when doing error improvement or lose accuracy when saving area and power, HEALM can improve accuracy and resource consumption at the same time. Our study shows that 8-bit HEALM can achieve up to 2.92%, 9.30%, 16.08%, 17.61% improvement in mean error, peak error, area, power consumption respectively over REALM, which is the state of art work with the same number of bits truncated. We also propose a single error coefficient mode named HEALM-TA-S, which improves the ALM design with a truncation adder (TA) for mantissa summation. Furthermore, we evaluate the proposed HEALM design in a discrete cosine transformation (DCT) application. The result shows that with different values of k, HEALM-TA can improve the image quality upon the ALM baseline by 7.8~17.2dB in average and HEALM-SOA can improve 2.9~15.8dB in average, respectively. Besides, HEALM-TA and HEALM-SOA outperform all the state of art works with k = 2, 3, 4 on the image quality. And the single coefficient mode, HEALM-TA-S, can improve the image quality upon the baseline up to 4.1dB in average with extremely low resource consumption.
Shuyuan Yu, Maliha Tasnim, Sheldon X.-D. Tan
ASP-DAC1
2022 Conceptual Prerequisites for Proportional Analogy
Shuyuan Yu, Ho-Chieh Lin, John Opfer
CogSci1
2022 Scaled-CBSC: scaled counting-based stochastic computing multiplication for improved accuracy
abstract
Stochastic computing (SC) can lead area-efficient implementation of logic designs. Existing SC multiplication, however, suffers a long-standing problem: large multiplication error with small inputs due to its intrinsic nature of bit-stream based computing. In this article, we propose a new scaled counting-based SC multiplication approach, called Scaled-CBSC, to mitigate this issue by introducing scaling bits to ensure the bit '1' density of the stochastic number is sufficiently large. The idea is to convert the "small" inputs to "large" inputs, thus improve the accuracy of SC multiplication. But different from an existing stream-bit based approach, the new method uses the binary format and does not require stochastic addition as the SC multiplication always starts with binary numbers. Furthermore, Scaled-CBSC only requires all the numbers to be larger than 0.5 instead of arbitrary defined threshold, which leads to integer numbers only for the scaling term. The experimental results show that the 8-bit Scaled-CBSC multiplication with 3 scaling bits can achieve up to 46.6% and 30.4% improvements in mean error and standard deviation, respectively; reduce the peak relative error from 100% to 1.8%; and improve 12.6%, 51.5%, 57.6%, 58.4% in delay, area, area-delay product, energy consumption, respectively, over the state of art work. Furthermore, we evaluate the proposed multiplication approach in a discrete cosine transformation (DCT) application. The results show that with 3 scaling bits, 8-bit scaled counting-based SC multiplication can improve the image quality with 5.9dB upon the state of art work in average.
Shuyuan Yu, Sheldon X.-D. Tan
DAC1
2021 Cognitive Supports for Objective Numeracy
Shuyuan Yu, John Opfer
CogSci1
2021 COSAIM: Counter-based Stochastic-behaving Approximate Integer Multiplier for Deep Neural Networks
abstract
In this work, we propose a new counter-based stochastic-behaving approximate integer unsigned multiplier, called COSAIM, for many emerging error tolerant application workloads such as deep neural networks. Unlike existing approximate multipliers, which are based on some deterministic ad-hoc methods or mathematical formula, the new design is an improved stochastic multiplier, which performs improved sequential counting for multiplication operation in a deterministic way. In this work, we further improve the counting efficiency by introducing approximate schemes to significantly speed up the counting process, which leads to significant clock cycle reduction with no accuracy loss. COSAIM bears all the advantages of stochastic computing such as built-in configurability for progressive performance-accuracy trade-off. At the same time, it shows very small latency and high energy efficiency. Our evaluation shows that the COSAIM with error improvement operation can achieve very low error bias (0.06%), along with lower mean error (0.30% to 3.49%), and low peak errors (around 1.81%) with variance of 1. $47\times 10^{-4}$ %. Experimental results obtained from Xilinx ISE show that compared with the 8-bit exact multiplier baseline, COSAIM can save up to 53.95%, 32.84%, 52.24%, 21.05% in area, power, energy and the product Area. 1/Throughput, respectively. Furthermore, by doing shared parallel design, COSAIM can further lead to improvements in area, power and energy reduction by 60.44%, 53.33% and 68.54%, respectively compared to the baseline. We also implement COSAIM in a Convolution Neural Network (CNN) and test it on CIFAR10 dataset and find that CNN with COSAIM delivers similar inference accuracy compared to some state of art approximate multipliers.
Shuyuan Yu, Sheldon X.-D. Tan
DAC1
2020 Reliable Power Grid Network Design Framework Considering EM Immortalities for Multi-Segment Wires
abstract
This paper presents a new power grid network design and optimization technique that considers the new EM immortality constraint due to EM void saturation volume for multi-segment interconnects. Void may grow to its saturation volume without changing the wire resistance significantly. However, this phenomenon was ignored in existing EM-aware optimization methods. By considering this new effect, we can remove more conservativeness in the EM-aware on-chip power grid design. Along with recently proposed nucleation phase immortality constraint for multi-segment wires, we show that both EM immortality constraints can be naturally integrated into the existing programming based power grid optimization framework. To further mitigate the overly conservative problem of existing immortality-constrained optimization methods, we further explore two strategies: first we size up failed wires to meet one of the immorality conditions subject to design rules; second, we consider the EM-induced aging effects on power supply networks for a target lifetime, which allows some short-lifetime wires to fail and optimizes the rest of the wires. Numerical results on a number of IBM-format power grid networks demonstrate that the new method can reduce more power grid area compared to the existing EM-immortality constrained optimizations. Furthermore, the new method can optimize power grids with nucleated wires, which would not be possible with the existing methods.
Han Zhou 0002, Shuyuan Yu, Zeyu Sun 0001, Sheldon X.-D. Tan
ASP-DAC2
2020 Run-Time Accuracy Reconfigurable Stochastic Computing for Dynamic Reliability and Power Management: Work-in-Progress
abstract
In this paper, we propose a novel accuracy-reconfigurable stochastic computing (ARSC) framework for dynamic reliability and power management. Different than the existing stochastic computing works, where the accuracy versus power/energy trade-off is carried out in the design time, the new ARSC design can change accuracy or bit-width of the data in the run-time so that it can accommodate the long-term aging effects by slowing the system clock frequency at the cost of accuracy while maintaining the throughput of the computing. We validate the ARSC concept on a discrete cosine transformation (DCT) and inverse DCT designs for image compressing/decompressing applications, which are implemented on Xilinx Spartan-6 family XC6SLX45 platform. Experimental results show that the new design can easily mitigate the long-term aging induced effects by accuracy trade-off while maintaining the throughput of the whole computing process using simple frequency scaling. We further show that one-bit precision loss for the input data, which translated to 3.44dB of the accuracy loss in term of Peak Signal to Noise Ratio (PSNR) for images, we can sufficiently compensate the NBTI induced aging effects in 10 years while maintaining the pre-aging computing throughput of 7.19 frames per second. At the same time, we can save 74% power consumption by 10.67dB of accuracy loss. The proposed ARSC computing framework also allows much aggressive frequency scaling, which can lead to order of magnitude power savings compared to the traditional dynamic voltage and frequency scaling (DVFS) techniques.
Shuyuan Yu, Han Zhou 0002, Shaoyi Peng, Hussam Amrouch, Jörg Henkel, Sheldon X.-D. Tan
CASES1
2020 From Integers to Fractions: Developing a Coherent Understanding of Proportional Magnitude
Shuyuan Yu, Dan Kim, Marta K. Mielicki, Charles J. Fitzsimmons, Clarissa A. Thompson, John Opfer
CogSci1
2018 Source Retrieval Cues Facilitate Transfer in Fraction Learning
Shuyuan Yu, John Opfer
CogSci1
2013 Vehicle logo recognition based on Bag-of-Words
abstract
The recognition of vehicle manufacturer logo is a crucial and very challenging problem, which is still an area with few published effective methods. This paper proposes a new fast and reliable system for Vehicle Logo Recognition (VLR) based on Bag-of-Words (BoW). In our system, vehicle logo images are represented as histograms of visual words and classified by SVM in three steps: firstly, extract dense-SIFT features; secondly, quantize features into visual words by `Soft-assignment' thirdly, build histograms of visual words with spatial information. Compared with traditional VLR methods, experiment results show that our proposed system achieves higher recognition accuracy with less processing time. The proposed system is evaluated on a dataset of 840 low-resolution vehicle logo images with about 30×30 pixels, which verifies that our system is practical and effective.
Shuyuan Yu, Shibao Zheng, Hua Yang 0001, Longfei Liang
AVSS1