EDBT 2026 Demo / reviewers in the wild / expert
Yu Bai 0004
dblp:03/6325-4
· DBLP profile ↗
20ranked-venue papers
9as first author
5since 2021 · last 2025
0000-0002-2303-1120ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 9 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LID-Drug: A Localized Interactive Domain-Aggregated (LID) Framework for Protein Drug Editing
Mingshuo Liu, Yunduan Lou, Shiyi Luo, Shangping Ren, Yu Bai 0004 |
PRICAI | 6 |
| 2024 | Educational Tool-spaces for Convolutional Neural Network FPGA Design Space Exploration Using High-Level SynthesisabstractThere is significant demand and urgency to prepare electrical and computer engineering students regarding the operational and performance characteristics of machine learning (ML) hardware accelerators. Convolutional Neural Networks (CNNs), which are utilized for real-time and large dataset image classification tasks, are appropriate targets for hardware acceleration. Designing accelerators for CNNs necessitates understanding the manipulation of CNN parameters. We introduce a hands-on pedagogy whereby learners can identify, modify, and appreciate the interaction of the CNN parameters within an interactive GUI. CASCADE (Computer Aided Student's CNN Analyzer for Design Exploration), a simulation-based framework for Design Space Exploration (DSE) of CNN FPGA-based accelerators is developed, including datapath synthesis, simulation, training, and testbench steps. We offer a case study of High-Level Synthesis (HLS) based CNN implementations targeting the MNIST dataset and present simulation results, namely hardware utilization, accuracy, and operating frequency, and offer insight into potential design trade-offs facing modern engineers. Richard C. Yarnell, Mousam Hossain, Raul Graterol, Ayush Pindoria, Sujan Ghimire, Md Muhtasim Alam Chowdhury, Soheil Salehi, Yu Bai 0004, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 8 |
| 2024 | Swin-MSP: A Shifted Windows Masked Spectral Pretraining Model for Hyperspectral Image ClassificationabstractDeep learning has found widespread application in the hyperspectral image (HSI) classification, where transformer architectures based on self-attention have emerged as state-of-the-art (SOTA). The Swin-MAE framework utilizes a masked autoencoder approach with a shifted windows transformer as its backbone, demonstrating strong representational power and performance. This study proposes a shifted windows masking spectral pretraining (Swin-MSP) model, which achieves hierarchical modeling of hyperspectral data from local to global scales by introducing spectral masking pretraining techniques and a hierarchical architecture. To fit with this pretraining, we introduce the uniaxial continuous cross correlation layer (UC3L), a straightforward yet effective solution tailored for hyperspectral imagery masking. We design the shift frequency band transformer (SFBT) to hierarchically characterize spectral features. Experiments with publicly available datasets establish that our pretrained network significantly improves classification efficiency compared with SOTA networks. Furthermore, we systematically investigate the sensitivity of various datasets to pretraining hyper-parameters. The results underscore that the universal spectral representation acquired during the pretraining phase serves as a robust initialization for subsequent task-specific fine-tuning. It is noted that this work breaks from traditional vision transformer (ViT) approaches, offering a new perspective on hyperspectral dataset pretraining. The code is available athttps://github.com/teaRRe/Swin-MSP. Danqing Liu, Yu Bai 0004, Guanliang Wan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | An Efficient Real-Time Object Detection Framework on Resource-Constricted Hardware Devices via Software and Hardware Co-designabstractThe fast development of object detection techniques has attracted attention to developing efficient Deep Neural Networks (DNNs). However, the current state-of-the-art DNN models can not provide a balanced solution among accuracy, speed, and model size. This paper proposes an efficient real-time object detection framework on resource-constricted hardware devices through hardware and software co-design. The Tensor Train (TT) decomposition is proposed for compressing the YOLOv5 model. By unitizing the unique characteristics given by the TT decomposition, we develop an efficient hardware accelerator based on FPGA devices. Experimental results show that the proposed method can significantly reduce the model size and improve the execution time. Mingshuo Liu, Shiyi Luo, Kevin Han, Bo Yuan 0001, Ronald F. DeMara, Yu Bai 0004 |
ASAP | 6 |
| 2021 | An Efficient Video Prediction Recurrent Network using Focal Loss and Decomposed Tensor Train for Imbalance DatasetabstractNowadays, from companies to academics, researchers across the world are interested in developing recurrent neural networks due to their incredible feats in various applications, such as speech recognition, video detection, predictions, and machine translation. However, the advantages of recurrent neural networks accompanied by high computational and power demands, which are a major design constraint for electronic devices with limited resources used in such network implementations. Optimizing the recurrent neural networks, such as model compression, is crucial to ensure the broad deployment of recurrent neural networks and promote recurrent neural networks for implementing most resource-constrained scenarios. Among many techniques, tensor train (TT) decomposition is considered an up-and-coming technology. Although our previous efforts have achieved 1) expanding limits of many multiplications within eliminating all redundant computations; and 2) decomposing into multi-stage processing to reduce memory traffic, this work still faces some limitations. In particular, current TT decomposition on recurrent neural networks leads to a complex computation sensitive to the quality of training datasets. In this paper, we investigate a new method for TT decomposition on recurrent neural networks for constructing an efficient model within imbalance datasets to overcome this issue. Experimental results show that the proposed new training method can achieve significant improvements in accuracy, precision, recall, F1-score, False Negative Rate (FNR), and False Omission Rate (FOR). Mingshuo Liu, Kevin Han, Shiyi Luo, Mingze Pan, Mousam Hossain, Bo Yuan 0001, Ronald F. DeMara, Yu Bai 0004 |
ACM Great Lakes Symposium on VLSI | 8 |
| 2019 | Compressing Deep Neural Networks Using Toeplitz Matrix: Algorithm Design and Fpga ImplementationabstractDeep neural networks (DNNs) have emerged as an important artificial intelligence technique. However, the computation-intensive and storage-intensive DNNs pose severe challenges on efficient execution over the underlying hardware platform. In this paper we propose to impose Toeplitz structure on DNN models to achieve high compression ratio with negligible performance loss. Accordingly, the hardware performance can be significantly improved after performing model compression. We evaluate the proposed approach on speech recognition and implement the corresponding compressed model on FPGA. Experimental results show that our approach enables high hardware performance while retaining high task performance. Siyu Liao, Ashkan Samiee, Chunhua Deng, Yu Bai 0004, Bo Yuan 0001 |
ICASSP | 4 |
| 2019 | Information, knowledge, and semantics for interacting with Internet-of-Things
Yunchuan Sun, Xiuzhen Cheng, Yu Bai 0004, Jiguo Yu |
Comput. Networks | 3 |
| 2018 | FRLDM: Empowering K-nearest Neighbor (KNN) through FPGA-based Reduced-rank Local Distance MetricabstractWhile fast and accurate data classification techniques are vital to many applications, K-Nearest Neighbor algorithm (KNN) is considered the most important algorithm used in data mining, text categorization, and image recognition. However, conventional KNN for computing distances may not necessarily perform well for all problems. In this paper, we propose a new framework named FRLDM to empower KNN through FPGA-based reduced-rank local distance metric. Experimental results on the collection of classification problems and hardware measurement imply that the FRLDM offers notable performance advantages over other approaches on CPU. Ashkan Samiee, Yinjie Huang, Yu Bai 0004 |
IEEE BigData | 3 |
| 2018 | Leveraging Spintronic Devices for Efficient Approximate Logic and Stochastic Neural NetworksabstractITRS has identified nano-magnet based spintronic devices as promising post-CMOS technologies for information processing and data storage due to their ultra-low switching energy, non-volatility, superior endurance, excellent retention time, high integration density and compatibility with CMOS technology. As for data storage, spintronic memory has been widely accepted as a universal high performance next-generation non-volatile memory candidate. As for information processing, spintronic computing remains complementary in its features to CMOS technology. In this paper, we present two innovative spintronic computing primitives, i.e. spintronic approximate logic and spintronic stochastic neural network, which both leverage the intrinsic spintronic device physics to achieve much more compact and efficient designs than CMOS counterparts. In spintronic approximate logic, we employ the intrinsic current-mode thresholding operation to implement an accuracy-configurable adder and further demonstrate its application in approximate DSP applications. In spintronic stochastic neural networks, we leverage the stochastic properties of domain wall devices and magnetic tunnel junction to implement a low-power and robust artificial neural network design. Shaahin Angizi, Zhezhi He, Yu Bai 0004, Jie Han 0001, Mingjie Lin, Ronald F. DeMara, Deliang Fan |
ACM Great Lakes Symposium on VLSI | 3 |
| 2018 | Clockless Spintronic Logic: A Robust and Ultra-Low Power Computing ParadigmabstractAsynchronous logic offers the advantages of no clock tree, robust circuit operation, avoidance of worst-case timing margins, and a reduced emission spectrum. Thus, computational paradigms are sought to attain advantages of clockless logic by leveraging the complementary characteristics of emerging devices and CMOS transistors within novel circuit designs. This paper introduces Spin Torque Enabled NULL Convention Logic (STENCL), which exploits the physical characteristics of non-volatile Domain-Wall (DW) and memristive devices to realize the Quasi-Delay-Insensitive (QDI) NULL Convention Logic (NCL) asynchronous design methodology. First, a formal algorithm is developed to transform NCL-based threshold m-of-n gate realizations to STENCL, in order to generate the corresponding input memristance and NULL module memristance required for nominal currents achieving DW device biasing. Second, hysteresis and set/reset conditions are realized by determining the corresponding current fluctuations required to move the DW within each threshold logic gate to realize all 27 foundational NCL gate structures, which are then simulated to assess energy and delay metrics. Third, a case study of a four-stage pipelined 32-bit IEEE single-precision floating point co-processor implemented as a dual-rail STENCL architecture is compared to a conventional CMOS-based NCL design implemented by an IBM SOI1250 45nm CMOS process. Fourth, a sensitivity analysis is performed to assess the impact of write accuracy and drift on memristor and DW device operation. Results indicate that STENCL-based designs achieve between 2-fold to 20-fold reduction in energy consumption with up to 8-fold reduction in area, over an equivalent CMOS-based NCL design for 32-bit full adders. Comparisons for various four-stage pipelined 32-bit IEEE single-precision floating-point co-processors and ISCAS benchmarks further substantiate those benefits for operation within acceptable tolerances at identical process technology nodes. Yu Bai 0004, Ronald F. DeMara, Jia Di, Mingjie Lin |
IEEE Trans. Computers | 1 |
| 2017 | A Spin-Orbit Torque based Cellular Neural Network (CNN) ArchitectureabstractIn this paper, we propose a differential Spin Hall Effect(SHE) assisted domain wall synapse, which can generate either positive or negative synaptic weighting values without the significant cost of multiple power supply voltages, supply rails, or computationally-intensive digital hardware. The architecture of the proposed synapse utilizes reading currents flowing through two oppositely-oriented devices as weighted by device conductance. The conductance is used to encode synaptic weight and programmed by domain wall position through writing current. The ability to set the current as positively or negatively weighted results in highly-configurable functionality within a compact synapse design. The synapses are used with a soft-limiting nonlinear neuron to employ the relationship between positions and input current magnitude. We show through micro-magnetic simulation how the non-volatile physical characteristic of the domain wall calibrated synapse is used to implement a numerical integration function to realize a Cellular Neural Network(CNN). The performance of the proposed CNN design for isolated letter denoising at 0ns to 4ns demonstrates noise filtering functionality with total energy consumption during sensing of 24fJ. This compares favorably to existing spin CNN cell designs to provide a promising design approach for intrinsic neural computation. Yu Bai 0004, Xiaobo Sharon Hu, Ronald F. DeMara, Mingjie Lin |
ACM Great Lakes Symposium on VLSI | 1 |
| 2017 | CirCNN: accelerating and compressing deep neural networks using block-circulant weight matricesabstractLarge-scale deep neural networks (DNNs) are both compute and memory intensive. As the size of DNNs continues to grow, it is critical to improve the energy efficiency and performance while maintaining accuracy. For DNNs, the model size is an important factor affecting performance, scalability and energy efficiency. Weight pruning achieves good compression ratios but suffers from three drawbacks: 1) the irregular network structure after pruning, which affects performance and throughput; 2) the increased training complexity; and 3) the lack of rigirous guarantee of compression ratio and inference accuracy. Caiwen Ding, Siyu Liao, Yanzhi Wang 0001, Zhe Li 0001, Ning Liu 0007, Youwei Zhuo, Chao Wang 0051, Xuehai Qian, Yu Bai 0004, Geng Yuan, Jian Tang 0008, Qinru Qiu, Xue Lin 0001, Bo Yuan 0001 |
MICRO | 9 |
| 2016 | Stochastic-Based Spin-Programmable Gate Array with Emerging MTJ Device Technology (Abstract Only)abstractThis paper describes the stochastic-based Spin-Programmable Gate Array (SPGA), an innovative architecture attempting to exploit the stochastic switching behavior newly found in emerging spintronic devices for reconfigurable computing. While many recently studies have investigated using Spin Transfer Torque Memory (STTM) devices to replace configuration memory in FPGAs, our study, for the first time, attempts to use the quantum-induced stochastic property exhibited by spintronic devices directly for reconfiguration and logic computation. Specifically, the SPGA was designed from scratch for high performance, routability, and ease-of-use. It supports variable granularity multiple-input-multiple-output (MIMO) logic blocks and variable-length bypassing interconnects with a symmetrical structure. Due to its unconventional architectural features, the SPGA requires several major modifications to be made in the standard VPR placement/routing CAD flow, which include a new technology mapping algorithm based on computing (k, l)-cut, a new placement algorithm, and a modified delay-based routing procedure. Our mixed mode simulation results have shown that, with FPGA architecture innovations, on average, a SPGA can further achieve more than 10x improvement in logic density, about 5x improvement in average net delay, and about 5x improvement in the critical path delay for the largest 12 MCNC benchmark circuits over an island-style baseline FPGA with spintronic configuration bits. Yu Bai 0004, Mingjie Lin |
FPGA | 1 |
| 2016 | Ultra-Robust Null Convention Logic Circuit with Emerging Domain Wall DevicesabstractDespite many attractive advantages, Null Convention Logic (NCL) remains to be a niche largely due to its high imple- mentation costs. Using emerging spintronic devices, this paper proposes a Domain-Wall-Motion-based NCL circuit design methodology that achieves approximately 30x and 8x improvements in energy efficiency and chip layout area, respectively, over its equivalent CMOS design, while main- taining similar delay performance for a 32-bit full adder. These advantages are made possible mostly by exploiting the domain wall motion physics to natively realize the hys- teresis critically needed in NCL. More Interestingly, this de- sign choice achieves ultra-high robustness by allowing spin- tronic device parameters to vary within a predetermined range while still achieving correct operations. Yu Bai 0004, Weidong Kuang, Mingjie Lin |
ACM Great Lakes Symposium on VLSI | 1 |
| 2015 | Energy-Efficient Discrete Signal Processing with Field Programmable Analog Arrays (FPAAs)abstractLarge-scale field programmable analog array (FPAA) devices have made analog and analog-digital signal processing techniques accessible to a much wider community. However, largely due to its severe resource constraints, high noise sensitivity, and enormous design space, reconfigurable analog computing remains a niche in the DSP application space. In this paper, we develop a probabilistic-based methodology for designing and implementing the analog computing engines that specifically target at energy-efficient signal processing systems. We will first demonstrate how to decompose a given DSP application into various functional modules within the framework of probabilistic-based processing. Furthermore, we will show how these individual functional modules can be easily mapped to the limited selection of analog blocks found in an commercially available FPAA device: the PSoC chip platform from Cypress. To keep our study concrete, our implementation example focuses on the 1-D convolution module, a fundamental algorithmic building block in many applications of computer vision and artificial intelligence. In the end, we construct a complete image processing system based on the PSoC chip platform, and use the application of image key point extraction to demonstrate that our proposed approach to reconfigurable analog computing has considerable advantages in hardware usage, energy efficiency, and computing robustness over the traditional DSP approaches. Yu Bai 0004, Mingjie Lin |
FPGA | 1 |
| 2014 | Energy-efficient multiplier-less discrete convolver through probabilistic domain transformationabstractEnergy efficiency and algorithmic robustness typically are conflicting circuit characteristics, yet with CMOS technology scaling towards 10-nm feature size, both become critical design metrics simultaneously for modern logic circuits. This paper propose a novel computing scheme hinged on probabilistic domain transformation aiming for both low power operation and fault resilience. In such a computing paradigm, algorithm inputs are first encoded through probabilistic means, which translates the input values into a number of random samples. Subsequently, light-weight operations, such as sim- ple additions will be performed onto these random samples in order to generate new random variables. Finally, the resulting random samples will be decoded probabilistically to give the final results. Mohammed Alawad, Yu Bai 0004, Ronald F. DeMara, Mingjie Lin |
FPGA | 2 |
| 2014 | Optimally mitigating BTI-induced FPGA device aging with discriminative voltage scaling (abstract only)abstractWith the CMOS technology aggressively scaling towards the 22nm node, modern FPGA devices face tremendous aging- induced reliability challenges due to Bias Temperature In- stability (BTI) and Hot Carrier Injection (HCI). This paper presents a novel antiaging technique at logic level that is both scalable and applicable for VLSI digital circuits implemented with FPGA devices. The key idea is to prolong the lifetime of FPGA-mapped designs by strategically elevating the VDD values of some LUTs based on their modular criticality values. Although the idea of scaling VDD in order to improve either energy efficiency or circuit reliability has been explored extensively, our study distinguishes itself by approaching this challenge through analytical procedure, therefore able to maximize the overall reliability of target FPGA design by rigorously modelling the BTI-induce de- vice reliability and optimally solving the VDD assignment problem. Yu Bai 0004, Mohammed Alawad, Mingjie Lin |
FPGA | 1 |
| 2013 | Boosting Memory Performance of Many-Core FPGA Device through Dynamic Precedence GraphabstractEmerging FPGA device, integrated with abundant RAM blocks and high-performance processor cores, offers an unprecedented opportunity to effectively implement single-chip distributed logic-memory (DLM) architectures [1]. Being “memory-centric”, the DLM architecture can significantly improve the overall performance and energy efficiency of many memory-intensive embedded applications, especially those that exhibit irregular array data access patterns at algorithmic level. However, implementing DLM architecture poses unique challenges to an FPGA designer in terms of 1) organizing and partitioning diverse on-chip memory resources, and 2) orchestrating effective data transmission between on-chip and off-chip memory. In this paper, we offer our solutions to both of these challenges. Specifically, 1) we propose a stochastic memory partitioning scheme based on the well-known simulated annealing algorithm. It obtains memory partitioning solutions that promote parallelized memory accesses by exploring large solution space; 2) we augment the proposed DLM architecture with a reconfigure hardware graph that can dynamically compute precedence relationship between memory partitions, thus effectively exploiting algorithmic level memory parallelism on a per-application basis. We evaluate the effectiveness of our approach (A3) against two other DLM architecture synthesizing methods: an algorithmic-centric reconfigurable computing architectures with a single monolithic memory (A1) and the heterogeneous distributed architectures synthesized according to [1] (A2). To make our comparison fair, in all three architectures, the data path remains the same while local memory architecture differs. For each of ten benchmark applications from SPEC2006 and MiBench [2], we break down the performance benefit of using A3 into two parts: the portion due to stochastic local memory partitioning and the portion due to the dynamic graph-based memory arbitration. All experiments have been conducted with a Virtex-5 (XCV5LX155T-2) FPGA. On average, our experimental results show that our proposed A3 architecture outperforms A2 and A1 by 34% and 250%, respectively. Within the performance improvement of A3 over A2, more than 70% improvement comes from the hardware graph-based memory scheduling. Yu Bai 0004, Abigail Fuentes-Rivera, Michael Riera, Mohammed Alawad, Mingjie Lin |
FCCM | 1 |
| 2013 | Exploiting algorithmic-level memory parallelism in distributed logic-memory architecture through hardware-assisted dynamic graph (abstract only)abstractEmerging FPGA device, integrated with abundant RAM blocks and high-performance processor cores, offers an unprecedented opportunity to effectively implement single-chip distributed logic-memory (DLM) architectures. Being "memory-centric", the DLM architecture can significantly improve the overall performance and energy efficiency of many memory-intensive embedded applications, especially those that exhibit irregular array data access patterns at algorithmic level. However, implementing DLM architecture poses unique challenges to an FPGA designer in terms of 1) organizing and partitioning diverse on-chip memory resources, and 2) orchestrating effective data transmission between on-chip and off-chip memory. In this paper, we offer our solutions to both of these challenges. Specifically, 1) we propose a stochastic memory partitioning scheme based on the well-known simulated annealing algorithm. It obtains memory partitioning solutions that promote parallelized memory accesses by exploring large solution space; 2) we augment the proposed DLM architecture with a reconfigure hardware graph that can dynamically compute precedence relationship between memory partitions, thus effectively exploiting algorithmic level memory parallelism on a per-application basis. We evaluate the effectiveness of our approach (A3) against two other DLM architecture synthesizing methods: an algorithmic-centric reconfigurable computing architectures with a single monolithic memory (A1) and the heterogeneous distributed architectures synthesized according to (A2). All experiments have been conducted with a Virtex-5 (XCV5LX155T-2) FPGA. On average, our experimental results show that our proposed A3 architecture outperforms A2 and A1 by 34% and 250%, respectively. Within the performance improvement of A3 over A2, more than 70% improvement comes from the hardware graph-based memory scheduling. Yu Bai 0004, Abigail Fuentes-Rivera, Mingjie Lin, Mike Riera |
FPGA | 1 |
| 2011 | Design of Asynchronous Circuits on FPGAs for Soft Error ToleranceabstractIn this paper, we investigate the mechanism of soft error generation, propagation in asynchronous circuits which are implemented on FPGA. We also proposed the circuit to detect the soft errors which propagate in asynchronous Pipelines. The effects of the soft errors on Quasi-delay-insensitive (QDI) asynchronous circuits are analyzed and detected. The simulation results show that the proposed detect circuit can detect the soft error in asynchronous circuits implemented on FPGAs easily so that FPGAs can be reprogrammed, compared with traditional synchronous circuits. Yu Bai 0004, Weidong Kuang |
DSD | 1 |