Miaoqing Huang

dblp:45/5897 · DBLP profile ↗
← Back
36ranked-venue papers
11as first author
14since 2021 · last 2025
0000-0001-7376-3744ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 32 · 11 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Optimized Coding and Parameter Selection for Efficient FPGA Design of Attention Mechanisms
abstract
Efficient utilization of on-chip computational and memory resources, along with optimized high-level synthesis (HLS) coding, is vital to maximize parallelism and minimize latency. This paper demonstrates the HLS algorithms to achieve high utilization of processing elements to enhance parallelism. It also analyzes how various parameters of an attention layer impact latency, employs an efficient tiling technique, and explains the process of selecting an optimized tile size (TS).
Ehsan Kabir, Austin R. J. Downey, Jason D. Bakos, David Andrews 0001, Miaoqing Huang
FCCM5
2025 N-TORC: Native Tensor Optimizer for Real-Time Constraints
abstract
Compared to overlay-based tensor architectures like VTA or Gemmini, compilers that directly translate machine learning models into a dataflow architecture as HLS code, such as HLS4ML and FINN, generally can achieve lower latency by generating customized matrix-vector multipliers and memory structures tailored to the specific fundamental tensor operations required by each layer. However, this approach has significant drawbacks: the compilation process is highly time-consuming and the resulting deployments have unpredictable area and latency, making it impractical to constrain the latency while simultaneously minimizing area. Currently, no existing methods address this type of optimization. In this paper, we present N-TORC (Native Tensor Optimizer for Real-Time Constraints), a novel approach that utilizes data-driven performance and resource models to optimize individual layers of a dataflow architecture. When combined with model hyperparameter optimization, N-TORC can quickly generate architectures that satisfy latency constraints while simultaneously optimizing for both accuracy and resource cost (i.e. offering a set of optimal trade-offs between cost and accuracy). To demonstrate its effectiveness, we applied this framework to a cyber-physical application, DROPBEAR (Dynamic Reproduction of Projectiles in Ballistic Environments for Advanced Research). N-TORC's HLS4ML performance and resource models achieve higher accuracy than prior efforts, and its Mixed Integer Program (MIP)-based solver generates equivalent solutions to a stochastic search in 1000X less time.
Suyash Vardhan Singh, Iftakhar Ahmad, David Andrews 0001, Miaoqing Huang, Austin R. J. Downey, Jason D. Bakos
FCCM4
2025 Resource Scheduling for Real-Time Machine Learning
Suyash Vardhan Singh, Iftakhar Ahmad, David Andrews 0001, Miaoqing Huang, Austin R. J. Downey, Jason D. Bakos
FPGA4
2025 DA-VinCi: A Deep-Learning Accelerator Overlay Using In-Memory Computing
abstract
The matrix operations that underpin today’s deep learning models are routinely implemented in Single Instruction Multiple Data (SIMD) domain specific accelerators. SIMD accelerators including GPUs and array processors can effectively leverage parallelism in models that are compute-bound, but their effectiveness can be diminished for models that are memory-bound. Processing-in-Memory (PIM) architectures are being explored to provide better energy efficiency and scalable performance for these memory-bound models. Modern Field Programmable Gate Arrays (FPGAs) feature hundreds of megabits of Static Random Access Memory (SRAM) distributed across the device as disaggregated memory resources. This makes FPGAs ideal programmable platforms for developing custom Processor In/Near Memory accelerators. Several PIM array-based accelerator designs have been proposed to leverage this substantial internal bandwidth. However, results reported to date show the FPGA based PIM architectures operating at system clock frequencies well below a chips Block-RAM (BRAM) Fmax clock frequency. Results also show that the compute densities of the designs do not scale linearly with BRAM densities. These results indicate that FPGA PIM architectures will never be competitive with their custom Application-Specific Integrated Circuit (ASIC) counterparts. In this article, we introduce DA-VinCi, a D eep-Learning A ccelerator O v erlay using In -Memory C omput i ng. DA-VinCi is the first scalable FPGA based PIM deep-learning accelerator overlay capable of clocking at the maximum frequency of a device’s BRAM. Further, the architecture of DA-VinCi allows the number of compute units to scale linearly up to the maximum capacity of a devices BRAM, and at the maximum clock frequency of the BRAM. The DA-VinCi overlay has a programmable Instruction Set Architecture (ISA) that allows the same synthesized design to provide low-latency inferencing of a range of memory-bound deep-learning models, including Multilayer Perceptrons, Recurrent Neural Network, Long Short-Term Memory, and Gated Recurrent Unit networks. The scalability and high clocking frequency of DA-VinCi is achieved through a new Processor In Memory (PIM) tile architecture and a highly scalable system-level framework. We present results showing DA-VinCi linearly scaling the number of Processing Elements (PEs) to 100% of the BRAM capacity (over 60K PEs) on an Alveo U55 clocking at 737 MHz, the chips BRAM Fmax. We provide comparative studies on inference latency across multiple deep-learning applications that show DA-VinCi achieves up to a 201 \(\times\) improvement over a state-of-the-art PIM overlay accelerator, up to 87 \(\times\) improvement over existing PIM-based FPGA accelerators, and up to 57 \(\times\) improvement over custom deep-learning accelerators on FPGAs.
M. D. Arafat Kabir, Nathaniel Fredricks, Tendayi Kamucheka, Joel Mandebi, Miaoqing Huang, Jason D. Bakos, David Andrews 0001
ACM Trans. Reconfigurable Technol. Syst.5
2024 The BRAM is the Limit: Shattering Myths, Shaping Standards, and Building Scalable PIM Accelerators
abstract
Many recent FPGA-based Processor-in-Memory (PIM) architectures have appeared with promises of impressive levels of parallelism but with performance that falls short of expectations due to reduced maximum clock frequencies, an inability to scale processing elements up to the maximum BRAM capacity, and minimal hardware support for large reduction operations. In this paper, we propose a “Standard” set of design objectives for PIM array-based FPGA designs. We then propose a PIM array-based GEMV accelerator architecture as a case study to show the proposed Standard can be realized in practice. The GEMV accelerator serves as existence proof that dispels several myths surrounding what is normally accepted as clocking and scaling FPGA performance limitations. Specifically, the proposed accelerator clocks at the maximum frequency of the BRAM and scales to 100% of the available BRAMs. Comparative analyses show execution speeds over existing PIM-based GEMV engines on FPGAs and achieving a 2.65Χ – 3.2Χ faster clock. An AMD Alveo U55 implementation achieves a system clock speed of 737 MHz, providing 64K bit serial multiply-accumulate (MAC) units for GEMV operation.
M. D. Arafat Kabir, Tendayi Kamucheka, Nathaniel Fredricks, Joel Mandebi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001
FCCM6
2024 IMAGine: An In-Memory Accelerated GEMV Engine Overlay
abstract
Processor-in-Memory (PIM) overlays and alternative reconfigurable tile fabrics have been proposed to eliminate the von Neumann bottleneck and enable processing performance to scale with BRAM capacity. The performance of these FPGA-based PIM architectures has been limited due to a reduction of the BRAMs maximum clock frequencies and less than ideal scaling of processing elements with increased BRAM capacity. This paper presents IMAGine, an In-Memory Accelerated GEMV engine, a PIM-array accelerator that clocks at the maximum frequency of the BRAM and scales to 100% of the available BRAMs. Comparative analyses are presented showing execution speeds over existing PIM-based GEMV engines on FPGAs and achieving a $2.65 \times-3.2 \times$ faster clock. An AMD Alveo U55 implementation is presented that achieves a system clock speed of 737 MHz, providing 64 K bit-serial multiply-accumulate (MAC) units for GEMV operation. This establishes IMAGine as the fastest PIM-based GEMV overlay, outperforming even the custom PIM-based FPGA accelerators reported to date. Additionally, it surpasses TPU v1-v2 and Alibaba Hanguang 800 in clock speed while offering an equal or greater number of multiply-accumulate (MAC) units.
M. D. Arafat Kabir, Tendayi Kamucheka, Nathaniel Fredricks, Joel Mandebi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001
FPL6
2024 A Reliable and Efficient Online Solution for Adaptive Voltage and Frequency Scaling on FPGAs
abstract
Adaptive voltage and frequency scaling (AVFS) technology adjusts the supply voltage and clock frequency based on the actual operating conditions of the circuit. It can significantly improve performance or reduce the power consumption of the device. Existing online field-programmable gate array (FPGA) AVFS solutions have relatively low adjustment efficiency. Many existing solutions rely on offline steps, which do not consider the runtime operating conditions. This article proposes a complete FPGA AVFS solution, which includes a versatile self-checking timing monitor (SCTM) with small resource overhead, efficient AVFS algorithms without any offline steps, and user-friendly comprehensive automation software. Compared with existing online solutions, the proposed solution improves scaling efficiency by reducing the number of configuration times for the clock generation unit. The effectiveness of the solution is evaluated by a set of pubic benchmarks. Experimental results indicate that it can set an appropriate voltage–frequency operating point for the application circuit within dozens of milliseconds. For power-oriented adjustment, the proposed solution can save power ranging from 33.93% to 43.46%, while keeping the frequency not slower than the one reported by the static timing analysis (STA). For performance-oriented adjustment, it can achieve a performance improvement ranging from 60.26% to 101.90% at the nominal voltage.
Jiacheng Cao, YaoZhang Liu, Jian Wang 0036, Jinmei Lai 0001, Miaoqing Huang
IEEE Trans. Very Large Scale Integr. Syst.6
2023 Making BRAMs Compute: Creating Scalable Computational Memory Fabric Overlays
abstract
The increasing density of distributed BRAMs diffused throughout modern Field Programmable Gate Arrays (FP-GAs) is ideal for forming processor in/near memory architectures. This breaks the traditional von Neumann memory bottleneck limiting concurrency and degrading energy efficiency. Ideally, processing density should scale linearly with BRAM capacity, and clock frequencies should be set by the read/write access times of the BRAM. In this paper, we present a PIM overlay that achieves these goals. We observe an improvement of performance by 2.25 x, logic resource utilization by 2 x, and accumulation delay by 17 x compared to prior published work.
M. D. Arafat Kabir, Joshua Hollis, Atiyehsadat Panahi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001
FCCM5
2023 Accelerating LSTM-Based High-Rate Dynamic System Models
abstract
In this paper, we evaluate the use of a trained Long Short-Term Memory (LSTM) network as a surrogate for a Euler-Bernoulli beam model, and then we describe and characterize an FPGA-based deployment of the model for use in real-time structural health monitoring applications. The focus of our efforts is the DROPBEAR (Dynamic Reproduction of Projectiles in Ballistic Environments for Advanced Research) dataset, which was generated as a benchmark for the study of real-time structural modeling applications. The purpose of DROPBEAR is to evaluate models that take vibration data as input and give the initial conditions of the cantilever beam on which the measurements were taken as output. DROPBEAR is meant to serve an exemplar for emerging high-rate “active structures” that can be actively controlled with feedback latencies of less than one microsecond. Although the Euler-Bernoulli beam model is a well-known solution to this modeling problem, its computational cost is prohibitive for the time scales of interest. It has been previously shown that a properly structured LSTM network can achieve comparable accuracy with less workload, but achieving sub-microsecond model latency remains a challenge. Our approach is to deploy the LSTM optimized specifically for latency on FPGA. We designed the model using both high-level synthesis (HLS) and hardware description language (HDL). The lowest latency of$1.42\ \mu\mathrm{S}$and the highest throughput of 7.87 Gops/s were achieved on Alveo U55C platform for HDL design.
Ehsan Kabir, Daniel Coble, Joud N. Satme, Austin R. J. Downey, Jason D. Bakos, David Andrews 0001, Miaoqing Huang
FPL7
2023 FPGA Processor In Memory Architectures (PIMs): Overlay or Overhaul ?
abstract
The dominance of machine learning and the ending of Moore's law have renewed interests in Processor in Memory (PIM) architectures. This interest has produced several recent proposals to modify an FPGA's BRAM architecture to form a next-generation PIM reconfigurable fabric [1], [2]. PIM architectures can also be realized within today's FPGAs as overlays without the need to modify the underlying FPGA architecture. To date, there has been no study to understand the comparative advantages of the two approaches. In this paper, we present a study that explores the comparative advantages between two proposed custom architectures and a PIM overlay running on a commodity FPGA. We created PiCaSO, a Processor in/near Memory Scalable and Fast Overlay architecture as a representative PIM overlay. The results of this study show that the PiCaSO overlay achieves up to 80% of the peak throughput of the custom designs with 2.56 x shorter latency and 25% - 43% better BRAM memory utilization efficiency. We then show how several key features of the PiCaSO overlay can be integrated into the custom PIM designs to further improve their throughput by 18%, latency by 19.5%, and memory efficiency by 6.2%.
M. D. Arafat Kabir, Ehsan Kabir, Joshua Hollis, Eli Levy-Mackay, Atiyehsadat Panahi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001
FPL7
2022 BrainVGAE: End-to-End Graph Neural Networks for Noisy fMRI Dataset
abstract
Graph Neural Networks (GNNs), a deep learning model for non-Euclidean data structures, have shown significant improvement in brain-related intelligent tasks (neuroimaging, brain clustering). A graph constructed from brain atlas is input of GNNs for various tasks. However, the choice of node connectivity (edges) receives little attention in current works, the performance thereby degrades when dealing with noisy dataset. In this paper, we propose an end-to-end framework to boost the performance robustness on noisy functional Magnetic Resonance Imaging (fMRI) dataset, using a Variational Graph Auto-Encoders (VGAE)-based edge predictor.
Quan Mai, Ukash Nakarmi, Miaoqing Huang
BIBM3
2022 High-Rate Machine Learning for Forecasting Time-Series Signals
abstract
"Active structures" are physical structures that incorporate real-time monitoring and control. Examples include active vibration damping or blast mitigation systems. Evaluating physics-based models in real-time is generally not feasible for such systems having high-rate dynamics which require microsecond response times, but data-driven machine-learning-based models can potentially offer a solution. This paper compares the cost and performance of two FPGA-based implementations of real-time, continuously-trained models for forecasting time-series signals with non-stationarities, with one using High-Level Synthesis (HLS) and the other a programmable overlay architecture. The proposed model accepts a uni-variate vibration signal and seeks to forecast future samples to inform high-rate controllers. The proposed forecasting method performs two concurrent neural inference operations. One inference forecasts the state of the signal f samples into the future as a function of the most recent h samples, while the other forecasts the current sample given h samples starting from h+f−1 samples into the past. The first forecast produces the forecast while the second forecast allows the system to calculate the model’s loss and perform an immediate model update before the next sample period.
Atiyehsadat Panahi, Ehsan Kabir, Austin R. J. Downey, David Andrews 0001, Miaoqing Huang, Jason D. Bakos
FCCM5
2022 A Masked Pure-Hardware Implementation of Kyber Cryptographic Algorithm
abstract
Quantum computing-specifically Shor's algorithm [1]-presents an existential threat to some standard cryptographic algorithms. In preparation, post-quantum cryptography (PQC) algorithms have been in development and are nearing mathematical and cryptanalytic maturity. Standardization efforts through the National Institute of Standards and Technology (NIST) PQC standardization process have chosen one PKE/KEM algorithm (i.e., CRYSTALS-Kyber) and three digital signature algorithms (i.e., CRYSTALS-Dilithium, Falcon, and SPHINCS+). CRYSTALS-Kyber is a lattice-based, IND-CCA2-secure, key-encapsulation mechanism (KEM) based on the learning-with-errors problem over module lattices. This paper presents a masked hardware implementation of Kyber that is demonstrably secure against side-channel power analysis methods.
Tendayi Kamucheka, Alexander Nelson 0001, David Andrews 0001, Miaoqing Huang
FPT4
2022 Self-Supervised Domain Adaptation in Crowd Counting
abstract
Self-training crowd counting has not been attentively explored though it is one of the important challenges in computer vision. In practice, the fully supervised methods usually require an intensive resource of manual annotation. In order to address this challenge, this work introduces a new approach to utilize existing datasets with ground truth to produce more robust predictions on unlabeled datasets, named domain adaptation, in crowd counting. While the network is trained with labeled data, samples without labels from the target domain are also added to the training process. In this process, the entropy map is computed and minimized in addition to the adversarial training process designed in parallel. Experiments on Shanghaitech, UCF_CC_50, and UCF-QNRF datasets prove a more generalized improvement of our method over the other state-of-the-arts in the cross-domain setting.
Pha A. Nguyen, Thanh-Dat Truong, Miaoqing Huang, T. Hoang Ngan Le, Khoa Luu
ICIP3
2019 Generic attribute revocation systems for attribute-based encryption in cloud storage
abstract
Attribute-based encryption (ABE) has been a preferred encryption technology to solve the problems of data protection and access control, especially when the cloud storage is provided by third-party service providers. ABE can put data access under control at each data item level. However, ABE schemes have practical limitations on dynamic attribute revocation. We propose a generic attribute revocation system for ABE with user privacy protection. The attribute revocation ABE (AR-ABE) system can work with any type of ABE scheme to dynamically revoke any number of attributes.
Genlang Chen, Zhiqian Xu 0001, Jiajian Zhang, Guojun Wang 0001, Hai Jiang 0003, Miaoqing Huang
Frontiers Inf. Technol. Electron. Eng.6
2017 PolyPC: Polymorphic parallel computing framework on embedded reconfigurable system
abstract
With the help of parallelism provided by the fine-grained architecture, hardware accelerators on Field Programmable Gate Arrays (FPGAs) can significantly improve the performance of many applications. However, designers are typically required to have excellent hardware programming skills and unique optimization techniques to fully explore the potential of FPGA resources. In this work, we propose the PolyPC (Polymorphic Parallel Computing) framework that aims to improve productivity while achieving performance speedup. The PolyPC framework implements a custom hardware platform on which the PolyPC framework extends vendor-provided tools to convert OpenCL-like programs into executables by using highlevel synthesis (HLS) tools. The PolyPC framework is evaluated regarding performance, area efficiency, and multitasking. The results show a maximum of 66 folds of speedup over a dual-core ARM processor, and 1,043 folds of speedup over a high-performance MicroBlaze soft processor, with 125 folds of area efficiency. In addition, it delivers a significant improvement in response time to high-priority PolyTasks with the priority-aware scheduling.
Hongyuan Ding, Miaoqing Huang
FPL2
2016 A hybrid parallel cellular automata model for urban growth simulation over GPU/CPU heterogeneous architectures
abstract
As an important spatiotemporal simulation approach and an effective tool for developing and examining spatial optimization strategies (e.g., land allocation and planning), geospatial cellular automata (CA) models often require multiple data layers and consist of complicated algorithms in order to deal with the complex dynamic processes of interest and the intricate relationships and interactions between the processes and their driving factors. Also, massive amount of data may be used in CA simulations as high-resolution geospatial and non-spatial data are widely available. Thus, geospatial CA models can be both computationally intensive and data intensive, demanding extensive length of computing time and vast memory space. Based on a hybrid parallelism that combines processes with discrete memory and threads with global memory, we developed a parallel geospatial CA model for urban growth simulation over the heterogeneous computer architecture composed of multiple central processing units (CPUs) and graphics processing units (GPUs). Experiments with the datasets of California showed that the overall computing time for a 50-year simulation dropped from 13,647 seconds on a single CPU to 32 seconds using 64 GPU/CPU nodes. We conclude that the hybrid parallelism of geospatial CA over the emerging heterogeneous computer architectures provides scalable solutions to enabling complex simulations and optimizations with massive amount of data that were previously infeasible, sometimes impossible, using individual computing approaches.
Qingfeng Guan 0001, Xuan Shi, Miaoqing Huang, Chenggang Lai
Int. J. Geogr. Inf. Sci.3
2015 Performance and Energy Optimization on MPSoCs by Enabling STT-MRAM LUTs
abstract
Increasing computation demands with limited power budget require more energy-efficient design without performance degradation in embedded systems and mobile computing platforms. Reconfigurable computing is an alternative to optimize both performance and power consumption. However, due to the complexity of hardware design, implementing dedicated accelerators usually lacks flexibility and productivity. In this work, our previous hybrid parallel co-design framework is extended. By using the partial reconfiguration technique, a dynamic reconfiguration scheme is presented to optimize both performance and power consumption without losing the programming flexibility. In addition, spin-torque transfer magneto resistive RAM (STT-MRAM) LUTs are exploited to replace traditional SRAM-based LUTs for further reducing the static power consumption. The results show that dynamic scheduling with hardware kernels implemented in STT-MRAM LUTs is 7.7 times better than the purely software implementation in terms of the product between the performance and the energy. Besides, when comparing STT-MRAM and SRAM hardware kernels, the STT-MRAM kernels have significant advantages over SRAM ones on the power consumption, especially the static power consumption.
Hongyuan Ding, Miaoqing Huang
FCCM2
2015 An Automatic Design Flow for Hybrid Parallel Computing on MPSoCs (Abstract Only)
abstract
State-of-the-art high-level synthesis (HLS) tools are able to lower the threshold for designers to exploit performance benefits of hardware accelerators. However, it is still a challenge to achieve parallelism on a hybrid multiprocessor system-on-chip (MPSoC). In this work, we present an automatic hybrid design flow. The hybrid hardware platform as well as both the hardware and software kernels can be generated through this flow. In addition, a hybrid OpenCL-like programming model is proposed to combine software and hardware kernels running on the unified hardware platform. Our results show that our automatic design flow can not only significantly minimize the development time, but also gain about 11 times speedup compared with pure software parallel implementation for a matrix multiplication benchmark.
Hongyuan Ding, Miaoqing Huang
FPGA2
2014 A Hierarchical Memory Architecture with NoC Support for MPSoC on FPGAs
abstract
This work presents a memory hierarchy with the support of network-on-chip (NoC) for MPSoC systems. The memory hierarchy consists of a shared global memory and private local memories as shown in Figure 1. Each core in the system is equipped with two local memories, one for instructions and one for data. The MicroBlaze soft core used in this work connects the main bus through the PLB interface and connects the local memory modules through the LMB interface. Further it connects to a 4x4 mesh NoC through the FSL interface, as shown in Figure 2(a). We built the generic NoC (NoC-g) using the open-source router designed by the Concurrent VLSI Architecture group at the Stanford University [2]. Each router has 5 input ports and 5 output ports. Each input physical channel and each output physical channel is connected to 4 input virtual channels and 4 output virtual channels, respectively. The 40 virtual channels are connected to an internal crossbar switch for routing. We designed the adapter to connect the MicroBlaze processor to the router.
Miaoqing Huang, Hongyuan Ding, Sen Ma
FCCM2
2014 Improve memory access for achieving both performance and energy efficiencies on heterogeneous systems
abstract
Hardware accelerators are capable of achieving significant performance improvement for many applications. In this work we demonstrate that it is critical to provide sufficient memory access bandwidth for accelerators to improve the performance and reduce energy consumption. We use the scale-invariant feature transform (SIFT) algorithm as a case study in which three bottleneck stages are accelerated on hardware logic. Based on different memory access patterns of SIFT algorithms, two different approaches are designed to accelerate different functions in SIFT on the Xilinx Zynq-7045 device. In the first approach, convolution is accelerated by designing fully customized hardware accelerator. On top of it, three interfacing methods are analyzed. In the second approach, a distributed multi-processor hardware system with its programming model is built to handle inconsecutive memory accesses. Furthermore, the last level cache (LLC) on the host processor is shared by all slaves to achieve better performance. Experiment results on the Zynq-7045 device show that the hybrid design in which two approaches are combined can achieve ~10 times and better improvement for both performance improvement and energy reduction compared with the pure software implementation for the convolution stage and the SIFT algorithm, respectively.
Hongyuan Ding, Miaoqing Huang
FPT2
2013 A Delay-based PUF Design Using Multiplexers on FPGA
abstract
Summary form only given. Physically unclonable functions (PUFs) have been a hot research topic in hardware-oriented security for many years. Given a challenge as an input to the PUF, it generates a corresponding response, which can be treated as a unique fingerprint or signature for authentication purpose. In this paper, a delay-based PUF design involving multiplexers on FPGA is presented. Due to the intrinsic difference of the switching latencies of two chained multiplexers, a positive pulse may be produced at the output of the downstream multiplexer. This pulse can be used to set the output of a D flip-flop to `1'. The proposed design improves the randomness of the outputs of the PUF.
Miaoqing Huang
FCCM1
2013 Improve Effective Capacity and Lifetime of Solid State Drives
abstract
Flash-based SSDs are becoming increasingly popular in modern storage systems, especially in high-performance computing infrastructures. However, several inherent technical limitations still remain to prevent their widespread deployment. One of the critical concerns is their limited lifetime, which is directly relevant to the total writes experienced by SSDs. In this paper, we present a Content and semantics Aware File System (CSA-FS) which is able to reduce write traffic to SSDs. It employs deduplication and delta-encoding techniques to file system data blocks and semantic blocks, respectively. It is motivated by two important observations: (1) there exists a huge amount of content redundancy within primary storage systems, and (2) semantic blocks are visited much more frequently than data blocks, with each update bringing very minimal changes. By separately deduplicating redundant data blocks and delta-encoding similar semantic blocks, CSA-FS can significantly reduce the total write traffic to SSDs and greatly improve their lifetime correspondingly, at an acceptable cost of at most 7% performance degradation across a variety of workloads.
Ping Huang 0001, Guangping Wan, Ke Zhou 0001, Miaoqing Huang, Chun-hua Li, Hua Wang 0008
NAS4
2013 Modular Design of Fully Pipelined Reduction Circuits on FPGAs
abstract
Fast and efficient reduction circuits are critical for a broad range of scientific and embedded system applications. High throughput reduction circuits are typically hand designed for specific vector lengths. These circuits need to be modified when the set lengths are changed. In this paper, we present a new design approach that can handle any set length or combination of different consecutive set lengths without stalling and generates in-order results. The flexibility of the design allows it to be used for any reduction operations, such as floating-point addition and multiplication. By providing a simple and efficient interface to the user and a modular architecture for the designer, the proposed technique has a broad impact across a wide range of custom hardware designs.
Miaoqing Huang, David Andrews 0001
IEEE Trans. Parallel Distributed Syst.1
2012 Automating the design of mLUT MPSoPC FPGAs in the cloud
abstract
Modern platform FPGAs are over the million-LUT level, large enough to support complete heterogeneous Multiprocessor System-On-Chips (MPSoCs). Constructing systems with 10's of processors is currently feasible using existing manual methods within vendor-specific CAD tools. However these manual, by-hand, approaches will not be feasible for constructing future systems with 100's to 1,000's of processors. Instead, new automated system assembly approaches will be required to handle these levels of system complexity and diversity. In this paper we present a new automated design flow for creating such next generation heterogeneous MPSoCs. An integral part of the MPSoPC system created is the inclusion of a general purpose PThreads-compliant HW/SW co-designed operating system and heterogeneous compiler. Our design flow has been placed in the cloud and is freely accessible across the Internet.
Eugene Cartwright, Azad Fahkari, Sen Ma, Christina Smith, Miaoqing Huang, David Andrews 0001, Jason Agron
FPL5
2012 Improving the Performance of On-Board Cache for Flash-Based Solid-State Drives
abstract
Flash-based Solid-State Drives (SSDs) are data storage devices that use flash memory to store persistent data. Previously we presented a centralized on-board cache for SSDs to improve the response time and reduce the physical writes to the flash media. An automatic periodic update (APU) feature with fixed period is used to write dirty and stable cache lines to flash media when the SSD is idle. In this work we propose a distributed on-board cache architecture to match the intrinsic parallelism of SSDs. Since there are N individual flash packages in one SSD device, N dirty cache lines are selected by the APU feature, each of which is written to a separate flash package. Simulation results show that this distributed on-board cache can significantly reduce the response time of SSDs by up to 90% compared with the SSDs without such a cache, while reducing the number of physical writes to the flash memory.
Miaoqing Huang, Liang Men
NAS1
2012 Efficient Mapping of Task Graphs onto Reconfigurable Hardware Using Architectural Variants
abstract
High-performance reconfigurable computing involves acceleration of significant portions of an application using reconfigurable hardware. Mapping application task graphs onto reconfigurable hardware is, therefore, of rising attention. In this work, we approach the mapping problem by incorporating multiple architectural variants for each hardware task; the variants reflect tradeoffs between the logic resources consumed and the task execution throughput. We propose a mapping approach based on the genetic algorithm, and show its effectiveness for random task graphs as well as an N-body simulation application, demonstrating improvements of up to 78.6 percent in the execution time compared with choosing a fixed implementation variant for all tasks. We then validate our methodology through experiments on real hardware, an SRC-6 reconfigurable computer.
Miaoqing Huang, Vikram K. Narayana, Mohamed Bakhouya, Jaafar Gaber, Tarek A. El-Ghazawi
IEEE Trans. Computers1
2011 New Hardware Architectures for Montgomery Modular Multiplication Algorithm
abstract
Montgomery modular multiplication is one of the fundamental operations used in cryptographic algorithms, such as RSA and Elliptic Curve Cryptosystems. At CHES 1999, Tenca and Koç proposed the Multiple-Word Radix-2 Montgomery Multiplication (MWR2MM) algorithm and introduced a now-classic architecture for implementing Montgomery multiplication in hardware. With parameters optimized for minimum latency, this architecture performs a single Montgomery multiplication in approximately 2n clock cycles, where n is the size of operands in bits. In this paper, we propose two new hardware architectures that are able to perform the same operation in approximately n clock cycles with almost the same clock period. These two architectures are based on precomputing partial results using two possible assumptions regarding the most significant bit of the previous word. These two architectures outperform the original architecture of Tenca and Koç in terms of the product latency times area by 23 and 50 percent, respectively, for several most common operand sizes used in cryptography. The architecture in radix-2 can be extended to the case of radix-4, while preserving a factor of two speedup over the corresponding radix-4 design by Tenca, Todorov, and Koç from CHES 2001. Our optimization has been verified by modeling it using Verilog-HDL, implementing it on Xilinx Virtex-II 6000 FPGA, and experimentally testing it using SRC-6 reconfigurable computer.
Miaoqing Huang, Kris Gaj, Tarek A. El-Ghazawi
IEEE Trans. Computers1
2010 Reaping the Processing Potential of FPGA on Double-Precision Floating-Point Operations: An Eigenvalue Solver Case Study
abstract
Many scientific applications such as electromagnetics require their operations carried out in double-precision floating-point format. The efficiency of these applications is mainly subject to the floating-point processing performance on the target processors. In this work, we use an eigenvalue solver application as a case study to demonstrate the processing potential of an FPGA device when dealing with floating-point operations. Relying on deep pipelines and large local memory directly accessible within the FPGA device, more than 20 × performance improvement has been achieved for this particular case whose computation is intrinsically sequential. The methodology described in this paper can be conveniently applied to other domain applications that share similar data processing characteristics, demonstrating an impact on a broad range of scientific fields, e.g., numerical linear algebra.
Miaoqing Huang, Özlem Kilic
FCCM1
2010 Modular design of fully pipelined accumulators
abstract
Fast and efficient accumulation arithmetic circuits are critical for a broad range of scientific and embedded system applications. High throughput accumulation circuits are typically hand designed for specific vector lengths requiring the circuit to be modified when the lengths are changed. In this work we present a new design approach that can achieve low latency and near optimal throughput for input data vectors of arbitrary length. The flexibility of the design allows it to be used for both integer and floating-point operations. By providing a simple and efficient interface to the user and a modular architecture for the designer, the proposed technique has broad impact across a wide range of custom hardware designs.
Miaoqing Huang, David Andrews 0001
FPT1
2010 Reconfiguration and Communication-Aware Task Scheduling for High-Performance Reconfigurable Computing
abstract
High-performance reconfigurable computing involves acceleration of significant portions of an application using reconfigurable hardware. When the hardware tasks of an application cannot simultaneously fit in an FPGA, the task graph needs to be partitioned and scheduled into multiple FPGA configurations, in a way that minimizes the total execution time. This article proposes the Reduced Data Movement Scheduling (RDMS) algorithm that aims to improve the overall performance of hardware tasks by taking into account the reconfiguration time, data dependency between tasks, intertask communication as well as task resource utilization. The proposed algorithm uses the dynamic programming method. A mathematical analysis of the algorithm shows that the execution time would at most exceed the optimal solution by a factor of around 1.6, in the worst-case. Simulations on randomly generated task graphs indicate that RDMS algorithm can reduce interconfiguration communication time by 11% and 44% respectively, compared with two other approaches that consider data dependency and hardware resource utilization only. The practicality, as well as efficiency of the proposed algorithm over other approaches, is demonstrated by simulating a task graph from a real-life application - N-body simulation - along with constraints for bandwidth and FPGA parameters from existing high-performance reconfigurable computers. Experiments on SRC-6 are carried out to validate the approach.
Miaoqing Huang, Vikram K. Narayana, Harald Simmler, Olivier Serres, Tarek A. El-Ghazawi
ACM Trans. Reconfigurable Technol. Syst.1
2009 Efficient Mapping of Hardware Tasks on Reconfigurable Computers Using Libraries of Architecture Variants
abstract
Scheduling and partitioning of task graphs on reconfigurable hardware needs to be carefully carried out in order to achieve the best possible performance. In this paper, we demonstrate that a significant improvement to the total execution time is possible by incorporating a library of hardware task implementations, which contains multiple architectural variants for each hardware task reflecting tradeoffs between the resources utilization and the task execution throughput. We develop a genetic algorithm based mapping approach, which considers both task graph and target platform, and present results for an N-body simulation application using estimated numbers for resource utilization for the constituent tasks and based on actual architectural constraints from different reconfigurable platforms. The results demonstrate improvements of up to 85.3% in the execution time, compared to choosing a fixed implementation variant for each task while keeping a reasonable searching time.
Miaoqing Huang, Vikram K. Narayana, Tarek A. El-Ghazawi
FCCM1
2009 RDMS: A hardware task scheduling algorithm for Reconfigurable Computing
abstract
Reconfigurable computers (RC) can provide significant performance improvement for domain applications. However, wide acceptance of today's RCs among domain scientist is hindered by the complexity of design tools and the required hardware design experience. Recent developments in HW/SW co-design methodologies for these systems provide the ease of use, but they are not comparable in performance to manual co-design. This paper aims at improving the overall performance of hardware tasks assigned to FPGA devices by minimizing both the communication overhead and configuration overhead, which are introduced by using FPGA devices. The proposed reduced data movement scheduling (RDMS) algorithm takes data dependency among tasks, hardware task resource utilization, and inter-task communication into account during the scheduling process and adopts a dynamic programming approach to reduce the communication between muP and FPGA co-processor and the number of FPGA configurations to a minimum. Compared to two other approaches that consider data dependency and hardware resource utilization only, RDMS algorithm can reduce inter-configuration communication time by 11% and 44% respectively based on simulation using randomly generated data flow graphs. The implementation of RDMS on a real-life application, N-body simulation, verifies the efficiency of RDMS algorithm against other approaches.
Miaoqing Huang, Harald Simmler, Olivier Serres, Tarek A. El-Ghazawi
IPDPS1
2008 Portable library development for reconfigurable computing systems: A case study
Proshanta Saha, Esam El-Araby, Miaoqing Huang, Mohamed Taher, Sergio López-Buedo, Tarek A. El-Ghazawi, Chang Shu 0003, Kris Gaj, Alan Michalski, Duncan A. Buell
Parallel Comput.3
2007 A Portable Memory Access Framework on Reconfigurable Computers
abstract
Current reconfigurable computers (RCs) do not share a unified architectural model, which presents a challenge to any developer who intends to port hardware designs across different RC platforms. In this paper, we propose a portable memory access framework that gives the user a unified memory view combining the host memory and the local memory of FPGA. Three memory access modes are provided, and the hardware cost and performance impact have been measured on three major RCs: SRC-6, SGI RC-100 and Cray XD1. Under current implementation, the penalty to hardware resource utilization and performance of applying this framework is reduced to minimum.
Miaoqing Huang, Iván González 0004, Tarek A. El-Ghazawi
FPT1
2005 Performance of Sorting Algorithms on the SRC 6 Reconfigurable Computer
John Harkins, Tarek A. El-Ghazawi, Esam El-Araby, Miaoqing Huang
FPT4