EDBT 2026 Demo / reviewers in the wild / expert
Hayden Kwok-Hay So
dblp:95/4575 · also Hayden K. H. So
· DBLP profile ↗
72ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0002-6514-0237ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 59 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 9 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher DistillationabstractHengyuan Zhang, Shiping Yang, Xiao Liang, Chenming Shang, Yuxuan Jiang, Chaofan Tao, Jing Xiong, Hayden Kwok-Hay So, Ruobing Xie, Angel X. Chang, Ngai Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chenming Shang, Chaofan Tao, Hayden Kwok-Hay So, Ruobing Xie, Angel X. Chang, Ngai Wong 0001 |
ACL (1) | 8 |
| 2026 | WireLightning: Harnessing Capacitances for In-Transit Massively Parallel Matrix MultiplicationabstractAnalog computing-in-memory accelerators promise ultra-low-power, on-device AI by reducing data transfer and energy usage. Yet inherent device variations and high energy consumption for analog-digital conversion continue to hinder their wide-scale adoption in mainstream systems. To address these issues, this paper introduces WireLightning, a novel capacitive-computing accelerator featuring a mixed-signal architecture that rethinks analog AI acceleration. Unlike conventional analog crossbars that encode weights in programmable devices, WireLightning exploits intrinsic charge dynamics in passive capacitors, encoding matrix multiplication through spike amplitude and timing. This design addresses critical limitations such as weight drift, stochasticity, and power-intensive ADC bottlenecks. Key innovations include: amplitude-temporal dual encoding that enables constant-time analog dot-products; time-based decoding scheme that significantly reduces reliance on power-intensive ADCs; row-wise parallel architecture for concurrent dot-product calculations across multiple rows to enhance throughput; and value repetition exploitation in low-bit quantized vectors to reduce multiplications to constant time complexity. A PCB pro totype achieved a range-normalized RMSE of 1.80%—69.2% of the error in RRAM crossbar implementations—and a normalized error of 9.18%, corresponding to 77.1% of that in leading PCM crossbars. Implemented in a 40-nm CMOS technology, WireLightning macro delivers 465.47 TOPS/W at 4-bit precision, outperforming state-of-the-art analog accelerators while maintaining 0.23% range-normalized RMSE. By integrating algorithm-circuit co-design with physical computing, this work establishes capacitive computing as a promising path toward combining digital precision and analog efficiency in next-generation edge AI. Song Wang 0023, Zhu Wang 0014, Can Li 0024, Hayden Kwok-Hay So |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer ReviewabstractYuan Chang, Ziyue Li, Hengyuan Zhang, Yuanbo Kong, Yanru Wu, Hayden Kwok-Hay So, Zhijiang Guo, Liya Zhu, Ngai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yuanbo Kong, Yanru Wu, Hayden Kwok-Hay So, Zhijiang Guo, Liya Zhu, Ngai Wong 0001 |
EMNLP | 6 |
| 2025 | A Hardware-Software Design Framework for SpMV Acceleration with Flexible Access Pattern PortfolioabstractSparse matrix-vector multiplications (SpMV) are notoriously challenging to accelerate due to their highly irregular data access pattern. Although a fully customized static accelerator design may be adequate for small problems that can fit entirely within an on-chip memory buffer, practical SpMV problems are large and have dynamic matrix structures that cannot easily be optimized at compile time. To address this need for trade-off between flexibility and performance, we present SPASM, a hardware-software framework that accelerates SpMV computation using a customizable portfolio of data access patterns as templates and a reconfigurable hardware to support their run-time execution. SPASM extracts local data access patterns of the input matrices and derives a set of template patterns to encode these inputs. Subsequently, a novel hardware computing structure is proposed to support vectorized computation and flexible switching between different template patterns for each tile computation. Furthermore, SPASM leverages the global compositions of input matrices to derive hardware configuration and workload schedules that improve load balancing among the parallel processing units. Importantly, although SPASM can optimize the pattern portfolio for a particular set of expected input matrices, the generated hardware can flexibly be used to accelerate SpMV of different input patterns albeit with reduced performance. Experimental results show that SPASM can achieve an average $2.81 \times$ speedup compared to the state-of-the-art SpMV accelerator while keeping a relatively low customization cost. Maolin Wang 0002, Hayden Kwok-Hay So |
HPCA | 3 |
| 2025 | Taijigraph: an Out-Of-Core Graph Processing System Enhanced with Computational StorageabstractOut-of-core graph processing systems are severely bottlenecked by I/O to the external storage because of the low compute-to-I/O ratio and the substantial amount of irregular data accesses. In order to alleviate the I/O bottleneck, prior works either focus on improving the bandwidth utilization by converting random I/O requests into sequential ones, or improving the data utilization by fetching only the required data to avoid the I/O redundancy. However, the former usually loads massive unused data, while the latter can induce frequent finegrained I/O requests, wasting the parallelism of the I/O channels and leading to under-utilization of the limited I/O bandwidth. Different from prior works, we systematically explore the use of computational storage devices (CSDs), which offer in-storage computing facilities with higher I/O bandwidth, to improve both the bandwidth utilization and data utilization for higher I/O efficiency. Specifically, we first introduce a graph-semanticaware data organization to enable the loading of only active graph partitions at the granularity of a physical page, reducing redundant I/O and enhancing data utilization. Additionally, we propose to coalesce parallel I/O requests of graph partitions distributed across different flash dies to maximize the parallelism of internal I/O channels, thereby fully utilizing the internal I/O bandwidth of CSDs. In addition, we capture the dynamic status of graph processing tasks across the iterations and partitions at runtime to dynamically offload I/O-intensive workloads into the instorage processors with restricted computing resources but higher I/O bandwidth to further improve the I/O efficiency. With the above techniques, we implement an out-of-core graph processing system prototype, namely TaijiGraph, on an open-channel CSD. According to our experiments on a set of representative graph datasets and algorithms, TaijiGraph achieves average speedups of$2.43 \times, 3.81 \times, 2.21 \times$and$7.89 \times$, respectively, when compared to state-of-the-art out-of-core graph processing systems including GridGraph, LUMOS, Blaze, and GraphSSD. Xinmiao Zhang 0004, Cheng Liu 0008, Shengwen Liang, Hayden Kwok-Hay So, Ying Wang 0001, Lei Zhang 0008, Huawei Li 0001, Xiaowei Li 0001 |
IPDPS | 4 |
| 2025 | SeerAttention: Self-distilled Attention Gating for Efficient Long-context PrefillingabstractAttention is the cornerstone of modern Large Language Models (LLMs). Yet its quadratic complexity hinders efficiency and scalability, especially for long-context processing. A promising approach is to leverage sparsity in attention. However, existing sparsity-based solutions predominantly rely on predefined patterns or heuristics at the attention head level, struggling to adapt dynamically to different contexts efficiently. We propose SeerAttention, a simple yet effective attention mechanism that directly learns the block-level attention sparsity from the LLM itself. Inspired by the gating mechanism in Mixture of Experts (MoE), SeerAttention augments the conventional attention with a **learnable gate** that **selectively activates important blocks** within the attention map. Specifically, the gate first pools the query (Q) and key (K) tensors along the sequence dimension and processes them through learnable linear layers. The resulting matrices are then multiplied together to produce the gating scores, which are used to predict block-level attention sparsity. Combined with our block-sparse FlashAttention kernel, SeerAttention can achieve significant speedup on GPUs. When applied to pre-trained LLMs, SeerAttention only requires training the gate parameters in a lightweight self-distillation manner, allowing rapid convergence. Our evaluation results demonstrate that SeerAttention achieves better model accuracy and lower latency for long-context pre-filling compared to prior methods. Code is available at: https://github.com/microsoft/SeerAttention. Yizhao Gao 0002, Zhichen Zeng 0002, Dayou Du, Shijie Cao, Peiyuan Zhou, Jiaxing Qi, Junjie Lai, Hayden Kwok-Hay So, Ting Cao 0003, Fan Yang 0024, Mao Yang 0004 |
NeurIPS | 8 |
| 2025 | TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic ArchitectureabstractModern transformer-based deep neural networks present unique technical challenges for effective acceleration in real-world applications. Apart from the vast amount of linear operations needed due to their sizes, modern transformer models are increasingly reliance on precise non-linear computations that make traditional low-bitwidth quantization methods and fixed-dataflow matrix accelerators ineffective for end-to-end acceleration. To address this need to accelerate both linear and non-linear operations in a unified and programmable framework, this article introduces TATAA. TATAA employs 8-bit integer ( int8 ) arithmetic for quantized linear layer operations through post-training quantization, while it relies on bfloat16 floating-point arithmetic to approximate non-linear layers of a transformer model. TATAA hardware features a transformable arithmetic architecture that supports both formats during runtime with minimal overhead, enabling it to switch between a systolic array mode for int8 matrix multiplications and a SIMD mode for vectorized bfloat16 operations. An end-to-end compiler is presented to enable flexible mapping from emerging transformer models to the proposed hardware. Experimental results indicate that our mixed-precision design incurs only 0.14% to 1.16% accuracy drop when compared with the pre-trained single-precision transformer models across a range of vision, language, and generative text applications. Our prototype implementation on the Alveo U280 FPGA currently achieves 2,935.2 GOPS throughput on linear layers and a maximum of 189.5 GFLOPS for non-linear operations, outperforming related works by up to \(1.45\times\) in end-to-end throughput and \(2.29\times\) in DSP efficiency, while achieving \(2.19\times\) higher power efficiency than modern NVIDIA RTX4090 GPU. Jiajun Wu 0006, Mo Song, Jingmin Zhao, Yizhao Gao 0002, Jia Li 0057, Hayden Kwok-Hay So |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2024 | A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGAabstractEvent-based vision represents a paradigm shift in how vision information is captured and processed. By only responding to dynamic intensity changes in the scene, event-based sensing produces far less data than conventional frame-based cameras, promising to springboard a new generation of high-speed, low-power machines for edge intelligence. However, processing such dynamically sparse input originated from event cameras efficiently in real time, particularly with complex deep neural networks (DNN), remains a formidable challenge. Existing solutions that employ GPUs and other frame-based DNN accelerators often struggle to efficiently process the dynamically sparse event data, missing the opportunities to improve processing efficiency with sparse data. To address this, we propose ESDA, a composable dynamic sparse dataflow architecture that allows customized DNN accelerators to be constructed rapidly on FPGAs for event-based vision tasks. ESDA is a modular system that is composed of a set of parametrizable modules for each network layer type. These modules share a uniform sparse token-feature interface and can be connected easily to compose an all-on-chip dataflow accelerator on FPGA for each network model. To fully exploit the intrinsic sparsity in event data, ESDA incorporates the use of submanifold sparse convolutions that largely enhance the activation sparsity throughout the layers while simplifying hardware implementation. Finally, a network architecture and hardware implementation co-optimizing framework that allows tradeoffs between accuracy and performance is also presented. Experimental results demonstrate that when compared with existing GPU and hardware-accelerated solutions, ESDA achieves substantial speedup and improvement in energy efficiency across different applications, and it allows much wider design space for real-world deployments. Yizhao Gao 0002, Baoheng Zhang, Yuhao Ding, Hayden Kwok-Hay So |
FPGA | 4 |
| 2024 | Multi-Issue Butterfly Architecture for Sparse Convex Quadratic ProgrammingabstractConvex quadratic optimization solvers are extensively utilized in various domains; however, achieving optimal performance in diverse situations remains a significant challenge due to the sparse nature of objective and constraint matrices. General-purpose architectures struggle with hardware utilization when performing critical sparse matrix operations, such as factorization and multiplication. To address this issue, we introduce a pipelined spatial architecture, Multi-Issue Butterfly (MIB), which supports all primitive scalar, vector, and matrix operations required by the Alternating Direction Method of Multipliers (ADMM) based solver algorithm. The proposed architecture features a butterfly computational network with innovative working modes for each node, controlled by runtime instructions. We developed a companion scheduling method for matrix operations based on their sparsity patterns. For factorization, an elimination tree guides the network instructions reordering to avoid data hazards caused by computation dependencies. For matrix-vector multiplication, data prefetching resolves structural hazards caused by read and write conflicts to register files. Instructions without hazards are issued simultaneously to increase pipeline throughput and function unit utilization. We evaluate the proposed architecture using FPGA prototypes, representing the first fully FPGA-based generic QP solver. Our assessment includes extensive performance and efficiency bench-marks across 100 QP problems from five application domains. Compared to the same algorithm variation running on CPU backends, our prototype achieves a geometric mean of$30.5\times$end-to-end speedup,$127.0 \times$greater energy efficiency, and$16.5\times$less runtime jitter. In comparison to GPU backends, the prototype attains a geometric mean of$4.3\times$faster end-to-end speedup,$21.7\times$higher energy efficiency, and$33.4\times$less runtime jitter. Maolin Wang 0002, Ian McInerney, Bartolomeo Stellato, Fengbin Tu, Stephen P. Boyd, Hayden Kwok-Hay So, Kwang-Ting Cheng |
MICRO | 6 |
| 2024 | DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network InferenceabstractTo accelerate the inference of deep neural networks (DNNs), quantization with low-bitwidth numbers is actively researched. A prominent challenge is to quantize the DNN models into low-bitwidth numbers without significant accuracy degradation, especially at very low bitwidths (< 8 bits). This work targets an adaptive data representation with variablelength encoding called DyBit. DyBit can dynamically adjust the precision and range of separate bit-fields to be adapted to the DNN weights/activations distribution. We also propose a hardware-aware quantization framework with a mixed-precision accelerator to trade-off the inference accuracy and speedup. Experimental results demonstrate that the ImageNet inference accuracy via DyBit is 1.97% higher than the state-of-the-art at 4-bit quantization, and the proposed framework can achieve up to 8.1× speedup compared with the original ResNet-50 model. Jiajun Zhou 0004, Jiajun Wu 0006, Yizhao Gao 0002, Yuhao Ding, Chaofan Tao, Fengbin Tu, Kwang-Ting Cheng, Hayden Kwok-Hay So, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2023 | DPACS: Hardware Accelerated Dynamic Neural Network Pruning through Algorithm-Architecture Co-designabstractBy eliminating compute operations intelligently based on the run time input, dynamic pruning (DP) promises to improve deep neural network inference speed substantially without incurring a major impact on their accuracy. Although many DP algorithms with good pruning performance have been proposed, it remains a challenge to translate these theoretical reductions in compute operations into satisfactory end-to-end speedups in practical real-world implementations. The overhead of identifying operations to be pruned during run time, the need to efficiently process the resulting dynamic dataflow, and the non-trivial memory I/O bottleneck that emerges as the number of compute operations reduces, have all contributed to the challenge of implementing practical DP systems. Yizhao Gao 0002, Baoheng Zhang, Xiaojuan Qi 0001, Hayden Kwok-Hay So |
ASPLOS (2) | 4 |
| 2023 | MSD: Mixing Signed Digit Representations for Hardware-efficient DNN Acceleration on FPGA with Heterogeneous ResourcesabstractBy quantizing weights with different precision for different parts of a network, mixed-precision quantization promises to reduce the hardware cost and improve the speed of deep neural network (DNN) accelerators that typically operate with a fixed quantization scheme. However, the additional control needed, and the decreased hardware efficiency arising from multi-precision operations have made mixed-precision quantization schemes challenging to deploy in practice. In this paper, a practical mixed-precision quantization framework called MSD that leverages the heterogeneous computing resources on FPGA to perform bit-serial and bit-parallel operations simultaneously is presented. MSD combines the use of a custom restricted signed digit (RSD) representation, which utilizes a limited number of effectual bits, and the conventional 2's complement representation to quantize DNN weights. Depending on the availability of fine-grained and coarse-grained resources, MSD encodes a subset of weights with RSD to allow highly efficient bit-serial multiply-accumulate implementation using LUT resources. Furthermore, the number of effectual bits used in RSD is optimized to match the bit-serial hardware latency to the bit-parallel operation on the coarse-grained resources to ensure the highest run-time utilization of all on-chip resources. Experiments show that MSD achieved a 1.36× speedup on the ResNet-18 model over the state-of-the-art, and a remarkable 4.91% higher accuracy on MobileNet-V2. Jiajun Wu 0006, Jiajun Zhou 0004, Yizhao Gao 0002, Yuhao Ding, Ngai Wong 0001, Hayden Kwok-Hay So |
FCCM | 6 |
| 2023 | Model-Platform Optimized Deep Neural Network Accelerator Generation through Mixed-Integer Geometric ProgrammingabstractAlthough there are distinct power-performance advantages in customizing an accelerator for a specific combination of FPGA platform and neural network model, developing such highly customized accelerators is a challenging task due to the massive design space spans from the range of network models to be accelerated, the target platform's compute capability, and its memory capacity and performance characteristics. To address this architectural customization problem, an automatic design space exploration (DSE) framework using a mixed-integer geometric programming (MIGP) approach is presented. Given the set of DNN models to be accelerated and a generic description of the target platform's compute and memory capabilities as input, the proposed framework automatically customizes an architectural template for the platform-model combination and produces the associated I/O schedule to maximize its end-to-end performance. By formulating DNN inference as a multi-level loop tiling problem, the proposed framework first customizes an accelerator template that consists of a parameterizable array architecture with SIMD execution cores and a customizable memory hierarchy using a MIGP to maximize the expected resource utilization. Subsequently, a second MIGP is used to schedule memory and compute operations as tiles to improve on-chip data reuse and memory bandwidth utilization. Experimental results from a wide range of neural network models and FPGA platform combinations show that the proposed scheme is able to produce accelerators with performance comparable to the state-of-the-art. The proposed DSE framework and the resulting hardware/software generator are available as an open-source package called AGNA with the hope that it may facilitate vendor-agnostic DNN accelerator development from the research community in the future. Yuhao Ding, Jiajun Wu 0006, Yizhao Gao 0002, Maolin Wang 0002, Hayden Kwok-Hay So |
FCCM | 5 |
| 2023 | RSQP: Problem-specific Architectural Customization for Accelerated Convex Quadratic OptimizationabstractConvex optimization is at the heart of many performance-critical applications across a wide range of domains. Although many high-performance hardware accelerators have been developed for specific optimization problems in the past, designing such accelerator is a challenging task and the resulting computing architecture is often so specific to the targeted application that they can hardly be reused even in a related application within the same domain. To accelerate general-purpose optimization solvers that must operate on diverse user input during run time, an ideal hardware solver should be able to adapt to the provided optimization problem dynamically while achieving high performance and power-efficiency. In this work, a hardware-accelerated general-purpose quadratic program solver, called RSQP, with reconfigurable functional units and data path that facilitate problem-specific customization is presented. RSQP uses a string-based encoding to describe the problem structure with fine granularity. Based on this encoding, functional units and datapath customized to the sparsity pattern of the problem are created by solving a dictionary-based lossless string compression problem and a mixed integer linear program respectively. RSQP has been integrated to accelerate the general-purpose quadratic programming solver OSQP and has been tested using an extensive benchmark with 120 optimization problems from 6 application domains. Through architectural customization, RSQP achieves up to 7× performance improvement over its baseline generic design. Furthermore, when compared with a CPU and a GPU-accelerated implementation, RSQP achieves up to 31.2× and 6.9× end-to-end speedup on these benchmark programs, respectively. Finally, the FPGA accelerator operates at up to 6.6× lower dynamic power consumption and up to 22.7× higher power efficiency over the GPU implementation, making it an attractive solution for power-conscious datacenter applications. Maolin Wang 0002, Ian McInerney, Bartolomeo Stellato, Stephen P. Boyd, Hayden Kwok-Hay So |
ISCA | 5 |
| 2023 | A Reconfigurable Architecture for Real-time Event-based Multi-Object TrackingabstractAlthough advances in event-based machine vision algorithms have demonstrated unparalleled capabilities in performing some of the most demanding tasks, their implementations under stringent real-time and power constraints in edge systems remain a major challenge. In this work, a reconfigurable hardware-software architecture called REMOT, which performs real-time event-based multi-object tracking on FPGAs, is presented. REMOT performs vision tasks by defining a set of actions over attention units (AUs). These actions allow AUs to track an object candidate autonomously by adjusting its region of attention and allow information gathered by each AU to be used for making algorithmic-level decisions. Taking advantage of this modular structure, algorithm-architecture codesign can be performed by implementing different parts of the algorithm in either hardware or software for different tradeoffs. Results show that REMOT can process 0.43–2.91 million events per second at 1.75–5.45 W. Compared with the software baseline, our implementation achieves up to 44 times higher throughput and 35.4 times higher power efficiency. Migrating the Merge operation to hardware further reduces the worst-case latency to be 95 times shorter than the software baseline. By varying the AU configuration and operation, a reduction of 0.59–0.77 mW per AU on the programmable logic has also been demonstrated. Yizhao Gao 0002, Song Wang 0023, Hayden Kwok-Hay So |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2022 | REMOT: A Hardware-Software Architecture for Attention-Guided Multi-Object Tracking with Dynamic Vision Sensors on FPGAsabstractIn contrast to conventional vision sensors that produce images of the entire field-of-view at a fixed frame rate, dynamic vision sensors (DVS) are neuromorphic devices that only produce sparse events in response to changes in light intensity local to each pixel, making them promising technologies for use in demanding edge scenarios where energy-efficient intelligent computations are needed. While several early research have demonstrated promising results in performing high-level machine vision tasks using vision events only, these algorithms are often too complex for real-time deployments in edge systems with limited processing and storage capabilities. In this work, a novel hardware-software architecture, called REMOT, is proposed to leverage the unique properties of DVS to perform real-time multi-object tracking (MOT) on FPGAs. REMOT incorporates a parallel set of reconfigurable hardware attention units (AUs) that work in tandem with a modular attention-guided software framework running in the attached processor. Each hardware AU autonomously adjusts its region of attention by processing each vision event as they are produced by the DVS. Using information aggregated by the AUs, high-level analyses are performed in software. To demonstrate the flexibility and modularity of REMOT, a family of MOT algorithms with different hardware-software configurations and tradeoffs have been implemented on 2 different edge reconfigurable systems. Experimental results show that REMOT is capable of processing 0.43-2.22 million events per second at 1.75-5.68 watts, making them suitable for real-time operations while maintaining good MOT accuracy in our target datasets. When compared with a software-only implementation using the same edge platforms, our HW-SW implementation results in up to 33.6 times higher event processing throughput and 25.9 times higher power efficiency. Yizhao Gao 0002, Song Wang 0023, Hayden Kwok-Hay So |
FPGA | 3 |
| 2022 | Low-Latency In Situ Image Analytics With FPGA-Based Quantized Convolutional Neural NetworkabstractReal-time in situ image analytics impose stringent latency requirements on intelligent neural network inference operations. While conventional software-based implementations on the graphic processing unit (GPU)-accelerated platforms are flexible and have achieved very high inference throughput, they are not suitable for latency-sensitive applications where real-time feedback is needed. Here, we demonstrate that high-performance reconfigurable computing platforms based on field-programmable gate array (FPGA) processing can successfully bridge the gap between low-level hardware processing and high-level intelligent image analytics algorithm deployment within a unified system. The proposed design performs inference operations on a stream of individual images as they are produced and has a deeply pipelined hardware design that allows all layers of a quantized convolutional neural network (QCNN) to compute concurrently with partial image inputs. Using the case of label-free classification of human peripheral blood mononuclear cell (PBMC) subtypes as a proof-of-concept illustration, our system achieves an ultralow classification latency of 34.2 [Formula: see text] with over 95% end-to-end accuracy by using a QCNN, while the cells are imaged at throughput exceeding 29 200 cells/s. Our QCNN design is modular and is readily adaptable to other QCNNs with different latency and resource requirements. Maolin Wang 0002, Kelvin C. M. Lee, Bob M. F. Chung, B. Sharat Chandra Varma 0001, Ho-Cheung Ng, Justin S. J. Wong, Ho Cheung Shum, Kevin K. Tsia, Hayden Kwok-Hay So |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2022 | NITI: Training Integer Neural Networks Using Integer-Only ArithmeticabstractLow bitwidth integer arithmetic has been widely adopted in hardware implementations of deep neural network inference applications. However, despite the promised energy-efficiency improvements demanding edge applications, the use of low bitwidth integer arithmetic for neural network training remains limited. Unlike inference, training demands high dynamic range and numerical accuracy for high quality results, making the use of low-bitwidth integer arithmetic particularly challenging. To address this challenge, we present a novel neural network training framework called NITI that exclusively utilizes low bitwidth integer arithmetic. NITI stores all parameters and accumulates intermediate values as 8-bit integers while using no more than 5 bits for gradients. To provide the necessary dynamic range during the training process, a per-layer block scaling exponentiation scheme is utilized. By deeply integrating with the rounding procedures and integer entropy loss calculation, the proposed scaling scheme incurs only minimal overhead in terms of storage and additional computation. Furthermore, a hardware-efficient pseudo-stochastic rounding scheme that eliminates the need for external random number generation is proposed to facilitate conversion from wider intermediate arithmetic results to lower precision for storage. Since NITI operates only with standard 8-bit integer arithmetic and storage, it is possible to accelerate it using existing low bitwidth operators originally developed for inference in commodity accelerators. To demonstrate this, an open-source software implementation of end-to-end training, using native 8-bit integer operations in modern GPUs is presented. In addition, experiments have been conducted on an FPGA-based training accelerator to evaluate the hardware advantage of NITI. When compared with an equivalent training setup implemented with floating point storage and arithmetic, NITI has no accuracy degradation on the MNIST and CIFAR10 datasets. On ImageNet, NITI achieves similar accuracy as state-of-the-art integer training frameworks without relying on full-precision floating-point first and last layers. Maolin Wang 0002, Seyedramin Rasoulinezhad, Philip H. W. Leong, Hayden Kwok-Hay So |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | HAO: Hardware-aware Neural Architecture Optimization for Efficient InferenceabstractAutomatic algorithm-hardware co-design for DNN has shown great success in improving the performance of DNNs on FPGAs. However, this process remains challenging due to the intractable search space of neural network architectures and hardware accelerator implementation. Differing from existing hardware-aware neural architecture search (NAS) algorithms that rely solely on the expensive learning-based approaches, our work incorporates integer programming into the search algorithm to prune the design space. Given a set of hardware resource constraints, our integer programming formulation directly outputs the optimal accelerator configuration for mapping a DNN subgraph that minimizes latency. We use an accuracy predictor for different DNN subgraphs with different quantization schemes and generate accuracy-latency pareto frontiers. With low computational cost, our algorithm can generate quantized networks that achieve state-of-the-art accuracy and hardware performance on Xilinx Zynq (ZU3EG) FPGA for image classification on ImageNet dataset. The solution searched by our algorithm achieves 72.5% top-1 accuracy on ImageNet at framerate 50, which is 60% faster than MnasNet [37] and 135% faster than FBNet [43] with comparable accuracy. Zhen Dong 0003, Yizhao Gao 0002, Qijing Huang 0001, John Wawrzynek, Hayden Kwok-Hay So, Kurt Keutzer |
FCCM | 5 |
| 2021 | Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization FrameworkabstractDeep Neural Networks (DNNs) have achieved extraordinary performance in various application domains. To support diverse DNN models, efficient implementations of DNN inference on edge-computing platforms, e.g., ASICs, FPGAs, and embedded systems, are extensively investigated. Due to the huge model size and computation amount, model compression is a critical step to deploy DNN models on edge devices. This paper focuses on weight quantization, a hardware-friendly model compression approach that is complementary to weight pruning.Unlike existing methods that use the same quantization scheme for all weights, we propose the first solution that applies different quantization schemes for different rows of the weight matrix. It is motivated by (1) the distribution of the weights in the different rows are not the same; and (2) the potential of achieving better utilization of heterogeneous FPGA hardware resources. To achieve that, we first propose a hardware-friendly quantization scheme named sum-of-power-of-2 (SP2) suitable for Gaussian-like weight distribution, in which the multiplication arithmetic can be replaced with logic shifter and adder, thereby enabling highly efficient implementations with the FPGA LUT resources. In contrast, the existing fixed-point quantization is suitable for Uniform-like weight distribution and can be implemented efficiently by DSP. Then to fully explore the resources, we propose an FPGA-centric mixed scheme quantization (MSQ) with an ensemble of the proposed SP2 and the fixed-point schemes. Combining the two schemes can maintain, or even increase accuracy due to better matching with weight distributions.For the FPGA implementations, we develop a parameterized architecture with heterogeneous Generalized Matrix Multiplication (GEMM) cores-one using LUTs for computations with SP2 quantized weights and the other utilizing DSPs for fixed-point quantized weights. Given the partition ratio among the two schemes based on resource characterization, MSQ quantization training algorithm derives an optimally quantized model for the FPGA implementation. We evaluate our FPGA-centric quantization framework across multiple application domains. With optimal SP2/fixed-point ratios on two FPGA devices, i.e., Zynq XC7Z020 and XC7Z045, we achieve performance improvement of 2.1 × -4.1 × compared to solely exploiting DSPs for all multiplication operations. In addition, the CNN implementations with the proposed MSQ scheme can achieve higher accuracy and comparable hardware utilization efficiency compared to the state-of-the-art designs. Sung-En Chang, Yanyu Li, Mengshu Sun, Runbin Shi, Hayden Kwok-Hay So, Xuehai Qian, Yanzhi Wang 0001, Xue Lin 0001 |
HPCA | 5 |
| 2021 | High-Dimensional Dense Residual Convolutional Neural Network for Light Field ReconstructionabstractWe consider the problem of high-dimensional light field reconstruction and develop a learning-based framework for spatial and angular super-resolution. Many current approaches either require disparity clues or restore the spatial and angular details separately. Such methods have difficulties with non-Lambertian surfaces or occlusions. In contrast, we formulate light field super-resolution (LFSR) as tensor restoration and develop a learning framework based on a two-stage restoration with 4-dimensional (4D) convolution. This allows our model to learn the features capturing the geometry information encoded in multiple adjacent views. Such geometric features vary near the occlusion regions and indicate the foreground object border. To train a feasible network, we propose a novel normalization operation based on a group of views in the feature maps, design a stage-wise loss function, and develop the multi-range training strategy to further improve the performance. Evaluations are conducted on a number of light field datasets including real-world scenes, synthetic data, and microscope light fields. The proposed method achieves superior performance and less execution time comparing with other state-of-the-art schemes. Nan Meng, Hayden Kwok-Hay So, Xing Sun 0001, Edmund Y. Lam |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | FTDL: A Tailored FPGA-Overlay for Deep Learning with High ScalabilityabstractFast inference is of paramount value to a wide range of deep learning applications. This work presents FTDL, a highly-scalable FPGA overlay framework for deep learning applications, to address the architecture and hardware mismatch faced by traditional efforts. The FTDL overlay is specifically optimized for the tiled structure of FPGAs, thereby achieving post-place-and-route operating frequencies exceeding 88 % of the theoretical maximum across different devices and design scales. A flexible compilation framework efficiently schedules matrix multiply and convolution operations of large neural network inference on the overlay and achieved over 80 % hardware efficiency on average. Taking advantage of both high operating frequency and hardware efficiency, FTDL achieves 402.6 and 151.2 FPS with GoogLeNet and ResNet50 on ImageNet, respectively, while operating at a power efficiency of 27.6 GOPS/W, making it up to 7.7× higher performance and 1.9× more power-efficient than the state-of-the-art. Runbin Shi, Yuhao Ding, Xuechao Wei, He Li 0008, Hang Liu 0001, Hayden Kwok-Hay So, Caiwen Ding |
DAC | 6 |
| 2020 | FTDL: An FPGA-tailored Architecture for Deep Learning SystemsabstractHardware acceleration of deep learning (DL) systems has been increasingly studied to achieve desirable performance and energy efficiency. The FPGA strikes a balance between high energy efficiency and fast development cycle and therefore is widely used as a DNN accelerator. However, there exists an architecture-layout mismatch in the current designs, which introduces scalability and flexibility issues, leading to irregular routing and resource imbalance problems. To address these limitations, in this work, we propose FTDL, an FPGA-tailored architecture with a parameterized and hierarchical hardware that is adaptive to different FPGA devices. FTDL has the following novelties: (i) At the architecture level, FTDL consists of Tiled Processing Elements (TPE) and super blocks, to achieve a near-to-theoretical digital signal processing (DSP) operating-frequency of 650 MHz. More importantly, FTDL is configurable and delivers good scalability, i.e., the timing is stabilized even when the design is scaled-up to 100% resource utilization for different deep learning systems. (ii) In workload compilation, FTDL provides a compiler that manages to map the DL workloads to the architecture level in an optimal manner. Experimental results show that for most benchmark layers in MLPerf, FTDL achieves an over 80% hardware efficiency. Runbin Shi, Yuhao Ding, Xuechao Wei, Hang Liu 0001, Hayden Kwok-Hay So, Caiwen Ding |
FPGA | 5 |
| 2020 | Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked Layers
Zhe Xu 0008, Runbin Shi, Ray C. C. Cheung, Hayden Kwok-Hay So |
ICLR | 5 |
| 2020 | Exploiting Elasticity in Tensor Ranks for Compressing Neural NetworksabstractElasticities in depth, width, kernel size and resolution have been explored in compressing deep neural networks (DNNs). Recognizing that the kernels in a convolutional neural network (CNN) are 4-way tensors, we further exploit a new elasticity dimension along the input-output channels. Specifically, a novel nuclear-norm rank minimization factorization (NRMF) approach is proposed to dynamically and globally search for the reduced tensor ranks during training. Correlation between tensor ranks across multiple layers is revealed, and a graceful tradeoff between model size and accuracy is obtained. Experiments then show the superiority of NRMF over the previous non-elastic variational Bayesian matrix factorization (VBMF) scheme. Jie Ran, Hayden Kwok-Hay So, Graziano Chesi, Ngai Wong 0001 |
ICPR | 3 |
| 2020 | CSB-RNN: a faster-than-realtime RNN acceleration framework with compressed structured blocksabstractRecurrent neural networks (RNNs) have been widely adopted in temporal sequence analysis, where realtime performance is often in demand. However, RNNs suffer from heavy computational workload as the model often comes with large weight matrices. Pruning (a model compression method) schemes have been proposed for RNNs to eliminate the redundant (close-to-zero) weight values. On one hand, the non-structured pruning methods achieve a high pruning rate but introducing computation irregularity (random sparsity), which is unfriendly to parallel hardware. On the other hand, hardware-oriented structured pruning suffers from low pruning rate due to restricted constraints on allowable pruning structure. Runbin Shi, Peiyan Dong, Tong Geng, Yuhao Ding, Hayden Kwok-Hay So, Martin C. Herbordt, Ang Li 0006, Yanzhi Wang 0001 |
ICS | 6 |
| 2020 | Vision Guided Crop Detection in Field Robots using FPGA-Based Reconfigurable ComputersabstractA case study in applying modern FPGAs as a platform to accelerate intelligent vision-guided crop detection in agricultural field robots is presented. A state-of-the-art YOLOv3 object detection neural network was adapted to detect broccoli and cauliflower in image dataset obtained from autonomous agricultural robots. A baseline floating point implementation achieved 96% mAP, and an efficient, quantized implementation suitable for FPGA implementation 92% mAP. The proposed FPGA solution has 136.86 ms inference latency while consuming 12.43W in a low latency configuration, and 28.48 frames per second while consuming 17.78W in a high throughput one. Compared to an embedded GPU implementation of the same task, the FPGA solution was 4.12 times more power-efficient and offers 6.85 times higher throughput, translating to faster and longer operation of a battery-powered field robot. Cyrus Wing-Hei Chan, Philip H. W. Leong, Hayden Kwok-Hay So |
ISCAS | 3 |
| 2020 | PARC: ultrafast and accurate clustering of phenotypic data of millions of single cellsabstractMOTIVATION: New single-cell technologies continue to fuel the explosive growth in the scale of heterogeneous single-cell data. However, existing computational methods are inadequately scalable to large datasets and therefore cannot uncover the complex cellular heterogeneity. RESULTS: We introduce a highly scalable graph-based clustering algorithm PARC-Phenotyping by Accelerated Refined Community-partitioning-for large-scale, high-dimensional single-cell data (>1 million cells). Using large single-cell flow and mass cytometry, RNA-seq and imaging-based biophysical data, we demonstrate that PARC consistently outperforms state-of-the-art clustering algorithms without subsampling of cells, including Phenograph, FlowSOM and Flock, in terms of both speed and ability to robustly detect rare cell populations. For example, PARC can cluster a single-cell dataset of 1.1 million cells within 13 min, compared with >2 h for the next fastest graph-clustering algorithm. Our work presents a scalable algorithm to cope with increasingly large-scale single-cell analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/ShobiStassen/PARC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shobana V. Stassen, Dickson M. D. Siu, Kelvin C. M. Lee, Joshua W. K. Ho, Hayden Kwok-Hay So, Kevin K. Tsia |
Bioinform. | 5 |
| 2019 | E-LSTM: Efficient Inference of Sparse LSTM on Embedded Heterogeneous SystemabstractVarious models with Long Short-Term Memory (LSTM) network have demonstrated prior art performances in sequential information processing. Previous LSTM-specific architectures set large on-chip memory for weight storage to alleviate the memory-bound issue and facilitate the LSTM inference in cloud computing. In this paper, E-LSTM is proposed for embedded scenarios with the consideration of the chip-area and limited data-access bandwidth. The heterogeneous hardware in E-LSTM tightly couples an LSTM co-processor with an embedded RISC-V CPU. The eSELL format is developed to represent the sparse weight matrix. With the proposed cell fusion optimization based on the inherent sparsity in computation, E-LSTM achieves up to 2.2× speedup of processing throughput. Runbin Shi, Hayden Kwok-Hay So, Shuo Wang 0009, Yun Liang 0001 |
DAC | 3 |
| 2019 | Design of quadruple precision multiplier architectures with SIMD single and double precision support
Manish Kumar Jaiswal, Hayden Kwok-Hay So |
Integr. | 2 |
| 2019 | Fringe Pattern Improvement and Super-Resolution Using Deep Learning in Digital HolographyabstractDigital holographic imaging is a powerful technique that can provide wavefront information of a three-dimensional object for biological and industrial applications. However, due to the constraint and cost of imaging sensors, the acquired digital hologram is limited in terms of pixel count, thus affecting the resolution in holographic reconstruction. To overcome this constraint, in this paper we propose a deep learning-based method to super-resolve holograms and to improve the quality of low-resolution holograms by training a convolutional neural network with large-scale data for resolution enhancement. Moreover, this algorithm can be broadly adapted to enhance the space-bandwidth product of a holographic imaging system without the need of any advanced hardware. We experimentally validate its capability using a lens-free off-axis holographic system, and compare the performance of various loss functions and interpolation methods in training such a network. Zhenbo Ren, Hayden Kwok-Hay So, Edmund Y. Lam |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | Large-Scale Multi-Class Image-Based Cell Classification With Deep LearningabstractRecent advances in ultra-high-throughput microscopy have enabled a new generation of cell classification methodologies using image-based cell phenotypes alone. In contrast to current single-cell analysis techniques that rely solely on slow and costly genetic/epigenetic analysis, these image-based analyses allow morphological profiling and screening of thousands or even millions of single cells at a fraction of the cost, and have been proven to demonstrate the statistical significance required for understanding the role of cell heterogeneity in diverse biological applications, ranging from cancer screening to drug candidate identification/validation processes. This paper examines the efficacies and opportunities presented by machine learning algorithms in processing large scale datasets with millions of label-free cell images. An automatic single-cell classification framework using convolutional neural network (CNN) has been developed. A comparative analysis of its efficiency in classifying large datasets against conventional k-nearest neighbors (kNN) and support vector machine (SVM) based methods are also presented. Experiments have shown that our proposed framework can efficiently identify multiple types cells with over 99% accuracy based on the phenotypic label-free bright-field images; and CNN-based models perform well and relatively stable against data volume compared with kNN and SVM. Nan Meng, Edmund Y. Lam, Kevin K. Tsia, Hayden Kwok-Hay So |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | GraVF-M: Graph Processing System Generation for Multi-FPGA PlatformsabstractDue to the irregular nature of connections in most graph datasets, partitioning graph analysis algorithms across multiple computational nodes that do not share a common memory inevitably leads to large amounts of interconnect traffic. Previous research has shown that FPGAs can outcompete software-based graph processing in shared memory contexts, but it remains an open question if this advantage can be maintained in distributed systems. In this work, we present GraVF-M, a framework designed to ease the implementation of FPGA-based graph processing accelerators for multi-FPGA platforms with distributed memory. Based on a lightweight description of the algorithm kernel, the framework automatically generates optimized RTL code for the whole multi-FPGA design. We exploit an aspect of the programming model to present a familiar message-passing paradigm to the user, while under the hood implementing a more efficient architecture that can reduce the necessary inter-FPGA network traffic by a factor equal to the average degree of the input graph. A performance model based on a theoretical analysis of the factors influencing performance serves to evaluate the efficiency of our implementation. With a throughput of up to 5.8GTEPS (billions of traversed edges per second) on a 4-FPGA system, the designs generated by GraVF-M compare favorably to state-of-the-art frameworks from the literature and reach 94% of the projected performance limit of the system. Nina Engelhardt, Hayden Kwok-Hay So |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2018 | Universal number posit arithmetic generator on FPGAabstractPosit number system format includes a run-time varying exponent component, defined by a combination of regime-bit (with run-time varying length) and exponent-bit (with size of up to ES bits, the exponent size). This also leads to a run-time variation in its mantissa field size and position. This run-time variation in posit format poses a hardware design challenge. Being a recent development, posit lacks for its adequate hardware arithmetic architectures. Thus, this paper is aimed towards the posit arithmetic algorithmic development and their generic hardware generator. It is focused on basic posit arithmetic (floating-point to posit conversion, posit to floating point conversion, addition/subtraction and multiplication). These are also demonstrated on a FPGA platform. Target is to develop an open-source solution for generating basic posit arithmetic architectures with parameterized choices. This contribution would enable further exploration and evaluation of posit system. Manish Kumar Jaiswal, Hayden Kwok-Hay So |
DATE | 2 |
| 2018 | Performance-Driven System Generation for Distributed Vertex-Centric Graph Processing on Multi-FPGA SystemsabstractIn this paper, we present a multi-FPGA graph processing framework and an accompanying performance model. Our framework emphasizes programmability, requiring minimal user input beyond providing the application kernel and the dataset. The framework predicts the performance of the system based on the algorithm characteristics and problem size and automatically selects the optimal FPGA configuration. We implement our system on an experimental 4-FPGA platform and compare the results to the predicted performance. Nina Engelhardt, C.-H. Dominic Hung, Hayden Kwok-Hay So |
FPL | 3 |
| 2018 | Architecture Generator for Type-3 Unum Posit Adder/SubtractorabstractThis paper is aimed towards the hardware architecture aspect of a recently proposed posit number system under type-3 unum (universal number system). Here, an algorithmic flow for the posit addition/subtraction arithmetic is developed and its hardware architecture is designed. Compare to floating point, posit provides better dynamic range and accuracy over same word size, along with more accurate and exact arithmetic support. Posit format includes a run-time varying exponent component, provided by a combination of regime-bits (of run-time varying length) and exponent-bits (of size up to ES bits). Thus, the mantissa precision also varies at run-time. This provides a combination of dynamic range and precision under a given word size (N). This possible variation in format along dynamic range and precision may attract various applications with different(accuracy and dynamic range) requirement. However, this run-time variation in posit format also poses a hardware design challenge. So, this paper is aimed towards the construction of an open-source parameterized Verilog HDL (Hardware Description Language) generator for posit adder/subtractor arithmetic, with parameterized N and ES. Manish Kumar Jaiswal, Hayden Kwok-Hay So |
ISCAS | 2 |
| 2017 | A Parameterizable Activation Function Generator for FPGA-Based Neural Network ApplicationsabstractNeural network applications on FPGAs at times require arithmetic operators that are either not available in the manufacturer's core library, or are complex operators made up of several elementary functions, requiring more resources than if they were built as single operators. In this work, we built an open-source, parameterized floating-point core generator named NnCore, for operators used as activation functions, and their derivatives. We propose a binary search algorithm to search for minimax-polynomial segments, with adjusting steps for ensuring monotonicity between different segments. Sam M. H. Ho, C.-H. Dominic Hung, Ho-Cheung Ng, Maolin Wang 0002, Hayden Kwok-Hay So |
FCCM | 5 |
| 2017 | OLAF'17: Third International Workshop on Overlay Architectures for FPGAs
Hayden Kwok-Hay So, John Wawrzynek |
FPGA | 1 |
| 2017 | NnCore: A parameterized non-linear function generator for machine learning applications in FPGAsabstractEfficient implementation of machine learning applications on FPGAs often requires non-linear numerical functions with a non-standard numerical precision that is not readily available from vendor provided standard libraries. While application-specific designs of such functions can result in superior numerical accuracy and area efficiency when compared to ad-hoc composition using vendor-provided primitives, the effort devoted to this challenging task can hardly be portable to other similar applications. In this work, we present an open source generator, NnCore, for floating-point non-linear operator cores built using fixed-point piecewise polynomial segments. The proposed framework takes advantage of properties such as oddness/evenness and intercept-at-origin, often found in the numerical functions commonly used in machine learning applications, and applied an improved segmentation algorithm that specifically handles “outlier” segments, to reduce the required memory size for storing polynomial coefficients. Experimental results show that, at single-precision setting, NnCore generated cores use up to 65% fewer BRAMs, 63% fewer shift-registers, and runs at up to 2.2 χ the clock speed, compare with cores generated from a previous generic function generator. At half-precision, cores can run around 1.2 χ higher clock speed while requiring higher resource usage, or use a comparable number of resource but run at 12% to 45% lower clock speed. The use of HLS C++ as output format allows core integration into modern high-level workflow such as Xilinx SDAccel. Sam M. H. Ho, Hayden Kwok-Hay So |
FPT | 2 |
| 2017 | Ultra-low latency continuous block-parallel stream windowing using FPGA on-chip memoryabstractIn this paper, we propose and demonstrate a real-time ultra-fast multi-data stream processing methodology on FPGA called “SWIM” (Stream Windowing on Interleaved Memory). The method exploits the flexible on-chip block memory fabric on existing FPGA architectures to achieve ultra-low-latency and fully pipelined continuous data flow while maintaining linear spatial locality of data for efficient data addressing and processing. The SWIM method is directly applicable to many practical applications such as real-time stencil computing, streaming image data processing, as well as closed loop-control systems that require ultra-low latency interleaved access and processing of high-speed sensor data. We demonstrate two practical cases on actual FPGA for generic 3-by-3 2-D convolution filter and image super-resolution method using pixel interleaving. Both memory usage and latency scales linearly with window height, or width of the 2-D input data set. The generic implementation of SWIM on FPGA showed impressive worst-case operation frequency of 410 MHz and uses 9.0χ and 5.6χ less Register and LUT resources respectively compared with a high-level synthesis solution. Justin S. J. Wong, Runbin Shi, Maolin Wang 0002, Hayden Kwok-Hay So |
FPT | 4 |
| 2017 | Computationally Efficient Hyperspectral Data Learning Based on the Doubly Stochastic Dirichlet ProcessabstractThe Dirichlet process (DP) prior is effective in modeling HSIs (HSI) and identifying land-cover classes. However, modeling a continuously varying intensity of these land covers elegantly and consistently is still a challenge. We propose a doubly stochastic DP (DSDP) as an efficient model of the global topic measurement space, which imposes a weaker assumption compared with the discrete Markov assumption, resulting in a lower computational cost than other DP-prior-based models. We also present a mixture model of DSDP, which is termed the marked sigmoidal Gaussian process (SGP) DSDP mixture model. It can be thinned from a DP mixture without massive auxiliary covariates, and the marked function prior makes the number of land-cover classes consistent, whereas the SGP function prior models the HSI land-cover variation globally. The consistency of the number of land covers is maintained for various HSIs with large-scale geographical areas. Experiments show that the model is robust and consistent on HSI identification with weak or even no supervision. Xing Sun 0001, Nelson H. C. Yung, Edmund Y. Lam, Hayden Kwok-Hay So |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2017 | The First 25 Years of the FPL Conference: Significant PapersabstractA summary of contributions made by significant papers from the first 25 years of the Field-Programmable Logic and Applications conference (FPL) is presented. The 27 papers chosen represent those which have most strongly influenced theory and practice in the field. Philip H. W. Leong, Hideharu Amano, Jason Helge Anderson, Koen Bertels, João M. P. Cardoso, Oliver Diessel, Guy Gogniat, Mike Hutton, Wayne Luk, Patrick Lysaght, Marco Platzner, Viktor Prasanna 0001, Tero Rissa, Cristina Silvano, Hayden Kwok-Hay So, Yu Wang 0002 |
ACM Trans. Reconfigurable Technol. Syst. | 16 |
| 2016 | Architecture for quadruple precision floating point division with multi-precision supportabstractThis paper proposes a FPGA based hardware architecture for quadruple precision (QP) division arithmetic which can also process a single, a double and a double-extended precision (SP, DP, DPE) computations. The mantissa division employs a series expansion methodology of division, integrated with a wide integer multiplier further optimized for FPGA implementations facilitating the built-in DSP blocks efficiently. The proposed division architecture is demonstrated using a Xilinx FPGA based implementation has shown a significant area saving and much improvement in latency with improved speed. Manish Kumar Jaiswal, Hayden Kwok-Hay So |
ASAP | 2 |
| 2016 | Vertex-Centric Graph Processing on FPGAabstractPast research and implementation efforts have shown that FPGAs are efficient at processing many graph algorithms. However, they are notoriously hard to program, leading to impractically long development times even for simple applications. We propose a vertex-centric framework for graph processing on FPGAs, providing a base execution model and distributed architecture so that developers need only write very small application kernels. Nina Engelhardt, Hayden Kwok-Hay So |
FCCM | 2 |
| 2016 | OLAF'16: Second International Workshop on Overlay Architectures for FPGAsabstractThe Second International Workshop on Overlay Architec- tures for FPGAs is held in Monterey, California, USA, on February 21, 2016 and co-located with FPGA 2016: The 24th ACM/SIGDA International Symposium on Field Pro- grammable Gate Arrays. The main objective of the work- shop is to address how overlay architectures can help address the challenges and opportunities provided by FPGA-based reconfigurable computing. The workshop provides a venue for researchers to present and discuss the latest develop- ments in FPGA overlay architecture and related areas. We have assembled a program of six refereed papers and a panel discussion with prominent experts in the field. Hayden Kwok-Hay So, John Wawrzynek |
FPGA | 1 |
| 2016 | GraVF: A vertex-centric distributed graph processing framework on FPGAsabstractFPGAs are promising platforms to efficiently execute distributed graph algorithms. Unfortunately, they are notoriously hard to program, especially when the problem size and system complexity increases. In this paper, we propose GraVF, a high-level design framework for distributed graph processing on FPGAs. It leverages the vertex-centric paradigm, which is naturally distributed and requires the user to define only very small kernels and their associated message semantics for the target application. The user design may subsequently be elaborated and compiled to the target system automatically by the framework. To demonstrate the flexibility and capabilities of the proposed framework, 4 graph algorithms with distinct requirements have been implemented, namely breadth-first search, PageRank, single source shortest path, and connected component. Results show that the proposed framework is capable of producing FPGA designs with performance comparable to similar custom designs while requiring only minimal input from the user. Nina Engelhardt, Hayden Kwok-Hay So |
FPL | 2 |
| 2016 | Real-time object detection and classification for high-speed asymmetric-detection time-stretch optical microscopy on FPGAabstractA real-time object detection and classification system using FPGA developed for high-speed asymmetric time-stretched optical microscopy (ATOM) framework is presented. Due to the massive amount of data generated by optical frontend, storing the raw data for offline post-processing is slow and impractical for the targeted single cell analysis applications. The proposed FPGA solution eliminates the need to transfer and persist the entire raw data by processing low-level signals and forming high-level images in real-time. Objects of interest are detected and segmented from the image stream and a classifier subsequently performs high-level analysis on the segmented images. When compared with existing software-based post-processing workflow, this FPGA-based approach will improve both the number of objects captured per experiment and the overall end-to-end object classification performance. The system also allows co-optimization between optical system, low-level signal processing and image analytic in a unified environment that enables new scientific discoveries previously unachievable. Maolin Wang 0002, Ho-Cheung Ng, Bob M. F. Chung, B. Sharat Chandra Varma 0001, Manish Kumar Jaiswal, Kevin K. Tsia, Ho Cheung Shum, Hayden Kwok-Hay So |
FPT | 8 |
| 2016 | Sparse Hierarchical Nonparametric Bayesian learning for light field representation and denoisingabstractIn this paper, we present a sparse hierarchical non-parametric Bayesian (SHNB) model, which is used to represent the data captured by the light field cameras. Specifically, a light field can be represented as a set of sub-aperture views. In order to capture the visual variations of these viewpoints, we propose the so-called “depth flow” features. Then based on the depth flow features, we model these views statistically with a sparse representation in a fully unsupervised manner. While local dictionaries are learned based on each sub-aperture view, all the views with different perspectives share one global dictionary. To show the effectiveness of the proposed model, we apply our model to denoise the light field data. In the experiments, we demonstrate that our method outperforms several state-of-the-art light field denoising approaches. Xing Sun 0001, Nan Meng, Edmund Y. Lam, Hayden Kwok-Hay So |
IJCNN | 5 |
| 2016 | Data-driven light field depth estimation using deep Convolutional Neural NetworksabstractThis paper presents a data-driven approach to estimate the object depths from light field data using Convolutional Neural Networks (CNN). By exploring the relationship between the epipolar-plane images (EPI) and the corresponding depth map, we propose an enhanced EPI feature that encodes the depth information of each physical point in the light field and obtains the disparity map of the whole scene in a supervised manner. This work covers two major contributions, namely the extraction of the enhanced EPI features and the light field depth estimation with CNN. The proposed features augment the depth information of the corresponding points in the light field, and then our CNN architecture differentiates them into different depth layers. Forward propagation step of the CNN model allows rapid recognition of the disparity map of the test light field data. In the experiments, we apply our method on the HCI (Heidelberg Col-laboratory for Image Processing) benchmark dataset and demonstrate that it is significantly faster than the state-of-the-art light field depth estimation approaches while achieving satisfactory performance. Xing Sun 0001, Nan Meng, Edmund Y. Lam, Hayden Kwok-Hay So |
IJCNN | 5 |
| 2015 | Automatic Soft CGRA Overlay Customization for High-Productivity Nested Loop Acceleration on FPGAsabstractCompiling high level compute intensive kernels to FPGAs via an abstract overlay architecture has been demonstrated to be an effective way to improve designers' productivity. However, achieving the desired performance and overhead constraints requires exploration in a complex design space involving multiple architectural parameters and counteracts the benefit of utilizing an overlay as a productivity enhancer. In this work, a soft CGRA (SCGRA) which provides unique opportunity to improve the power-performance of the resulting accelerators is used an FPGA overlay. With the observation that the loop unrolling factor and SCGRA size typically have monotonic impact on the loop compute time and the loop performance benefit degrades with the increase of the two design parameters, we took a marginal performance revenue metric to prune the design space to a small feasible design space (FDS) and then performed an intensive customization on the FDS by using analytical models of various design metrics such as power and overhead. Cheng Liu 0008, Hayden Kwok-Hay So |
FCCM | 2 |
| 2015 | Significant papers from the first 25 years of the FPL conferenceabstractThe list of significant papers from the first 25 years of the Field-Programmable Logic and Applications conference (FPL) is presented in this paper. These 27 papers represent those which have most strongly influenced theory and practice in the field. Philip H. W. Leong, Hideharu Amano, Jason Helge Anderson, Koen Bertels, João M. P. Cardoso, Oliver Diessel, Guy Gogniat, Mike Hutton, Wayne Luk, Patrick Lysaght, Marco Platzner, Viktor Prasanna 0001, Tero Rissa, Cristina Silvano, Hayden Kwok-Hay So, Yu Wang 0002 |
FPL | 16 |
| 2015 | QuickDough: A rapid FPGA loop accelerator design framework using soft CGRA overlayabstractThe use of FPGAs as compute accelerators has been demonstrated by numerous researchers as an effective solution to meet the performance requirement across many application domains. However, the design productivity of developing FPGA accelerators remains much lower compared to the use of a typical software development flow. Although the use of the high-level design tools may partly alleviate this shortcoming, the lengthy low-level FPGA implementation process including synthesis, placing and routing still dramatically limits the number of compile-debug-edit cycles per day and hinders the widespread adoption of FPGAs. To address this design productivity problem, we have developed a rapid FPGA loop accelerator generation framework called QuickDough. By utilizing a soft coarse-grained reconfigurable array (SCGRA) overlay built on top of off-the-shelf FPGAs, it compiles a high-level loop to the overlay through a rapid operation scheduling first and then generates the FPGA accelerator bitstream through a rapid integration of the scheduling result and a pre-built overlay bitstream. According to the experiments, QuickDough is able to produce accelerators in the order of seconds while achieving up to 9X performance speedup over the execution of the same software running on a hard ARM processor. Cheng Liu 0008, Ho-Cheung Ng, Hayden Kwok-Hay So |
FPT | 3 |
| 2015 | Accelerated cell imaging and classification on FPGAs for quantitative-phase asymmetric-detection time-stretch optical microscopyabstractWith the fundamental trade-off between speed and sensitivity, existing quantitative phase imaging (QPI) systems for diagnostics and cell classification are often limited to batch processing only small amount of offline data. While quantitative asymmetric-detection time-stretch optical microscopy (Q-ATOM) offers a unique optical platform for ultrafast and high-sensitivity quantitative phase cellular imaging, performing the computationally demanding backend QPI phase retrieval and image classification in real-time remains a major technical challenge. In this paper, we propose an optimized architecture for QPI on FPGA and compare its performance against CPU and GPU implementations in terms of speed and power efficiency. Results show that our implementation on single FPGA card demonstrates a speedup of 9.4 times over an optimized C implementation running on a 6-core CPU, and 3.47 times over the GPU implementation. It is also 24.19 and 4.88 times more power-efficient than the CPU and GPU implementation respectively. Throughput increase linearly when four FPGA cards are used to further improve the performance. We also demonstrate an increased classification accuracy when phase images instead of single-angle ATOM images are used. Overall, one FPGA card is able to process and categorize 2497 cellular images per second, making it suitable for real-time single-cell analysis applications. Junyi Xie, Xinyu Niu, Andy K. S. Lau, Kevin K. Tsia, Hayden Kwok-Hay So |
FPT | 5 |
| 2015 | Dual-mode double precision / two-parallel single precision floating point multiplier architectureabstractFloating point multiplication is an integral part of any contemporary computing system. This paper presents a configurable dual-mode double precision floating point multiplier architecture, which can also process two-parallel single precision multiplication. This unified, double precision dual (two-parallel) single precision, architecture is named as DPdSP multiplier. The proposed architecture is based on the standard state-of-the-art flow of floating point multiplication, which can process normal and sub-normal operands along with exceptional case handling. The proposed architecture is aimed for a ASIC (UMC 90nm) implementation. The key single-mode design units in the computational flow (like mantissa multiplier, dynamic right/left shifters, leading one detector, etc) are re-designed for configurable dual-mode operation to enable efficient resource sharing. The proposed architecture is compared with the best available literature in terms of area, period and area × period / throughput complexity metric. The proposed dual mode architecture shows a significant improvement in design metrics and also provides more computation support. Manish Kumar Jaiswal, Hayden Kwok-Hay So |
VLSI-SoC | 2 |
| 2014 | Map-reduce processing of k-means algorithm with FPGA-accelerated computer clusterabstractThe design and implementation of the k-means clustering algorithm on an FPGA-accelerated computer cluster is presented. The implementation followed the Map-Reduce programming model, with both the map and reduce functions executing autonomously to the CPU on multiple FPGAs. A hardware/software framework was developed to manage gateware execution on multiple FPGAs across the cluster. Using this k-means implementation as an example, system-level tradeoff study between computation and I/O performance in the target multi-FPGA execution environment was performed. When compared to a similar software implementation executing over the Hadoop MapReduce framework, 15.5× to 20.6× performance improvement has been achieved across a range of input data sets. Yuk-Ming Choi, Hayden Kwok-Hay So |
ASAP | 2 |
| 2014 | Scheduling Mixed-Architecture Processes in Tightly Coupled FPGA-CPU Reconfigurable ComputersabstractThe design and implementation of a multitasking run-time system on a tightly coupled FPGA-CPU platform is presented. Using a mix of CPU and FPGA programmable logic for computing, user applications are executed as mixed-architecture processes from the perspective of the OS. Context switching mechanisms with hybrid scheduling containing both blocking and preemption support were implemented to support concurrent execution of multiple mixed-architecture processes, and evaluated under a synthetic workload. Brandon Kyle Hamilton, Michael R. Inggs, Hayden Kwok-Hay So |
FCCM | 3 |
| 2014 | Mixed-architecture process scheduling on tightly coupled reconfigurable computersabstractThe design and implementation of a multitasking runtime system for mixed-architecture applications on a tightly coupled FPGA-CPU platform is presented. The runtime environment and the user applications assume an underlying machine that encompasses multiple computing architectures within a unified machine model. Using this model, a unified process scheduling mechanism was developed that enables concurrent execution of multiple mixed-architecture processes. Scheduling and allocation strategies, including blocking and preemption, were implemented and evaluated with respect to performance and fairness on a Xilinx Zynq platform using a mix of synthetic workloads. Brandon Kyle Hamilton, Michael R. Inggs, Hayden Kwok-Hay So |
FPL | 3 |
| 2013 | A Soft Coarse-Grained Reconfigurable Array Based High-level Synthesis Methodology: Promoting Design Productivity and Exploring Extreme FPGA FrequencyabstractCompared to the use of a typical software development flow, the productivity of developing FPGA-based compute applications remains much lower. Although the use of high-level synthesis (HLS) tools may partly alleviate this shortcoming, the lengthy low-level FPGA implementation process remains a major obstacle to high productivity computing, limiting the number of compile-debug-edit cycles per day. Furthermore, high-level application developers often lack the intimate hardware engineering experience that is needed to achieve high performance on FPGAs, therefore undermining their usefulness as accelerators. To address the productivity and performance problems, a HLS methodology that utilizes soft coarse-grained reconfigurable arrays (SCGRAs) as an intermediate compilation step is presented. Instead of compiling high-level applications directly to circuits, the compilation process is reduced to an operation scheduling task targeting the SCGRA. Cheng Liu 0008, Colin Yu Lin, Hayden Kwok-Hay So |
FCCM | 3 |
| 2013 | Direct virtual memory access from FPGA for high-productivity heterogeneous computingabstractHeterogeneous computing utilizing both CPU and FPGA requires access to data in the main memory from both devices. While a typical system relies on software executing on the CPU to orchestrate all data movements between the FPGA and the main memory, our demo presents a complementary FPGA-centric approach that allows gateware to directly access the virtual memory space as part of the executing process without involving the CPU. A caching address translation buffer was implemented alongside the user FPGA gateware to provide runtime mapping between virtual and physical memory addresses. The system was implemented on a commercial off-the-shelf FPGA add-on card to demonstrate the viability of such approach in low-cost systems. Experiment demonstrated reasonable performance improvement when compared to a typical software-centric implementation; while the number of context switches between FPGA and CPU in both kernel and user mode was significantly reduced, freeing the CPU for other concurrent user tasks. Ho-Cheung Ng, Yuk-Ming Choi, Hayden Kwok-Hay So |
FPT | 3 |
| 2012 | Operation scheduling and architecture co-synthesis for energy-efficient dataflow computations on FPGAs (abstract only)abstractCompiling high-level user applications for execution on FPGAs often involves synthesizing dataflow graphs beyond the size of the available on-chip computational resources. One way to address this is by folding the execution of the given dataflow graphs onto an array of directly connected simple configurable processing elements (CPEs). Under this scenario, the performance and energy-efficiency of the resulting system depends not only on the mapping schedule of the compute operations on the CPEs, but also on the topology of the interconnect array that connects the CPEs. This paper presents a framework in which the operation scheduler and the underlying CPE interconnect network topology are co-optimized on a per-application basis for energy-efficient FPGA computation. Given the same application, more than 2.5x difference in energy-efficiency was achievable by the use of different common regular array topologies to connect the CPEs. Moreover, by using irregular application-specific interconnect topologies derived from a genetic algorithm, up to 50% improvement in energy-delay-product was achievable when compared to the use of even the best regular topology. The use of such framework is anticipated to serve as part of a rapid high-level FPGA application compiler since minimum hardware place-and-route is needed to generate the optimal schedule and topology. Colin Yu Lin, Ngai Wong 0001, Hayden Kwok-Hay So |
FPGA | 3 |
| 2012 | Extending BORPH for shared memory reconfigurable computersabstractWe extend BORPH for shared memory reconfigurable computers in this paper. BORPH is an operating system designed for FPGA based reconfigurable computers. BORPH introduced the concept of hardware process in contrast to software process. With our extension, hardware processes are supported to communicate with other processes based on shared memory. In our system, the program of hardware process is not just hardware design, but the software program running on embedded processor in FPGA. Our experiment shows the overhead of shared memory segments management is acceptable. And with independent virtual memory access, bandwidth of repeated shared memory access is high. Changqing Xun, Mei Wen, Nan Wu 0003, Chunyuan Zhang, Hayden Kwok-Hay So |
FPL | 5 |
| 2012 | Design considerations of real-time adaptive beamformer for medical ultrasound research using FPGA and GPUabstractAdaptive beamforming has been well considered as a potential solution for improving the imaging quality of medical ultrasound systems. Despite the promised improvement in lateral resolution, image contrast and imaging penetration, the use of adaptive beamforming is substantially more computationally demanding than conventional delay-and-sum beamformers. While a dedicated hardware solution may be able to address the computational demand of one particular design, the need for an efficient algorithm exploration framework demands a platform solution that is high-performance and easily reprogrammable. To that end, the use of FPGA and GPU for implementing real-time adaptive beamforming on such platform has been explored. The results are evaluated quantitatively in terms of performance and image quality, and qualitatively with respect to ease of system integration and ease of use. In our test cases, both FPGA- and GPU-based solutions achieved real-time throughput exceeding 80 frames-per-second, and over 38x improvement when compared to our baseline CPU implementation. While the development time on GPU platform remains much lower than its FPGA counterpart, the FPGA solution is effective in providing the necessary I/O bandwidth to enable an end-to-end real-time reconfigurable image formation system. Alfred C. H. Yu, Hayden Kwok-Hay So |
FPT | 3 |
| 2011 | A Model for Peak Matrix Performance on FPGAs
Colin Yu Lin, Hayden Kwok-Hay So, Philip H. W. Leong |
FCCM | 2 |
| 2011 | A Model for Matrix Multiplication Performance on FPGAsabstractComputations involving matrices form the kernel of a large spectrum of computationally demanding applications for which FPGAs have been utilized as accelerators. Their performance is related to their underlying architectural and system parameters such as computational resources, memory and I/O bandwidth. A simple analytic model that gives an estimate of the performance of FPGA-based sparse matrix-vector and matrix-matrix multiplication is presented, dense matrix multiplication being a special case. The efficiency of existing implementations are compared to the model and performance trends for future technologies examined. Colin Yu Lin, Hayden Kwok-Hay So, Philip H. W. Leong |
FPL | 2 |
| 2010 | Design space exploration for sparse matrix-matrix multiplication on FPGAsabstractThe design and implementation of a sparse matrix-matrix multiplication architecture on FPGAs is presented. Performance of the design, in terms of computational latency, as well as the associated power-delay and energy-delay tradeoff are studied. Taking advantage of the sparsity of the input matrices, the proposed design allows user-tunable power-delay and energy-delay tradeoffs by employing different number of processing elements (PEs) in the architecture design and different block size in the blocking decomposition. Such ability allows designers to employ different on-chip computational architecture for different system power-delay and energy-delay requirements. It is in contrast to conventional dense matrix-matrix multiplication architectures that always favor the maximum number of PEs and largest block size. In our implementation, the better energy consumption and power-delay product favors less PEs and smaller block size for the 90%-sparsity matrix-matrix multiplications. While in order to achieve better energy-delay product, more PEs and larger block size are preferred. Colin Yu Lin, Zheng Zhang 0005, Ngai Wong 0001, Hayden Kwok-Hay So |
FPT | 4 |
| 2009 | Operation scheduling for FPGA-based reconfigurable computersabstractMany high-performance applications involve large data sets that are impossible to fit entirely within on-chip memories of even the largest FPGAs. As a result, they must be stored in off-chip SDRAMs and loaded onto the FPGAs as computations progress. Because of the high latency and energy consumption associated with off-chip memory accesses, it is important to develop efficient operation schedules that not only minimize latency of computations, but also the amount of data I/Os. We formulate this problem as a modified resource-constrained job scheduling problem. The problem is then solved using a list scheduling algorithm that takes advantage of the fast burst-mode access of SDRAMs. Results have shown that for large problem sizes, the performance of our algorithm is within 1% of a hand-optimized matrix-matrix multiplication implementation, with no memory overhead, and is within 0.03% of the theoretical minimum latency of an 8-by-8 cofactor matrix computation. Colin Yu Lin, Ngai Wong 0001, Hayden Kwok-Hay So |
FPL | 3 |
| 2008 | Runtime Filesystem Support for Reconfigurable FPGA Hardware Processes in BORPHabstractThis paper presents the design of BORPH's file system layer for FPGA-based reconfigurable computers. BORPH provides user FPGA designs that execute as hardware processes access to the general file system using familiar UNIX file I/O semantics. Such capability provides FPGA designers an intuitive interface not only for regular file I/O, but also for representing streaming hardware/software and hardware/hardware communication using UNIX pipes. Design trade-offs among system manageability, user usability and application performance are explored. A case of mixed hardware/software video processing is presented as a proof-of-concept. Hayden Kwok-Hay So, Robert W. Brodersen |
FCCM | 1 |
| 2008 | Direct sigma-delta modulated signal processing in FPGAabstractThe effectiveness of implementing bit-stream signal processing (BSSP) multiplier circuits in FPGAs, in terms of hardware resources and clock frequency, is presented. In particular, the result of realizing BSSP multipliers on FPGA architectures that utilize 6-input lookup tables (LUTs) is compared against architectures that utilize 4-input LUTs. It is found that architectures featuring 6-input LUTs suit well in BSSP applications where wide combinatorial paths are common. Furthermore, the performance of a BSSP multiplier is compared against conventional parallel multipliers in terms of LUT resource requirements. For a given resource requirement, it is found that an over-sampling ratio of less than 32 is required for a BSSP multiplier to outperform its parallel counterpart. Chiu-Wah Ng, Ngai Wong 0001, Hayden Kwok-Hay So, Tung-Sang Ng |
FPL | 3 |
| 2008 | File system access from reconfigurable FPGA hardware processes in BORPHabstractThis paper presents the design and implementation of BORPH’s kernel file system layer that provides FPGA processes direct access to the general file system. Using a semantics resembling that of conventional UNIX file I/Os, an FPGA accesses the file system through a special hardware system call interface. By extending the semantics of a UNIX pipe, a single file system access mechanism is used for both regular file I/O, as well as for hardware/software and hardware/hardware data streaming. An FPGA design may switch between different communication modes dynamically during run time by means of file redirection. Design trade-offs among system manageability, user usability and application performance are explored. An example of constructing a video processing system during run time using commodity software and FPGA applications connected by pipes is used to demonstrate the feasibility and potential of such FPGA-centric file system access capability. Hayden Kwok-Hay So, Robert W. Brodersen |
FPL | 1 |
| 2008 | Quad-level bit-stream signal processing on FPGAsabstractQuad-level bit-stream signal processing (BSSP) circuits are implemented and their performances are compared with previously published tri-level and bi-level BSSP implementations on FPGAs. BSSP refers to the process of performing computation directly on over-sampled delta-sigma modulated signals to eliminate the need of resource consuming decimators and interpolators. Quad-level BSSP offers better performance than their bi-and tri-level counterparts at the expense of higher resource utilization. Using a digital phase locked loop (DPLL) and a quadrature phase-shift keying (QPSK) demodulator as application examples, the effectiveness of quad-level BSSP on FPGAs is studied. The BSSP approach will be contrasted with conventional multi-bit implementations using built-in digital signal processing blocks in modern FPGAs. Chiu-Wah Ng, Ngai Wong 0001, Hayden Kwok-Hay So, Tung-Sang Ng |
FPT | 3 |
| 2008 | A unified hardware/software runtime environment for FPGA-based reconfigurable computers using BORPHabstractThis paper explores the design and implementation of BORPH, an operating system designed for FPGA-based reconfigurable computers. Hardware designs execute as normal UNIX processes under BORPH, having access to standard OS services, such as file system support. Hardware and software components of user designs may, therefore, run as communicating processes within BORPH's runtime environment. The familiar language independent UNIX kernel interface facilitates easy design reuse and rapid application development. To develop hardware designs, a Simulink-based design flow that integrates with BORPH is employed. Performances of BORPH on two on-chip systems implemented on a BEE2 platform are compared. Hayden Kwok-Hay So, Robert W. Brodersen |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2006 | Improving Usability of FPGA-Based Reconfigurable Computers Through Operating System SupportabstractAdvances in FPGA-based reconfigurable computers have made them a viable computing platform for a vast variety of computation demanding areas such as bioinformatics, speech recognition, and high-end digital signal processing. The lack of common, intuitive operating system support, however, hinders their wide deployment. This paper presents BORPH, an operating system framework for FPGA-based reconfigurable computers with a goal to ease and accelerate development of high-level applications to run on these computers. It provides kernel support for FPGA resources by extending a standard Linux operating system. Users therefore compile and execute hardware processes on FPGA resources the same way they run software processes on conventional processor-based systems. The operating system offers run-time general file system support to hardware processes as if they were software. Furthermore, a virtual file system is built to allow access to memories and registers defined in the FPGA, which provides communication links with running hardware processes. Increased productivities have been observed for high-level application developers, who have few previous experiences in hardware design, to implement complex mixed software/hardware designs on a FPGA-based reconfigurable computer running BORPH Hayden Kwok-Hay So, Robert W. Brodersen |
FPL | 1 |