EDBT 2026 Demo / reviewers in the wild / expert
Thomas C. P. Chau
dblp:18/5848 · also Thomas Chau 0001, Thomas Chun-Pong Chau
· DBLP profile ↗
26ranked-venue papers
8as first author
4since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 7 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Efficient and distributed learning · 75% Graph learning · 12% Speech recognition and synthesis · 7% | |
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Hardware accelerators and domain-specific architectures · 47% Reconfigurable computing and FPGAs · 22% Performance modeling and evaluation · 19% |
Topics — the 25 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
2.6 | 5 | 2023 | Zero-Cost Operation Scoring in Differentiable Architecture Search · AAAI 2023 BLOX: Macro Neural Architecture Search Benchmark and Algorithms · NeurIPS 2022 NAS-Bench-ASR: Reproducible Neural Architecture Search for Speech Recognition · ICLR 2021 |
Reconfigurable computing and FPGAs
FPGA accelerator |
1.0 | 2 | 2022 | Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design · MICRO 2022 Codesign-NAS: Automatic FPGA/CNN Codesign Using Neural Architecture Search · FPGA 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
1.0 | 2 | 2022 | Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design · MICRO 2022 Best of Both Worlds: AutoML Codesign of a CNN and its Hardware Accelerator · DAC 2020 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
hardware-aware neural architecture search |
0.9 | 2 | 2020 | BRP-NAS: Prediction-based NAS using GCNs · NeurIPS 2020 Best of Both Worlds: AutoML Codesign of a CNN and its Hardware Accelerator · DAC 2020 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search › one-shot neural architecture search
differentiable architecture search |
0.7 | 1 | 2023 | Zero-Cost Operation Scoring in Differentiable Architecture Search · AAAI 2023 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
zero-cost proxy |
0.7 | 1 | 2023 | Zero-Cost Operation Scoring in Differentiable Architecture Search · AAAI 2023 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
attention accelerator |
0.6 | 1 | 2022 | Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design · MICRO 2022 |
Performance modeling and evaluation
benchmarking |
0.6 | 1 | 2022 | BLOX: Macro Neural Architecture Search Benchmark and Algorithms · NeurIPS 2022 |
Performance modeling and evaluation › benchmarking › parallel benchmark suites
NAS parallel benchmarks |
0.6 | 1 | 2022 | BLOX: Macro Neural Architecture Search Benchmark and Algorithms · NeurIPS 2022 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.5 | 1 | 2021 | NAS-Bench-ASR: Reproducible Neural Architecture Search for Speech Recognition · ICLR 2021 |
Machine learning › Graph learning › graph neural network
graph convolutional network |
0.4 | 1 | 2020 | BRP-NAS: Prediction-based NAS using GCNs · NeurIPS 2020 |
Machine learning › Graph learning
graph neural network |
0.4 | 1 | 2020 | BRP-NAS: Prediction-based NAS using GCNs · NeurIPS 2020 |
Machine learning › Efficient and distributed learning
latency prediction |
0.4 | 1 | 2020 | BRP-NAS: Prediction-based NAS using GCNs · NeurIPS 2020 |
Machine learning › Efficient and distributed learning
model compression |
0.4 | 1 | 2020 | BRP-NAS: Prediction-based NAS using GCNs · NeurIPS 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator |
0.4 | 1 | 2020 | Codesign-NAS: Automatic FPGA/CNN Codesign Using Neural Architecture Search · FPGA 2020 |
Hardware accelerators and domain-specific architectures › neural architecture search
hardware-aware neural architecture search |
0.4 | 1 | 2020 | Codesign-NAS: Automatic FPGA/CNN Codesign Using Neural Architecture Search · FPGA 2020 |
Electronic design automation
hardware/software co-design |
0.4 | 1 | 2020 | Best of Both Worlds: AutoML Codesign of a CNN and its Hardware Accelerator · DAC 2020 |
Hardware accelerators and domain-specific architectures
neural architecture search |
0.4 | 1 | 2020 | Codesign-NAS: Automatic FPGA/CNN Codesign Using Neural Architecture Search · FPGA 2020 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.2 | 1 | 2022 | Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design · MICRO 2022 |
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention |
0.2 | 1 | 2022 | Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design · MICRO 2022 |
Reconfigurable computing and FPGAs
dynamic reconfiguration |
0.2 | 1 | 2013 | Automating resource optimisation in reconfigurable design (abstract only) · FPGA 2013 |
Cloud and datacenter computing › resource management
resource optimization |
0.2 | 1 | 2013 | Automating resource optimisation in reconfigurable design (abstract only) · FPGA 2013 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2020 | Codesign-NAS: Automatic FPGA/CNN Codesign Using Neural Architecture Search · FPGA 2020 |
Integrated circuit design
digital circuit design |
0.1 | 1 | 2009 | A comparison of via-programmable gate array logic cell circuits · FPGA 2009 |
Integrated circuit design
low-power circuit design |
0.1 | 1 | 2009 | A comparison of via-programmable gate array logic cell circuits · FPGA 2009 |
Methods — techniques the papers use, named apart from their topics
neural architecture search · 2.0multi-objective optimization · 1.7butterfly sparsity · 1.1zero-cost operation scoring · 0.7perturbation-based scoring · 0.7hardware/algorithm co-design · 0.6hardware-algorithm co-design · 0.6blockwise search · 0.6block-wise search · 0.6reinforcement learning · 0.4performance prediction · 0.4pareto frontier · 0.4iterative data selection · 0.4graph convolutional network · 0.4run-time solution generation · 0.2function analysis · 0.2configuration organization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Zero-Cost Operation Scoring in Differentiable Architecture SearchabstractWe formalize and analyze a fundamental component of dif- ferentiable neural architecture search (NAS): local “opera- tion scoring” at each operation choice. We view existing operation scoring functions as inexact proxies for accuracy, and we find that they perform poorly when analyzed empir- ically on NAS benchmarks. From this perspective, we intro- duce a novel perturbation-based zero-cost operation scor- ing (Zero-Cost-PT) approach, which utilizes zero-cost prox- ies that were recently studied in multi-trial NAS but de- grade significantly on larger search spaces, typical for dif- ferentiable NAS. We conduct a thorough empirical evalu- ation on a number of NAS benchmarks and large search spaces, from NAS-Bench-201, NAS-Bench-1Shot1, NAS- Bench-Macro, to DARTS-like and MobileNet-like spaces, showing significant improvements in both search time and accuracy. On the ImageNet classification task on the DARTS search space, our approach improved accuracy compared to the best current training-free methods (TE-NAS) while be- ing over 10× faster (total searching time 25 minutes on a single GPU), and observed significantly better transferabil- ity on architectures searched on the CIFAR-10 dataset with an accuracy increase of 1.8 pp. Our code is available at: https://github.com/zerocostptnas/zerocost operation score. Lichuan Xiang, Lukasz Dudziak, Mohamed S. Abdelfattah, Thomas C. P. Chau, Nicholas D. Lane, Hongkai Wen 0001 |
AAAI | 4 |
| 2022 | Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-designabstractAttention-based neural networks have become pervasive in many AI tasks. Despite their excellent algorithmic performance, the use of the attention mechanism and feedforward network (FFN) demands excessive computational and memory resources, which often compromises their hardware performance. Although various sparse variants have been introduced, most approaches only focus on mitigating the quadratic scaling of attention on the algorithm level, without explicitly considering the efficiency of mapping their methods on real hardware designs. Furthermore, most efforts only focus on either the attention mechanism or the FFNs but without jointly optimizing both parts, causing most of the current designs to lack scalability when dealing with different input lengths. This paper systematically considers the sparsity patterns in different variants from a hardware perspective. On the algorithmic level, we propose FABNet, a hardware-friendly variant that adopts a unified butterfly sparsity pattern to approximate both the attention mechanism and the FFNs. On the hardware level, a novel adaptable butterfly accelerator is proposed that can be configured at runtime via dedicated hardware control to accelerate different butterfly layers using a single unified hardware engine. On the Long-Range-Arena dataset, FABNet achieves the same accuracy as the vanilla Transformer while reducing the amount of computation by 10$\sim66\times$ and the number of parameters 2$\sim22\times$. By jointly optimizing the algorithm and hardware, our FPGA-based butterfly accelerator achieves 14.2$\sim23.2\times$ speedup over state-of-the-art accelerators normalized to the same computational budget. Compared with optimized CPU and GPU designs on Raspberry Pi 4 and Jetson Nano, our system is up to $273.8\times$ and $15.1\times$ faster under the same power budget Hongxiang Fan, Thomas C. P. Chau, Stylianos I. Venieris, Royson Lee, Alexandros Kouris, Wayne Luk, Nicholas D. Lane, Mohamed S. Abdelfattah |
MICRO | 2 |
| 2022 | BLOX: Macro Neural Architecture Search Benchmark and AlgorithmsabstractNeural architecture search (NAS) has been successfully used to design numerous high-performance neural networks. However, NAS is typically compute-intensive, so most existing approaches restrict the search to decide the operations and topological structure of a single block only, then the same block is stacked repeatedly to form an end-to-end model. Although such an approach reduces the size of search space, recent studies show that a macro search space, which allows blocks in a model to be different, can lead to better performance. To provide a systematic study of the performance of NAS algorithms on a macro search space, we release Blox – a benchmark that consists of 91k unique models trained on the CIFAR-100 dataset. The dataset also includes runtime measurements of all the models on a diverse set of hardware platforms. We perform extensive experiments to compare existing algorithms that are well studied on cell-based search spaces, with the emerging blockwise approaches that aim to make NAS scalable to much larger macro search spaces. The Blox benchmark and code are available at https://github.com/SamsungLabs/blox. Thomas C. P. Chau, Lukasz Dudziak, Hongkai Wen 0001, Nicholas D. Lane, Mohamed S. Abdelfattah |
NeurIPS | 1 |
| 2021 | NAS-Bench-ASR: Reproducible Neural Architecture Search for Speech Recognition
Abhinav Mehrotra, Alberto Gil C. P. Ramos, Sourav Bhattacharya, Lukasz Dudziak, Ravichander Vipperla, Thomas C. P. Chau, Mohamed S. Abdelfattah, Samin Ishtiaq, Nicholas D. Lane |
ICLR | 6 |
| 2020 | Best of Both Worlds: AutoML Codesign of a CNN and its Hardware AcceleratorabstractNeural architecture search (NAS) has been very successful at outperforming human-designed convolutional neural networks (CNN) in accuracy, and when hardware information is present, latency as well. However, NAS-designed CNNs typically have a complicated topology, therefore, it may be difficult to design a custom hardware (HW) accelerator for such CNNs. We automate HW-CNN codesign using NAS by including parameters from both the CNN model and the HW accelerator, and we jointly search for the best model-accelerator pair that boosts accuracy and efficiency. We call this Codesign-NAS. In this paper we focus on defining the Codesign-NAS multiobjective optimization problem, demonstrating its effectiveness, and exploring different ways of navigating the codesign search space. For CIFAR-10 image classification, we enumerate close to 4 billion model-accelerator pairs, and find the Pareto frontier within that large search space. This allows us to evaluate three different reinforcement-learning-based search strategies. Finally, compared to ResNet on its most optimal HW accelerator from within our HW design space, we improve on CIFAR-100 classification accuracy by 1.3% while simultaneously increasing performance/area by 41% in just ~1000 GPU-hours of running Codesign-NAS. Mohamed S. Abdelfattah, Lukasz Dudziak, Thomas C. P. Chau, Royson Lee, Hyeji Kim, Nicholas D. Lane |
DAC | 3 |
| 2020 | Codesign-NAS: Automatic FPGA/CNN Codesign Using Neural Architecture SearchabstractField-programmable gate arrays (FPGAs) have become a popular compute platform for convolutional neural network (CNN) inference; however, the design of a CNN model and its FPGA accelerator has been inherently sequential. A CNN is first prototyped with no-or-little hardware awareness to attain high accuracy; subsequently, an FPGA accelerator is tuned to that specific CNN to maximize its efficiency. Instead, we formulate a neural architecture search (NAS) optimization problem that contains parameters from both the CNN model and the FPGA accelerator, and we jointly search for the best CNN model-accelerator pair that boosts accuracy and efficiency -we call this Codesign-NAS. In this paper we focus on defining the Codesign-NAS multiobjective optimization problem, demonstrating its effectiveness, and exploring different ways of navigating the codesign search space. For Cifar-10 image classification, we enumerate close to 4 billion model-accelerator pairs, and find the Pareto frontier within that large search space. Next we propose accelerator innovations that improve the entire Pareto frontier. Finally, we compare to ResNet on a highly-tuned accelerator, and show that using codesign, we can improve on Cifar-100 classification accuracy by 1.8% while simultaneously increasing performance/area by 41% in just 1000 GPU-hours of running Codesign-NAS, thus demonstrating that our automated codesign approach is superior to sequential design of a CNN model and accelerator. Mohamed S. Abdelfattah, Lukasz Dudziak, Thomas C. P. Chau, Royson Lee, Hyeji Kim, Nicholas D. Lane |
FPGA | 3 |
| 2020 | BRP-NAS: Prediction-based NAS using GCNsabstractNeural architecture search (NAS) enables researchers to automatically explore broad design spaces in order to improve efficiency of neural networks. This efficiency is especially important in the case of on-device deployment, where improvements in accuracy should be balanced out with computational demands of a model. In practice, performance metrics of model are computationally expensive to obtain. Previous work uses a proxy (e.g., number of operations) or a layer-wise measurement of neural network layers to estimate end-to-end hardware performance but the imprecise prediction diminishes the quality of NAS. To address this problem, we propose BRP-NAS, an efficient hardware-aware NAS enabled by an accurate performance predictor-based on graph convolutional network (GCN). What is more, we investigate prediction quality on different metrics and show that sample efficiency of the predictor-based NAS can be improved by considering binary relations of models and an iterative data selection strategy. We show that our proposed method outperforms all prior methods on NAS-Bench-101, NAS-Bench-201 and DARTS. Finally, to raise awareness of the fact that accurate latency estimation is not a trivial task, we release LatBench -- a latency dataset of NAS-Bench-201 models running on a broad range of devices. Lukasz Dudziak, Thomas C. P. Chau, Mohamed S. Abdelfattah, Royson Lee, Hyeji Kim, Nicholas D. Lane |
NeurIPS | 2 |
| 2019 | Transparent Heterogeneous Cloud AccelerationabstractThis work proposes a cloud computing platform (PaaS) with a novel micro-service architecture designed to support transparent acceleration on large-scaled heterogeneous cloud infrastructures with hardware accelerators such as FPGAs. Jessica Vandebon, José Gabriel F. Coutinho, Wayne Luk, Thomas C. P. Chau |
ASAP | 4 |
| 2018 | Towards Hardware Accelerated Reinforcement Learning for Application-Specific Robotic ControlabstractReinforcement Learning (RL) is an area of machine learning in which an agent interacts with the environment by making sequential decisions. The agent receives reward from the environment based on how good the decisions are and tries to find an optimal decision-making policy that maximises its longterm cumulative reward. This paper presents a novel approach which has showon promise in applying accelerated simulation of RL policy training to automating the control of a real robot arm for specific applications. The approach has two steps. First, design space exploration techniques are developed to enhance performance of an FPGA accelerator for RL policy training based on Trust Region Policy Optimisation (TRPO), which results in a 43% speed improvement over a previous FPGA implementation, while achieving 4.65 times speed up against deep learning libraries running on GPU and 19.29 times speed up against CPU. Second, the trained RL policy is transferred to a real robot arm. Our experiments show that the trained arm can successfully reach to and pick up predefined objects, demonstrating the feasibility of our approach. Shengjia Shao, Jason Tsai, Michal Mysior, Wayne Luk, Thomas C. P. Chau, Alexander Warren, B. P. Jeppesen |
ASAP | 5 |
| 2016 | An FPGA-based platform for integrated power and motion controlabstractMotor control applications using small fast-spinning motors such as e-turbos, UAVs, surgical instruments and high-speed pumps, and the advent of SiC and GaN switching transistors, are driving a need for compact drives with high-frequency control loop updates. This paper describes a multi-axis motor and power control platform suitable for academic or commercial research and development. It includes motor and power control kit, FPGA development board and FPGA hardware and software design. It supports integrated power and motor control with PWM time resolution up to 300MHz and control update frequencies of 100kHz or more. Ben P. Jeppesen, Andrew Crosland, Thomas C. P. Chau |
IECON | 3 |
| 2015 | Recursive pipelined genetic propagation for bilevel optimisationabstractThe bilevel optimisation problem (BLP) is a subclass of optimisation problems in which one of the constraints of an optimisation problem is another optimisation problem. BLP is widely used to model hierarchical decision making where the leader and the follower correspond to the upper level and lower level optimisation problem, respectively. In BLP, the optimal solutions to the lower level optimisation problem are the feasible solutions to the upper level problem, which makes it particularly difficult to solve. This paper proposes a novel hardware architecture known as Recursive Pipelined Genetic Propagation (RPGP), to solve BLP efficiently on FPGA. RPGP features a graph of genetic operation nodes which can be scaled to exploit hardware resources. In addition, the topology of the RPGP graph can be changed at run-time to escape from local optima. We evaluate the proposed architecture on an Altera Stratix-V FPGA, using a benchmark bilevel optimisation problem set. Our experimental results show that RPGP can achieve a significant speed-up against previous work. Shengjia Shao, Liucheng Guo, Ce Guo 0002, Thomas C. P. Chau, David B. Thomas, Wayne Luk, Stephen Weston |
FPL | 4 |
| 2015 | Mapping Adaptive Particle Filters to Heterogeneous Reconfigurable SystemsabstractThis article presents an approach for mapping real-time applications based on particle filters (PFs) to heterogeneous reconfigurable systems, which typically consist of multiple FPGAs and CPUs. A method is proposed to adapt the number of particles dynamically and to utilise runtime reconfigurability of FPGAs for reduced power and energy consumption. A data compression scheme is employed to reduce communication overhead between FPGAs and CPUs. A mobile robot localisation and tracking application is developed to illustrate our approach. Experimental results show that the proposed adaptive PF can reduce up to 99% of computation time. Using runtime reconfiguration, we achieve a 25% to 34% reduction in idle power. A 1U system with four FPGAs is up to 169 times faster than a single-core CPU and 41 times faster than a 1U CPU server with 12 cores. It is also estimated to be 3 times faster than a system with four GPUs. Thomas C. P. Chau, Xinyu Niu, Alison Eele, Jan M. Maciejowski, Peter Y. K. Cheung, Wayne Luk |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2015 | Automating Elimination of Idle Functions by Runtime ReconfigurationabstractA design approach is proposed to automatically identify and exploit runtime reconfiguration opportunities with optimised resource utilisation by eliminating idle functions. We introduce Reconfiguration Data Flow Graph, a hierarchical graph structure enabling reconfigurable designs to be synthesised in three steps: function analysis, configuration organisation, and runtime solution generation. The synthesised reconfigurable designs are dynamically evaluated and selected under various runtime conditions. Three applications—barrier option pricing, particle filter, and reverse time migration—are used in evaluating the proposed approach. The runtime solutions approximate their theoretical performance by eliminating idle functions and are 1.31 to 2.19 times faster than optimised static designs. FPGA designs developed with the proposed approach are up to 43.8 times faster than optimised CPU reference designs and 1.55 times faster than optimised GPU designs. Xinyu Niu, Thomas C. P. Chau, Qiwei Jin, Wayne Luk, Qiang Liu 0011, Oliver Pell |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2014 | SMCGen: Generating Reconfigurable Design for Sequential Monte Carlo ApplicationsabstractThe Sequential Monte Carlo (SMC) method is a simulation-based approach to compute posterior distributions. SMC methods often work well on applications considered intractable by other methods due to high dimensionality, but they are computationally demanding. While SMC has been implemented efficiently on FPGAs, design productivity remains a challenge. This paper introduces a design flow for generating efficient implementation of reconfigurable SMC designs. Through templating the SMC structure, the design flow enables efficient mapping of SMC applications to multiple FPGAs. The proposed design flow consists of a parametrisable SMC computation engine, and an open-source software template which enables efficient mapping of a variety of SMC designs to reconfigurable hardware. Design parameters that are critical to the performance and to the solution quality are tuned using a machine learning algorithm based on surrogate modelling. Experimental results for three case studies show that design performance is substantially improved after parameter optimisation. The proposed design flow demonstrates its capability of producing reconfigurable implementations for a range of SMC applications that have significant improvement in speed and in energy efficiency over optimised CPU and GPU implementations. Thomas C. P. Chau, Maciej Kurek, James Stanley Targett, Jake Humphrey, Georgios Skouroupathis, Alison Eele, Jan M. Maciejowski, Benjamin Cope, Kathryn Cobden, Philip H. W. Leong, Peter Y. K. Cheung, Wayne Luk |
FCCM | 1 |
| 2014 | Automating Optimization of Reconfigurable DesignsabstractWe present Automatic Reconfigurable Design Efficient Global Optimization (ARDEGO), a new algorithm based on the existing Efficient Global Optimization (EGO) methodology for automating optimization of reconfigurable designs targeting Field-Programmable Gate Array (FPGA) technology. It is a potentially disruptive design approach: instead of manually improving designs repeatedly but without understanding the design space as a whole, ARDEGO users follow a novel approach that: (a) automates the manual optimization process, significantly reducing optimization time and (b) does not require the user to calibrate or understand the inner workings of the algorithm. We evaluate ARDEGO using two case studies: financial option pricing and seismic imaging. Maciej Kurek, Tobias Becker, Thomas C. P. Chau, Wayne Luk |
FCCM | 3 |
| 2013 | Automating Elimination of Idle Functions by Run-Time ReconfigurationabstractA design approach is proposed to automatically identify and exploit run-time reconfiguration opportunities while optimising resource utilisation. We introduce Reconfiguration Data Flow Graph, a hierarchical graph structure enabling reconfigurable designs to be synthesised in three steps: function analysis, configuration organisation, and run-time solution generation. Three applications, based on barrier option pricing, particle filter, and reverse time migration are used in evaluating the proposed approach. The run-time solutions approximate the theoretical performance by eliminating idle functions, and are 1.31 to 2.19 times faster than optimised static designs. FPGA designs developed with the proposed approach are up to 28.8 times faster than optimised CPU reference designs and 1.55 times faster than optimised GPU designs. Xinyu Niu, Thomas C. P. Chau, Qiwei Jin, Wayne Luk, Qiang Liu 0011 |
FCCM | 2 |
| 2013 | Automating resource optimisation in reconfigurable design (abstract only)abstractA design approach is proposed to automatically identify and exploit run-time reconfiguration opportunities while optimising resource utilisation. We introduce Configuration Data Flow Graph, a hierarchical graph structure enabling reconfigurable designs to be synthesised in three steps: function analysis, configuration organisation, and run-time solution generation. Three applications, based on barrier option pricing, particle filter, and reverse time migration are used in evaluating the proposed approach. The run-time solutions approximate the theoretical performance by eliminating idle functions, and are 1.61 to 2.19 times faster than optimised static designs. FPGA designs developed with the proposed approach are up to 28.8 times faster than optimised CPU reference designs and 1.55 times faster than optimised GPU designs. Xinyu Niu, Thomas C. P. Chau, Qiwei Jin, Wayne Luk, Qiang Liu 0011 |
FPGA | 2 |
| 2013 | Acceleration of real-time Proximity Query for dynamic active constraintsabstractProximity Query (PQ) is a process to calculate the relative placement of objects. It is a critical task for many applications such as robot motion planning, but it is often too computationally demanding for real-time applications, particularly those involving human-robot collaborative control. This paper derives a PQ formulation which can support non-convex objects represented by meshes or cloud points. We optimise the proposed PQ for reconfigurable hardware by function transformation and reduced precision, resulting in a novel data structure and memory architecture for data streaming while maintaining the accuracy of results. Run-time reconfiguration is adopted for dynamic precision optimisation. Experimental results show that our optimised PQ implementation on a reconfigurable platform with four FPGAs is 58 times faster than an optimised CPU implementation with 12 cores, 9 times faster than a GPU, and 3 times faster than a double precision implementation with four FPGAs. Thomas C. P. Chau, Ka-Wai Kwok, Gary C. T. Chow, Kuen Hung Tsoi, Kit-Hang Lee, Zion Tsz Ho Tse, Peter Y. K. Cheung, Wayne Luk |
FPT | 1 |
| 2013 | Architecture and Design Flow for a Highly Efficient Structured ASICabstractAs fabrication process technology continues to advance, mask set costs have become prohibitively expensive. Structured application specific integrated circuits (sASICs) offer a middle ground in price and performance between ASICs and field-programmable gate arrays (FPGAs) by sharing masks across different designs. In this paper, two sASIC architectures are proposed, the first being based on three-input lookup-tables, and the second on AOI22 gates. The sASICs are programmed using a standard-cell compatible design flow. They are customized using a minimum of three masks, i.e., two metals and one via. The area and delay of the sASIC are compared with ASICs and FPGAs. Results over a set of benchmark circuits show that our AOI22-based sASIC had an average of 1.76x/1.41x increase in area/delay compared to ASICs, a considerable improvement compared with the 26.56x/5.09x increase for FPGAs. This is, to the best of our knowledge, the best performance reported in the literature for a practical sASIC. A prototype using the sASIC was fabricated using a universal machine control 0.13-μm mixed-mode/RF process. It was fully verified using scan and functional tests, and used in a demonstration system. S. Man Ho Ho, Yanqing Ai, Thomas C. P. Chau, Steve C. L. Yuen, Oliver Chiu-sing Choy, Philip H. W. Leong, Kong-Pang Pun |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | Adaptive Sequential Monte Carlo approach for real-time applicationsabstractThis paper presents an adaptive Sequential Monte Carlo approach for real-time applications. Sequential Monte Carlo method is employed to estimate the states of dynamic systems using weighted particles. The proposed approach reduces the run-time computation complexity by adapting the size of the particle set. Multiple processing elements on FPGAs are dynamically allocated for improved energy efficiency without violating real-time constraints. A robot localisation application is developed based on the proposed approach. Compared to a non-adaptive implementation, the dynamic energy consumption is reduced by up to 70% without affecting the quality of solutions. Thomas C. P. Chau, Wayne Luk, Peter Y. K. Cheung, Alison Eele, Jan M. Maciejowski |
FPL | 1 |
| 2010 | Rapid prototyping on a structured ASIC fabricabstractWe describe the architecture of a structured ASIC fabric in which the logic and routing can be customized using three masks. A standard Cadence based design flow is employed, and using an active dynamic backlight controller as an example, performance is compared to that of an ASIC implementation in the same technology. Steve C. L. Yuen, Yanqing Ai, Brian P. W. Chan, Thomas C. P. Chau, Sam M. H. Ho, Oscar K. L. Lau, Kong-Pang Pun, Philip H. W. Leong, Oliver Chiu-sing Choy |
ASP-DAC | 4 |
| 2010 | Design of a single layer programmable Structured ASIC libraryabstractA Structured Application-specific Integrated Circuit (SASIC) is a programmable fabric in which a small set of masks are customized for a particular application, serving to reduce the associated non-recurring engineering cost (NRE). In this paper we describe the implementation of a SASIC logic cell which is programmable via a single metal layer. A SASIC fabric prototype is fabricated and all implemented functions are verified on silicon. Experimental measurement verifies correct operation of our SASIC with a clock frequency of over 250 MHz. Thomas C. P. Chau, David W. L. Wu, Yanqing Ai, Brian P. W. Chan, Sam M. H. Ho, Oscar K. L. Lau, Steve C. L. Yuen, Kong-Pang Pun, Oliver Chiu-sing Choy, Philip H. W. Leong |
DDECS | 1 |
| 2010 | Structured ASIC: Methodology and comparisonabstractAs fabrication process technology continues to advance, mask set costs have become prohibitively expensive. Structured ASICs can offer price and performance between ASICs and FPGAs. They are attractive for mid-volume production and offer good intellectual property security. In this paper, a structured ASIC methodology, where 2 metal- and 1 via-mask are customised, is described. The CAD tools are fully compatible with conventional ASIC design flows and a comparison of area and delay performance with ASICs and FPGAs is given. A prototype structured ASIC implementing an LED-backlit LCD controller was fabricated in a 0.13 μm CMOS process. It was verified and power consumption compared with an ASIC design. Sam M. H. Ho, Steve C. L. Yuen, Hiu Ching Poon, Thomas C. P. Chau, Yanqing Ai, Philip H. W. Leong, Oliver Chiu-sing Choy, Kong-Pang Pun |
FPT | 4 |
| 2009 | A comparison of via-programmable gate array logic cell circuitsabstractVia-programmable gate arrays (VPGAs) offer a middle ground between application specific integrated circuits and field programmable gate arrays in terms of flexibility, manufactuing cost, speed, power and area. In this paper, we present a novel VPGA logic cell, the complementary universal logic gate (CULG) which can be used to implement both sequential and combinatorial elements. Its performance is compared with a number of other designs including transmission gate, differential cascode voltage switch with pass gate, and standard cell. The CULG is found to have comparable power-delay product and process variation sensitivity to the other designs while offering the lowest power consumption. Thomas C. P. Chau, Philip H. W. Leong, Sam M. H. Ho, Brian P. W. Chan, Steve C. L. Yuen, Kong-Pang Pun, Oliver Chiu-sing Choy, Xinan Wang |
FPGA | 1 |
| 2009 | A detailed delay path model for FPGAsabstractA complete circuit-level description of a representative FPGA is presented in this paper, from which a simple RC delay model as a function of architectural and technology parameters is derived. Using this model, the expression for the optimal delay of any path through the FPGA can be formulated. We distill our model into being purely architecture dependent, and use it to capture new insight into how FPGA parameters can directly affect its delay. Several applications of this model are: (1) to gain better intuition of how architecture and process parameters affect the delay path in an FPGA, (2) for initial studies into new circuit designs and integrated circuit technologies, (3) in CAD tools for optimisation and sensitivity analysis. The technique described can be applied to arbitrary circuits, and simulations show that our closed form equations give delay values that are accurate to approximately 10% when compared to HSPICE simulation. Eddie Hung, Steve Wilton, Haile Yu, Thomas C. P. Chau, Philip H. W. Leong |
FPT | 4 |
| 2009 | Generation of Synthetic Floating-Point benchmark circuitsabstractSynthetic Floating-Point (SFP), a synthetic benchmark generator program for floating-point circuits is presented. SFP consists of two independent modules for characterisation and generation. The characterisation module extracts key dataflow statistics of an arbitrary software program. Generation involves producing randomised circuits with desired statistics which are either the output of the characterisation module or directly generated by the user. Using the basic linear algebra subprograms (BLAS) library, Whetstone benchmark and LINPACK benchmark, it is demonstrated that SFP can be used to generate floating-point benchmarks with different user-specified properties as well as benchmarks that mimic real computational programs. Thomas C. P. Chau, S. Man Ho Ho, Philip H. W. Leong, Peter Zipf, Manfred Glesner |
IPDPS | 1 |