VLDB 2026 Research / reviewers in the wild / expert
Ce Guo 0002
dblp:21/8636-2
· DBLP profile ↗
30ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0002-0272-9175ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 9 first-author · 14 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPAC: Automating FPGA-based Network Switches with Protocol Adaptive CustomizationabstractWith network requirements diverging across emerging applications, latency-critical services demand minimal logic delay, while hyperscale training and collectives require sustained line-rate throughput for synchronized bulk transfers. This divergence creates an urgent need for custom network switches tailored to specialized protocols and application-specific traffic patterns. This paper presents SPAC (Switch and Protocol Adaptive Customization), a novel approach that automates the generation of FPGA-based network switches co-optimized for custom protocols and application-specific traffic patterns. SPAC introduces a unified workflow with a domain-specific language (DSL) for protocol-architecture co-design, a library of modular HLS-based adaptive switch components, and a trace-aware Design Space Exploration (DSE) engine. By providing a multi-fidelity simulation stack, SPAC enables rapid identification of Pareto-optimal designs prior to deployment. We demonstrate the efficacy of the domain-specific adaptation of SPAC across a spectrum of real-world scenarios, spanning from latency-sensitive sensor and HFT networks to hyperscale datacenter fabrics. Experimental results show that by tailoring the micro-architecture and protocol to the specific workload, SPAC-generated designs reduce LUT and BRAM usage by 55% and 53%, respectively. Compared to fixed-architecture counterparts, SPAC delivers latency reductions ranging from 7.8% to 38.4% across various tasks while maintaining adequate resource consumption and packet drop rate.1 Lucas H. L. Ng, Alexander Charlton, Qianzhou Wang, Will Punter, Philippos Papaphilippou, Ce Guo 0002, Hongxiang Fan, Wayne Luk, Saman P. Amarasinghe, Ajay Brahmakshatriya |
FCCM | 8 |
| 2026 | EDSSC: An Efficient FPGA-based Accelerator for Dynamic Sparse Spectral ClusteringabstractThis paper presents a hardware-software co-design for efficient dynamic sparse spectral clustering. We introduce a heterogeneous streaming architecture that replaces dense SVD with a sparse iterative solver. Key contributions include: (1) dynamic graph generation to minimize memory footprint; (2) tile-centric dataflow to maximize sparse matrix reuse; and (3) an INT8 mixed-precision datapath. Implemented on Xilinx VCK190, our design achieves up to 22× speedup and 39× higher energy efficiency than an RTX 3090 GPU. Zhengyan Liu, Ce Guo 0002, Zehuan Zhang, Qiang Liu 0011, Wayne Luk |
FCCM | 2 |
| 2026 | CODESCA: Co-Design for Spectral Clustering AccelerationabstractAbstract: Spectral clustering is powerful but limited by O(N³) complexity. We present CODESCA, a co-design on Xilinx VCK190. By offloading sparse graph construction to the host and utilizing a quantization-aware block power iteration engine on FPGA, CODESCA achieves 22× speedup over CPU and 9× better throughput-per-watt than RTX 3080 GPU, enabling efficient edge data mining. Zhengyan Liu, Ce Guo 0002, Zehuan Zhang, Qiang Liu 0011, Wayne Luk |
FPGA | 2 |
| 2026 | Advancing Full-Stack Acceleration for SchröDinger-Style Quantum SimulationabstractRecent developments in quantum hardware, including the scaling of physical qubits and advanced quantum error correction techniques, have increased the number of reliable logical qubits. However, this progress has introduced new challenges for quantum algorithm developers. Limited access to physical quantum machines and the insufficient performance of classical quantum simulators for near-term scales ($\sim 30$logical qubits) hinder the simulation and validation of quantum algorithms. To address this urgent need for improving simulation performance, we propose a novel end-to-end full-stack solution for Schrödingerstyle simulation that jointly explores algorithm, software, and hardware optimizations. At the algorithmic level, by identifying the inefficiency in executing complex signed permutations and complex unitary permutation gates, we introduce index redirection and pre-compute merging that significantly reduce data movement and computational complexity. At the hardware level, we propose a reconfigurable dataflow architecture with adaptive memory scheduling and swapping optimizations. At the software level, an end-to-end toolchain is introduced to jointly explore both algorithmic and hardware optimizations. A comprehensive evaluation across a large suite of quantum circuits demonstrates that our work achieves a maximum speedup exceeding$50 \times$over the GPU-based Qiskit baseline. Shuang Liang 0012, Yuncheng Lu, Ce Guo 0002, Paul H. J. Kelly, Wayne Luk, Hongxiang Fan |
HPCA | 3 |
| 2026 | VCDF: A Validated Consensus-Driven Framework for Time Series Causal Discovery
Gene Yu, Ce Guo 0002, Wayne Luk |
PAKDD (3) | 2 |
| 2026 | MetaML-Pro: Cross-Stage Design Flow Automation for Efficient Deep Learning AccelerationabstractThis paper presents a unified framework for codifying and automating optimization strategies to efficiently deploy deep neural networks (DNNs) on resource-constrained hardware, such as FPGAs, while maintaining high performance, accuracy, and resource efficiency. Deploying DNNs on such platforms involves addressing the significant challenge of balancing performance, resource usage (e.g., DSPs and LUTs), and inference accuracy, which often requires extensive manual effort and domain expertise. Our novel approach addresses two core issues: (i) encoding custom optimization strategies and (ii) enabling cross-stage optimization search. In particular, our proposed framework seamlessly integrates programmatic DNN optimization techniques with high-level synthesis (HLS)-based metaprogramming, leveraging advanced design space exploration (DSE) strategies like Bayesian optimization to automate both top-down and bottom-up design flows. Hence, we reduce the need for manual intervention and domain expertise. In addition, the framework introduces customizable optimization, transformation, and control blocks to enhance DNN accelerator performance and resource efficiency. We further formalize a cross-stage, constrained Bayesian optimization procedure that couples predicate–action bottom-up feedback (via BRANCH) with FORK/REDUCE order search, enabling automated selection, ordering, and tuning across software and HLS tasks. Experimental results demonstrate up to a 92% DSP and 89% LUT usage reduction for select networks, while preserving accuracy, along with a 15.6-fold reduction in optimization time compared to grid search. These results highlight the potential for automating the generation of resource-efficient DNN accelerator designs with minimal effort, resulting in large resource savings with bounded exploration cost. Zhiqiang Que, Jose G. F. Coutinho, Ce Guo 0002, Hongxiang Fan, Wayne Luk |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2025 | Versatile Cross-platform Compilation Toolchain for Schrödinger-style Quantum Circuit SimulationabstractWhile existing quantum hardware resources have limited availability and reliability, there is a growing demand for exploring and verifying quantum algorithms. Efficient classical simulators for high-performance quantum simulation are critical to meeting this demand. However, due to the vastly varied characteristics of classical hardware, implementing hardware-specific optimizations for different hardware platforms is challenging. To address such needs, we propose CAST (Cross-platform Adaptive Schrödinger-style Simulation Toolchain), a novel compilation toolchain with cross-platform (CPU and Nvidia GPU) optimization and high-performance backend supports. CAST exploits a novel sparsity-aware gate fusion algorithm that automatically selects the best fusion strategy and backend configuration for targeted hardware platforms. CAST also aims to offer versatile and high-performance backend for different hardware platforms. To this end, CAST provides an LLVM IR-based vectorization optimization for various CPU architectures and instruction sets, and a PTX-based code generator for Nvidia GPU support. We benchmark CAST against IBM Qiskit, Google QSimCirq, Nvidia cuQuantum backend, and other high-performance simulators. On various 32-qubit CPU-based benchmarks, CAST achieves up to 8.03x speedup than Qiskit. On various 30-qubit GPU-based benchmarks, CAST achieves up to 39.3x speedup than Nvidia cuQuantum backend. Yuncheng Lu, Shuang Liang 0012, Hongxiang Fan, Ce Guo 0002, Wayne Luk, Paul H. J. Kelly |
DAC | 4 |
| 2025 | Aspo: Constraint-Aware Bayesian Optimization for Fpga-Based Soft ProcessorsabstractBayesian Optimization (BO) has shown promise in tuning processor design parameters. However, standard BO does not support constraints involving categorical parameters such as types of branch predictors and division circuits. In addition, optimization time of BO grows with processor complexity, which becomes increasingly significant especially for FPGA-based soft processors. This paper introduces ASPO, an approach that leverages disjunctive form to enable BO to handle constraints involving categorical parameters. Unlike existing methods that directly apply standard BO, the proposed ASPO method, for the first time, customizes the mathematical mechanism of BO to address challenges faced by soft-processor designs on FPGAs. Specifically, ASPO supports categorical parameters using a novel customized BO covariance kernel. It also accelerates the design evaluation procedure by penalizing the BO acquisition function with potential evaluation time and by reusing FPGA synthesis checkpoints from previously evaluated configurations. ASPO targets three soft processors: RocketChip, BOOM, and EL2 VeeR. The approach is evaluated based on seven RISC-V benchmarks. Results show that ASPO can reduce execution time for the “multiply” benchmark on the BOOM processor by up to 35 % compared to the default configuration. Furthermore, it reduces design time for the BOOM processor by up to 74 % compared to Boomerang, a state-of-the-art hardware-oriented BO approach. Ce Guo 0002, Wayne Luk, Robert Mullins 0001 |
FPL | 2 |
| 2024 | PCQ: Parallel Compact Quantum Circuit SimulationabstractSince quantum computers are not readily available, much quantum computing research such as quantum algorithm verification has to be conducted on classical computer platforms. While many quantum circuit simulators have been developed on CPUs and GPUs, the potential of FPGAs as a platform with parallel computing capabilities and high energy efficiency has not been fully explored. This paper describes a novel approach with two modes of data movement optimization for an FPGA-based parallel pipelined dataflow architecture targeting a compact computation format. A data decoupling method is adapted to partition computing tasks and data into non-interacting sub- sets, significantly reducing external data interaction overhead. The proposed approach shows significant promise in improving performance and energy efficiency compared with existing state vector based CPU, GPU, and FPGA implementations. Shuang Liang 0012, Yuncheng Lu, Ce Guo 0002, Wayne Luk, Paul H. J. Kelly |
FCCM | 3 |
| 2023 | Co-Design of Algorithm and FPGA Accelerator for Conditional Independence TestabstractConditional independence (CI) testing is a critical statistical method that determines conditional independence between variables using data. It is useful for various data mining applications, such as causal discovery, Bayesian inference, and agent-based model validation. However, the high volume of CI test queries and the large data sizes make CI testing computationally intensive. This paper proposes a hardware-oriented residual-based CI testing algorithm, co-designed with an FPGA accelerator, to address this issue. Our system accelerates CI tests by skipping least-squares computations algorithmically, enabling fixed-point operations in correlation evaluation and parallelization of permutation tests. Our experimental evaluation demonstrates that our method is as accurate as state-of-the-art CI testing approaches. Furthermore, our experimental implementation on an Intel Arria 10 FPGA delivers up to 32 times higher performance compared to state-of-the-art CI test tools running on eight Intel Xeon Silver 4110 CPU cores. Ce Guo 0002, Wayne Luk, Alexander Warren, Joshua M. Levine, Peter Brookes |
ASAP | 1 |
| 2023 | FPGA-Accelerated Causal Discovery with Conditional Independence Test PrioritizationabstractCausal discovery is a data mining approach that finds causal relations between variables from data. Causal discovery algorithms are computationally demanding when the data set has a high dimensionality or a large sample size. A promising way to expedite causal discovery is by utilizing FPGAs, but a significant drawback is that FPGA designs become inefficient when the on-chip memory cannot store the entire data set. This paper proposes Conditional Independence Test Prioritization (CITP), a novel approach that overcomes this limitation and enables fast FPGA-based causal discovery for large datasets with comparable speed and adequate accuracy to state-of-the-art methods. The main idea behind CITP is to design a workflow that allows a small subset of data to be stored in on-chip memory for prioritizing conditional independence tests. The paper provides experimental results that demonstrate the effectiveness of CITP in terms of both accuracy and speed. Our experiments show that for specific datasets, the proposed approach can respectively be 79 times, 2.6 times and 2.1 times faster than current CPU, GPU and FPGA designs. Ce Guo 0002, Diego Cupello, Wayne Luk, Joshua M. Levine, Alexander Warren, Peter Brookes |
FPL | 1 |
| 2023 | MetaML: Automating Customizable Cross-Stage Design-Flow for Deep Learning AccelerationabstractThis paper introduces a novel optimization framework for deep neural network (DNN) hardware accelerators, enabling the rapid development of customized and automated design flows. More specifically, our approach aims to automate the selection and configuration of low-level optimization techniques, encompassing DNN and FPGA low-level optimizations. We introduce novel optimization and transformation tasks for building design-flow architectures, which are highly customizable and flexible, thereby enhancing the performance and efficiency of DNN accelerators. Our results demonstrate considerable reductions of up to 92% in DSP usage and 89% in LUT usage for two networks, while maintaining accuracy and eliminating the need for human effort or domain expertise. In comparison to state-of-the-art approaches, our design achieves higher accuracy and utilizes three times fewer DSP resources, underscoring the advantages of our proposed framework. Zhiqiang Que, Markus Rognlien, Ce Guo 0002, José Gabriel F. Coutinho, Wayne Luk |
FPL | 4 |
| 2022 | Optimizing quantum circuit placement via machine learningabstractQuantum circuit placement (QCP) is the process of mapping the synthesized logical quantum programs on physical quantum machines, which introduces additional SWAP gates and affects the performance of quantum circuits. Nevertheless, determining the minimal number of SWAP gates has been demonstrated to be an NP-complete problem. Various heuristic approaches have been proposed to address QCP, but they suffer from suboptimality due to the lack of exploration. Although exact approaches can achieve higher optimality, they are not scalable for large quantum circuits due to the massive design space and expensive runtime. By formulating QCP as a bilevel optimization problem, this paper proposes a novel machine learning (ML)-based framework to tackle this challenge. To address the lower-level combinatorial optimization problem, we adopt a policy-based deep reinforcement learning (DRL) algorithm with knowledge transfer to enable the generalization ability of our framework. An evolutionary algorithm is then deployed to solve the upper-level discrete search problem, which optimizes the initial mapping with a lower SWAP cost. The proposed ML-based approach provides a new paradigm to overcome the drawbacks in both traditional heuristic and exact approaches while enabling the exploration of optimality-runtime trade-off. Compared with the leading heuristic approaches, our ML-based method significantly reduces the SWAP cost by up to 100%. In comparison with the leading exact search, our proposed algorithm achieves the same level of optimality while reducing the runtime cost by up to 40 times. Hongxiang Fan, Ce Guo 0002, Wayne Luk |
DAC | 2 |
| 2022 | Accelerating Constraint-Based Causal Discovery by Shifting Speed BottleneckabstractCausal discovery is a technique to find the causal relationship between variables using data. This technique has many applications in data mining and knowledge discovery. However, the high data dimensionality results in a significant computational efficiency problem. A common speed bottleneck in conventional causal discovery methods is the execution of conditional independence (CI) tests. This paper proposes, analyzes, and evaluates a novel acceleration strategy for causal discovery, which has low communication costs and can effectively exploit FPGA on-chip memory and parallelism. First, we propose an algorithmic method to shift the speed bottleneck from CI test execution to CI test generation. Second, we design a hardware accelerator for CI test generation on FPGAs. Third, we evaluate the proposed approach by comparing the accuracy-speed trade-off against four state-of-the-art accelerated causal discovery tools on CPUs and GPUs. Our accelerated implementation running on an Intel Arria 10 GX FPGA shows a superior accuracy-speed trade-off in 12 causal discovery problems. The implementation achieves up to 8.8 times speedup over the cuPC software running on an NVIDIA GeForce RTX 2080 Ti GPU. It also achieves up to 155.7 times speedup over the stable.fast software running on an Intel Xeon Silver 4110 octa-core CPU. To the best of our knowledge, the proposed approach is the first FPGA-based acceleration approach for constraint-based causal discovery. Ce Guo 0002, Wayne Luk |
FPGA | 1 |
| 2021 | Analytical Performance Estimation for Large-Scale Reconfigurable Dataflow PlatformsabstractNext-generation high-performance computing platforms will handle extreme data- and compute-intensive problems that are intractable with today’s technology. A promising path in achieving the next leap in high-performance computing is to embrace heterogeneity and specialised computing in the form of reconfigurable accelerators such as FPGAs, which have been shown to speed up compute-intensive tasks with reduced power consumption. However, assessing the feasibility of large-scale heterogeneous systems requires fast and accurate performance prediction. This article proposes Performance Estimation for Reconfigurable Kernels and Systems (PERKS), a novel performance estimation framework for reconfigurable dataflow platforms. PERKS makes use of an analytical model with machine and application parameters for predicting the performance of multi-accelerator systems and detecting their bottlenecks. Model calibration is automatic, making the model flexible and usable for different machine configurations and applications, including hypothetical ones. Our experimental results show that PERKS can predict the performance of current workloads on reconfigurable dataflow platforms with an accuracy above 91%. The results also illustrate how the modelling scales to large workloads, and how performance impact of architectural features can be estimated in seconds. Ryota Yasudo, José Gabriel F. Coutinho, Ana Lucia Varbanescu, Wayne Luk, Hideharu Amano, Tobias Becker, Ce Guo 0002 |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2020 | Accelerating Simulation for Agent-based Epidemic Models using FPGAsabstractAgent-based models (ABMs) play a critical role in the containment and mitigation of epidemics. An ABM for epidemics involves a population of agents that represent healthy and infected individuals. Users can simulate the interactions of the agents to analyse spreading patterns, make predictions and conduct what-if tests. A major drawback that limits the application of ABMs is the long simulation time for large populations. This paper proposes new techniques to speed up ABM-based simulation. Specifically, we present an agent-based model for epidemic spreading that enables effective acceleration using reconfigurable hardware. The key idea is to compute the infection probability in a novel way so that the amount of on-chip memory usage is independent of the population size. Also, we propose a parameter pre-computation algorithm, a parallel simulation algorithm and an efficient collaboration method between the host computer and the accelerator. We simulate the proposed model on an Intel Arria 10 FPGA and compare it with a software reference that simulates a conventional ABM on an Intel Core i5-9400F CPU with six cores. The two systems produce similar results, but the FPGA-based system achieves 14 times speedup compared to the software reference. Ce Guo 0002, Wayne Luk, Stephen Weston |
AICCSA | 1 |
| 2020 | Fast and Accurate Training of Ensemble Models with FPGA-based SwitchabstractRandom projection is gaining more attention in large scale machine learning. It has been proved to reduce the dimensionality of a set of data whilst approximately preserving the pairwise distance between points by multiplying the original dataset with a chosen matrix. However, projecting data to a lower dimension subspace typically reduces the training accuracy. In this paper, we propose a novel architecture that combines an FPGA-based switch with the ensemble learning method. This architecture enables reducing training time while maintaining high accuracy. Our initial result shows a speedup of 2.12-6.77 times using four different high dimensionality datasets. Jiuxi Meng, Ce Guo 0002, Nadeen Gebara, Wayne Luk |
ASAP | 2 |
| 2019 | Customisable Control Policy Learning for RoboticsabstractDeep reinforcement learning algorithms integrate deep neural networks with traditional reinforcement learning methodologies. These techniques have been developed and used for various applications to produce exciting results in many fields, including robotics. However, physical robots require a large amount of training episodes which can damage the robot if directed by immature policies. Training using simulations can serve as a viable alternative before a robot is deployed in the field. This study addresses a computational challenge of deep reinforcement learning by developing a hardware architecture for the Deep Deterministic Policy Gradient (DDPG) algorithm. Additionally, we identify the customisation opportunities for a full-stack development framework with reinforcement learning to discover control policies for robotic arms. Finally, we transfer policies encoded in fixed-point numbers from our FPGA DDPG implementation to a robotic arm to evaluate the feasibility of our learning platform. Ce Guo 0002, Wayne Luk, Stanley Qing Shui Loh, Alexander Warren, Joshua M. Levine |
ASAP | 1 |
| 2019 | Towards Efficient Deep Neural Network Training by FPGA-Based Batch-Level ParallelismabstractTraining Deep Neural Networks (DNNs) requires a significant amount of time and resources to obtain acceptable results, which severely limits its deployment in resource-limited platforms. This paper proposes DarkFPGA, a novel customizable framework to efficiently accelerate the entire DNN training on a single FPGA platform. First, we explore batch-level parallelism to enable efficient training on FPGAs. Second, we devise a novel hardware architecture optimised by a batch-oriented data pattern and tiling techniques to effectively exploit parallelism. Moreover, an analytical model is developed to determine the optimal design parameters for the DarkFPGA accelerator with respect to a specific network specification and FPGA resource constraints. Our results show that the accelerator is able to perform about 11 times faster than CPU training and about a third of the energy consumption than GPU training using 8-bit integers for training VGG-like networks on the CIFAR dataset for the Maxeler MAX5 platform. Man-Kit Sit, Hongxiang Fan, Shuanglong Liu, Wayne Luk, Ce Guo 0002 |
FCCM | 6 |
| 2017 | A Fully-Pipelined Hardware Design for Gaussian Mixture ModelsabstractGaussian Mixture Models (GMMs) are widely used in many applications such as data mining, signal processing and computer vision, for probability density modeling and soft clustering. However, the parameters of a GMM need to be estimated from data by, for example, the Expectation-Maximization algorithm for Gaussian Mixture Models (EM-GMM), which is computationally demanding. This paper presents a novel design for the EM-GMM algorithm targeting reconfigurable platforms, with five main contributions. First, a pipeline-friendly EM-GMM with diagonal covariance matrices that can easily be mapped to hardware architectures. Second, a function evaluation unit for Gaussian probability density based on fixed-point arithmetic. Third, our approach is extended to support a wide range of dimensions or/and components by fitting multiple pieces of smaller dimensions onto an FPGA chip. Fourth, we derive a cost and performance model that estimates logic resources. Fifth, our dataflow design targeting the Maxeler MPCX2000 with a Stratix-5SGSD8 FPGA can run over 200 times faster than a 6-core Xeon E5645 processor, and over 39 times faster than a Pascal TITAN-X GPU. Our design provides a practical solution to applications for training and explores better parameters for GMMs with hundreds of millions of high dimensional input instances, for low-latency and high-performance applications. Conghui He, Haohuan Fu, Ce Guo 0002, Wayne Luk, Guangwen Yang 0002 |
IEEE Trans. Computers | 3 |
| 2015 | Pipelined Genetic PropagationabstractGenetic Algorithms (GAs) are a class of numerical and combinatorial optimisers which are especially useful for solving complex non-linear and non-convex problems. However, the required execution time often limits their application to small-scale or latency-insensitive problems, so techniques to increase the computational efficiency of GAs are needed. FPGA-based acceleration has significant potential for speeding up genetic algorithms, but existing FPGA GAs are limited by the generational approaches inherited from software GAs. Many parts of the generational approach do not map well to hardware, such as the large shared population memory and intrinsic loop-carried dependency. To address this problem, this paper proposes a new hardware-oriented approach to GAs, called Pipelined Genetic Propagation (PGP), which is intrinsically distributed and pipelined. PGP represents a GA solver as a graph of loosely coupled genetic operators, which allows the solution to be scaled to the available resources, and also to dynamically change topology at run-time to explore different solution strategies. Experiments show that pipelined genetic propagation is effective in solving seven different applications. Our PGP design is 5 times faster than a recent FPGA-based GA system, and 90 times faster than a CPU-based GA system. Liucheng Guo, Ce Guo 0002, David B. Thomas, Wayne Luk |
FCCM | 2 |
| 2015 | Recursive pipelined genetic propagation for bilevel optimisationabstractThe bilevel optimisation problem (BLP) is a subclass of optimisation problems in which one of the constraints of an optimisation problem is another optimisation problem. BLP is widely used to model hierarchical decision making where the leader and the follower correspond to the upper level and lower level optimisation problem, respectively. In BLP, the optimal solutions to the lower level optimisation problem are the feasible solutions to the upper level problem, which makes it particularly difficult to solve. This paper proposes a novel hardware architecture known as Recursive Pipelined Genetic Propagation (RPGP), to solve BLP efficiently on FPGA. RPGP features a graph of genetic operation nodes which can be scaled to exploit hardware resources. In addition, the topology of the RPGP graph can be changed at run-time to escape from local optima. We evaluate the proposed architecture on an Altera Stratix-V FPGA, using a benchmark bilevel optimisation problem set. Our experimental results show that RPGP can achieve a significant speed-up against previous work. Shengjia Shao, Liucheng Guo, Ce Guo 0002, Thomas C. P. Chau, David B. Thomas, Wayne Luk, Stephen Weston |
FPL | 3 |
| 2014 | Pipelined reconfigurable accelerator for ordinal pattern encodingabstractOrdinal analysis is a statistical method for analysing the complexity of time series. This method has been used in characterising dynamic changes in time series, with various applications such as financial risk modelling and biomedical signal processing. Ordinal pattern encoding is a fundamental calculation in ordinal analysis. It is computationally demanding particularly for high query orders and large time series data. This paper presents the first reconfigurable accelerator for this encoding calculation, with four main contributions. First, we propose a two-level hardware-oriented ordinal pattern encoding scheme to avoid sequence sorting operations in the accelerator, enabling theoretically best code compactness. Second, we develop a hardware mapping method by promoting data reuse, by parallelising arithmetic operations, and by pipelining the data path. Third, we conduct an experimental implementation of the proposed system, showing promising accelerated performance compared to software solutions. Finally, we apply the proposed accelerator to the computation of permutation entropy, demonstrating the significant potential for acceleration that would benefit such computation. Ce Guo 0002, Wayne Luk, Stephen Weston |
ASAP | 1 |
| 2014 | Accelerating parameter estimation for multivariate self-exciting point processesabstractSelf-exciting point processes are stochastic processes capturing occurrence patterns of random events. They offer powerful tools to describe and predict temporal distributions of random events like stock trading and neurone spiking. A critical calculation in self-exciting point process models is parameter estimation, which fits a model to a data set. This calculation is computationally demanding when the number of data points is large and when the data dimension is high. This paper proposes the first reconfigurable computing solution to accelerate this calculation. We derive an acceleration strategy in a mathematical specification by eliminating complex data dependency, by cutting hardware resource requirement, and by parallelising arithmetic operations. In our experimental evaluation, an FPGA-based implementation of the proposed solution is up to 79 times faster than one CPU core, and 13 times faster than the same CPU with eight cores. Ce Guo 0002, Wayne Luk |
FPGA | 1 |
| 2014 | Automated framework for FPGA-based parallel genetic algorithmsabstractParallel genetic algorithms (pGAs) are a variant of genetic algorithms which can promise substantial gains in both efficiency of execution and quality of results. pGAs have attracted researchers to implement them in FPGAs, but the implementation always needs large human effort. To simplify the implementation process and make the hardware pGA designs accessible to potential non-expert users, this paper proposes a general-purpose framework, which takes in a high-level description of the optimisation target and automatically generates pGA designs for FPGAs. Our pGA system exploits the two levels of parallelism found in GA instances and genetic operations, allowing users to tailor the architecture for resource constraints at compile-time. The framework also enables users to tune a subset of parameters at run-time without time-consuming recompilation. Our pGA design is more flexible than previous ones, and has an average speedup of 26 times compared to the multi-core counterparts over five combinatorial and numerical optimisation problems. When compared with a GPU, it also shows a 6.8 times speedup over a combinatorial application. Liucheng Guo, David B. Thomas, Ce Guo 0002, Wayne Luk |
FPL | 3 |
| 2014 | Accelerating transfer entropy computationabstractTransfer entropy is a measure of information transfer between two time series. It is an asymmetric measure based on entropy change which only takes into account the statistical dependency originating in the source series, but excludes dependency on a common external factor. Transfer entropy is able to capture system dynamics that traditional measures cannot, and has been successfully applied to various areas such as neuroscience, bioinformatics, data mining and finance. When time series becomes longer and resolution becomes higher, computing transfer entropy is demanding. This paper presents the first reconfigurable computing solution to accelerate transfer entropy computation. The novel aspects of our approach include a new technique based on Laplace's Rule of Succession for probability estimation; a novel architecture with optimised memory allocation, bit-width narrowing and mixed-precision optimisation; and its implementation targeting a Xilinx Virtex-6 SX475T FPGA. In our experiments, the proposed FPGA-based solution is up to 111.47 times faster than one Xeon CPU core, and 18.69 times faster than a 6-core Xeon CPU. Shengjia Shao, Ce Guo 0002, Wayne Luk, Stephen Weston |
FPT | 2 |
| 2014 | Collaborative processing of Least-Square Monte Carlo for American optionsabstractAmerican options are popularly traded in the financial market, so pricing those options becomes crucial in practice. In reality, many popular pricing models do not have analytical solutions. Hence techniques such as Monte Carlo are often used in practice. This paper presents a CPU-FPGA collaborative accelerator using state-of-the-art Least-Square Monte Carlo method, for pricing American options. We provide a new sequence of generating the Monte Carlo paths, and a precalculation strategy for the regression process. Our design is customisable for different pricing models, discretisation schemes, and regression functions. The Heston model is used as a case study for evaluating our strategy. Experimental results show that an FPGA-based solution could provide 22 to 64.5 times faster than a single-core CPU implementation. Jinzhe Yang, Ce Guo 0002, Wayne Luk, Terence Nahar |
FPT | 2 |
| 2013 | Accelerating HAC estimation for multivariate time seriesabstractHeteroskedasticity and autocorrelation consistent (HAC) covariance matrix estimation, or HAC estimation in short, is one of the most important techniques in time series analysis and forecasting. It serves as a powerful analytical tool for hypothesis testing and model verification. However, HAC estimation for long and high-dimensional time series is computationally expensive. This paper describes a novel pipeline-friendly HAC estimation algorithm derived from a mathematical specification, by applying transformations to eliminate conditionals, to parallelize arithmetic, and to promote data reuse in computation. We then develop a fully-pipelined hardware architecture based on the proposed algorithm. This architecture is shown to be efficient and scalable from both theoretical and empirical perspectives. Experimental results show that an FPGA-based implementation of the proposed architecture is up to 111 times faster than an optimised CPU implementation with one core, and 14 times faster than a CPU with eight cores. Ce Guo 0002, Wayne Luk |
ASAP | 1 |
| 2013 | Accelerating maximum likelihood estimation for Hawkes point processesabstractHawkes processes are point processes that can be used to build probabilistic models to describe and predict occurrence patterns of random events. They are widely used in high-frequency trading, seismic analysis and neuroscience. A critical numerical calculation in Hawkes process models is parameter estimation, which is used to fit a Hawkes process model to a data set. The parameter estimation problem can be solved by searching for a parameter set that maximises the log-likelihood. A core operation of this search process, the log-likelihood evaluation, is computationally demanding if the number of data points is large. To accelerate the computation, we present a log-likelihood evaluation strategy which is suitable for hardware acceleration. We then design and optimise a pipelined engine based on our proposed strategy. In the experiments, an FPGA-based implementation of the proposed engine is shown to be up to 72 times faster than a single-core CPU, and 10 times faster than an 8-core CPU. Ce Guo 0002, Wayne Luk |
FPL | 1 |
| 2012 | A fully-pipelined expectation-maximization engine for Gaussian Mixture ModelsabstractGaussian Mixture Models (GMMs) are powerful tools for probability density modeling and soft clustering. They are widely used in data mining, signal processing and computer vision. In many applications, we need to estimate the parameters of a GMM from data before working with it. This task can be handled by the Expectation-Maximization algorithm for Gaussian Mixture Models (EM-GMM), which is computationally demanding. In this paper we present our FPGA-based solution for the EM-GMM algorithm. We propose a pipeline-friendly EM-GMM algorithm, a variant of the original EM-GMM algorithm that can be converted to a fully-pipelined hardware architecture. To further improve the performance, we design a Gaussian probability density function evaluation unit that works with fixed-point arithmetic. In the experiments, our FPGA-based solution generates fairly accurate results while achieving a maximum of 517 times speedup over a CPU-based solution, and 28 times speedup over a GPU-based solution. Ce Guo 0002, Haohuan Fu, Wayne Luk |
FPT | 1 |