Terrence S. T. Mak

dblp:46/5280 · also Sui-Tung Mak · DBLP profile ↗
← Back
77ranked-venue papers
11as first author
3since 2021 · last 2022
0000-0003-1945-8292ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 68 · 11 first-author · 3 since 2021Software engineering, systems software and programming languages · 11Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Computer networks · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2022 Thermal and Performance Efficient On-Chip Surface-Wave Communication for Many-Core Systems in Dark Silicon Era
abstract
Due to the exceedingly high integration density of VLSI circuits and the resulting high power density, thermal integrity became a major challenge. One way to tackle this problem is Dark silicon. Dark silicon is the amount of circuitry in a chip that is forced to switch off to insure thermal integrity of the system and prevent permanent thermal-related faults. In many-core systems, the presence of Dark Silicon adds new design constraints, in general, and on the communication fabric of such systems, in particular. This is due to the fact that system-level thermal-management systems tend to increase the distance between high activity cores to insure better thermal balancing and integrity. Consequently, a designing dilemma is created where a compromise has to be made between interconnect performance and power consumption. This study proposes a hybrid wire and surface-wave interconnect (SWI) based Network-on-Chip (NoC) to address the dark silicon challenge. Through efficient utilization of one-hop cross the chip communication SWI links, the proposed architecture is able to offer an efficient and scalable communication platform in terms of performance, power, and thermal impact. As a result, evaluations of the proposed architecture compared to baseline architecture under dark silicon scenarios show reduction in maximum temperature by 15∘C, average delay up to 73.1%, and energy-saving up to ∼3X. This study explores the promising potential of the proposed architecture in extending the utilization wall for current and future many-core systems in dark silicon era.
Ammar Karkar, Nizar Dahir, Terrence S. T. Mak, Kin-Fai Tong
ACM J. Emerg. Technol. Comput. Syst.3
2021 Power density aware application mapping in mesh-based network-on-chip architecture: An evolutionary multi-objective approach
Nizar Dahir, Ammar Karkar, Maurizio Palesi, Terrence S. T. Mak, Alexandre Yakovlev
Integr.4
2021 On Performance Optimization and Quality Control for Approximate-Communication-Enabled Networks-on-Chip
abstract
For many applications showing error forgiveness, approximate computing is a new design paradigm that trades application output accuracy for mitigating computation/communication effort, which results in performance/energy benefit. Since networks-on-chip (NoCs) are one of the major contributors to system performance and power consumption, the underlying communication is approximated to achieve time/energy improvement. However, performing approximation blindly causes unacceptable quality loss. In this article, first, an optimization problem to maximize NoC performance is formulated with the constraint of application quality requirement, and the application quality loss is studied. Second, a congestion-aware quality control method is proposed to improve system performance by aggressively dropping network data, which is based on flow prediction and a lightweight heuristic. In the experiments, two recent approximation methods for NoCs are augmented with our proposed control method to compare with their original ones. Experimental results show that our proposed method can speed up execution by as much as 29.42% over the two state-of-the-art works.
Siyuan Xiao, Xiaohang Wang 0001, Maurizio Palesi, Amit Kumar Singh 0002, Liang Wang 0020, Terrence S. T. Mak
IEEE Trans. Computers6
2020 On hardware-trojan-assisted power budgeting system attack targeting many core systems
Xiaohang Wang 0001, Yingtao Jiang, Liang Wang 0020, Mei Yang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
J. Syst. Archit.7
2020 An Active Silicon Interposer With Low-Power Hybrid Wireless-Wired Clock Distribution Network for Many-Core Systems
abstract
Due to the increasing interconnect delay caused by shrinking wiring dimensions, modern synchronous many-core systems are now facing critical issues. Particularly, the power budget to propagate high-frequency clock signals across the chip is limited. It becomes more challenging using conventional metallic interconnects to deliver a clock with low uncertainties across active dies. This article proposes a novel hybrid wireless-wired clock distribution network, which improves the performance of ON-chip clock distribution significantly. By using embedded wireless clock transmitter and receiver designs, because of the high fan-out feature of the wireless clock transmission, the overall clock delay, skew, and power have been reduced. The proposed hybrid clock distribution scheme is verified through a novel test circuit by using Arm Mali-G77 GPU as an example. Experimental results indicate that the proposed clock distribution network exhibits a significant global delay reduction of up to 28.8%. Also, for the best case scenario, a maximum of 46.7% and 17.7% reduction in clock skew and power consumption, are identified, respectively. Thus, our proposed approach offers a promising solution to clock distribution for many-core integrated circuits, especially for high-performance systems.
Graham Knight, Terrence S. T. Mak
IEEE Trans. Very Large Scale Integr. Syst.3
2019 CoDAPT: A Concurrent Data And Power Transceiver for Fully Wireless 3D-ICs
abstract
Three dimensional system integration is a promising enabling technology for realising heterogeneous ICs, facilitating stacking of disparate elements such as MEMS, sensors, analogue components, memories and digital processing. Recently, research has looked to contactless 3D integration using inductive coupling links (ICLs) to provide a low-cost alternative to conventional contact-based approaches (e.g. through silicon vias) for 3D integration. In this paper, we present a novel, fully wireless, ICL architecture for Concurrent Data and Power Transfer (CoDAPT) between tiers of a 3D-IC. The proposed CoDAPT architecture uses only a single inductor for simultaneous power transmission and data communication, resulting in high area efficiency, whilst facilitating low-cost, straightforward die stacking. The proposed design is experimentally validated through full wave EM and SPICE simulation and demonstrates capability to communicate data vertically at a rate of 1.3Gbps/channel (utilising an area of only 0.052mm2) whilst simultaneously achieving power delivery of 0.83mW, under standard operating conditions. A case study is also presented, demonstrating that CoDAPT achieves an area reduction greater than 1.7× when compared with existing works, representing an important progression towards ultra low-cost 3D-ICs through fully wireless stacking.
Benjamin J. Fletcher, Shidhartha Das, Terrence S. T. Mak
DATE3
2019 ACDC: An Accuracy- and Congestion-aware Dynamic Traffic Control Method for Networks-on-Chip
abstract
Many applications exhibit error forgiving features. For these applications, approximate computing provides the opportunity of accelerating the execution time or reducing power consumption, by mitigating computation effort to get an approximate result. Among the components on a chip, network-on-chip (NoC) contributes a large portion to system power and performance. In this paper, we exploit the opportunity of aggressively reducing network congestion and latency by selectively dropping data. Essentially, the importance of the dropped data is measured based on a quality model. An optimization problem is formulated to minimize the network congestion with constraint of the result quality. A lightweight online algorithm is proposed to solve this problem. Experiments show that on average, our proposed method can reduce the execution time by as much as 12.87% and energy consumption by 12.42% under strict quality requirement, speed up execution by 19.59% and reduce energy consumption by 21.20% under relaxed requirement, compared to a recent work on approximate computing approach for NoCs.
Siyuan Xiao, Xiaohang Wang 0001, Maurizio Palesi, Amit Kumar Singh 0002, Terrence S. T. Mak
DATE5
2019 Type-Spread Molecular Communications: Principles and Inter-Symbol Interference Mitigation
abstract
Diffusion-based Molecular Communication (DMC) is a feasible method for information transmission in some nano-networks operated in gas or liquid environments. In this paper, we first propose an information modulation scheme for DMC, which is referred to as the Type-Spread Molecular Shift Keying (TS-MoSK). Considering that DMC signals usually experience severe inter-symbol interference (ISI), our TS-MoSK is characterized by introducing extra types of molecules for ISI mitigation (ISIM). Furthermore, we introduce two ISIM methods to the TS-MoSK modulated DMC systems, which are the active ISIM and passive ISIM. We detail their operation principles, and investigate as well as compare their achievable performance. Our studies show that, aided by the extra types of molecules, TS-MoSK outperforms the MoSK without spreading. Both ISIM approaches are effective for further improving the performance of TS-MoSK.
Weidong Gao 0004, Terrence S. T. Mak, Lie-Liang Yang
ICC2
2019 A Low-Energy Inductive Transceiver using Spike-Latency Encoding for Wireless 3D Integration
abstract
Recently, the use of wireless (or contactless) 3D integration has been proposed as a low-cost method of stacking disparate processing and sensor dies into singular, small form-factor ICs. Whilst such devices would be ideally suited for the Internet of Things (IoT), in the IoT, maintaining low-power consumption is of paramount importance. Contactless intertier links use significant energy when forming a magnetic field which can penetrate multiple silicon dies, and hence are often criticised for their poor power efficiency when compared to wired alternatives such as through silicon vias (TSVs). To address this, in this paper we present a novel, neuro-inspired, inductive transceiver (for transmitting data between tiers of a 3D-IC) that maintains low power consumption by encoding frames of data in terms of the latency between pulses, thereby reducing the number of transmit pulses and energy required per bit. The proposed approach is validated using commercial electromagnetic and electrical circuit simulators in 65nm CMOS technology. Results demonstrate an energy consumption of 0.79pJ/bit, representing a reduction of 31% when compared to existing state-of-the-art transceivers, or an increased communication distance of up to 1.8× for the same energy budget.
Benjamin J. Fletcher, Shidhartha Das, Terrence S. T. Mak
ISLPED3
2019 A Lifetime Reliability-Constrained Runtime Mapping for Throughput Optimization in Many-Core Systems
abstract
Due to technology scaling, lifetime reliability is becoming one of the major design constraints in the performance optimization of future many-core systems. Given a lifetime reliability constraint, the existing lifetime-constrained runtime mapping schemes often lead to low throughput because of the requirement to map all applications to compact regions. In this paper, we propose a runtime application mapping scheme that exploits a borrowing strategy to improve the throughput of many-core systems given a lifetime constraint. First, we propose using different strategies for mapping communication-intensive applications and computation-intensive applications. The lifetime reliability constraint can be relaxed in the local time scale when the communication requirement is high. The throughput is improved because the communication distance of communication-intensive applications is optimized while the waiting time of computation-intensive application is reduced. Then, we propose a method to effectively classify applications depending on the communication-to-computation ratio. A dynamic threshold is determined according to the current locations of available cores. Finally, we propose an improved neighborhood allocation scheme to reduce the communication cost in the task mapping. The experimental results show that compared to the state-of-the-art lifetime-constrained mapping, the proposed mapping scheme improves the throughput of many-core systems by 26% on average for synthetic task graphs and by 20% on average for realistic task graphs while the lifetime reliability is maintained within a constraint.
Liang Wang 0020, Ping Lv, Leibo Liu, Jie Han 0001, Ho-fung Leung, Xiaohang Wang 0001, Shouyi Yin, Shaojun Wei, Terrence S. T. Mak
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2019 A Non-Minimal Routing Algorithm for Aging Mitigation in 2D-Mesh NoCs
abstract
Due to technology scaling, aging issue is becoming one of major concerns in the design of network-on-chip (NoC). The imbalanced workload distribution and routing algorithm cause aging hotspots, where a certain group of routers have higher aging effect than others. This can possibly lead to shorter lifetime of NoC. Most existing aging-aware routing algorithms are based on minimal routing, which suffers from less degree of adaptiveness compared to non-minimal routing. Thus, they are inefficient to mitigate the aging effect of routers. In this paper, we propose to use a non-minimal routing scheme to detour the traffic away from the aging hotspots, with the objective of mitigating the aging effect for NoCs. The problem is formulated as a bottleneck shortest path problem and solved using a dynamic programming approach. Finally, the experimental results show that compared to the state-of-the-art aging-aware routing algorithm, the non-minimal routing algorithm has up to 20% lifetime improvement for hotspot traffic patterns and realistic workload traces.
Liang Wang 0020, Xiaohang Wang 0001, Ho-fung Leung, Terrence S. T. Mak
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 On Runtime Communication and Thermal-Aware Application Mapping and Defragmentation in 3D NoC Systems
abstract
Many-core systems connected by 3D Networks-on-Chip (NoC) are emerging as a promising computation engine for systems like cloud computing servers, big data systems, etc. Mapping applications at runtime to 3D NoCs is the key to maintain high throughput of the overall chip under a thermal/power constraint. However, the goals of optimizing both the communication latency and chip peak temperature are contradicting due to several reasons. First, exploiting the vertical TSV links can accelerate communications, while low peak temperature prefers that the tasks to be mapped closer to the heat sink, instead of using the vertical links. Second, mapping tasks in close proximity can reduce communication latency, but at the cost of poor heat dissipation. To address these issues, in this paper, we propose an efficient runtime mapping algorithm to reduce both communication latency and overall application running time under thermal constraint. In essence, this algorithm first selects a 3D cuboid core region of a specific shape for each incoming application by setting the region's number of occupied vertical layers and its distance to the heat sink, in order to optimize its communication performance and peak temperature. Next, the exact locations of the core regions in the chip are determined, followed by a task-to-core mapping. A defragmentation algorithm is also proposed to keep free core regions contiguous. The experimental results have confirmed that, compared to two recently proposed runtime mapping algorithms, our proposed approach can reduce the total running time by up to 48% and communication cost by up to 44%, with a low runtime overhead.
Xiaohang Wang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
IEEE Trans. Parallel Distributed Syst.4
2019 Design and Optimization of Inductive-Coupling Links for 3-D-ICs
abstract
Recent research in the field of 3-D system integration has looked to the use of inductive-coupling links (ICLs) to provide vertical connectivity without incurring the inflated fabrication and testing costs associated with through-silicon vias. For power-efficient ICL design, optimization of the utilized physical inductor geometries is essential, but currently must be performed manually in a process that can take several hours. As a result, the generation of optimized inductor designs poses a significant challenge. In this paper, we address this challenge in three main contributions: 1) a novel, nonuniform planar inductor layout that exhibits enhanced performance when compared with conventional uniform inductors; 2) a rapid solver for evaluating inductor layouts; and 3) a high-speed optimization algorithm for determining best performing coil pairs. These three contributions are combined as a CAD tool for optimization of ICLs for 3-D-ICs (COIL-3-D). Results demonstrate that COIL-3-D achieves an average accuracy within 7.8% of finite-element tools consuming a small fraction of the time (1.5 × 10-3%), significantly ameliorating the design of ICL-based 3-D-ICs. We also demonstrate that using COIL-3-D to optimize ICL inductor layouts can yield significant performance (up to 41.5% bandwidth improvement) and power (up to 8.1% power improvement) benefits, when compared with layouts used in prior ICL implementations. For these reasons, this paper unlocks new potential for low-cost, power-efficient 3-D integration using ICLs.
Benjamin J. Fletcher, Shidhartha Das, Terrence S. T. Mak
IEEE Trans. Very Large Scale Integr. Syst.3
2018 A high-speed design methodology for inductive coupling links in 3D-ICs
abstract
Inductive coupling links (ICLs) are gaining traction as an alternative to through silicon vias (TSVs) for 3D integration, promising high-bandwidth connectivity without the inflated fabrication costs associated with TSV-enabled processes. For power-efficient ICL design, optimisation of the utilised physical inductor geometries is essential, however typically necessitates the use of finite element analysis (FEA) in addition to manual parameter fitting, a process that can take several hours even for a single geometry. As a result, the generation of optimised inductor designs poses a significant challenge. In this paper, we address this challenge, presenting a CAD-tool for Optimisation of Inductive coupling Links for 3D-ICs (COIL-3D1). COIL-3D uses a rapid solver based upon semi-empirical expressions to quickly and accurately characterise a given link, in conjunction with a high-speed refined optimisation flow to find optimal inductor geometries for use in ICL-based 3D-ICs. The proposed solver achieves an average accuracy within 9.1% of commercial FEA software tools, and the proposed optimisation flow reduces the search time by 26 orders of magnitude. This work unlocks new potential for power-efficient 3D integration using inductive coupling links.
Benjamin J. Fletcher, Shidhartha Das, Terrence S. T. Mak
DATE3
2018 Low-power 3D integration using inductive coupling links for neurotechnology applications
abstract
Three dimensional system integration offers the ability to stack multiple dies, fabricated in disparate technologies, within a single IC. For this reason, it is gaining popularity for use in sensor devices which perform concurrent analogue and digital processing, as both analogue and digital dies can be coupled together. One such class of devices are closed-loop neuromodulators; neurostimulators which perform real-time digital signal processing (DSP) to deliver bespoke treatment. Due to their implantable nature, these devices are inherently governed by very strict volume constraints, power budgets, and must operate with high reliability. To address these challenges, this paper presents a low-power inductive coupling link (ICL) transceiver for 3D integration of digital CMOS and analogue BiCMOS dies for use in closed-loop neuromodulators. The use of an ICL, as opposed to through silicon vias (TSVs), ensures high reliability and fabrication yield in addition to circumventing the use of voltage level conversion between disparate dies, improving power efficiency. The proposed transceiver is experimentally evaluated using SPICE as well as nine traditional TSV baseline solutions. Results demonstrate that, whilst the achievable bandwidth of the TSV-based approaches is much higher, for the typical data rates demanded by neuromodulator applications (0.5-1 Gbps) the ICL design consumes on average 36.7% less power through avoiding the use of voltage level shifters.
Benjamin J. Fletcher, Shidhartha Das, Chi-Sang Poon, Terrence S. T. Mak
DATE4
2018 Improving the efficiency of thermal covert channels in multi-/many-core systems
abstract
In many-core chips seen in mobile computing, data center, AI, and elsewhere, thermal covert channels could be established to transmit data (e.g., passwords), supposedly to be kept secret and private. Effectiveness of a thermal covert channel, measured by its transmission rate and bit error rate (BER), is so much dependent on the thermal noise/interference imposed on the channel. In this paper, we present a few techniques to improve the capacity of thermal covert channel by overcoming the thermal interference. In particular, data in a thermal covert channel are encoded and represented following a new thermal signaling scheme where logic value, 0 or 1, modules the thermal signals duty cycle. Next, we show in this study that proper selection of transmission frequency can significantly minimize thermal interference. In addition, we propose a robust end-to-end communication protocol for reliable communications. Our experiments have confirmed that, compared to an existing thermal covert channel attack [1] [2], a thermal covert channel enhanced with all the improvements proposed in this study is seeing significant BER reduction (by as much as 75%), and transmission rate boost (by more than threefold). Building such a strong thermal covert channel is the key step towards developing robust defense and countermeasures against information leaking over thermal covert channel.
Zijun Long, Xiaohang Wang 0001, Yingtao Jiang, Guofeng Cui, Terrence S. T. Mak
DATE6
2018 Globally Wireless Locally Wired (GloWiLoW): A Clock Distribution Network for Many-Core Systems
abstract
Modern high-performance systems are now facing critical issues on delivering power-efficient and globally interconnected clock networks. Conventional metal-based interconnect has gradually reached its bottleneck with the technology scaling which limits system performance as the interconnect delay has already overweighed circuit gate delay. A typical clock distribution network (CDN) might consume up to 50% of the total chip power and could generate large phase delay and clock uncertainties due to its RC characteristic and unbalanced load. This paper proposes a novel hybrid wire-wireless CDN, which improves the performance of on-chip clock distribution significantly using embedded wireless clock transmitter and receiver fabrics. In particular, On-Off-Keying (OOK) transceivers are implemented for overall system simplicity and has achieved significant power-efficiency. Results indicate that the total propagation delay can be reduced to 39.4ps, which is 57 times lower than the conventional H-tree. Besides, the system clock skew can be predictable and limited only by the displacement of clock receivers. Less than 26.9ps of clock skew (6.7% of a clock period at 2.5GHz) could be found within the proposed CDN and hence, shows the promising potential of future high-performance on-chip clock distribution.
Benjamin J. Fletcher, Terrence S. T. Mak
ISCAS3
2018 Special session on bringing cores closer together: The wireless revolution in on-chip communication
abstract
The emerging field of NoC with Wireless interconnects is actively being pursued by a number of researchers worldwide, from a variety of different perspectives, ranging from very high levels of abstraction (e.g., system architecture) to very low levels (physical layer and transceiver design). Successful solutions will likely adopt and encompass elements from all or at least several levels of abstraction and rely on interdisciplinary concepts from multi-core architectures, integrated circuits, 3D ICs, digital communications, complex networks, and optimization techniques. This special session will provide a timely and insightful journey into various challenges and emerging solutions regarding the design of future NoC architectures. By scope and contents, this special session represents an engaging proposition to attendees belonging to both academia and industry.
Terrence S. T. Mak, Hiroki Matsutani, Partha Pratim Pande
VTS1
2018 An integrated web-based air pollution decision support system - a prototype
abstract
To efficiently and effectively monitor and mitigate air pollution in the urban environment, it is of paramount importance to integrate into a unified whole air pollutant concentration databases coming from different sources including the ground-based stations, mobile sensors, remote sensing, atmospheric-chemical-transport models and social media for the analysis and unraveling of the complex air pollution processes in space and time. This study constructs and implements for the first time a prototype of the fully integrated air pollution decision support system (APDSS) that put together in an integrated manner all relevant multi-scale, multi-type and multi-source data for decision-making on urban air pollution. The prototype contains the main system that handles the multi-source, multi-type and multi-scale databases, queries, visualization and data mining algorithms and the integrated modules that individually and holistically capitalize on the power of the ground-based stations, ground and aerial mobile sensors, satellite-borne remote-sensing technologies, atmospheric-chemical-transport models and social media. It renders a solid scientific foundation and system development methodology for the study of the spatiotemporal air pollution profiles crucial to the mitigation of urban air pollution. Real-life applications of the prototype are employed to illustrate the functionality of the APDSS.
Yee Leung, Kwong-Sak Leung, Man Hon Wong 0001, Terrence S. T. Mak, Kwan-Yau Cheung, Leung-Yau Lo, Wei Ying Yi, Yuan-Lin Dong
Int. J. Geogr. Inf. Sci.4
2018 Effectiveness of HT-assisted sinkhole and blackhole denial of service attacks targeting mesh networks-on-chip
Xiaohang Wang 0001, Yingtao Jiang, Mei Yang 0001, Terrence S. T. Mak, Amit Kumar Singh 0002
J. Syst. Archit.5
2018 A neuro-inspired visual tracking method based on programmable system-on-chip platform
Shufan Yang, KongFatt Wong-Lin, James Andrew, Terrence S. T. Mak, T. Martin McGinnity
Neural Comput. Appl.4
2018 Bubble Budgeting: Throughput Optimization for Dynamic Workloads by Exploiting Dark Cores in Many Core Systems
abstract
All the cores of a many-core chip cannot be active at the same time, due to reasons like low CPU utilization in server systems and limited power budget in dark silicon era. These free cores (referred to as bubbles) can be placed near active cores for heat dissipation so that the active cores can run at a higher frequency level, boosting the performance of applications that run on active cores. Budgeting inactive cores (bubbles) to applications to boost performance has the following three challenges. First, the number of bubbles varies due to open workloads. Second, communication distance increases when a bubble is inserted between two communicating tasks (a task is a thread or process of a parallel application), leading to performance degradation. Third, budgeting too many bubbles as coolers to running applications leads to insufficient cores for future applications. In order to address these challenges, in this paper, a bubble budgeting scheme is proposed to budget free cores to each application so as to optimize the throughput of the whole system. Throughput of the system depends on the execution time of each application and the waiting time incurred for newly arrived applications. Essentially, the proposed algorithm determines the number and locations of bubbles to optimize the performance and waiting time of each application, followed by tasks of each application being mapped to a core region. A Rollout algorithm is used to budget power to the cores as the last step. Experiments show that our approach achieves 50 percent higher throughput when compared to state-of-the-art thermal-aware runtime task mapping approaches. The runtime overhead of the proposed algorithm is in the order of 1M cycles, making it an efficient runtime task management method for large-scale many-core systems.
Xiaohang Wang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
IEEE Trans. Computers6
2017 Runtime task mapping for lifetime budgeting in many-core systems
abstract
Due to technology scaling, lifetime reliability is becoming one of major design constraints in the design of future many-core systems. In this paper, we propose a novel runtime mapping scheme which can dynamically map the applications given a lifetime reliability constraint. A borrowing strategy is adopted to manage the lifetime in a long-term scale, and the lifetime constraint can be relaxed in short-term scale when the communication performance requirement is high. The through-put can be improved because the communication performance of communication intensive applications is optimized, and mean-while the waiting time of computation intensive application is reduced. An improved neighborhood allocation method is proposed for the runtime mapping scheme. Moreover, we propose a method to effectively classify communication intensive applications and computation intensive applications. The experimental results show that compared to the state-of-the-art lifetime-constrained mapping, the proposed scheme has more than 20% throughput improvement in average.
Liang Wang 0020, Xiaohang Wang 0001, Ho-fung Leung, Terrence S. T. Mak
FDL4
2017 Throughput Optimization for Lifetime Budgeting in Many-Core Systems
abstract
Due to technology scaling, lifetime reliability is becoming one of major design constraints in the design of future many-core systems. In this paper, we propose a novel runtime mapping scheme which could dynamically map the applications given a lifetime reliability constraint. A borrowing strategy is adopted to manage the lifetime in a long-term scale, and the lifetime constraint could be relaxed in short-term scale when the communication performance requirement is high. The throughput could be improved because the communication performance of communication intensive applications is optimized, and meanwhile the waiting time of computation intensive application is reduced. Furthermore, an improved neighborhood allocation method is proposed for the runtime mapping scheme. The experimental results show that compared to the state-of-the-art lifetime-constrained mapping, the proposed mapping scheme could have over 20% throughput improvement.
Liang Wang 0020, Xiaohang Wang 0001, Ho-fung Leung, Terrence S. T. Mak
ACM Great Lakes Symposium on VLSI4
2017 On Runtime Communication- and Thermal-aware Application Mapping in 3D NoC
abstract
Many-core systems connected by 3D Network-on-Chips (NoC) are emerging as a promising computation engine for systems like cloud computing servers, big data systems, etc. Mapping applications at runtime to 3D NoCs is the key to maintain high throughput of the overall chip under a thermal/power constraint. However, the goals of optimizing both the communication latency and chip peak temperature are contradicting due to several reasons. Firstly, exploiting the vertical TSV links can accelerate communications, while low peak temperature prefers that the tasks to be mapped closer to the heat sink, instead of using the vertical links. Secondly, mapping tasks in close proximity can reduce communication latency, but at the cost of poor heat dissipation. To address these issues, in this paper, we propose an efficient runtime mapping algorithm to reduce both communication latency and overall application running time under thermal constraint. In essence, this algorithm first selects a 3D cuboid core region of a specific shape for each incoming application by setting the region's number of occupied vertical layers and its distance to the heat sink, in order to optimize its communication performance and peak temperature. Next, the exact locations of the core regions in the chip are determined, followed by a task-to-core mapping. The experimental results have confirmed that, compared to two recently proposed runtime mapping algorithms, our proposed approach can reduce the total running time by up to 48% and communication cost by up to 44%, with a low runtime overhead.
Xiaohang Wang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
NOCS4
2017 HRC: A 3D NoC Architecture with Genuine Support for Runtime Thermal-Aware Task Management
abstract
In spite of escalating thermal challenges imposed by high power consumption, most reported 3D Network-on-chip (NoC) systems that adopt classic 3D cube (mesh) topology are unable to tackle the thermal management issues directly at the architectural level. Rather, to avoid chip being overheated, tasks running in a “hot” node have to be migrated to a “cooler” one, resulting in increased distance between communicating nodes and ultimately poor performance. In this paper, we propose a new 3D NoC architecture that genuinely supports runtime thermal-aware task management. Dubbed Hierarchical Ring Cluster (HRC), this new hierarchical 3D NoC architecture has three levels across its entire network hierarchy: 1) nodes are grouped as rings, 2) rings are then grouped into cubes, and 3) multiple cubes are connected to form the whole network. Routing in a HRC system is also performed in a hierarchical manner: Paths are set up within rings using low latency circuit switching, and data that need to cross the rings or cubes are routed following dimension-order routing supported by wormhole switching. In this organization, “hot” tasks that need to migrate can move along the rings without incurring increased communication distances. Our experimental results have confirmed that the proposed HRC architecture has a much lower network latency than other known 3D NoC architectures. When working with runtime thermal-aware task migration approaches, HRC can help reduce latency by as much as 80 percent compared to thermal-aware task migration approaches applied to 3D mesh NoC topologies.
Xiaohang Wang 0001, Yingtao Jiang, Mei Yang 0001, Terrence S. T. Mak
IEEE Trans. Computers5
2017 A Resilient 2-D Waveguide Communication Fabric for Hybrid Wired-Wireless NoC Design
abstract
Hybrid wired-wireless Network-on-Chip (WiNoC) has emerged as an alternative solution to the poor scalability and performance issues of conventional wireline NoC design for future System-on-Chip (SoC). Existing feasible wireless solution for WiNoCs in the form of millimeter wave (mm-Wave) relies on free space signal radiation which has high power dissipation with high degradation rate in the signal strength per transmission distance. Moreover, over the lossy wireless medium, combining wireless and wireline channels drastically reduces the total reliability of the communication fabric. Surface wave has been proposed as an alternative wireless technology for low power on-chip communication. With the right design considerations, the reliability and performance benefits of the surface wave channel could be extended. In this paper, we propose a surface wave communication fabric for emerging WiNoCs that is able to match the reliability of traditional wireline NoCs. First, we propose a realistic channel model which demonstrates that existing mm-Wave WiNoCs suffers from not only free-space spreading loss (FSSL) but also molecular absorption attenuation (MAA), especially at high frequency band, which reduces the reliability of the system. Consequently, we employ a carefully designed transducer and commercially available thin metal conductor coated with a low cost dielectric material to generate surface wave signals with improved transmission gain. Our experimental results demonstrate that the proposed communication fabric can achieve a 5 dB operational bandwidth of about 60 GHz around the center frequency (60 GHz). By improving the transmission reliability of wireless layer, the proposed communication fabric can improve maximum sustainable load of NoCs by an average of 20:9 and 133:3 percent compared to existing WiNoCs and wireline NoCs, respectively.
Michael Opoku Agyeman, Quoc-Tuan Vien, Ali Ahmadinia, Alexandre Yakovlev, Kin-Fai Tong, Terrence S. T. Mak
IEEE Trans. Parallel Distributed Syst.6
2016 Bubble budgeting: throughput optimization for dynamic workloads by exploiting dark cores in many core systems
abstract
All the cores of a many-core chip cannot be active at the same time, due to reasons like low CPU utilization in server systems and limited power budget in dark silicon era. These free cores (referred to as bubbles) can be placed near active cores for heat dissipation so that the active cores can run at a higher frequency level, boosting the performance of active cores and applications. Budgeting inactive cores (bubbles) to workloads to boost performance has the following three challenges. First, the number of bubbles varies due to dynamic workloads. Second, communication distance increases when a bubble is inserted between two communicating tasks, leading to performance degradation. Third, budgeting too many bubbles as cooler to running applications leads to insufficient cores for future applications. In order to address these challenges, in this paper, a bubble budgeting scheme is proposed to budget free cores to each application so as to optimize the throughput of the whole system, including the execution time of each application and the waiting time incurred for newly arrived applications. Essentially, the proposed algorithm determines the number and locations of bubbles to optimize the performance and waiting time of each application, followed by tasks of each application being mapped to a core region. Experiments show that our approach achieves 50% higher throughput when compared to state-of-the-art thermal-aware runtime task mapping approaches.
Xiaohang Wang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
NOCS5
2016 Adaptive Routing Algorithms for Lifetime Reliability Optimization in Network-on-Chip
abstract
Technology scaling leads to the reliability issue as a primary concern in Network-on-Chip (NoC) design. We observe that due to routing algorithm some routers age much faster than others which becomes a bottleneck for NoC lifetime. In this paper, lifetime is modeled as a resource consumed over time. A metric lifetime budget is associated with each router, indicating the maximum allowed workload for current period. Since the heterogeneity in router lifetime reliability has strong correlation with the routing algorithm, we define a problem to optimize the lifetime by routing packets along the path with maximum lifetime budgets. The problem is then extended for both performance and lifetime reliability optimization. The lifetime is optimized in long-term time scale while performance is optimized in short-term time scale. Two dynamic programming-based adaptive routing algorithms (lifetime aware routing and multi-objective routing) are proposed to solve the problems. In the experiments, the lifetime aware routing and multi-objective routing algorithms are evaluated with synthetic traffic and real benchmarks respectively. The experimental results show that the lifetime aware routing has around 20, 45 and 55 percent minimal lifetime improvement than XY routing, NoP routing and Oddeven routing, respectively. In addition, the multi-objective adaptive routing algorithm can effectively improve both performance and lifetime.
Liang Wang 0020, Xiaohang Wang 0001, Terrence S. T. Mak
IEEE Trans. Computers3
2016 On Fine-Grained Runtime Power Budgeting for Networks-on-Chip Systems
abstract
Power budgeting is an essential aspect of networks-on-chip (NoC) to meet the power constraint for on-chip communications while assuring the best possible overall system performance. For simplicity and ease of implementation, existing NoC power budgeting schemes treat all the individual routers uniformly when allocating power to them. However, such homogeneous power budgeting schemes ignore the fact that the workloads of different NoC routers may vary significantly, and thus may provide excess power to routers with low workloads, whereas insufficient power to those with high workloads. In this paper, we formulate the NoC power budgeting problem in order to optimize the network performance over a power budget through per-router frequency scaling. We take into account of heterogeneous workloads across different routers as imposed by variations in traffic. Correspondingly, we propose a fine-grained solution using an agile algorithm with low time complexity. Frequency of each router is set individually according to its contribution to the average network latency while meeting the power budget. Experimental results have confirmed that with fairly low runtime and hardware overhead, the proposed scheme can help save up to$50$percent application execution time when compared with the latest proposed methods.
Xiaohang Wang 0001, Baoxin Zhao, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab
IEEE Trans. Computers3
2016 IP Protection of Mesh NoCs Using Square Spiral Routing
abstract
Intellectual property (IP) core reuse is essential for the design process of system-on-chip (SoC). Network-on-chip (NoC) has been used as an independent IP core during SoC design. However, the NoC has not been protected via IP protection and paid attention on its innovations. This paper proposes the first known approach to protect the authorship and the usage legitimacy of NoCs using specially designed routing, square spiral routing. The special routing algorithm exploits routing redundancy inherent in the mesh NoCs and transports packets along the paths, which have very low probability to be taken under commonly used routing algorithms. These unique and diverse paths are exploited in this paper to embed information of the author and identify the legal buyer of NoCs, showing high robustness and credibility. The hardware implementation of an IP-protected mesh NoC shows that the area overhead is small, which is ~0.74%, and the power overhead is ~0.52%, while the functionality and performance of the network is not affected. In this paper, the approach is presented for the mesh NoC, but the idea is equally applicable to other NoC topologies where the unique and diverse paths also inherently exist.
Qiang Liu 0011, Wenqing Ji, Terrence S. T. Mak
IEEE Trans. Very Large Scale Integr. Syst.4
2016 Defragmentation for Efficient Runtime Resource Management in NoC-Based Many-Core Systems
abstract
Efficient runtime resource allocation is critical to the overall performance and energy consumption of many-core systems. A region of free cores is allocated for each newly launched application. The cores are deallocated when the corresponding applications finish execution. The frequent allocations and deallocations of the cores might leave free cores scattered (not forming a contiguous region). This situation is referred to as fragmentation. Fragmentation could cause the inefficient mapping of the incoming applications, i.e., long communication distance between communicating cores. This further leads to poor performance and high energy consumption. In this paper, we propose a runtime defragmentation scheme that collects and reallocates the scattered cores in close proximity. We first define a fragmentation metric that is able to evaluate the scatteredness level of the free cores. Based on this, the proposed algorithm is executed to bring the scattered free cores together when the fragmentation metric is over a certain predefined threshold. In this way, the contiguous free core region is formed to facilitate the efficient mapping of the incoming applications. Moreover, the proposed algorithm also aims to minimize the negative impact on the performance of existing applications. Experimental results show that the proposed defragmentation scheme reduces the overall execution time and the energy consumption by 42% and 41%, respectively, when it is augmented to existing runtime mapping algorithms. Moreover, a negligible overhead, accounting for only less than 2.6% of the overall execution time, is required for the proposed defragmentation process. The proposed defragmentation scheme is an effective resource management enhancement to existing runtime mapping algorithms for many-core systems.
Jim Ng, Xiaohang Wang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
IEEE Trans. Very Large Scale Integr. Syst.4
2015 Fine-grained runtime power budgeting for networks-on-chip
abstract
Power budgeting for NoC needs to be performed to meet limited power budget while assuring the best possible overall system performance. For simplicity and ease of implementation, existing NoC power budgeting schemes, irrespective of the fact that the packet arrival rates of different NoC routers may vary significantly, treat all the individual routers indiscriminately when allocating power to them. However, such homogeneous power allocation may provide excess power to routers with low packet arrival rates whereas insufficient power to those with high arrival rates. In this paper, we formulate the NoC power budgeting problem as to optimize the network performance over a power budget through per-router frequency scaling, taking into account of heterogeneous packet arrival rates across different routers as imposed by run time traffic dynamics. Correspondingly, we propose a fine-grained solution using an agile dynamic programming network with a linear time complexity. In essence, frequency of a router is set individually according to its contribution to the average network latency while meeting the power budget. Experimental results have confirmed that with fairly low runtime and hardware overhead, the proposed scheme can help save up to 50% application execution time when compared with the best existing methods.
Xiaohang Wang 0001, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab
ASP-DAC3
2015 Mixed wire and surface-wave communication fabrics for decentralized on-chip multicasting
Ammar Karkar, Kin-Fai Tong, Terrence S. T. Mak, Alexandre Yakovlev
DATE3
2015 Novel Hybrid Wired-Wireless Network-on-Chip Architectures: Transducer and Communication Fabric Design
abstract
Existing wireless communication interface of Hybrid Wired-Wireless Network-on-Chip (WiNoC) has 3-dimensional free space signal radiation which has high power dissipation and drastically affects the received signal strength. In this paper, we propose a CMOS based 2-dimensional (2-D) waveguide communication fabric that is able to match the channel reliability of traditional wired NoCs as the wireless communication fabric. Our experimental results demonstrate that, the proposed communication fabric can achieve a 5dB operational bandwidth of about 60GHz around the center frequency (60GHz). Compared to existing WiNoCs, the proposed communication fabric can improve the reliability of WiNoCs with average gains of 21.4%, 13.8% and 10.6% performance efficiencies in terms of maximum sustainable load, throughput and delay, respectively.
Michael Opoku Agyeman, Wen Zong, Ji-Xiang Wan, Alexandre Yakovlev, Kenneth Tong, Terrence S. T. Mak
NOCS6
2015 Unbiased Regional Congestion Aware Selection Function for NoCs
abstract
Adaptive routing in Network-on-Chip (NoC) selects paths for packets according to network state to reduce packet latency and balance network load. Existing adaptive routing schemes can degrade network performance due to their dependency on either inadequate or outdated network information. We present an adaptive routing scheme in which a router is provided adequate and timely congestion information of the network. A low-complexity routing selection function that considers regional congestion status is proposed. The selection function is unbiased as it considers the same amount of congestion information on both admissible directions. Proposed selection function achieves 18% lower packet latency than local congestion aware selection under realistic workloads. It also reduces regional congestion aware selection logic area and power overhead by 73% and 35% on an 8×8 mesh network.
Wen Zong, Michael Opoku Agyeman, Xiaohang Wang 0001, Terrence S. T. Mak
NOCS4
2015 DeFrag: Defragmentation for Efficient Runtime Resource Allocation in NoC-Based Many-core Systems
abstract
Efficient runtime resource allocation is critical to the overall performance and energy consumption of many-core systems. However, due to the applications' unknown arrival and departure time under dynamic workloads, the runtime system resource management is challenging. The frequent allocations and deal locations of the applications might leave on-chip free cores scattered due to the lack of design-time knowledge of their finishing time. This situation is referred to as fragmentation. In order to optimize the performance and energy consumption of the system in such situations, in this paper, we propose a runtime defragmentation approach that collects and reshapes the scattered cores in close proximity. We also propose a fragmentation metric which is able to evaluate the scatteredness of the free cores. Based on this, the proposed algorithm will be executed to bring the scattered free cores together when the metric is over a certain predefined threshold. In this way, the contiguous free core region is formed to facilitate efficient mapping of the incoming applications. Moreover, the proposed algorithm is also aware of the existing applications and minimizes their performance impact. Experimental results demonstrated that the proposed defragmentation approach reduces the overall execution time and energy consumption by 42% and 41%, respectively when compared to some of the existing approaches. Moreover, a negligible overhead, accounting for only less than 2.6% of the overall execution time, is required for the defragmentation process.
Jim Ng, Xiaohang Wang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
PDP4
2015 An efficient runtime power allocation scheme for many-core systems inspired from auction theory
Xiaohang Wang 0001, Baoxin Zhao, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab
Integr.3
2015 Power-Adaptive Computing System Design for Solar-Energy-Powered Embedded Systems
abstract
Through energy harvesting system, new energy sources are made available immediately for many advanced applications based on environmentally embedded systems. However, the harvested power, such as the solar energy, varies significantly under different ambient conditions, which in turn affects the energy conversion efficiency. In this paper, we propose an approach for designing power-adaptive computing systems to maximize the energy utilization under variable solar power supply. Using the geometric programming technique, the proposed approach can generate a customized parallel computing structure effectively. Then, based on the prediction of the solar energy in the future time slots by a multilayer perceptron neural network, a convex model-based adaptation strategy is used to modulate the power behavior of the real-time computing system. The developed power-adaptive computing system is implemented on the hardware and evaluated by a solar harvesting system simulation framework for five applications. The results show that the developed power-adaptive systems can track the variable power supply better. The harvested solar energy utilization efficiency is 2.46 times better than the conventional static designs and the rule-based adaptation approaches. Taken together, the present thorough design approach for self-powered embedded computing systems has a better utilization of ambient energy sources.
Qiang Liu 0011, Terrence S. T. Mak, Tao Zhang 0025, Xinyu Niu, Wayne Luk, Alexandre Yakovlev
IEEE Trans. Very Large Scale Integr. Syst.2
2014 Agile frequency scaling for adaptive power allocation in many-core systems powered by renewable energy sources
abstract
As low-power electronics and miniaturization conspire to populate the world with emerging devices, one appealing approach is to power these multi-core/many-core-based devices with energy harvested from various environments. Of the most important issues concerning these devices is how to effectively allocate power budget among the cores competing for power, which is formulated as one specific type of power-performance optimization problem in this paper. We attempt to solve this problem by proposing an Adaptive Power Allocation Technique (APAT) that uses a dynamic programming network. Our goal here is to maximize the overall system performance, taking into account a unique yet challenging fact that, available power budget might have to undergo a significant change when a renewable energy source is scavenging. APAT has a linear time complexity and low hardware overhead. Experiments have confirmed that APAT can reduce 20 ~ 30% of execution time compared to other state-of-the-art power allocation algorithms. In addition, as APAT is quite insensitive to the changing rate of the power, lending itself well for power management in many-core systems powered by energy-harvesting sources.
Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab, Terrence S. T. Mak
ASP-DAC6
2014 Hybrid wire-surface wave architecture for one-to-many communication in networks-on-chip
abstract
Network-on-chip (NoC) is a communication paradigm that has emerged to tackle different on-chip challenges and has satisfied different demands in terms of high performance and economical interconnect implementation. However, merely metal based NoC pursuit offers limited scalability with the relentless technology scaling, especially in one-to-many (1-to-M) communication. To meet the scalability demand, this paper proposes a new hybrid architecture empowered by both metal interconnects and Zenneck surface wave interconnects (SWI). This architecture, in conjunction with newly proposed routing and global arbitration schemes, avoids overloading the NoC and alleviates traffic hotspots compared to the trend of handling 1-to-M traffic as unicast. This work addresses the system level challenges for intra chip multicasting. Evaluation results, based on a cycle-accurate simulation and hardware description, demonstrate the effectiveness of the proposed architecture in terms of power reduction ratio of 4 to 12X and average delay reduction of 25X or more, compared to a regular NoC. These results are achieved with negligible hardware overheads.
Ammar Karkar, Nizar Dahir, Ra'ed Al-Dujaily, Kenneth Tong, Terrence S. T. Mak, Alexandre Yakovlev
DATE5
2014 Adaptive power allocation for many-core systems inspired from multiagent auction model
abstract
Scaling of future many-core chips is hindered by the challenge imposed by ever-escalating power consumption. At its worst, an increasing fraction of the chips will have to be shut down, as power supply is inadequate to simultaneously switch all the transistors. This so-called dark silicon problem brings up a critical issue regarding how to achieve the maximum performance within a given limited power budget. This issue is further complicated by two facts. First, high variation in power budget calls for wide range power control capability, whereas most current frequency/voltage scaling techniques cannot effectively adjust power over such a wide range. Second, as the applications' behavior becomes more complicated, there is a pressing need for scalability and global coordination, rendering heuristic-based centralized or fully distributed control schemes inefficient. To address the aforementioned problems, in this paper, a power allocation method employing multiagent auction models is proposed, referred as Hierarchal MultiAgent based Power allocation (HiMAP). Tiles act the role of consumers to bid for power budget and the whole process is modeled by a combinatorial auction, whereas HiMAP finds the Walrasian equilibria. Experimental results have confirmed that HiMAP can reduce the execution time by as much as 45% compared to three competing methods. The runtime overhead and cost of HiMAP are also small, which makes it suitable for adaptive power allocation in many-core systems.
Xiaohang Wang 0001, Baoxin Zhao, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab, Maurizio Palesi
DATE3
2014 Design and Implementation of Dynamic Thermal-Adaptive Routing Strategy for Networks-on-Chip
abstract
Technology scaling is leading to extreme thermal challenges that make worst-case cooling system design unfavourable. On the other hand, on-chip communication, in terms of Network-on-Chip (NoC) workload, is expected to dominate Systems-on-Chip as a major heat source. In this paper a Runtime Thermal Management (RTM) design implementation for NoCs is proposed. Dynamic Programming Network (DPN) is introduced to implement the adaptive routing control logic and Ring Oscillators (ROs) are used for temperature sensing. Various challenges associated with DPN convergence and sensor accuracy and precision, such as isolating the IR drops and intra-chip process variations, are addressed. An FPGA implementation of the proposed strategy demonstrates promising results in terms of both thermal regulation and functionality with a variety of traffics. In terms of functionality, the proposed scheme is shown to be highly flexible in manoeuvring the packets away from hot regions. This results in up to 16% reduction in the maximum chip temperature and lowers chip thermal gradient by up to 51% compared with performance-driven routing. Moreover, the proposed scheme results in significantly slower chip heating which is reflected as up to 100% higher performance when the chip works under a thermal limit. These results imply that the proposed technique would improve thermal reliability and performance for future many-core VLSI systems.
Nizar Dahir, Ghaith Tarawneh, Terrence S. T. Mak, Ra'ed Al-Dujaily, Alexandre Yakovlev
PDP3
2014 Dynamic programming-based lifetime aware adaptive routing algorithm for Network-on-Chip
abstract
Technology scaling leads to the reliability issue as a primary concern in Network-on-Chip (NoC) design. Due to the routing algorithms, some routers may age much faster than others, which becomes a bottleneck for system lifetime. In this paper, lifetime is modeled as a resource consumed over time. A metric lifetime budget is associated with each router, indicating the maximum allowed workload for current period. Since the heterogeneity in router lifetime reliability has strong correlation with the routing algorithm, we define a problem to optimize the lifetime by routing flits along the path with maximum lifetime budgets. A dynamic programming-based lifetime aware routing algorithm is proposed based on the lifetime budget metric. The dynamic programming network approach is employed to solve this problem with linear complexity. The experimental results show that the lifetime aware routing has around 20%, 45%, 55% minimal MTTF improvement than XY routing, NoP routing, oddeven routing, respectively.
Liang Wang 0020, Xiaohang Wang 0001, Terrence S. T. Mak
VLSI-SoC3
2014 Modeling and Tools for Power Supply Variations Analysis in Networks-on-Chip
abstract
Power supply integrity has become a critical concern with the rapid shrinking feature size and the ever increasing power consumption in nanometre scale integration. In particular, on-chip communication in platforms such as networks-on-chip (NoC) dictates the power dissipation and overall system performance in multicore systems and embedded computing architectures. These architectures require a dedicated tool for analyzing the power supply noise which must embed distinctive communication characteristics and spatial parameters. In this paper, we present a tool dedicated to determining the on-chip VDDdrops due to communication workload in NoCs. This tool integrates a fast power grid model, an NoC simulator, an on-chip link model, and a microarchitectural power model for router. The model has been rigorously verified using SPICE simulations. The proposed model and tools are further exemplified through analyzing the impact of power supply noise for NoC links. Statistical timing analysis of NoC links in the presence of power supply noise was performed to evaluate the bit error rates (BERs). This work would enable better understanding of the tradeoffs existing in the design of NoCs, and the induced power supply noise due to on-chip communication. This understanding is crucial for the analysis of the quality of service (QoS) of communication fabrics in NoCs at the early design stages.
Nizar Dahir, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev
IEEE Trans. Computers2
2014 Thermal Optimization in Network-on-Chip-Based 3D Chip Multiprocessors Using Dynamic Programming Networks
abstract
The substantial silicon density in 3D VLSI, albeit its numerous advantages, introduces serious thermal threats that would lead to faults and system failures. This article introduces a new strategy to effectively diffuse heat from NoC-based 3D CMPs. Runtime Dynamic Programming Network (DPN) is proposed to optimize routing directions and provide silicon temperature moderation. Both on-chip reliability and computational performance have been improved by 63% and 27%, respectively, with the DPN approach. This work enables a new avenue to explore the adaptability for future large-scale 3D integration.
Nizar Dahir, Ra'ed Al-Dujaily, Terrence S. T. Mak, Alexandre Yakovlev
ACM Trans. Embed. Comput. Syst.3
2014 On self-tuning networks-on-chip for dynamic network-flow dominance adaptation
abstract
Modern network-on-chip (NoC) systems are required to handle complex runtime traffic patterns and unprecedented applications. Data traffics of these applications are difficult to fully comprehend at design time so as to optimize the network design. However, it has been discovered that the majority of dataflows in a network are dominated by less than 10% of the specific pathways. In this article, we introduce a method that is capable of identifying critical pathways in a network at runtime and can then dynamically reconfigure the network to optimize for network performance subject to the identified dominated flows. An online learning and analysis scheme is employed to quickly discover the emerging dominated traffic flows and provides a statistical traffic prediction using regression analysis. The architecture of a self-tuning network is also discussed which can be reconfigured by setting up the identified point-to-point paths for the dominance dataflows in large traffic volumes. The merits of this new approach are experimentally demonstrated using comprehensive NoC simulations. Compared to the conventional network architectures over a range of realistic applications, the proposed self-tuning network approach can effectively reduce the latency and power consumption by as much as 25% and 24%, respectively. We also evaluated the configuration time and additional hardware cost. This new approach demonstrates the capability of an adaptive NoC to handle more complex and dynamic applications.
Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Peng Liu 0016, Masoud Daneshtalab, Maurizio Palesi, Terrence S. T. Mak
ACM Trans. Embed. Comput. Syst.7
2014 Eliminating Synchronization Latency Using Sequenced Latching
abstract
Modern multicore systems have a large number of components operating in different clock domains and communicating through asynchronous interfaces. These interfaces use synchronizer circuits, which guard against metastability failures but introduce latency in processing the asynchronous input. We propose a speculative method that hides synchronization latency by overlapping it with computation cycles. We verify the correctness of our approach through a field programmable gate array implementation and apply it to a number of synthesized benchmarks. Synthesis results reveal that our approach achieves average savings of 135% and 204% in area costs and nearly 100% in power costs compared to two similar speculative techniques
Ghaith Tarawneh, Alexandre Yakovlev, Terrence S. T. Mak
IEEE Trans. Very Large Scale Integr. Syst.3
2013 Novel Multi-Layer Network Decomposition boosting acceleration of multi-core algorithms
abstract
Complex networks are a technique for the modeling and analysis of large data sets in many scientific and engineering disciplines. Due to their excessive size conventional algorithms and single core processors struggle with the efficient processing of such networks. Employing multi-core graphic processing units (GPUs) could provide sufficient processing power for the analysis of such networks. However, commonly designed algorithms cannot exploit these massively parallel processing power for the analysis of such networks. In this paper, we present the Multi Layer Network Decomposition (MLND) approach which provides a general approach for parallel network analysis using multi-core processors via efficient partitioning and mapping of networks onto GPU architectures. Evaluation using a 336 core GPU graphic card demonstrated a 16x speed-up in complex network analysis relative to a CPU based approach.
Athanasios K. Grivas, Terrence S. T. Mak, Alexandre Yakovlev, Jonny Wray
ASAP2
2013 A Fault-Tolerant Routing Algorithm for NoC Using Farthest Reachable Routers
abstract
As technology scaling, reliability has became one of the key challenges of Network-on-Chip (NoC). Many faulttolerant routing algorithms for NoC are developed to overcome fault components and provide reliable transmission. But proposed routing algorithms do not pay enough attention to find the shortest paths, which increases latency and power consumption. In this paper, a fault-tolerant routing algorithm using new component states diffusion method based on Farthest Reachable Router (FRR) is proposed. This algorithm can reduce latency by finding the shortest paths between source and destination routers. Experiment results verify that FRR routing algorithm can tolerate 79% fault patterns within 3 × 3 and reduce latency by 16-44% compared with FON.
Junshi Wang, Xiaohang Wang 0001, Letian Huang, Terrence S. T. Mak, Guangjun Li
DASC4
2013 Towards reliable hybrid bio-silicon integration using novel adaptive control system
abstract
Hybrid bio-silicon networks are difficult to implement in practice due to variations of biological neuron bursting frequency. This causes the hybrid network to have inaccuracies and unreliability. The network may produce irregular bursts or incorrect spiking phase relationships if the electrical neuron bursting frequency is not suitable for biological neurons. To solve this potentially vital problem, a novel adaptive control system based on dynamic clamp is proposed. Biological measurement is combined with an adaptive controller to control to silicon neuron bursting periods in real time. We use a hybrid pyloric network which contains three real neurons and one electronic neuron as a case study. Simulation results indicate that the silicon neuron can follow the biological neuron bursting frequency in real time to achieve hybrid network functionalities. System settling time can be achieved in 303 milliseconds and percentage overshoot kept to 1%. We believe that our methodology is scalable to various larger bio-silicon hybrid neural networks.
Patrick Degenaar, Graeme Coapes, Alexandre Yakovlev, Terrence S. T. Mak, Peter Andras 0001
ISCAS5
2013 On self-tuning networks-on-chip for dynamic network-flow dominance adaptation
abstract
Modern networks-on-chip (NoC) systems are required to handle complex run-time traffic patterns and unprecedented applications. Data traffics of these applications are difficult to be fully comprehended at design-time so as to optimize the network design. However, it has been discovered that the majority data flows in a network are dominated by less than 10% of the specific pathways. In this paper, we introduce a method that is capable of identifying critical pathways in a network at run-time and, then, can dynamically reconfigure the network to optimize for the network performance subjected to the identified dominated flows. An online learning and analysis scheme is employed to quickly discover the emerged dominated traffic flows and provides a statistical traffic prediction using regression analysis. The architecture of a self-tuning network is also discussed which can be reconfigured by setting up the identified point-to-point paths for the dominance data flows in large traffic volumes. The merits of this new approach are experimentally demonstrated using comprehensive NoC simulators. Compared to the conventional network architectures over a range of realistic applications, the proposed self-tuning network approach can effectively reduce the latency and power consumption by as much as 25% and 24%, respectively. We also evaluated the configuration time and additional hardware cost. This new approach demonstrates the capability of an adaptive NoC to handle more complex and dynamic applications.
Xiaohang Wang 0001, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab, Maurizio Palesi
NOCS2
2013 Dynamic On-Chip Thermal Optimization for Three-Dimensional Networks-On-Chip
abstract
The complex thermal behaviour prohibits the advancement of three-dimensional (3D) very-large-scale integration system. Particularly, the high-density through-silicon via based integration could lead to ultra-high temperature hotspots and permanent silicon device damage. In this paper, we introduce an adaptive strategy to effectively diffuse heat throughout the 3D geometry. This strategy employs a dynamic programming network to select and optimize the direction of data manoeuvre in a network-on-chip (NoC). We also developed a tool, which is based on the accurate HotSpot thermal model and SystemC cycle accurate model, to simulate the thermal system and evaluate our approach. We found that the proposed approach can significantly diffuse the hotspots from a 3D geometry and overall temperature can be significantly reduced. Given the same thermal constraints, the throughput performance of an adaptive NoC can also be improved. This work enables a new avenue to explore the on-chip adaptability for the future large-scale 3D integration.
Ra'ed Al-Dujaily, Terrence S. T. Mak, Kai-Pui Lam, Fei Xia 0001, Alexandre Yakovlev, Chi-Sang Poon
Comput. J.2
2013 Efficient multicast schemes for 3-D Networks-on-Chip
Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Maurizio Palesi, Peng Liu 0016, Terrence S. T. Mak, Nader Bagherzadeh
J. Syst. Archit.6
2013 Dynamic programming-based runtime thermal management (DPRTM): An online thermal control strategy for 3D-NoC systems
abstract
Complex thermal behavior inhibits the advancement of three-dimensional (3D) very-large-scale-integration (VLSI) system designs, as it could lead to ultra-high temperature hotspots and permanent silicon device damage. This article introduces a new runtime thermal management strategy to effectively diffuse and manage heat throughout 3D chip geometry for a better throughput performance in networks on chip (NoC). This strategy employs a dynamic programming-based runtime thermal management (DPRTM) policy to provide online thermal regulation. Reactive and proactive adaptive schemes are integrated to optimize the routing pathways depending on the critical temperature thresholds and traffic developments. Also, when the critical system thermal limit is violated, an urgent throttling will take place. The proposed DPRTM is rigorously evaluated through cycle-accurate simulations, and results show that the proposed approach outperforms conventional approaches in terms of computational efficiency and thermal stability. For example, the system throughput using the DPRTM approach can be improved by 33% when compared to other adaptive routing strategies for a given thermal constraint. Moreover, the DPRTM implementation presented in this article demonstrates that the hardware overhead is insignificant. This work opens a new avenue for exploring the on-chip adaptability and thermal regulation for future large-scale and 3D many-core integrations.
Ra'ed Al-Dujaily, Nizar Dahir, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev
ACM Trans. Design Autom. Electr. Syst.3
2012 A scalable FPGA-based design for field programmable large-scale ion channel simulations
abstract
The design of systems to replicate complex neural functionality is a requirement for the development of next-generation prosthetic devices. The demands of such neural models are growing exponentially as we discover more about how brain systems function. It is therefore important for the electronic architectures involved to scale effectively in terms of latency, area and power usage in order to be able to process more advanced neural models. Within this paper a design is proposed that utilises the parallel nature and the resources available upon modern FPGAs to achieve a scalable and efficient method for the implementation of complex neural models, allowing for the simulation of 150000 ion channels concurrently.
Graeme Coapes, Terrence S. T. Mak, Alexandre Yakovlev, Chi-Sang Poon
FPL2
2012 Intra-chip physical parameter sensor for FPGAS using flip-flop metastability
abstract
We present a novel intra-chip physical parameter sensor that exploits the clock-to-q delay response of flip-flops. The proposed design relies on deliberately violating the setup and hold time conditions of a flip-flop to bring it into metastable states and increase its clock-to-q delay. Traditionally, this is an undesired effect because it can result in unpredictable system failures. In this work, this phenomenon is exploited to quantify variations in intra-chip physical parameters. Our design has three benefits over conventional ring-oscillator-based sensors; it consumes less device resources, has a higher precision and does not require a high clock frequency. We present a small-signal model of the proposed sensor and compare its performance with ring oscillators by conducting voltage and temperature-controlled experiments on an Altera Cyclone II FPGA device.
Ghaith Tarawneh, Terrence S. T. Mak, Alexandre Yakovlev
FPL2
2012 Embedded Transitive Closure Network for Runtime Deadlock Detection in Networks-on-Chip
abstract
Interconnection networks with adaptive routing are susceptible to deadlock, which could lead to performance degradation or system failure. Detecting deadlocks at runtime is challenging because of their highly distributed characteristics. In this paper, we present a deadlock detection method that utilizes runtime transitive closure (TC) computation to discover the existence of deadlock-equivalence sets, which imply loops of requests in networks-on-chip (NoCs). This detection scheme guarantees the discovery of all true deadlocks without false alarms in contrast with state-of-the-art approximation and heuristic approaches. A distributed TC-network architecture, which couples with the NoC infrastructure, is also presented to realize the detection mechanism efficiently. Detailed hardware realization architectures and schematics are also discussed. Our results based on a cycle-accurate simulator demonstrate the effectiveness of the proposed method. It drastically outperforms timing-based deadlock detection mechanisms by eliminating false detections and, thus, reducing energy wastage in retransmission for various traffic scenarios including real-world application. We found that timing-based methods may produce two orders of magnitude more deadlock alarms than the TC-network method. Moreover, the implementations presented in this paper demonstrate that the hardware overhead of TC-networks is insignificant.
Ra'ed Al-Dujaily, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev, Maurizio Palesi
IEEE Trans. Parallel Distributed Syst.2
2011 Run-time deadlock detection in networks-on-chip using coupled transitive closure networks
abstract
Interconnection networks with adaptive routing are susceptible to deadlock, which could lead to performance degradation or system failure. Detecting deadlocks at run-time is challenging because of their highly distributed characteristics. In this paper, we present a deadlock detection method that utilizes run-time Transitive Closure (TC) computation to discover the existence of deadlock-equivalence sets, which imply loops of requests in networks-on-chip (NoC). This detection scheme guarantees the discovery of all true deadlocks without false alarms unlike state-of-the-art approximation and heuristic approaches. A distributed TC-network architecture which couples with the NoC architecture is also presented to realize the detection mechanism efficiently. Our results based on a cycle-accurate simulator demonstrate the effectiveness of the TC-network method. It drastically outperforms timing-based deadlock detection mechanisms by eliminating false detections and thus reducing energy dissipation in various traffic scenarios. For example, timing based methods may produce two orders of magnitude more deadlock alarms than the TC-network method. Moreover, the implementations presented in this paper demonstrate that the hardware overhead of TC-networks is insignificant.
Ra'ed Al-Dujaily, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev, Maurizio Palesi
DATE2
2011 Redressing timing issues for speed-independent circuits in deep submicron age
abstract
The class of speed independent (SI) circuits opens a promising way towards tolerating process variations. However, the fundamental assumption of speed independent circuit is that forks in some wires (usually, large percentage of wires) in such circuits are isochronic; this assumption is more and more challenged by the shrinking technology. This paper suggests a method to generate the weakest timing constraints for a SI circuit to work correctly under bounded delays in wires. The method works for all SI circuits and the generated timing constraints are significantly weaker than those suggested in the current literature claiming the weakest formally proved conditions.
Terrence S. T. Mak, Alexandre Yakovlev
DATE2
2011 Communication centric on-chip power grid models for networks-on-chip
abstract
Adverse effects of unreliable on-chip power supply delivery are exacerbated due to the rapid shrinking of device dimensions and the ever increasing power consumptions in nanometre-scale integration. Power supply integrity becomes a critical concern. Particularly, on-chip communication networks, such as networks-on-chip (NoC), dictates power dissipations and overall system performance in multi-core systems and emerging embedded computing architectures. These new communication centric architectures require dedicated power grid model that embeds distinctive communication characteristics and spatial parameters for analysing impacts of power supply voltage drop and noise. In this paper, we present a new on-chip power delivery model that captures the on-chip communication patterns and power grid dynamics. This model integrates cycle-accurate simulation of networks-on-chip to analyze the impact of different design entities on power supply noise. The model has been rigorously evaluated. Novel observations of power delivery integrity due to communication network design are presented. This model provides a unique and communication-centric perspective to analyse power supply integrity that leads to future robust and reliable multi-core system design.
Nizar Dahir, Terrence S. T. Mak, Alexandre Yakovlev
VLSI-SoC2
2011 Comparative ODE benchmarking of unidirectional and bidirectional DP networks for 3D-IC
abstract
There has been great technological stride in 3D-IC on its design, analysis, and fabrication, with prediction that they will eventually lead to significant advances in multicore, multiprocessor, and network-on-chip (NoC) systems. A dynamic programming (DP) network is well suited for the grid stack architecture, because of its capability to achieve global optimality using only local computational units with short inter-grid communication links. In this paper we extend the transitive closure and shortest path unidirectional networks to bidirectional networks, with the development of an effective simulation tool for such type of DP networks. In addition to helping to construct real DP networks on 3D-IC, ODE (ordinary differential equation) simulation methodology for solving an average shortest path length problem provides new insights for comparative bench-marking very large-scale 2D/3D networks for different design considerations in application.
Kai-Pui Lam, Terrence S. T. Mak, Chi-Sang Poon
VLSI-SoC2
2011 Cycle avoidance in 2D/3D bidirectional graphs using shortest-path dynamic programming network
abstract
An ordinary-differential-equation (ODE) simulation model has recently been proposed for an N-node dynamic programming (DP) network, which solves the transitive closure and shortest path problems on an architecturally equivalent N-node 2D/3D grid stack. For large-scale randomly generated bidirectional network, where N is large and the inter-grid paths may take either direction, cycles commonly occur leading to a high percentage of nodes with unbound path lengths. The detection of such cycle nodes can be readily found using a shortest-path DP network. In this work we address several issues on the cycle avoidance problem, by first defining the 〈H〉-index and 〈V〉-index and hence its product 〈HV〉 as the two-dimensional turn ability. A regression model was then proposed and obtained empirically to relate the cycle-node ratio, which is the percentage of cycle nodes over N, with 〈HV〉 for several random networks of sizes N = 10×10×2, 10×10×5, 10×10×8. By reducing 〈HV〉 from 0.8 to 0.2, the cycle-node ratio can be reduced from close to 60% to 20% and indicates a significant avoidance of cycle nodes.
Kai-Pui Lam, Terrence S. T. Mak, Chi-Sang Poon
VLSI-SoC2
2010 A Reconfigurable Hebbian Eigenfilter for Neurophysiological Spike Train Analysis
abstract
The emergence of multi-electrode array enables the study of real-time neurophysiological activities across multiple regions of the brain. However, the real-time extracellular action potentials recorded on any electrode represent the simultaneous electrical activity of an unknown number of neurons which present a critical challenge to the accuracy of interpretation and identification of the neural circuitry in the subsequent analysis. In this paper, we present a principal component analysis approach utilizing Hebbian eigenfilter to identify the corresponding electrical activities of each neuron, namely spike sorting. The Hebbian eigenfilter greatly simplifies the computational complexity of eigen-projection. An efficient FPGA-based Hebbian eigenfilter is proposed. The performance, accuracy and power consumption of our Hebbian eigenfilter are thoroughly evaluated through synthetic spike trains. The proposal enables real-time spike sorting and analysis, and leads the way towards future motor and cognitive neuroprosthetics.
Bo Yu 0014, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev, Yihe Sun, Chi-Sang Poon
FPL2
2010 Wave-pipelined intra-chip signaling for on-FPGA communications
Terrence S. T. Mak, N. Pete Sedcole, Peter Y. K. Cheung, Wayne Luk
Integr.1
2009 Throughput Maximization for Wave-pipelined Interconnects using Cascaded Buffers and Transistor Sizing
abstract
This paper presents two new design methodologies for throughput-centric wave-pipelined interconnects: cascaded buffers insertion and transistor sizing. Experimental results show that up to 185% throughput improvement can be achieved by applying the new proposed approaches compared with conventional interconnect optimization techniques, such as buffer insertion. Moreover, with the combination of cascaded buffers insertion and adequate techniques in supply voltage scaling, up to 60% dynamic power reduction can be gained compared to the conventional design.
Terrence S. T. Mak, N. Pete Sedcole, Peter Y. K. Cheung
ISCAS2
2008 High-throughput interconnect wave-pipelining for global communication in FPGAs
abstract
Global interconnection is fundamental to high bandwidth links for inter-module communication in FPGAs. The long range global interconnections at high clock frequencies are becoming more problematic. This is due to the large circuit delay and leakage power caused by interconnect switches along the line. The delay is worsen by the global interconnect deterioration in technology scaling. In this paper, we address this problem by presenting an interconnect wave-pipelining strategy by using the existing programmable interconnects fabrics to provide high-throughput global communication in FPGA. A novel global interconnection circuit model is presented and, from the model, interconnection throughput can be derived. The model has been verified using SPICE simulation and delay results from the Xilinx FPGA Editor. We demonstrate the feasibility of our proposal by implementing a wave-pipelined interconnect circuit in a Xilinx Virtex-5 FPGA device. The circuit is able to achieve a throughput that is 3 times faster than a conventional synchronous approach. We conclude this paper by having a discussion about two strategies to further enhance the wave-pipelining throughput
Terrence S. T. Mak, N. Pete Sedcole, Peter Y. K. Cheung, Wayne Luk
FPGA1
2008 Wave-pipelined signaling for on-FPGA communication
abstract
On-FPGA communication is becoming more problematic as the long interconnection performance is deteriorating in technology scaling. In this paper, we address this issue by presenting a new wave-pipelined signaling scheme to achieve high-throughput communication in FPGA. The throughput and power consumption of a wave-pipelined link have been derived analytically and compared to the conventional synchronous link. Two circuit designs are proposed to realize wave-pipelined link using FPGA fabrics. The proposed approaches are also compared with conventional synchronous and asynchronous pipelining techniques. It is shown that, the wave-pipelined approach can achieve up to 5.66 times improvement in throughput versus the synchronous link and 13% improvement in power consumption and 35% improvement in delay versus the synchronous register-pipelining. Also, trade-offs between power, speed and area between the proposed and conventional designs are studied.
Terrence S. T. Mak, N. Pete Sedcole, Peter Y. K. Cheung, Wayne Luk
FPT1
2008 Implementation of Wave-Pipelined Interconnects in FPGAs
Terrence S. T. Mak, Crescenzo D'Alessandro, N. Pete Sedcole, Peter Y. K. Cheung, Alexandre Yakovlev, Wayne Luk
NOCS1
2007 A Current-Mode Analog Circuit for Reinforcement Learning Problems
abstract
Reinforcement learning is important for machine-intelligence and neurophysiological modelling applications to provide time-critical decision making. Analog circuit implementation has been demonstrated as a powerful computational platform for power-efficient, bio-implantable and real-time applications. This paper presents a current-mode analog circuit design for solving reinforcement learning problem with simple and efficient computational network architecture. The design has been fabricated and a new procedure to validate the fabricated reinforcement learning circuit will also be presented. This work provides a preliminary study for future biomedical application using CMOS VLSI reinforcement learning model.
Terrence S. T. Mak, Kai-Pui Lam, H. S. Ng, Guy Rachmuth, Chi-Sang Poon
ISCAS1
2007 A Hybrid Analog-Digital Routing Network for NoC Dynamic Routing
abstract
Dynamic routing can substantially enhance the quality of service for multiprocessor communication, and can provide intelligent adaptation of faulty links during run time. Implementing dynamic routing on a network-on-chip (NoC) platform requires a design that provides highly efficient optimal path computation coupled with reduced area and power consumption. In this paper, we present a hybrid analog-digital routing network design that enables efficient dynamic routing on an NoC architecture. The digital part provides accurate real-time traffic estimation using a temporal cost evaluation and adaptation scheme. The analog network, which is distributed within the digital communication network, provides an efficient implementation for the optimal routing algorithm with extremely low power consumption. Our results demonstrate the effectiveness of the hybrid analog-digital design, with a significant improvement in latency over the static routing for random hot spot traffics
Terrence S. T. Mak, N. Pete Sedcole, Peter Y. K. Cheung, Wayne Luk, Kai-Pui Lam
NOCS1
2006 On-FPGA Communication Architectures and Design Factors
abstract
The recent development of Platform-FPGA or Field-Programmable System-on-Chip architectures, with immersed coarse-grain processors, embedded memories and IP cores, offers the potential for immense computing power as well as opportunities for rapid system prototyping. These platforms require high-performance on-chip communication architectures for efficient and reliable inter-processor communication. However, as the number of embedded processors increases, communication bandwidth between embedded components becomes a limiting factor to overall system performance. In this paper, we survey the state-of-the-art on-FPGA communication architectures and methodologies. Salient factors, which include quantitative performance metrics and qualitative factors, relevant to design are identified and used to analyze and classify the on-FPGA communication architectures. This survey aims to facilitate innovation in and development of future on-FPGA communication architectures.
Terrence S. T. Mak, N. Pete Sedcole, Peter Y. K. Cheung, Wayne Luk
FPL1
2004 FPGA-Based Computation for Maximum Likelihood Phylogenetic Tree Evaluation
Terrence S. T. Mak, Kai-Pui Lam
FPL1
2004 On Computing Maximum Likelihood Phylogeny Using FPGA p
Terrence S. T. Mak, Kai-Pui Lam
FPL1
2003 An FPGA-based eigenfilter using fast Hebbian learning
abstract
We present a high-gain, multiple learning/decay rate, "cooling off" annealing strategy to a modified generalized Hebbian algorithm (GHA) that gives good approximate solution within one training epoch, and with fast convergence to accurate principal components within a few more epochs. A novel bit-shifting normalization procedure is shown to bound the weight vector norm effectively and eliminates the need for performing division. This leads to an FPGA-based computational framework using only fixed point arithmetic instead of more complicated floating point design. Simulation results on Xilinx DSP System Generator tool indicate the practicality of the approach, where real-time eigenfilter can. be readily implemented on field programmable gate arrays with limited resources.
Kai-Pui Lam, Terrence S. T. Mak
ICASSP (2)2
2002 On Computing Transitive-Closure Equivalence Sets Using a Hybrid GA-DP Approach
Kai-Pui Lam, Terrence S. T. Mak
FPL2
2002 Serial-parallel tradeoff analysis of all-pairs shortest path algorithms in reconfigurable computing
abstract
Implementation of shortest path algorithm in FPGA has been recently proposed for solving the network routing problem. This paper discusses the architecture and implementation of shortest path algorithms for Floyd-Warshall algorithm and the parallel implementation of Bellman-Ford algorithm in the Binary Relation Inference Network architecture. There are significant differences in the performance of computing shortest paths for these two different approaches. The computation speed and resource consumption issues are discussed. An alternative, serial implementation of the synchronized inference network for single-destination problem is also explored, with emphasis on computation time, resource consumption, and scaling problem size.
Terrence S. T. Mak, Kai-Pui Lam
FPT1