EDBT 2026 Demo / reviewers in the wild / expert
Fei Xia 0001
dblp:79/1081-1
· DBLP profile ↗
36ranked-venue papers
3as first author
6since 2021 · last 2023
0000-0002-3426-8406ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 1 since 2021Theory of computation · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Asynchronous Control for Tsetlin Machine with Binary Memristor-Transistor ArrayabstractTsetlin machines (TMs) are a novel machine learning paradigm based on learning automata and Boolean logic inference, with better energy-efficiency and explainability than neural networks. This work exploits non-volatile ReRAM-transistor memory arrays to perform efficient in-memory TM computing. To accommodate the large timing variability of ReRAM devices and enhance energy efficiency and speed, the control path is implemented with quasi delay-insensitive (QDI) asynchronous circuits. The design of these circuits are derived and synthesized from their signal-transition graph specifications using the Workcraft tool. The resulting circuits offer high event-driven controllability and high variation tolerance for the mixed-signal ReRAM data path. Compared to state of the art TM hardware, the new TM design uses less than 5% of the power to achieve better than$4\times$the performance. Omar Ghazal, Gang Mao, Jesse Ojukwu, Fei Xia 0001, Alexandre Yakovlev, Rishad A. Shafik |
ISCAS | 5 |
| 2023 | IMBUE: In-Memory Boolean-to-CUrrent Inference ArchitecturE for Tsetlin MachinesabstractIn-memory computing for Machine Learning (ML) applications remedies the von Neumann bottlenecks by organizing computation to exploit parallelism and locality. Non-volatile memory devices such as Resistive RAM (ReRAM) offer integrated switching and storage capabilities showing promising performance for ML applications. However, ReRAM devices have design challenges, such as nonlinear digital-analog conversion and circuit overheads. This paper proposes an In-Memory Boolean-to-Current Inference Architecture (IMBUE) that uses ReRAM-transistor cells to eliminate the need for such conversions. IMBUE processes Boolean feature inputs expressed as digital voltages and generates parallel current paths based on resistive memory states. The proportional column current is then translated back to the Boolean domain for further digital processing. The IMBUE architecture is inspired by the Tsetlin Machine (TM), an emerging ML algorithm based on intrinsically Boolean logic. The IMBUE architecture demonstrates significant performance improvements over binarized convolutional neural networks and digital TM in-memory implementations, achieving up to a 12.99x and 5.28x increase, respectively. Omar Ghazal, Simranjeet Singh, Tousif Rahman, Shengqi Yu, Yujin Zheng, Domenico Balsamo, Sachin B. Patkar, Farhad Merchant, Fei Xia 0001, Alexandre Yakovlev, Rishad A. Shafik |
ISLPED | 9 |
| 2023 | Approximate digital-in analog-out multiplier with asymmetric nonvolatility and low energy consumptionabstractMany modern compute-intensive applications require arithmetic results (usually multiplication) to be represented as analog signals. Using digital multipliers followed by digital-to-analog conversion (DAC) results in high energy and performance costs. This is because digital multipliers have costly carry propagation, and DAC circuits add associated conversion costs. Another concern, especially for arithmetic on the edge, is the need for nonvolatile operands in the face of power uncertainty. To deal with this, nonvolatile memory technologies have been combined with in-memory computing. This paper proposes a mixed-signal multiplier which directly generates an analog product based on two digital input operands. Fundamental to the design are transistor-memristor cells, organized in a crossbar structure. Using analog resistive partial product accumulation in the crossbar, the approximate multiplier eliminates the need for carry propagation and an explicit DAC. It also provides asymmetric nonvolatility making memristor writing a rare event, extending the application significance of the method. The design is shown to be functionally correct up to 4-bit, and achieves 8× to over 300× speedup, competitive peak-power and orders of magnitude energy reduction, compared with existing full-digital memristor-based multipliers and low-power multiplication DAC solutions. Shengqi Yu, Fei Xia 0001, Rishad A. Shafik, Domenico Balsamo, Alexandre Yakovlev |
Integr. | 2 |
| 2022 | Runtime Energy Minimization of Distributed Many-Core Systems using Transfer LearningabstractThe heterogeneity of computing resources continues to permeate into many-core systems making energy-efficiency a challenging objective. Existing rule-based and model-driven methods return sub-optimal energy-efficiency and limited scalability as system complexity increases to the domain of distributed systems. This is exacerbated further by dynamic variations of workloads and quality-of-service (QoS) demands. This work presents a QoS-aware runtime management method for energy minimization using a transfer learning (TL) driven exploration strategy. It enhances standard Q-learning to improve both learning speed and operational optimality (i.e., QoS and energy). The core to our approach is a multi-dimensional knowledge transfer across a task's state-action space. It accelerates the learning of dynamic voltage/frequency scaling (DVFS) control actions for tuning power/performance trade-offs. Firstly, the method identifies and transfers already learned policies between explored and behaviorally similar states referred to as Intra-Task Learning Transfer (ITLT). Secondly, if no similar “expert” states are available, it accelerates exploration at a local state's level through what's known as Intra-State Learning Transfer (ISLT). A comparative evaluation of the approach indicates faster and more balanced exploration. This is shown through energy savings ranging from 7.30% to 18.06%, and improved QoS from 10.43% to 14.3%, when compared to existing exploration strategies. This method is demonstrated under WordPress and TensorFlow workloads on a server cluster. Dainius Jenkus, Fei Xia 0001, Rishad A. Shafik, Alexandre Yakovlev |
DATE | 2 |
| 2022 | Editable asynchronous control logic for SAR ADCsabstractThis paper presents a novel design method for asynchronous control logic targeting successive approximation register (SAR) analog-to-digital converters (ADCs). This work is based on modeling the control logic for SAR ADCs using signal transition graphs (STGs). Different from conventional synchronous controllers, the proposed method results in asynchronous controllers driven by the causality of signals rather than relying on clocks to control the conversion process. Moreover, the proposed asynchronous control logic can be modularized through the handshake protocol, making it possible to build ADCs of arbitrary precision based on single-bit control units. This work results in a formal, model-based asynchronous design flow for SAR ADC control, which is shown to produce resulting circuits of similar speeds but great power efficiency improvements. Fei Xia 0001, Gang Mao, Shengqi Yu, Rishad A. Shafik, Alexandre Yakovlev |
ISCAS | 2 |
| 2021 | Run-time Configurable Approximate Multiplier using Significance-Driven Logic CompressionabstractDesigning energy-efficient hardware continues to be challenging due to arithmetic complexities. The problem is further exacerbated in systems powered by energy harvesters as variable power levels can limit their computation capabilities. In this work, we propose a run-time configurable adaptive approximation method for multiplication that is capable of managing the energy and performance tradeoffs — ideally suited in these systems. Central to our approach is a Significance-Driven Logic Compression (SDLC) multiplier architecture that can dynamically adjust the level of approximation depending on the run-time power/accuracy constraints. The architecture can be configured to operate in the exact mode (no approximation) or in progressively higher approximation modes (i.e. 2 to 4-bit SDLC). Our method is implemented in both ASIC and FPGA. The implementation results indicate that our design has only a 2.3% silicon overhead, on top of what is required by a traditional exact multiplier. We evaluate the efficiency of the proposed design through a number of case studies. We show that our method achieves similar image fidelity as in the existing approximate methods, without a delay penalty. Further, the inclusion of the dynamic approximation techniques is justified by up to 62.6% energy savings when processing an image with a multiplier using 4-bit SDLC and 35% energy savings when using 2-bit SDLC. In addition, case study results show that the proposed approach incurs negligible loss in output quality with the worst PSNR of 30dB when using the 4-bit SDLC multiplier. Ibrahim Haddadi, Issa Qiqieh, Rishad A. Shafik, Fei Xia 0001, Mohammed A. Noaman Al-Hayanni, Alexandre Yakovlev |
ICCD | 4 |
| 2020 | Current-Mode Carry-Free Multiplier Design using a Memristor-Transistor Crossbar ArchitectureabstractMultipliers are a major energy and delay contributor in modern compute-intensive applications due to their complex logic architecture. As such, designing multipliers with reduced energy and faster speed has remained a thoroughgoing challenge. This paper presents a novel, carry-free multiplier, which is suitable for a new-generation of energy-constrained applications. The multiplier circuit consists of an array of memristor-transistor cells that can be selected (i.e., turned ON or OFF) using a combination of DC bias voltages based on the operand values. When a cell is selected it contributes to current in the array path, which is then amplified by current mirrors with variable transistor gate sizes. The different current paths are connected to a node for analogously accumulating the currents to produce the multiplier output directly. This removes the need for latency-sensitive carry propagation stages, typically seen in traditional multipliers. We conduct a number of experiments to validate the functional and parametric properties. Our experiments showed that proposed multiplier achieves 51.44% savings in energy at a similar accuracy when compared with recently proposed approaches. Shengqi Yu, Ahmed Soltan, Rishad A. Shafik, Thanasin Bunnam, Fei Xia 0001, Domenico Balsamo, Alexandre Yakovlev |
DATE | 5 |
| 2020 | Dynamics of Time-Domain Power-Elastic Circuits for Pervasive Machine LearningabstractTime-domain data encoding, in the form of the duty cycle of a pulse width modulated (PWM) signal, has recently shown promising ways of building Machine Learning (ML) circuits. As the temporal signals approximately retain their “pseudo-analog” capacitive charging rates under voltage/ frequency variations, the circuits designed are inherently power elastic, offering the crucial leverage of energy autonomy for pervasive applications. This paper focuses on the analysis of dynamic parametric variations and their impact on the temporally encoded Machine Learning circuits. The aim is to investigate and suitably optimize these parameters for robustness, power elasticity and energy efficiency. Our study of dynamics includes how the selection of passive (R and C) components affects the dynamic range of operating frequency, which we term as “PWM carrier frequency”. We investigate how RC values define the performance and energy in terms of computation latency and energy per operation. Additionally, we demonstrates how the dynamic range of voltage and frequency variations affect functional and non-functional parameters of the PWM-based neural network solutions. Sergey Mileiko, Thanasin Bunnam, Fei Xia 0001, Rishad A. Shafik, Alexandre Yakovlev |
ISCAS | 3 |
| 2020 | PARMA: Parallelization-Aware Run-Time Management for Energy-Efficient Many-Core SystemsabstractPerformance and energy efficiency considerations have shifted computing paradigms from single-core to many-core architectures. At the same time, traditional speedup models such as Amdahl's Law face challenges in the run-time reasoning for system performance and energy efficiency, because these models typically assume limited variations of the parallel fraction. Moreover, the parallel fraction, which varies dynamically in workloads, is generally unknown at run-time without application-level instrumentation. This article describes novel performance/energy trade-off models based on realistic architectural considerations, which describe the parallel fraction and speedup as functions of performance counter values available in modern processors, removing the need for application-level instrumentation. These are then used to develop a Parallelization-Aware Run-time Management (PARMA) approach. PARMA aims at controlling core allocations and operating voltage/frequency points for energy efficiency, according to the varying workload parallel fractions. The efficacy of our models and the PARMA approach is extensively validated using a number of PARSEC benchmark applications, involving two performance/energy trade-off metrics: energy-delay-product (EDP), typically used in high-performance applications and energy per instruction (EPI), suitable for energy-aware applications. Up to 48 and 68 percent improvements in EDP and EPI have been observed using the PARMA approach compared with parallelization-agnostic methods. Mohammed A. Noaman Al-Hayanni, Ashur Rafiev, Fei Xia 0001, Rishad A. Shafik, Alexander B. Romanovsky, Alexandre Yakovlev |
IEEE Trans. Computers | 3 |
| 2016 | Low power voltage sensing through capacitance to digital conversionabstractCapacitance sensors are widely used for sensing physical parameters. Conventional capacitance to digital methods use complex analog ADC techniques which are power hungry. Recently a fully digital solution was proposed with improved power consumption. This paper describes a number of problems in that solution, analyzes these problems, and proposes a new design free of these problems. A voltage senor as an example was designed based on the proposed capacitance to digital conversion in this paper. The new method achieves the same accuracy with less than half the circuit size, and 25% and 33% savings on power and energy consumption. Delong Shang, Yuqing Xu, Kaiyuan Gao, Fei Xia 0001, Alexandre Yakovlev |
DDECS | 4 |
| 2016 | Selective abstraction and stochastic methods for scalable power modelling of heterogeneous systemsabstractWith the increase of system complexity in both platforms and applications, power modelling of heterogeneous systems is facing grand challenges from the model scalability issue. To address these challenges, this paper studies two systematic methods: selective abstraction and stochastic techniques. The concept of selective abstraction via black-boxing is realised using hierarchical modelling and cross-layer cuts, respecting the concepts of boxability and error contamination. The stochastic aspect is formally underpinned by Stochastic Activity Networks (SANs). The proposed method is validated with experimental results from Odroid XU3 heterogeneous 8-core platform and is demonstrated to maintain high accuracy while improving scalability. Ashur Rafiev, Fei Xia 0001, Alexei Iliasov, Rem Gensh, Ali Aalsaud, Alexander B. Romanovsky, Alexandre Yakovlev |
FDL | 2 |
| 2016 | Power-Aware Performance Adaptation of Concurrent Applications in Heterogeneous Many-Core SystemsabstractModern embedded systems execute multiple applications, both sequentially and concurrently. These applications are exercised on heterogeneous platforms generating varying power consumption and system workloads (CPU or memory intensive or both). As a result, determining the most energy-efficient system configuration (i.e. the number of parallel threads, their core allocations and operating frequencies) tailored for each kind of workload and application scenario is extremely challenging. In this paper, we propose a novel runtime optimization approach with the aim of achieving maximized power normalized performance considering dynamic variation of workload and application scenarios. Fundamental to this approach is a comprehensive study to investigate the tradeoffs between inter-application concurrency with performance and power consumption under different system configurations. Using real experimental measurements on an Odroid XU-3 heterogeneous platform with a number of PARSEC benchmark applications, we model power normalized performance (in terms of IPS/Watt) underpinning analytical power and performance models, derived through multivariate linear regression (MLR). Using these models, we show that with increasing number of concurrent CPU intensive applications show variable gains in IPS/Watt compared to the memory intensive applications in both sequential and concurrent application scenarios. Furthermore, we demonstrate that it is possible to continuously adapt system configuration through a low-cost and linear-complexity runtime algorithm, which can improve the IPS/Watt by up to 125% compared to the existing approach. Ali Aalsaud, Rishad A. Shafik, Ashur Rafiev, Fei Xia 0001, Sheng Yang 0003, Alexandre Yakovlev |
ISLPED | 4 |
| 2015 | A Formal Specification and Prototyping Language for Multi-core System ManagementabstractWe relate the experience of a defining a formal domain specific language (DSL) for the construction and reasoning about OS-level management logic of multi-core systems. The approach is based on a novel, iterative development principle where results of prototyping studies feed back into the next language revision. We illustrate the DSL with several examples of executable scripts. Alexei Iliasov, Ashur Rafiev, Fei Xia 0001, Rem Gensh, Alexander B. Romanovsky, Alexandre Yakovlev |
PDP | 3 |
| 2014 | Asynchronous design for new on-chip wide dynamic range power electronicsabstractAsynchronous circuits will play an important role in microelectronic systems in the future, especially in energy harvesting and autonomous (EHA) systems where such circuits will be able to offer robustness and deliver high efficiency in a wide range of power-energy conditions. The concept of Capacitor Bank Block (CBB) mechanisms was proposed to form the basis of electronics for powering asynchronous loads. These mechanisms will benefit EHA systems by enabling effective co-scheduling of computational tasks and energy supply. This paper demonstrates how the CBB mechanisms can themselves be controlled by asynchronous circuits, thereby forming a new type of power delivery units (PDU) that will be able to deliver power to intelligent digital logic in future EHA systems. These PDUs are superior to traditional power converters largely because the latter can only regulate sufficiently high power and energy levels (regular and periodic) as well as their controllers require stable power levels themselves. This makes them unsuitable for intermittent and sporadic conditions inherent to EHA systems. In this paper, a novel asynchronous control for the CBB is described. Experiments and analysis of the new PDUs, comprising CBBs and asynchronous control, are presented and discussed in detail. Delong Shang, Xuefu Zhang, Fei Xia 0001, Alexandre Yakovlev |
DATE | 3 |
| 2014 | Asynchronously assisted FPGA for variabilityabstractThe effect of variability has become increasingly significant as a result of technology geometry scaling. This paper describes Asynchronous Assisting Logic (AAL) blocks and the method of introducing them into modern FPGA architecture, in order to increase tolerance of the wide range latency variations caused by parametric variation, and temperature and supply voltage fluctuations. The proposed method leverages the availability of variation maps and suggests deploying configurable AAL blocks only into the variation critical paths - reinforcing rather rerouting/remapping. This method reduces the size overhead significantly which normally will be incurred by fully asynchronous designs. The proposed technique maintains the existing FPGA architecture allowing potential reuse of design flow. Simulations show correct functionality given regularly variable, randomly variable and capacitor switching energy harvester voltage supplies. Hock Soon Low, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
FPL | 3 |
| 2014 | Modeling and Tools for Power Supply Variations Analysis in Networks-on-ChipabstractPower supply integrity has become a critical concern with the rapid shrinking feature size and the ever increasing power consumption in nanometre scale integration. In particular, on-chip communication in platforms such as networks-on-chip (NoC) dictates the power dissipation and overall system performance in multicore systems and embedded computing architectures. These architectures require a dedicated tool for analyzing the power supply noise which must embed distinctive communication characteristics and spatial parameters. In this paper, we present a tool dedicated to determining the on-chip VDDdrops due to communication workload in NoCs. This tool integrates a fast power grid model, an NoC simulator, an on-chip link model, and a microarchitectural power model for router. The model has been rigorously verified using SPICE simulations. The proposed model and tools are further exemplified through analyzing the impact of power supply noise for NoC links. Statistical timing analysis of NoC links in the presence of power supply noise was performed to evaluate the bit error rates (BERs). This work would enable better understanding of the tradeoffs existing in the design of NoCs, and the induced power supply noise due to on-chip communication. This understanding is crucial for the analysis of the quality of service (QoS) of communication fabrics in NoCs at the early design stages. Nizar Dahir, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev |
IEEE Trans. Computers | 3 |
| 2013 | Wide-range, reference free, on-chip voltage sensor for variable Vdd operationsabstractIn future systems with relatively unreliable and unpredictable energy sources such as harvesters, the system Vdd may become non-deterministic. Reliable and accurate on-chip voltage sensors are therefore indispensible for the power and computation management of such systems. Stable and known references are also difficult to obtain in this environment. This paper describes a reference-free voltage sensor implemented using a speed independent (SI) SRAM cell and an inverter chain. It can work under a wide range of Vdd, and provides accurate measurements of Vdd over this operating range with a precision range from 50mV to 10mV. Unlike existing methods, the voltage information is directly generated as a digital code without any analog circuits. This is realized by exploiting the inherently different latency behaviors of different types of circuits under different Vdd. Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 2 |
| 2013 | Dynamic On-Chip Thermal Optimization for Three-Dimensional Networks-On-ChipabstractThe complex thermal behaviour prohibits the advancement of three-dimensional (3D) very-large-scale integration system. Particularly, the high-density through-silicon via based integration could lead to ultra-high temperature hotspots and permanent silicon device damage. In this paper, we introduce an adaptive strategy to effectively diffuse heat throughout the 3D geometry. This strategy employs a dynamic programming network to select and optimize the direction of data manoeuvre in a network-on-chip (NoC). We also developed a tool, which is based on the accurate HotSpot thermal model and SystemC cycle accurate model, to simulate the thermal system and evaluate our approach. We found that the proposed approach can significantly diffuse the hotspots from a 3D geometry and overall temperature can be significantly reduced. Given the same thermal constraints, the throughput performance of an adaptive NoC can also be improved. This work enables a new avenue to explore the on-chip adaptability for the future large-scale 3D integration. Ra'ed Al-Dujaily, Terrence S. T. Mak, Kai-Pui Lam, Fei Xia 0001, Alexandre Yakovlev, Chi-Sang Poon |
Comput. J. | 4 |
| 2013 | Concurrent Multiresource Arbiter: Design and ApplicationsabstractThis paper presents a novel type of asynchronous arbiter that allocates M interchangeable resources among N clients. This arbiter enables the concurrent utilization of multiple resources and is a useful device in various load-balancing circuits. Dedicated request signals from the resources and the clients are used in pairs to form each new grant. The 2 × 2 arbiter is examined as an accessible special case of the N × M arbiter. A concurrent implementation is compared to fully sequential design. It is shown that the sequential design can be more practical when the time between a grant and the withdrawal of the initial request is small. The concurrent design provides higher performance in a system with a longer resource utilization time. A scalable tiled structure is developed to extend the arbiter structure beyond 2 × 2 to support N clients and M resources. Models and subsequent implementations of the tiles are presented. The tiles can be repeated without the use of additional connecting logic, enabling the construction of arbiters of larger sizes. Several examples demonstrate the usage of the arbiter. Stanislavs Golubcovs, Delong Shang, Fei Xia 0001, Andrey Mokhov, Alexandre Yakovlev |
IEEE Trans. Computers | 3 |
| 2013 | Dynamic programming-based runtime thermal management (DPRTM): An online thermal control strategy for 3D-NoC systemsabstractComplex thermal behavior inhibits the advancement of three-dimensional (3D) very-large-scale-integration (VLSI) system designs, as it could lead to ultra-high temperature hotspots and permanent silicon device damage. This article introduces a new runtime thermal management strategy to effectively diffuse and manage heat throughout 3D chip geometry for a better throughput performance in networks on chip (NoC). This strategy employs a dynamic programming-based runtime thermal management (DPRTM) policy to provide online thermal regulation. Reactive and proactive adaptive schemes are integrated to optimize the routing pathways depending on the critical temperature thresholds and traffic developments. Also, when the critical system thermal limit is violated, an urgent throttling will take place. The proposed DPRTM is rigorously evaluated through cycle-accurate simulations, and results show that the proposed approach outperforms conventional approaches in terms of computational efficiency and thermal stability. For example, the system throughput using the DPRTM approach can be improved by 33% when compared to other adaptive routing strategies for a given thermal constraint. Moreover, the DPRTM implementation presented in this article demonstrates that the hardware overhead is insignificant. This work opens a new avenue for exploring the on-chip adaptability and thermal regulation for future large-scale and 3D many-core integrations. Ra'ed Al-Dujaily, Nizar Dahir, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2012 | Ultra-low power transmitterabstractThis paper presents a design of an ultra-low power UWB transmitter based on 4thand 5thderivative Gaussian pulse shapes implemented in UMC 90nm CMOS technology. The simulations show 119mV peak to peak pulse amplitude and the pulse width of 240 ps for the 5thderivative Gaussian pulse and 99.71mV pulse amplitude and 190 ps pulse width for the 4thderivative Gaussian pulse. Power consumption of the pulse generators are calculated 30.11 uW and 21.5 uW for the 5thand 4thderivative Gaussian pulse respectively at a 100MHz pulse repeating frequency (PRF). Ultra-low power radio transmission is important in such application contexts as wireless network nodes and sensors powered by energy harvesters. Mohsen Ghasempour, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 3 |
| 2012 | Embedded Transitive Closure Network for Runtime Deadlock Detection in Networks-on-ChipabstractInterconnection networks with adaptive routing are susceptible to deadlock, which could lead to performance degradation or system failure. Detecting deadlocks at runtime is challenging because of their highly distributed characteristics. In this paper, we present a deadlock detection method that utilizes runtime transitive closure (TC) computation to discover the existence of deadlock-equivalence sets, which imply loops of requests in networks-on-chip (NoCs). This detection scheme guarantees the discovery of all true deadlocks without false alarms in contrast with state-of-the-art approximation and heuristic approaches. A distributed TC-network architecture, which couples with the NoC infrastructure, is also presented to realize the detection mechanism efficiently. Detailed hardware realization architectures and schematics are also discussed. Our results based on a cycle-accurate simulator demonstrate the effectiveness of the proposed method. It drastically outperforms timing-based deadlock detection mechanisms by eliminating false detections and, thus, reducing energy wastage in retransmission for various traffic scenarios including real-world application. We found that timing-based methods may produce two orders of magnitude more deadlock alarms than the TC-network method. Moreover, the implementations presented in this paper demonstrate that the hardware overhead of TC-networks is insignificant. Ra'ed Al-Dujaily, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev, Maurizio Palesi |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2011 | Run-time deadlock detection in networks-on-chip using coupled transitive closure networksabstractInterconnection networks with adaptive routing are susceptible to deadlock, which could lead to performance degradation or system failure. Detecting deadlocks at run-time is challenging because of their highly distributed characteristics. In this paper, we present a deadlock detection method that utilizes run-time Transitive Closure (TC) computation to discover the existence of deadlock-equivalence sets, which imply loops of requests in networks-on-chip (NoC). This detection scheme guarantees the discovery of all true deadlocks without false alarms unlike state-of-the-art approximation and heuristic approaches. A distributed TC-network architecture which couples with the NoC architecture is also presented to realize the detection mechanism efficiently. Our results based on a cycle-accurate simulator demonstrate the effectiveness of the TC-network method. It drastically outperforms timing-based deadlock detection mechanisms by eliminating false detections and thus reducing energy dissipation in various traffic scenarios. For example, timing based methods may produce two orders of magnitude more deadlock alarms than the TC-network method. Moreover, the implementations presented in this paper demonstrate that the hardware overhead of TC-networks is insignificant. Ra'ed Al-Dujaily, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev, Maurizio Palesi |
DATE | 3 |
| 2011 | Variation tolerant asynchronous FPGA (abstract only)abstractThis paper describes the realization of an interconnect Delay Insensitive (DI) FPGA architecture with distributed asynchronous control. This architecture maintains the basic block structure of traditional FPGAs allowing the potential use of existing FPGA design tools in block design. This asynchronous FPGA architecture is mainly aimed at tolerating the unpredictable delay variations caused by process and environment variations in current and future VLSI technology nodes and also targets low power operations, including modes such as dynamic voltage scaling and variable Vdd, as in applications featuring energy harvesting. This is achieved by making the longer inter-block interconnects DI, keeping the computational logic single-rail, and removing global clocks. Hock Soon Low, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
FPGA | 3 |
| 2011 | A Novel Power Delivery Method for Asynchronous Loads in Energy Harvesting SystemsabstractFor systems depending on power harvesting, a fundamental contradiction in the power delivery chain has existed between conventional synchronous computational loads requiring relatively stable Vdd and power harvesters unable to supply it. DC/DC conversion has therefore been an integral part of such systems to resolve this contradiction. On the other hand, asynchronous computational loads, in addition to their potential power-saving capabilities, can be made tolerant to a much wider range of Vdd variance. This may open up opportunities for much more energy efficient methods of power delivery. This article presents in-depth investigations into the behavior and performance of different on-chip power delivery methods driving both asynchronous and synchronous loads directly from a harvester source. A novel power delivery method, which employs a capacitor bank for adaptively storing the energy from power harvesters depending on load and source conditions, is developed. Its advantages, especially when driving asynchronous loads, are demonstrated through comprehensive comparative analysis. Xuefu Zhang, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2010 | A Reconfigurable Hebbian Eigenfilter for Neurophysiological Spike Train AnalysisabstractThe emergence of multi-electrode array enables the study of real-time neurophysiological activities across multiple regions of the brain. However, the real-time extracellular action potentials recorded on any electrode represent the simultaneous electrical activity of an unknown number of neurons which present a critical challenge to the accuracy of interpretation and identification of the neural circuitry in the subsequent analysis. In this paper, we present a principal component analysis approach utilizing Hebbian eigenfilter to identify the corresponding electrical activities of each neuron, namely spike sorting. The Hebbian eigenfilter greatly simplifies the computational complexity of eigen-projection. An efficient FPGA-based Hebbian eigenfilter is proposed. The performance, accuracy and power consumption of our Hebbian eigenfilter are thoroughly evaluated through synthetic spike trains. The proposal enables real-time spike sorting and analysis, and leads the way towards future motor and cognitive neuroprosthetics. Bo Yu 0014, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev, Yihe Sun, Chi-Sang Poon |
FPL | 4 |
| 2010 | Stochastic analysis of power, latency and the degree of concurrencyabstractConcurrent processing has become the default mode of operation in on-chip systems. Silicon has become cheap enough for having hardware facilities to support very large scale concurrent processing on chip. As a result the availability and applicability of power is becoming more of a limiting factor than logic. However, the advantage of parallelism in reducing power consumption will soon become unrealistic because of the limited scope of reducing Vdd beyond threshold voltage, leaving the reduction of concurrency (through the partial shut-down of system blocks) as a realistic means of reducing power consumption when needed. A stochastic modelling approach is presented in this paper which can integrate the degree of concurrency as a parameter into power and latency analysis. This will facilitate a system design and management regime where the degree of concurrency is used as a means of control to achieve power and performance goals. Yuan Chen 0002, Isi Mitrani, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 4 |
| 2010 | Asynchronous FPGA architecture with distributed controlabstractAsynchronous techniques have become more significant with continued scaling of VLSI technologies. This paper proposes an asynchronous FPGA architecture. Different from previous methods of introducing asynchrony into FPGAs, our method seeks to preserve the current FPGA cell structure as much as possible, whilst achieving delay insensitivity in the inter-cell interconnects. By using David Cells as the central technique in the delay insensitive clock replacement, this method is conducive to the establishment of an automatic design and synthesis flow. It also particularly caters for low power designs, where current FPGA solutions are not effective yet. Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 2 |
| 2010 | Highly parallel multi-resource arbitersabstractMulti-resource multi-client arbiters are becoming more important in on-chip systems because of the increasing significance of dynamic, run-time, allocation of various system performance resources such as power and computation and communication facilities. Arbiters, for example, can be used to limit the amount of concurrency for regulating voltage droops, and for balancing load and traffic. This paper describes the design of multi-resource arbiters with high degrees of concurrency. By using freezing logic, this design method guarantees correct computation whilst simplifies the implementation. Quick release mechanisms and the implementation of the multi-token concept through the duplication of the client requests help improve the efficiency. Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 2 |
| 2007 | Automating Synthesis of Asynchronous Communication Mechanisms
Kyller Costa Gorgônio, Jordi Cortadella, Fei Xia 0001, Alexandre Yakovlev |
Fundam. Informaticae | 3 |
| 2006 | Low-Cost Online Testing of Asynchronous HandshakesabstractA new low-cost low-complexity checker for online testing of asynchronous interfaces in globally-asynchronous locally-synchronous circuits is proposed. The solution is fully based upon the standard gate libraries. The checker itself is fully offline testable. It also provides a fault-locating functionality, which is achieved by combining the online mode with scan techniques Delong Shang, Alexandre Yakovlev, Frank P. Burns, Fei Xia 0001, Alexandre V. Bystrov |
ETS | 4 |
| 2006 | Virtual self-timed blocks for systems-on-chipabstractIntellectual properties (IP cores) are widely used as pre-designed and reusable units in various system-on-chip (SOC) designs, but their integration has presented difficulties for system designers. In this paper, we propose an approach to better reuse IP cores while maintain energy efficiency for SOC systems. Here we employ a so called self-timed event processor (STEP) to make each IP core into a virtual self-timed block. Much of the IP cores' pre-designed properties can be preserved and the new SOC systems that use virtual self-timed blocks can be more energy efficient. A MATLAB based investigation is carried out on an example STEP processor. Yuan Chen 0002, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 2 |
| 2006 | Buffered Asynchronous Communication Mechanisms
Fei Xia 0001, Ian G. Clark, Alexandre Yakovlev, E. Graeme Chester |
Fundam. Informaticae | 1 |
| 2004 | MATLAB Models of ACMS in Control Systems
Fei Xia 0001, E. Graeme Chester, Alexandre Yakovlev, Ian G. Clark |
ICINCO (3) | 2 |
| 2002 | Algorithms for Signal and Message Asynchronous Communication Mechanisms and their Analysis
Fei Xia 0001, Ian G. Clark |
Fundam. Informaticae | 1 |
| 1994 | A Parallel Simulation of Multiple Mobile Robots Using the DORIS Design MethodabstractThis paper introduces the data oriented requirements implementation scheme (DORIS) design strategy and the data interaction architecture demonstrator (DIADEM), by describing their use for a multiple-robot workspace simulation. The design process, the hardware and the software for the DIADEM are outlined, followed by a discussion of the design used for the simulation.> Fei Xia 0001, Sergio A. Velastin, Anthony C. Davies |
ICRA | 1 |