EDBT 2026 Demo / reviewers in the wild / expert
Alexandre Yakovlev
dblp:83/4837 · also Alex Yakovlev
· DBLP profile ↗
179ranked-venue papers
12as first author
31since 2021 · last 2026
0000-0003-0826-9330ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 136 · 8 first-author · 19 since 2021Software engineering, systems software and programming languages · 43 · 2 first-author · 9 since 2021Theory of computation · 17 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ANIMATE: Automated Framework for Scalable Design of Tsetlin Machines Using 1-Safe Petri Nets
Alex Chan, Mohamed Tarraf, Rishad A. Shafik, Alexandre Yakovlev |
PETRI NETS | 4 |
| 2026 | AURORA - AUtomated 8T SRAM Wired-OR Logic Array for Boolean-Based Machine Learning
Komal Krishnamurthy, Marcos L. L. Sartori, Shengyu Duan, Alexandre Yakovlev, Rishad A. Shafik |
DATE | 4 |
| 2026 | Learning dynamics, pattern recognition capability and interpretability of the Tsetlin MachineabstractThe inability to trace an AI’s reasoning process and understand why it makes each decision is known as the black box problem. This remains one of the major barriers to the trusted and widespread use of machine learning in many application domains. The paper explores pattern recognition performance and learning dynamics of the Tsetlin Machine – a new explainable logic-based machine-learning approach. Tsetlin Machine uses a collection of finite-state automata with a unique logic-based learning mechanism and provides a promising alternative to Artificial Neural Networks with several advantages, such as interpretability, low complexity, suitability for hardware implementation and high performance. This work investigates Tsetlin Machine’s mechanism for constructing conjunctive clauses from data and their interpretation for pattern recognition on several datasets. We demonstrate that during training the logical clauses learn persistent sub-patterns within the class. Each clause creates a class template by clustering a certain number of similar class samples, combining them through literal-wise logical conjunction (i.e., AND-ing). The number of class samples that each clause combines depends on Tsetlin Machine’s hyperparameters. The more class samples that are combined, the more general the clauses become. The paper aims at uncovering how Tsetlin Machine’s hyperparameters influence the balance between clause generalization and specialization and how this affects the accuracy of pattern recognition. It also studies the evolution of the machine’s internal state, its convergence and training completion. Olga Tarasyuk, Anatoliy Gorbenko, Tousif Rahman, Lei Jiao 0001, Ole-Christoffer Granmo, Rishad A. Shafik, Alexandre Yakovlev |
Pattern Recognit. | 7 |
| 2026 | An All-Digital 8.6-nJ/Frame 65-nm Tsetlin Machine Image Classification AcceleratorabstractWe present an all-digital programmable machine learning accelerator chip for image classification, underpinning on the Tsetlin machine (TM) principles. The TM is an emerging machine learning algorithm founded on propositional logic, utilizing sub-pattern recognition expressions called clauses. The accelerator implements the coalesced TM version with convolution, and classifies booleanized images of$28\times 28$pixels with 10 categories. A configuration with 128 clauses is used in a highly parallel architecture. Fast clause evaluation is achieved by keeping all clause weights and Tsetlin automata (TA) action signals in registers. The chip is implemented in a 65 nm low-leakage CMOS technology, and occupies an active area of 2.7 mm2. At a clock frequency of 27.8 MHz, the accelerator achieves 60.3 k classifications per second, and consumes 8.6 nJ per classification. This demonstrates the energy-efficiency of the TM, which was the main motivation for developing this chip. The latency for classifying a single image is$25.4~\mu $s which includes system timing overhead. The accelerator achieves 97.42%, 84.54% and 82.55% test accuracies for the datasets MNIST, Fashion-MNIST and Kuzushiji-MNIST, respectively, matching the TM software models. Svein Anders Tunheim, Yujin Zheng, Lei Jiao 0001, Rishad A. Shafik, Alexandre Yakovlev, Ole-Christoffer Granmo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | Distributed Places and Safe Net Reduction
Victor Khomenko, Maciej Koutny, Alexandre Yakovlev |
Petri Nets | 3 |
| 2025 | Dynamic Tsetlin Machine Accelerators for On-Chip Training Using FPGAsabstractThe increased demand for data privacy and security in machine learning (ML) applications has put impetus on effective edge training on Internet-of-Things (IoT) nodes. Edge training aims to leverage speed, energy efficiency and adaptability within the resource constraints of the nodes. Deploying and training Deep Neural Networks (DNNs)-based models at the edge, although accurate, posit significant challenges from the back-propagation algorithm’s complexity, bit precision trade-offs, and heterogeneity of DNN layers. This paper presents a Dynamic Tsetlin Machine (DTM) training accelerator as an alternative to DNN implementations. DTM utilizes logic-based on-chip inference with finite-state automata-driven learning within the same Field Programmable Gate Array (FPGA) package. Underpinned on the Vanilla and Coalesced Tsetlin Machine algorithms, the dynamic aspect of the accelerator design allows for a run-time reconfiguration targeting different datasets, model architectures, and model sizes without resynthesis. This makes the DTM suitable for targeting multivariate sensor-based edge tasks. Compared to DNNs, DTM trains with fewer multiply-accumulates, devoid of derivative computation. It is a data-centric ML algorithm that learns by aligning Tsetlin automata with input data to form logical propositions enabling efficient Look-up-Table (LUT) mapping and frugal Block RAM usage in FPGA training implementations. The proposed accelerator offers 2.54x more Giga operations per second per Watt (GOP/s per W) and uses 6x less power than the next-best comparable design. Gang Mao, Tousif Rahman, Sidharth Maheshwari, Bob Pattison, Rishad A. Shafik, Alexandre Yakovlev |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | Tsetlin Machine-Based Image Classification FPGA Accelerator With On-Device TrainingabstractThe Tsetlin Machine (TM) is a novel machine learning algorithm that uses Tsetlin automata (TAs) to define propositional logic expressions (clauses) for classification. This paper describes a field-programmable gate array (FPGA) accelerator for image classification based on the Convolutional Coalesced Tsetlin Machine. The accelerator classifies booleanized images of$28\times 28$pixels into 10 classes, and is configured with 128 clauses in a highly parallel architecture. To achieve fast clause evaluation and class prediction, the TA action signals and the clause weights per class are available from registers. Full on-device training is included, and the TAs are implemented with 34 Block RAM (BRAM) instances which operate in parallel. Each BRAM is addressed by the clause number and has a 72-bit word width that supports 8 TAs. The design is implemented in a Xilinx Zynq Ultrascale+ XCZU7 FPGA. Running at 50 MHz, the accelerator core achieves 134k image classifications per second, with an energy consumption per classification of$13.3~\mu $J. A single training epoch of 60k samples requires a processing time of 1.5 seconds. The accelerator obtains a test accuracy of 97.6% on MNIST, 84.1% on Fashion-MNIST and 82.8% on Kuzushiji-MNIST. Svein Anders Tunheim, Lei Jiao 0001, Rishad A. Shafik, Alexandre Yakovlev, Ole-Christoffer Granmo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Design of Event-Driven Tsetlin Machines Using Safe Petri Nets
Alex Chan, Adrian Wheeldon, Rishad A. Shafik, Alexandre Yakovlev |
Petri Nets | 4 |
| 2024 | Bridging the Design Methodologies of Burst-Mode Specifications and Signal Transition GraphsabstractAsynchronous circuits are a promising type of digital circuit that still see moderate usage in today’s commercial products, which has often been linked to the adaptation challenges that are posed within industry, e.g. time required to develop new tools and train designers versus using existing synchronous tools to quickly meet market demands. Several formal models were introduced to aid with the design of asynchronous circuits, including Burst-Mode (BM) Specifications and Signal Transition Graphs (STGs). BM specifications resemble synchronous Finite State Machines (FSMs) allowing circuit designers to easily adapt and use them, however their circuit implementations may be limited due to declining tool support. STGs have access to well-established tools that produce optimal hazard-free circuit implementations, but they are seen as too different by the industry. In this paper, we present a new ‘co-design’ methodology that bridges the gap between BM specifications and STGs by using a formal model called Burst Automaton (BA). BA is a generic FSM-like model that acts as a framework for enabling interoperability between many different formal models, and offers several benefits that BM specifications and STGs can leverage. Our ‘co-design’ methodology is implemented in Workcraft, and is evaluated on several benchmarks showing an improved synthesis flow. Alex Chan, Danil Sokolov, Victor Khomenko, Alexandre Yakovlev |
ASPDAC | 4 |
| 2024 | MATADOR: Automated System-on-Chip Tsetlin Machine Design Generation for Edge ApplicationsabstractSystem-on-Chip Field-Programmable Gate Arrays (SoC-FPGAs) offer significant throughput gains for machine learning (ML) edge inference applications via the design of co-processor accelerator systems. However, the design effort for training and translating ML models into SoC-FPGA solutions can be substantial and requires specialist knowledge aware trade-offs between model performance, power consumption, latency and resource utilization. Contrary to other ML algorithms, Tsetlin Machine (TM) performs classification by forming logic proposition between boolean actions from the Tsetlin Automata (the learning elements) and boolean input features. A trained TM model, usually, exhibits high sparsity and considerable overlapping of these logic propositions both within and among the classes. The model, thus, can be translated to RTL-level design using a miniscule number of AND and NOT gates. This paper presents MATADOR, an automated boolean-to-silicon tool with GUI interface capable of implementing optimized accelerator design of the TM model onto SoC-FPGA for inference at the edge. It offers automation of the full development pipeline: model training, system level design generation, design verification and deployment. It makes use of the logic sharing that ensues from propositional overlap and creates a compact design by effectively utilizing the TM model's sparsity. MATADOR accelerator designs are shown to be up to 13.4x faster, up to 7x more resource frugal and up to 2x more power efficient when compared to the state-of-the-art Quantized and Binary Deep Neural Network implementations. Tousif Rahman, Gang Mao, Sidharth Maheshwari, Rishad A. Shafik, Alexandre Yakovlev |
DATE | 5 |
| 2024 | An Event-Driven Approach to Genotype Imputation on a Custom RISC-V ClusterabstractThis article proposes an event-driven solution to genotype imputation, a technique used to statistically infer missing genetic markers in DNA. The work implements the widely accepted Li and Stephens model, primary contributor to the computational complexity of modern x86 solutions, in an attempt to determine whether further investigation of the application is warranted in the event-driven domain. The model is implemented using graph-based Hidden Markov Modeling and executed as a customized forward/backward dynamic programming algorithm. The solution uses an event-driven paradigm to map the algorithm to thousands of concurrent cores, where events are small messages that carry both control and data within the algorithm. The design of a single processing element is discussed. This is then extended across multiple cores and executed on a custom RISC-V NoC cluster called POETS. Results demonstrate how the algorithm scales over increasing hardware resources and a multi-core run demonstrates a 270X reduction in wall-clock processing time when compared to a single-threaded x86 solution. Optimisation of the algorithm via linear interpolation is then introduced and tested, with results demonstrating a wall-clock reduction time of ∼ 5 orders of magnitude when compared to a similarly optimised x86 solution. Jordan Morris, Ashur Rafiev, Graeme M. Bragg, Mark Vousden, David B. Thomas, Alexandre Yakovlev, Andrew D. Brown |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2023 | A Rapid Reset 8-Transistor Physically Unclonable Function Utilising Power GatingabstractPhysically Unclonable Functions (PUFs) need error correction whilst regenerating Secret Keys in cryptography. The proposed 8-Transistor (8T) PUF, which coordinates with the power gating technique, can significantly accelerate a single evaluation cycle 1000 times faster than$\mathbf{6}\mathbf{T}$-SRAM PUF does with a 12.8% area increase. This design enables multiple evaluations even in the key regeneration phase in field, hence greatly reducing the number of errors and the hardware penalty for error correction. The$\mathbf{8T}$PUF derives from the$\mathbf{6}\mathbf{T}$SRAM. It is built to eliminate data retention swiftly and maximise physical mismatches. And a two-phase power gating module is designed to provide controllable power-on/off cycles rapidly for the chosen PUF clusters in order to facilitate statistical measurements and curb the in-rush current, thereby enhancing PUF entropy and security. An architecture of the power-gated PUF is developed to accommodate fast multiple evaluations. Post-layout Monte Carlo simulations were performed with Cadence, and the extracted PUF Responses were processed with Matlab to evaluate the 8T PUF performance and statistical metrics for subsequent inclusion into PUF Responses. Yujin Zheng, Alexandre V. Bystrov, Alexandre Yakovlev |
DATE | 3 |
| 2023 | Asynchronous Control for Tsetlin Machine with Binary Memristor-Transistor ArrayabstractTsetlin machines (TMs) are a novel machine learning paradigm based on learning automata and Boolean logic inference, with better energy-efficiency and explainability than neural networks. This work exploits non-volatile ReRAM-transistor memory arrays to perform efficient in-memory TM computing. To accommodate the large timing variability of ReRAM devices and enhance energy efficiency and speed, the control path is implemented with quasi delay-insensitive (QDI) asynchronous circuits. The design of these circuits are derived and synthesized from their signal-transition graph specifications using the Workcraft tool. The resulting circuits offer high event-driven controllability and high variation tolerance for the mixed-signal ReRAM data path. Compared to state of the art TM hardware, the new TM design uses less than 5% of the power to achieve better than$4\times$the performance. Omar Ghazal, Gang Mao, Jesse Ojukwu, Fei Xia 0001, Alexandre Yakovlev, Rishad A. Shafik |
ISCAS | 6 |
| 2023 | IMBUE: In-Memory Boolean-to-CUrrent Inference ArchitecturE for Tsetlin MachinesabstractIn-memory computing for Machine Learning (ML) applications remedies the von Neumann bottlenecks by organizing computation to exploit parallelism and locality. Non-volatile memory devices such as Resistive RAM (ReRAM) offer integrated switching and storage capabilities showing promising performance for ML applications. However, ReRAM devices have design challenges, such as nonlinear digital-analog conversion and circuit overheads. This paper proposes an In-Memory Boolean-to-Current Inference Architecture (IMBUE) that uses ReRAM-transistor cells to eliminate the need for such conversions. IMBUE processes Boolean feature inputs expressed as digital voltages and generates parallel current paths based on resistive memory states. The proportional column current is then translated back to the Boolean domain for further digital processing. The IMBUE architecture is inspired by the Tsetlin Machine (TM), an emerging ML algorithm based on intrinsically Boolean logic. The IMBUE architecture demonstrates significant performance improvements over binarized convolutional neural networks and digital TM in-memory implementations, achieving up to a 12.99x and 5.28x increase, respectively. Omar Ghazal, Simranjeet Singh, Tousif Rahman, Shengqi Yu, Yujin Zheng, Domenico Balsamo, Sachin B. Patkar, Farhad Merchant, Fei Xia 0001, Alexandre Yakovlev, Rishad A. Shafik |
ISLPED | 10 |
| 2023 | A multi-step finite-state automaton for arbitrarily deterministic Tsetlin Machine learningabstractAbstract Due to the high arithmetic complexity and scalability challenges of deep learning, there is a critical need to shift research focus towards energy efficiency. Tsetlin Machines (TMs) are a recent approach to machine learning (ML) that has demonstrated significantly reduced energy compared to neural networks alike, while providing comparable accuracy on several benchmarks. However, TMs rely heavily on energy‐costly random number generation to stochastically guide a team of Tsetlin Automata (TA) in TM learning. In this paper, we propose a novel finite‐state learning automaton that can replace the TA in the TM, for increased determinism. The new automaton uses multi‐step deterministic state jumps to reinforce sub‐patterns, without resorting to randomization. A determinism parameter finely controls trading off the energy consumption of random number generation, against randomization for increased accuracy. Randomization is controlled by flipping a coin before every state jump, ignoring the state jump on tails. For example, makes every update random and makes the automaton completely deterministic. Both theoretically and empirically, we establish that the proposed automaton converges to the optimal action almost surely. Further, used together with the TM, only substantial degrees of determinism reduce accuracy. Energy‐wise, random number generation constitutes switching energy consumption of the TM, saving up to 11 mW power for larger datasets with high values. Our new learning automaton approach thus facilitates low‐energy ML. Kuruge Darshana Abeyrathna, Ole-Christoffer Granmo, Rishad A. Shafik, Lei Jiao 0001, Adrian Wheeldon, Alexandre Yakovlev, Jie Lei 0007, Morten Goodwin |
Expert Syst. J. Knowl. Eng. | 6 |
| 2023 | Approximate digital-in analog-out multiplier with asymmetric nonvolatility and low energy consumptionabstractMany modern compute-intensive applications require arithmetic results (usually multiplication) to be represented as analog signals. Using digital multipliers followed by digital-to-analog conversion (DAC) results in high energy and performance costs. This is because digital multipliers have costly carry propagation, and DAC circuits add associated conversion costs. Another concern, especially for arithmetic on the edge, is the need for nonvolatile operands in the face of power uncertainty. To deal with this, nonvolatile memory technologies have been combined with in-memory computing. This paper proposes a mixed-signal multiplier which directly generates an analog product based on two digital input operands. Fundamental to the design are transistor-memristor cells, organized in a crossbar structure. Using analog resistive partial product accumulation in the crossbar, the approximate multiplier eliminates the need for carry propagation and an explicit DAC. It also provides asymmetric nonvolatility making memristor writing a rare event, extending the application significance of the method. The design is shown to be functionally correct up to 4-bit, and achieves 8× to over 300× speedup, competitive peak-power and orders of magnitude energy reduction, compared with existing full-digital memristor-based multipliers and low-power multiplication DAC solutions. Shengqi Yu, Fei Xia 0001, Rishad A. Shafik, Domenico Balsamo, Alexandre Yakovlev |
Integr. | 5 |
| 2023 | REDRESS: Generating Compressed Models for Edge Inference Using Tsetlin MachinesabstractInference at-the-edge using embedded machine learning models is associated with challenging trade-offs between resource metrics, such as energy and memory footprint, and the performance metrics, such as computation time and accuracy. In this work, we go beyond the conventional Neural Network based approaches to explore Tsetlin Machine (TM), an emerging machine learning algorithm, that uses learning automata to create propositional logic for classification. We use algorithm-hardware co-design to propose a novel methodology for training and inference of TM. The methodology, called REDRESS, comprises independent TM training and inference techniques to reduce the memory footprint of the resulting automata to target low and ultra-low power applications. The array of Tsetlin Automata (TA) holds learned information in the binary form as bits: {0,1}, called excludes and includes, respectively. REDRESS proposes a lossless TA compression method, called the include-encoding, that stores only the information associated with includes to achieve over 99% compression. This is enabled by a novel computationally minimal training procedure, called the Tsetlin Automata Re-profiling, to improve the accuracy and increase the sparsity of TA to reduce the number of includes, hence, the memory footprint. Finally, REDRESS includes an inherently bit-parallel inference algorithm that operates on the optimally trained TA in the compressed domain, that does not require decompression during runtime, to obtain high speedups when compared with the state-of-the-art Binary Neural Network (BNN) models. In this work, we demonstrate that using REDRESS approach, TM outperforms BNN models on all design metrics for five benchmark datasets viz. MNIST, CIFAR2, KWS6, Fashion-MNIST and Kuzushiji-MNIST. When implemented on an STM32F746G-DISCO microcontroller, REDRESS obtained speedups and energy savings ranging 5-5700× compared with different BNN models. Sidharth Maheshwari, Tousif Rahman, Rishad A. Shafik, Alexandre Yakovlev, Ashur Rafiev, Lei Jiao 0001, Ole-Christoffer Granmo |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Burst Automaton: Framework for Speed-Independent Synthesis Using Burst-Mode SpecificationsabstractBurst-mode (BM) formalism is a variant of an asynchronous finite-state machine (FSM) that operates in “BM” timing assumption and offers simple entry into the asynchronous circuit design. However, some of BM’s well-formedness properties, while useful for implementing BM specifications as circuits, are rather restrictive in some important contexts, e.g., BM’s maximal set property (or its analog, extended BM (XBM) formalism’s distinguishability constraint) forbids nondeterministic specifications that are inherent in some design approaches, input and output bursts must alternate meaning BMs are not a proper extension of FSMs with arcs labeled by single events, and BMs cannot express input-output concurrency whereas FSMs can with interleaving. The latter limitation is particularly problematic when interoperability between several formalisms is desirable. In this article, we propose the burst automation (BA) model that is more powerful and yet simpler than BM, by relaxing BM’s well-formedness properties. BA is a proper extension of FSMs, and can express input-output concurrency and nondeterminism. We define BA’s interleaving semantics via its asynchronous reachability graph that is an FSM, and develop three translations from BAs to signal transition graphs (STGs) that preserve strong bisimulation, weak bisimulation, or the language. Former two translations may be exponential, whereas the latter translation is linear. The resulting STG can then be used for verification and synthesis into speed-independent (SI) or quasi-delay-insensitive (QDI) circuits, or for composition with other STGs. The proposed workflow was implemented in Workcraft, and experimental results show an improved synthesis rate and a significant reduction in the literal count. Alex Chan, Danil Sokolov, Victor Khomenko, Alexandre Yakovlev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Avoiding Exponential Explosion in Petri Net Models of Control Flows
Victor Khomenko, Maciej Koutny, Alexandre Yakovlev |
Petri Nets | 3 |
| 2022 | Slimming down Petri Boxes: Compact Petri Net Models of Control Flows
Victor Khomenko, Maciej Koutny, Alexandre Yakovlev |
CONCUR | 3 |
| 2022 | Runtime Energy Minimization of Distributed Many-Core Systems using Transfer LearningabstractThe heterogeneity of computing resources continues to permeate into many-core systems making energy-efficiency a challenging objective. Existing rule-based and model-driven methods return sub-optimal energy-efficiency and limited scalability as system complexity increases to the domain of distributed systems. This is exacerbated further by dynamic variations of workloads and quality-of-service (QoS) demands. This work presents a QoS-aware runtime management method for energy minimization using a transfer learning (TL) driven exploration strategy. It enhances standard Q-learning to improve both learning speed and operational optimality (i.e., QoS and energy). The core to our approach is a multi-dimensional knowledge transfer across a task's state-action space. It accelerates the learning of dynamic voltage/frequency scaling (DVFS) control actions for tuning power/performance trade-offs. Firstly, the method identifies and transfers already learned policies between explored and behaviorally similar states referred to as Intra-Task Learning Transfer (ITLT). Secondly, if no similar “expert” states are available, it accelerates exploration at a local state's level through what's known as Intra-State Learning Transfer (ISLT). A comparative evaluation of the approach indicates faster and more balanced exploration. This is shown through energy savings ranging from 7.30% to 18.06%, and improved QoS from 10.43% to 14.3%, when compared to existing exploration strategies. This method is demonstrated under WordPress and TensorFlow workloads on a server cluster. Dainius Jenkus, Fei Xia 0001, Rishad A. Shafik, Alexandre Yakovlev |
DATE | 4 |
| 2022 | Formal Modelling of Burst-Mode Specifications in a Distributed EnvironmentabstractGeneralised fundamental mode is an important timing assumption for implementing digital circuits, where the environment is assumed to wait for the circuit to stabilise before producing new inputs. In particular, Burst-Mode (BM) timing assumption states that the circuit must wait until a complete input burst has arrived and the environment must wait until a complete output burst is produced. However, this timing assumption may be difficult to enforce in a distributed environment, if each part only observes a subset of the circuit’s output burst.In this paper, we address the above by proposing two formal modelling methodologies: 1) Design by Signal Transition Graphs (STGs), and 2) Design by our new model called Burst Automata (BAs). STGs are flexible as they express many behaviours, while BAs extends the BM methodology and enables interoperability between many different models. Our experimental results show improved synthesis success rates and significant reduction in literal count. Alex Chan, Danil Sokolov, Victor Khomenko, Alexandre Yakovlev |
FDL | 4 |
| 2022 | Editable asynchronous control logic for SAR ADCsabstractThis paper presents a novel design method for asynchronous control logic targeting successive approximation register (SAR) analog-to-digital converters (ADCs). This work is based on modeling the control logic for SAR ADCs using signal transition graphs (STGs). Different from conventional synchronous controllers, the proposed method results in asynchronous controllers driven by the causality of signals rather than relying on clocks to control the conversion process. Moreover, the proposed asynchronous control logic can be modularized through the handshake protocol, making it possible to build ADCs of arbitrary precision based on single-bit control units. This work results in a formal, model-based asynchronous design flow for SAR ADC control, which is shown to produce resulting circuits of similar speeds but great power efficiency improvements. Fei Xia 0001, Gang Mao, Shengqi Yu, Rishad A. Shafik, Alexandre Yakovlev |
ISCAS | 6 |
| 2021 | Synthesis of SI Circuits from Burst-Mode SpecificationsabstractIn this paper, we present a new workflow that is based on the conversion of Extended Burst-Mode (XBM) specifications to Signal Transition Graphs (STGs). While XBMs offer a simple design entry to specify asynchronous circuits, they cannot be synthesised into speed-independent (SI) circuits, due to the ‘burst mode’ timing assumption inherent in the model. Furthermore, XBM synthesis tools are no longer supported, and there are no dedicated tools for formal verification of XBMs. Our approach addresses these issues, by granting the XBMs access to sophisticated synthesis and verification tools available for STGs, as well as the possibility to synthesise SI circuits. Experimental results show that our translation only linearly increases the model size and that our workflow achieves a much improved synthesis success rate, with a 33% average reduction in the literal count. Alex Chan, Danil Sokolov, Victor Khomenko, Alexandre Yakovlev |
DATE | 5 |
| 2021 | PLEDGER: Embedded Whole Genome Read Mapping using Algorithm-HW Co-design and Memory-aware ImplementationabstractWith over 6000 known genetic disorders, genomics is a key driver to transform the current generation of healthcare from reactive to personalized, predictive, preventive and participatory (P4) form. High throughput sequencing technologies produce large volumes of genomic data, making genome reassembly and analysis computationally expensive in terms of performance and energy. In this paper, we propose an algorithm-hardware co-design driven acceleration approach for enabling translational genomics. Core to our approach is a Pyopencl based tooL for gEnomic workloaDs tarGeting Embedded platforms (PLEDGER). PLEDGER is a scalable, portable and energy-efficient solution to genomics targeting low-cost embedded platforms. It is a read mapping tool to reassemble genome, which is a crucial prerequisite to genomics. Using bit-vectors and variable level optimisations, we propose a low-memory footprint, dynamic programming based filtration and verification kernel capable of accelerated parallel heterogeneous executions. We demonstrate, for the first time, mapping of real reads to whole human genome on a memory-restricted embedded platform using novel memory-aware preprocessed data structures. We compare the performance and accuracy of PLEDGER with state-of-the-art RazerS3, Hobbes3, CORAL and REPUTE on two systems: 1) Intel i7-8750H CPU + Nvidia GTX 1050 Ti, 2) Odroid N2 with 6 cores: 4xCortex-A73 + 2xCortex-A53 and Mali GPU. PLEDGER demonstrates persistent energy and accuracy advantages compared to state-of-the-art read mappers producing up to 11× speedups and 5.9× energy savings compared to state-of-the-art hardware resources. Sidharth Maheshwari, Rishad A. Shafik, Ian Wilson 0006, Alexandre Yakovlev, Venkateshwarlu Y. Gudur, Amit Acharyya |
DATE | 4 |
| 2021 | Low-Latency Asynchronous Logic Design for Inference at the EdgeabstractModern internet of things (IoT) devices leverage machine learning inference using sensed data on-device rather than offloading them to the cloud. Commonly known as inference at-the-edge, this gives many benefits to the users, including personalization and security. However, such applications demand high energy efficiency and robustness. In this paper we propose a method for reduced area and power overhead of self-timed early-propagative asynchronous inference circuits, designed using the principles of learning automata. Due to natural resilience to timing as well as logic underpinning, the circuits are tolerant to variations in environment and supply voltage whilst enabling the lowest possible latency. Our method is exemplified through an inference datapath for a low power machine learning application. The circuit builds on the Tsetlin machine algorithm further enhancing its energy efficiency. Average latency of the proposed circuit is reduced by 10× compared with the synchronous implementation whilst maintaining similar area. Robustness of the proposed circuit is proven through post-synthesis simulation with 0.25 V to 1.2 V supply. Functional correctness is maintained and latency scales with gate delay as voltage is decreased. Adrian Wheeldon, Alexandre Yakovlev, Rishad A. Shafik, Jordan Morris |
DATE | 2 |
| 2021 | Optimized Multi-Memristor Model based Low Energy and Resilient Current-Mode Multiplier DesignabstractMultipliers are central to modern compute-intensive applications, such as signal processing and artificial intelligence (AI).However, the complex logic chain in conventional multipliers, particularly due to cascaded carry propagation circuits, contributes to high energy and performance costs.This paper proposes a novel current-mode multiplier design that reduces the carry propagation chain and improves the current amplification.Fundamental to this design is a one transistor multi-memristor (1TxM) cell architecture.In each cell, transistor can be switched ON/OFF to determine the cell selection, while the high/low resistive states of memristors determine the corresponding cell output current when selected.The memristor states as well as biasing configurations in each memristor are suitably optimized through a new memristor model.The number of memristors implementing this model in each cell is suitably determined depending on the cell significance to achieve the required amplification.Consequently, the design reduces the need to have current mirror circuits in each current path, while also ensuring high resilience in transitional bias voltages.Parallel cell currents are then directed to a common current accumulation path to generate the multiplier output without requiring any carry propagation chain.We carried out a wide range of experiments to extensively validate our multiplier design in Cadence Virtuoso analogue design environment for functional and parametric properties.The results show that the proposed multiplier reduces up to 85% latency and 99% energy cost when compared with the recently proposed approaches. Shengqi Yu, Rishad A. Shafik, Thanasin Bunnam, Kaiyun Chen, Alexandre Yakovlev |
DATE | 5 |
| 2021 | Run-time Configurable Approximate Multiplier using Significance-Driven Logic CompressionabstractDesigning energy-efficient hardware continues to be challenging due to arithmetic complexities. The problem is further exacerbated in systems powered by energy harvesters as variable power levels can limit their computation capabilities. In this work, we propose a run-time configurable adaptive approximation method for multiplication that is capable of managing the energy and performance tradeoffs — ideally suited in these systems. Central to our approach is a Significance-Driven Logic Compression (SDLC) multiplier architecture that can dynamically adjust the level of approximation depending on the run-time power/accuracy constraints. The architecture can be configured to operate in the exact mode (no approximation) or in progressively higher approximation modes (i.e. 2 to 4-bit SDLC). Our method is implemented in both ASIC and FPGA. The implementation results indicate that our design has only a 2.3% silicon overhead, on top of what is required by a traditional exact multiplier. We evaluate the efficiency of the proposed design through a number of case studies. We show that our method achieves similar image fidelity as in the existing approximate methods, without a delay penalty. Further, the inclusion of the dynamic approximation techniques is justified by up to 62.6% energy savings when processing an image with a multiplier using 4-bit SDLC and 35% energy savings when using 2-bit SDLC. In addition, case study results show that the proposed approach incurs negligible loss in output quality with the worst PSNR of 30dB when using the 4-bit SDLC multiplier. Ibrahim Haddadi, Issa Qiqieh, Rishad A. Shafik, Fei Xia 0001, Mohammed A. Noaman Al-Hayanni, Alexandre Yakovlev |
ICCD | 6 |
| 2021 | Power density aware application mapping in mesh-based network-on-chip architecture: An evolutionary multi-objective approach
Nizar Dahir, Ammar Karkar, Maurizio Palesi, Terrence S. T. Mak, Alexandre Yakovlev |
Integr. | 5 |
| 2021 | CORAL: Verification-Aware OpenCL Based Read Mapper for Heterogeneous SystemsabstractGenomics has the potential to transform medicine from reactive to a personalized, predictive, preventive, and participatory (P4) form. Being a Big Data application with continuously increasing rate of data production, the computational costs of genomics have become a daunting challenge. Most modern computing systems are heterogeneous consisting of various combinations of computing resources, such as CPUs, GPUs, and FPGAs. They require platform-specific software and languages to program making their simultaneous operation challenging. Existing read mappers and analysis tools in the whole genome sequencing (WGS) pipeline do not scale for such heterogeneity. Additionally, the computational cost of mapping reads is high due to expensive dynamic programming based verification, where optimized implementations are already available. Thus, improvement in filtration techniques is needed to reduce verification overhead. To address the aforementioned limitations with regards to the mapping element of the WGS pipeline, we propose a Cross-platfOrm Read mApper using opencL (CORAL). CORAL is capable of executing on heterogeneous devices/platforms, simultaneously. It can reduce computational time by suitably distributing the workload without any additional programming effort. We showcase this on a quadcore Intel CPU along with two Nvidia GTX 590 GPUs, distributing the workload judiciously to achieve up to 2× speedup compared to when, only, the CPUs are used. To reduce the verification overhead, CORAL dynamically adapts k-mer length during filtration. We demonstrate competitive timings in comparison with other mappers using real and simulated reads. CORAL is available at: https://github.com/nclaes/CORAL. Sidharth Maheshwari, Venkateshwarlu Y. Gudur, Rishad A. Shafik, Ian Wilson 0006, Alexandre Yakovlev, Amit Acharyya |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | Asynchrony and persistence in reaction systems
Maciej Koutny, Marta Pietkiewicz-Koutny, Alexandre Yakovlev |
Theor. Comput. Sci. | 3 |
| 2020 | REPUTE: An OpenCL based Read Mapping Tool for Embedded GenomicsabstractGenomics is transforming medicine from reactive to personalized, predictive, preventive and participatory (P4). The massive amount of data produced by genomics is a major challenge as it requires extensive computational capabilities, consuming large amounts of energy. A crucial prerequisite for computational genomics is genome assembly but the existing mapping tools used are predominantly software based, optimized for homogeneous high-performance systems. In this paper, we propose an OpenCL based REad maPper for heterogeneoUs sysTEms (REPUTE), which can use diverse and parallel compute and storage devices effectively. Core to this tool are dynamic programming based filtration and verification kernel to map the reads on multiple devices, concurrently. We show hardware/ software co-design and implementations of REPUTE across different platforms, and compare it with state-of-the-art mappers. We demonstrate the performance of mappers on two systems: 1) Intel CPU + 2×Nvidia GPUs; 2) HiKey970 embedded SoC with ARM Cortex-A73/A53 cores. The results show that REPUTE outperforms other read mappers in most cases producing up to 13× speedup with better or comparable accuracy. We also demonstrate that the embedded implementation can achieve up to 27× energy savings, enabling low-cost genomics. Sidharth Maheshwari, Rishad A. Shafik, Ian Wilson 0006, Alexandre Yakovlev, Amit Acharyya |
DATE | 4 |
| 2020 | Current-Mode Carry-Free Multiplier Design using a Memristor-Transistor Crossbar ArchitectureabstractMultipliers are a major energy and delay contributor in modern compute-intensive applications due to their complex logic architecture. As such, designing multipliers with reduced energy and faster speed has remained a thoroughgoing challenge. This paper presents a novel, carry-free multiplier, which is suitable for a new-generation of energy-constrained applications. The multiplier circuit consists of an array of memristor-transistor cells that can be selected (i.e., turned ON or OFF) using a combination of DC bias voltages based on the operand values. When a cell is selected it contributes to current in the array path, which is then amplified by current mirrors with variable transistor gate sizes. The different current paths are connected to a node for analogously accumulating the currents to produce the multiplier output directly. This removes the need for latency-sensitive carry propagation stages, typically seen in traditional multipliers. We conduct a number of experiments to validate the functional and parametric properties. Our experiments showed that proposed multiplier achieves 51.44% savings in energy at a similar accuracy when compared with recently proposed approaches. Shengqi Yu, Ahmed Soltan, Rishad A. Shafik, Thanasin Bunnam, Fei Xia 0001, Domenico Balsamo, Alexandre Yakovlev |
DATE | 7 |
| 2020 | Explainability and Dependability Analysis of Learning Automata based AI HardwareabstractExplainability remains the holy grail in designing the next-generation pervasive artificial intelligence (AI) systems. Current neural network based AI design methods do not naturally lend themselves to reasoning for a decision making process from the input data. A primary reason for this is the overwhelming arithmetic complexity.Built on the foundations of propositional logic and game theory, the principles of learning automata are increasingly gaining momentum for AI hardware design. The lean logic based processing has been demonstrated with significant advantages of energy efficiency and performance. The hierarchical logic underpinning can also potentially provide opportunities for by-design explainable and dependable AI hardware. In this paper, we study explainability and dependability using reachability analysis in two simulation environments. Firstly, we use a behavioral SystemC model to analyze the different state transitions. Secondly, we carry out illustrative fault injection campaigns in a low-level SystemC environment to study how reachability is affected in the presence of hardware stuck-at 1 faults. Our analysis provides the first insights into explainable decision models and demonstrates dependability advantages of learning automata driven AI hardware design. Rishad A. Shafik, Adrian Wheeldon, Alexandre Yakovlev |
IOLTS | 3 |
| 2020 | Toward Designing Thermally-Aware Memristance DecoderabstractMemristors are intensively proposed in many applications, such as biosensors and machine learning. Regarding their analog characteristics, memristance decoder is, therefore, an essential part for every memristor-based system. As memristor is a temperature sensitive device, this work proposes a memristance decoder circuit with self-temperature calibration. Its main building block is a comparator which is based on current mode circuit to achieve high performance at low power. The design provides configurable precision based on the available energy and supports both synchronous and asynchronous schemes. Moreover, the VTEAM model is modified to include the temperature effect on the memristance in the analysis. The simulation results, based on UMC 65nm low-leakage CMOS technology, show the following comparator's characteristics: 1.70% maximum offset, 2.91ns worst case latency, 343MHz maximum frequency and 48.79fJ maximum energy per comparison. Monte Carlo simulation shows the metastable state in determining the memristor value. This can be solved by extending the clock period or applying a metastability resolver. The proposed memristor model reveals that memristance at high resistive state degrades quadratically with the rise of the temperature and at 85C nearly reaches the memristance of low resistive state. Thanasin Bunnam, Ahmed Soltan, Danil Sokolov, Alexandre Yakovlev, Oleg V. Maevsky |
ISCAS | 4 |
| 2020 | Dynamics of Time-Domain Power-Elastic Circuits for Pervasive Machine LearningabstractTime-domain data encoding, in the form of the duty cycle of a pulse width modulated (PWM) signal, has recently shown promising ways of building Machine Learning (ML) circuits. As the temporal signals approximately retain their “pseudo-analog” capacitive charging rates under voltage/ frequency variations, the circuits designed are inherently power elastic, offering the crucial leverage of energy autonomy for pervasive applications. This paper focuses on the analysis of dynamic parametric variations and their impact on the temporally encoded Machine Learning circuits. The aim is to investigate and suitably optimize these parameters for robustness, power elasticity and energy efficiency. Our study of dynamics includes how the selection of passive (R and C) components affects the dynamic range of operating frequency, which we term as “PWM carrier frequency”. We investigate how RC values define the performance and energy in terms of computation latency and energy per operation. Additionally, we demonstrates how the dynamic range of voltage and frequency variations affect functional and non-functional parameters of the PWM-based neural network solutions. Sergey Mileiko, Thanasin Bunnam, Fei Xia 0001, Rishad A. Shafik, Alexandre Yakovlev |
ISCAS | 5 |
| 2020 | PARMA: Parallelization-Aware Run-Time Management for Energy-Efficient Many-Core SystemsabstractPerformance and energy efficiency considerations have shifted computing paradigms from single-core to many-core architectures. At the same time, traditional speedup models such as Amdahl's Law face challenges in the run-time reasoning for system performance and energy efficiency, because these models typically assume limited variations of the parallel fraction. Moreover, the parallel fraction, which varies dynamically in workloads, is generally unknown at run-time without application-level instrumentation. This article describes novel performance/energy trade-off models based on realistic architectural considerations, which describe the parallel fraction and speedup as functions of performance counter values available in modern processors, removing the need for application-level instrumentation. These are then used to develop a Parallelization-Aware Run-time Management (PARMA) approach. PARMA aims at controlling core allocations and operating voltage/frequency points for energy efficiency, according to the varying workload parallel fractions. The efficacy of our models and the PARMA approach is extensively validated using a number of PARSEC benchmark applications, involving two performance/energy trade-off metrics: energy-delay-product (EDP), typically used in high-performance applications and energy per instruction (EPI), suitable for energy-aware applications. Up to 48 and 68 percent improvements in EDP and EPI have been observed using the PARMA approach compared with parallelization-agnostic methods. Mohammed A. Noaman Al-Hayanni, Ashur Rafiev, Fei Xia 0001, Rishad A. Shafik, Alexander B. Romanovsky, Alexandre Yakovlev |
IEEE Trans. Computers | 6 |
| 2020 | Advance Interconnect Circuit Modeling Design Using Fractional-Order ElementsabstractNowadays, the interconnect circuits' conduct plays a crucial role in determining the performance of the CMOS systems, especially those related to nano-scale technology. Modeling the effect of such an influential component has been widely studied from many perspectives. In this article, we propose a new general formula for RLC interconnect circuit model in CMOS technology using the fractional-order elements approach. The study is based on approximating an infinite transfer function of the CMOS circuit with a noninteger distributed RLC load to a finite number of poles. It is accurate due to the effect of adding fractional-order variables and since these variables are utilized for tuning the model to match the design regardless of its complexity. As such, delay calculations employing our analytical model are within 0.4 absolute error of COMSOL-computed delay across a range of interconnect lengths. Furthermore, the effect of the interconnect conductivity G has been taken into account tacitly although the model included the resistance R, inductance L, and capacitance C of the interconnect. A number of analyses were set up at different levels of the design to evaluate the effectiveness. First, demonstrating the significant effects of generalizing parameters was gained by studying the fractional-order impedance and propagation constant of the transmission line for a range of frequencies. Second, using MATLAB we assessed the potential of the proposed approximated model besides the exact one, which shows similarity in the fundamental features of the system, such as stability and resonance. Third, the proposed approach showed that with a very small tuning reach 0.01 of the generalizing parameters can achieve up to 15% improvement in the model accuracy. Mohammed Al-daloo, Ahmed Soltan, Alexandre Yakovlev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Automating the Design of Asynchronous Logic Control for AMS ElectronicsabstractAnalog and mixed signal (AMS) electronics becomes increasingly complex and needs to be digitally enhanced by its own control circuitry. The RTL synthesis flow routinely used for digital logic is, however, optimized for synchronous data processing and produces inefficient control for AMS. In this paper, we demonstrate the evident benefits of asynchronous circuits in the context of AMS systems, and propose an asynchronous design for analog electronics (A4A) flow for their specification, synthesis, and formal verification. A library of specialized analog-to-asynchronous (A2A) components is developed for interfacing analog and asynchronous worlds. A4A flow is automated in the Workcraft framework and evaluated using a multiphase buck converter case study, where A2A components are employed to sanitize analog sensor readings. Timing analysis of asynchronous buck control shows improved response time: 4× reaction to high-load and 7× to under-voltage condition, compared with a 333 MHz clocked controller (to achieve a similar response time, a clocked controller would require ~3 GHz frequency). The simulation results of a 4-phase asynchronous buck demonstrate improved voltage ripple and peak current -16% and 12% reduction, respectively. These benefits lead to the higher efficiency of power conversion, and can be traded off for the cost of analog components, e.g., coils. Moreover, the use of the proposed design flow and tools helps to improve design productivity and overall robustness of AMS circuits. Danil Sokolov, Victor Khomenko, Andrey Mokhov, Vladimir Dubikhin, Alexandre Yakovlev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2019 | A Pulse Width Modulation based Power-elastic and Robust Mixed-signal Perceptron DesignabstractNeural networks are exerting burgeoning influence in emerging artificial intelligence applications at the micro-edge, such as sensing systems and image processing. As many of these systems are typically self-powered, their circuits are expected to be resilient and efficient in the presence of continuous power variations caused by the harvesters. In this paper, we propose a novel mixed-signal (i.e. analogue/digital) approach of designing a power-elastic perceptron using the principle of pulse width modulation (PWM). Fundamental to the design are a number of parallel inverters that transcode the input-weight pairs based on the principle of PWM duty cycle. Since PWM-based inverters are typically agnostic to amplitude and frequency variations, the perceptron shows a high degree of power elasticity and robustness under these variations. We show extensive design analysis in Cadence Analog Design Environment tool using a 3 × 3 perceptron circuit as a case study to demonstrate the resilience in the presence of parameric variations. Sergey Mileiko, Rishad A. Shafik, Alexandre Yakovlev, Jonathan Edwards |
DATE | 3 |
| 2019 | Ultra-low power m-sequence code generator for body sensor node applications
Ahmad N. Abdulfattah, Charalampos Tsimenidis, Alexandre Yakovlev |
Integr. | 3 |
| 2019 | Self-timed, minimum latency circuits for the internet of things
Adrian Wheeldon, Jordan Morris, Danil Sokolov, Alexandre Yakovlev |
Integr. | 4 |
| 2018 | An Excitation Time Model for General-purpose Memristance Tuning CircuitabstractThis paper presents a way to realize a simple yet accurate excitation time model of memristor-based circuits that are in the form of voltage dividers. According to the supported circuit structure, this model is compatible with a wide range of memristor applications, such as delay elements and analog memories. Memristance tuning based on excitation time estimation, i.e. pulse width, instead of using comparators, helps to save area and power. A general-purpose memristance tuning (GMET) circuit is proposed in order to demonstrate the simplicity of the proposed model and to evaluate its accuracy. Our model estimates are compared against the simulation results for the GMET circuit with VTEAM model of the memristor. The results show that the whole-range memristance shifts can be estimated with the worst case average error of 5.49%. They also show the worst case maximum error of 13.25%, which reduces to less than 7% when the operated memristance is higher than 2.5kΩ. Thanasin Bunnam, Ahmed Soltan, Danil Sokolov, Alexandre Yakovlev |
ISCAS | 4 |
| 2018 | Real-Power ComputingabstractThe traditional hallmark in embedded systems is to minimize energy consumption considering hard or soft real-time deadlines. The basic principle is to transfigure the uncertainties of task execution times in the real world into energy saving opportunities. The energy saving is achieved by suitably controlling the reliable power supply at circuit or system-level with the aim of minimizing the slack times, while meeting the specified performance requirements. Computing paradigm for emerging ubiquitous systems, particularly for the energy-harvested ones, has clearly shifted from the traditional systems. The energy supply of these systems can vary temporally and spatially within a dynamic range, essentially making computation extremely challenging. Such a paradigm shift requires disruptive approaches to design computing systems that can provide continued functionality under unreliable supply power envelope and operate with autonomous survivability (i.e., the ability to automatically guarantee retention and/or completion of a given computation task). In this paper, we introduce Real-Power Computing, inspired by the above trends and tenets. We show how computation systems must be designed with power-proportionality to achieve sustained computation and survivability when operating at extreme power conditions. We present extensive analysis of the need for this new computing approach using definitions, where necessary, coupled with detailed taxonomies, empirical observations, a review of relevant research works and example scenarios using three case studies representing the proposed paradigm. Rishad A. Shafik, Alexandre Yakovlev, Shidhartha Das |
IEEE Trans. Computers | 2 |
| 2018 | High-Level Asynchronous Concepts at the Interface Between Analog and Digital WorldsabstractAsynchronous circuits are becoming increasingly important in system design for Internet of Things, where they orchestrate the interface between big synchronous computation components and the analog environment, which is inherently asynchronous and has high uncertainty with respect to power supply, temperature, and long-term aging effects. However, wide adoption of asynchronous circuits by industrial users is hindered by a steep learning curve for asynchronous control models, such as signal transition graphs (STGs), that are developed by the academic community for specification, verification, and synthesis of asynchronous circuits. In this paper, we introduce a novel high-level description language for asynchronous circuits, which is based on behavioral concepts-high-level descriptions of asynchronous circuit requirements, that can be shared, reused, and extended by users, and can be automatically translated to STGs for further processing by conventional asynchronous and synchronous electronic design automation tools, such as Petrify and Mpsat. Our aim is to simplify the process of capturing system requirements in the form of a formal specification, and to promote behavioral concepts as a means for design reuse. The proposed design flow is fully automated in open-source toolsuite Workcraft, and is applied to the development of an asynchronous power regulator. Jonathan Beaumont, Andrey Mokhov, Danil Sokolov, Alexandre Yakovlev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | Advances in Formal Methods for the Design of Analog/Mixed-Signal Systems: InvitedabstractAnalog/mixed-signal (AMS) systems are rapidly expanding in all domains of information and communication technology. They are a critical part of the support for large-scale high-performance digital systems, provide important functionalities in medium-scale embedded and mobile systems, and act as a core organ of autonomous electronics such as sensor nodes. Analog and digital parts are closely intermixed, hence demanding AMS design methods and tools to be more holistic. In particular, the emergence of "little digital" electronics inside or near analog circuitry calls for the increasing use of asynchronous logic. To cope with the growing complexity of AMS designs, formal methods are required to complement traditional simulation approaches. This paper presents an overview of the state-of-the-art in AMS formal verification and asynchronous design that enables the development of analog/asynchronous co-design methods. One such co-design methodology is exemplified by the LEMA-Workcraft workflow currently under development by the authors. Vladimir Dubikhin, Chris J. Myers, Danil Sokolov, Ioannis Syranidis, Alexandre Yakovlev |
DAC | 5 |
| 2017 | Energy-efficient approximate multiplier design using bit significance-driven logic compressionabstractApproximate arithmetic has recently emerged as a promising paradigm for many imprecision-tolerant applications. It can offer substantial reductions in circuit complexity, delay and energy consumption by relaxing accuracy requirements. In this paper, we propose a novel energy-efficient approximate multiplier design using a significance-driven logic compression (SDLC) approach. Fundamental to this approach is an algorithmic and configurable lossy compression of the partial product rows based on their progressive bit significance. This is followed by the commutative remapping of the resulting product terms to reduce the number of product rows. As such, the complexity of the multiplier in terms of logic cell counts and lengths of critical paths is drastically reduced. A number of multipliers with different bit-widths (4-bit to 128-bit) are designed in SystemVerilog and synthesized using Synopsys Design Compiler. Post-synthesis experiments showed that up to an order of magnitude energy savings, and reductions of 65% in critical delay and almost 45% in silicon area can be achieved for a 128-bit multiplier compared to an accurate equivalent. These gains are achieved with low accuracy losses estimated at less than 0.00071 mean relative error. Additionally, we demonstrate the energy-accuracy trade-offs for different degrees of compression, achieved through configurable logic clustering. In evaluating the effectiveness of our approach, a case study image processing application showed up to 68.3% energy reduction with negligible losses in image quality expressed as peak signal-to-noise ratio (PSNR). Issa Qiqieh, Rishad A. Shafik, Ghaith Tarawneh, Danil Sokolov, Alexandre Yakovlev |
DATE | 5 |
| 2017 | Benefits of asynchronous control for analog electronics: Multiphase buck case studyabstractAnalog and mixed signal (AMS) electronics becomes increasingly complex and needs to be digitally enhanced by its own control circuitry. The RTL synthesis flow routinely used for digital logic is however optimized for synchronous data processing and produces inefficient control for AMS. In this paper we demonstrate the evident benefits of asynchronous circuits in the context of AMS systems, and propose an asynchronous design for analog electronics (A4A) flow for their specification, synthesis, and formal verification. A library of specialized analog-to-asynchronous (A2A) components is developed for interfacing analog signals to asynchronous control. A4A flow is automated in the Workcraft framework and evaluated using a multiphase buck converter case study. The simulation results show improved response time, voltage ripple, and peak current of the buck when controlled asynchronously. These benefits lead to the higher efficiency of power conversion, and can be traded off for the cost of analog components. A4A flow, A2A interfaces, and Workcraft tools are used for development of power converters at Dialog Semiconductor. Danil Sokolov, Vladimir Dubikhin, Victor Khomenko, Andrey Mokhov, Alexandre Yakovlev |
DATE | 6 |
| 2017 | Language and hardware acceleration backend for graph processingabstractGraphs are important in many applications however their analysis on conventional computer architectures is generally inefficient because it involves highly irregular access to memory when traversing vertices and edges. As an example, when finding a path from a source vertex to a target one the performance is typically limited by the memory bottleneck whereas the actual computation is trivial. This paper presents a methodology for embedding graphs into silicon, where graph vertices become finite state machines communicating via the graph edges. With this approach many common graph analysis tasks can be performed by propagating signals through the physical graph and measuring signal propagation time using the on-chip clock distribution network. This eliminates the memory bottleneck and allows thousands of vertices to be processed in parallel. We present a domain-specific language for graph description and transformation, and demonstrate how it can be used to translate application graphs into an FPGA board, where they can be analysed up to 1000× faster than on a conventional computer. Andrey Mokhov, Alessandro de Gennaro, Ghaith Tarawneh, Jonny Wray, Georgy Lukyanov, Sergey Mileiko, Joe Scott, Alexandre Yakovlev, Andrew D. Brown |
FDL | 8 |
| 2017 | A Structured Visual Approach to GALS Modeling and Verification of Communication CircuitsabstractIn this paper, a novel globally asynchronous locally synchronous (GALS) modeling and verification tool is introduced for xMAS circuits. The tool provides a structured environment for GALS in which organization of the modeling and verification enables it to handle a variety of implementation tasks facilitating a process which would otherwise be difficult for the end user. The tool provides verification techniques at different levels. A new unfolding algorithm is presented that uses structured occurrence nets. A novel representation for deadlocks is introduced using deadlock relations enabling the causality of local and global deadlocks to be visualized. This helps in the investigation of total or partial system shutdown. In particular, the approach enables the visualization of point-to-point causality of problems occurring between different parts of the system which are more difficult to analyze. In addition different types of deadlock related to the synchronizer can be detected. The work presented here provides structured visualization capability facilitating the analysis of complex communication systems. Frank P. Burns, Danil Sokolov, Alexandre Yakovlev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | A Resilient 2-D Waveguide Communication Fabric for Hybrid Wired-Wireless NoC DesignabstractHybrid wired-wireless Network-on-Chip (WiNoC) has emerged as an alternative solution to the poor scalability and performance issues of conventional wireline NoC design for future System-on-Chip (SoC). Existing feasible wireless solution for WiNoCs in the form of millimeter wave (mm-Wave) relies on free space signal radiation which has high power dissipation with high degradation rate in the signal strength per transmission distance. Moreover, over the lossy wireless medium, combining wireless and wireline channels drastically reduces the total reliability of the communication fabric. Surface wave has been proposed as an alternative wireless technology for low power on-chip communication. With the right design considerations, the reliability and performance benefits of the surface wave channel could be extended. In this paper, we propose a surface wave communication fabric for emerging WiNoCs that is able to match the reliability of traditional wireline NoCs. First, we propose a realistic channel model which demonstrates that existing mm-Wave WiNoCs suffers from not only free-space spreading loss (FSSL) but also molecular absorption attenuation (MAA), especially at high frequency band, which reduces the reliability of the system. Consequently, we employ a carefully designed transducer and commercially available thin metal conductor coated with a low cost dielectric material to generate surface wave signals with improved transmission gain. Our experimental results demonstrate that the proposed communication fabric can achieve a 5 dB operational bandwidth of about 60 GHz around the center frequency (60 GHz). By improving the transmission reliability of wireless layer, the proposed communication fabric can improve maximum sustainable load of NoCs by an average of 20:9 and 133:3 percent compared to existing WiNoCs and wireline NoCs, respectively. Michael Opoku Agyeman, Quoc-Tuan Vien, Ali Ahmadinia, Alexandre Yakovlev, Kin-Fai Tong, Terrence S. T. Mak |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | Formal verification of clock domain crossing using gate-level models of metastable flip-flops
Ghaith Tarawneh, Andrey Mokhov, Alexandre Yakovlev |
DATE | 3 |
| 2016 | Low power voltage sensing through capacitance to digital conversionabstractCapacitance sensors are widely used for sensing physical parameters. Conventional capacitance to digital methods use complex analog ADC techniques which are power hungry. Recently a fully digital solution was proposed with improved power consumption. This paper describes a number of problems in that solution, analyzes these problems, and proposes a new design free of these problems. A voltage senor as an example was designed based on the proposed capacitance to digital conversion in this paper. The new method achieves the same accuracy with less than half the circuit size, and 25% and 33% savings on power and energy consumption. Delong Shang, Yuqing Xu, Kaiyuan Gao, Fei Xia 0001, Alexandre Yakovlev |
DDECS | 5 |
| 2016 | Selective abstraction and stochastic methods for scalable power modelling of heterogeneous systemsabstractWith the increase of system complexity in both platforms and applications, power modelling of heterogeneous systems is facing grand challenges from the model scalability issue. To address these challenges, this paper studies two systematic methods: selective abstraction and stochastic techniques. The concept of selective abstraction via black-boxing is realised using hierarchical modelling and cross-layer cuts, respecting the concepts of boxability and error contamination. The stochastic aspect is formally underpinned by Stochastic Activity Networks (SANs). The proposed method is validated with experimental results from Odroid XU3 heterogeneous 8-core platform and is demonstrated to maintain high accuracy while improving scalability. Ashur Rafiev, Fei Xia 0001, Alexei Iliasov, Rem Gensh, Ali Aalsaud, Alexander B. Romanovsky, Alexandre Yakovlev |
FDL | 7 |
| 2016 | MEMS-based power delivery control for bursty applicationsabstractThe current trends for greater heterogeneity in future Systems-on-Chip (SoC) do not only concern their functionality but also their timing and power aspects. The increasing diversity of timing and power supply conditions, and associated concurrently operating modes, within an SoC calls for more efficient power delivery networks (PDN) for battery operated devices. This is especially important for systems with mixed duty cycling, where some parts are required to work regularly with low-throughput while other parts are activated spontaneously, i.e. in bursts. To improve their reaction time vs energy efficiency, this paper proposes to incorporate a power-switching network based on MEM relays to switch the SoC power-performance state (PPS) into an active mode while eliminating the leakage current when it is idle. Results show that even with todays large and high pull-in voltages, a MEM-relay-based power switching network (PSN) can achieve a 1000x saving in energy compared to its CMOS counterpart for low duty cycle. Finally, the paper presents, by means of multiphysics simulation results, design guidelines for idle time and energy-saving estimates as a function of MEM relay scaling parameters. Haider Alrudainy, Andrey Mokhov, Nizar Dahir, Alexandre Yakovlev |
ISCAS | 4 |
| 2016 | Power-Aware Performance Adaptation of Concurrent Applications in Heterogeneous Many-Core SystemsabstractModern embedded systems execute multiple applications, both sequentially and concurrently. These applications are exercised on heterogeneous platforms generating varying power consumption and system workloads (CPU or memory intensive or both). As a result, determining the most energy-efficient system configuration (i.e. the number of parallel threads, their core allocations and operating frequencies) tailored for each kind of workload and application scenario is extremely challenging. In this paper, we propose a novel runtime optimization approach with the aim of achieving maximized power normalized performance considering dynamic variation of workload and application scenarios. Fundamental to this approach is a comprehensive study to investigate the tradeoffs between inter-application concurrency with performance and power consumption under different system configurations. Using real experimental measurements on an Odroid XU-3 heterogeneous platform with a number of PARSEC benchmark applications, we model power normalized performance (in terms of IPS/Watt) underpinning analytical power and performance models, derived through multivariate linear regression (MLR). Using these models, we show that with increasing number of concurrent CPU intensive applications show variable gains in IPS/Watt compared to the memory intensive applications in both sequential and concurrent application scenarios. Furthermore, we demonstrate that it is possible to continuously adapt system configuration through a low-cost and linear-complexity runtime algorithm, which can improve the IPS/Watt by up to 125% compared to the existing approach. Ali Aalsaud, Rishad A. Shafik, Ashur Rafiev, Fei Xia 0001, Sheng Yang 0003, Alexandre Yakovlev |
ISLPED | 6 |
| 2015 | GALS synthesis and verification for xMAS models
Frank P. Burns, Danil Sokolov, Alexandre Yakovlev |
DATE | 3 |
| 2015 | Mixed wire and surface-wave communication fabrics for decentralized on-chip multicasting
Ammar Karkar, Kin-Fai Tong, Terrence S. T. Mak, Alexandre Yakovlev |
DATE | 4 |
| 2015 | Compositional design of asynchronous circuits from behavioural conceptsabstractAsynchronous circuits can be useful in many applications, however, they are yet to be widely used in industry. The main reason for this is a steep learning curve for concurrency models, such Signal Transition Graphs, that are developed by the academic community for specification and synthesis of asynchronous circuits. In this paper we introduce a compositional design flow for asynchronous circuits using concepts - a set of formalised descriptions for system requirements. Our aim is to simplify the process of capturing system requirements in the form of a formal specification, and promote the concepts as a means for design reuse. The proposed design flow is applied to the development of an asynchronous buck converter. Jonathan Beaumont, Andrey Mokhov, Danil Sokolov, Alexandre Yakovlev |
MEMOCODE | 4 |
| 2015 | Novel Hybrid Wired-Wireless Network-on-Chip Architectures: Transducer and Communication Fabric DesignabstractExisting wireless communication interface of Hybrid Wired-Wireless Network-on-Chip (WiNoC) has 3-dimensional free space signal radiation which has high power dissipation and drastically affects the received signal strength. In this paper, we propose a CMOS based 2-dimensional (2-D) waveguide communication fabric that is able to match the channel reliability of traditional wired NoCs as the wireless communication fabric. Our experimental results demonstrate that, the proposed communication fabric can achieve a 5dB operational bandwidth of about 60GHz around the center frequency (60GHz). Compared to existing WiNoCs, the proposed communication fabric can improve the reliability of WiNoCs with average gains of 21.4%, 13.8% and 10.6% performance efficiencies in terms of maximum sustainable load, throughput and delay, respectively. Michael Opoku Agyeman, Wen Zong, Ji-Xiang Wan, Alexandre Yakovlev, Kenneth Tong, Terrence S. T. Mak |
NOCS | 4 |
| 2015 | A Formal Specification and Prototyping Language for Multi-core System ManagementabstractWe relate the experience of a defining a formal domain specific language (DSL) for the construction and reasoning about OS-level management logic of multi-core systems. The approach is based on a novel, iterative development principle where results of prototyping studies feed back into the next language revision. We illustrate the DSL with several examples of executable scripts. Alexei Iliasov, Ashur Rafiev, Fei Xia 0001, Rem Gensh, Alexander B. Romanovsky, Alexandre Yakovlev |
PDP | 6 |
| 2015 | Persistent and Nonviolent Steps and the Design of GALS SystemsabstractA concurrent system is persistent if throughout its operation no activity which became enabled can subsequently be prevented from being executed by any other activity. This is often a highly desirable (or even necessary) property; in particular, if the system is to be implemented in hardware. Over the past 40 years, persistence has been investigated and applied in practical implementations assuming that each activity is a single atomic action which can be represented, for example, by a single transition of a Petri net. In this paper we investigate the behaviour of GALS (Globally Asynchronous Locally Synchronous) systems in the context of VLSI circuits. The specification of a system is given in the form of a Petri net. Our aim is to re-design the system to optimise signal management, by grouping together concurrent events. Looking at the concurrent reachability graph of the given Petri net, we are interested in discovering events that appear in ‘bundles’, so that they all can be executed in a single clock tick. The best candidates for bundles are sets of events that appear and re-appear over and over again in the same configurations, forming ‘persistent’ sets of events. Persistence was considered so far only in the context of sequential semantics. In this paper, we move to the realm of step based execution and consider not only steps which are persistent and cannot be disabled by other steps, but also steps which are nonviolent and cannot disable other steps. We then introduce a formal definition of a bundle and propose an algorithm to prune the behaviour of a system, so that only bundled steps remain. The pruned reachability graph represents the behaviour of a re-engineered system, which in turn can be implemented in a new Petri net using the standard techniques of net synthesis. The proposed algorithm prunes reachability graphs of persistent and safe nets leaving bundles that represent maximally concurrent steps. Johnson Fernandes, Maciej Koutny, Lukasz Mikulski, Marta Pietkiewicz-Koutny, Danil Sokolov, Alexandre Yakovlev |
Fundam. Informaticae | 6 |
| 2015 | Design of Self-Timed Reconfigurable Controllers for Parallel Synchronization via WaggingabstractSynchronization is an important issue in modern system design as systems-on-chips integrate more diverse technologies, operating voltages, and clock frequencies on a single substrate. This paper presents a methodology for the design and implementation of a self-timed reconfigurable control device suitable for a parallel cascaded flip-flop synchronizer based on a principle known as wagging, through the application of distributed feedback graphs. By modifying the endpoint adjacency of a common behavior graph via one-hot codes, several configurable modes can be implemented in a single design specification, thereby facilitating direct control over the synchronization time and the mean-time between failures of the parallel master-slave latches in the synchronizer. Therefore, the resulting implementation is resistant to process nonidealities, which are present in physical design layouts. This paper includes a discussion of the reconfiguration protocol, and implementations of both a sequential token ring control device, and an interrupt subsystem necessary for reconfiguration, all simulated in UMC 90-nm technology. The interrupt subsystem demonstrates operating frequencies between 505 and 818 MHz per module, with average power consumptions between 70.7 and 90.0 μW in the typical-typical case under a corner analysis. James Sebastian Guido, Alexandre Yakovlev |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Power-Adaptive Computing System Design for Solar-Energy-Powered Embedded SystemsabstractThrough energy harvesting system, new energy sources are made available immediately for many advanced applications based on environmentally embedded systems. However, the harvested power, such as the solar energy, varies significantly under different ambient conditions, which in turn affects the energy conversion efficiency. In this paper, we propose an approach for designing power-adaptive computing systems to maximize the energy utilization under variable solar power supply. Using the geometric programming technique, the proposed approach can generate a customized parallel computing structure effectively. Then, based on the prediction of the solar energy in the future time slots by a multilayer perceptron neural network, a convex model-based adaptation strategy is used to modulate the power behavior of the real-time computing system. The developed power-adaptive computing system is implemented on the hardware and evaluated by a solar harvesting system simulation framework for five applications. The results show that the developed power-adaptive systems can track the variable power supply better. The harvested solar energy utilization efficiency is 2.46 times better than the conventional static designs and the rule-based adaptation approaches. Taken together, the present thorough design approach for self-powered embedded computing systems has a better utilization of ambient energy sources. Qiang Liu 0011, Terrence S. T. Mak, Tao Zhang 0025, Xinyu Niu, Wayne Luk, Alexandre Yakovlev |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2014 | Network on Chip optimization based on surrogate model assisted evolutionary algorithmsabstractNetwork-on-Chip (NoC) design is attracting more and more attention nowadays, but there is a lack of design optimization method due to the computationally very expensive simulations of NoC. To address this problem, an algorithm, called NoC design optimization based on Gaussian process model assisted differential evolution (NDPAD), is presented. Using the surrogate model-aware evolutionary search (SMAS) framework with the tournament selection based constraint handling method, NDPAD can obtain satisfactory solutions using a limited number of expensive simulations. The evolutionary search strategies and training data selection methods are then investigated to handle integer design parameters in NoC design optimization problems. Comparison shows that comparable or even better design solutions can be obtained compared to standard EAs, and much less computation effort is needed. Mengyuan Wu, Ammar Karkar, Bo Liu 0003, Alexandre Yakovlev, Georges Gielen, Vic Grout |
IEEE Congress on Evolutionary Computation | 4 |
| 2014 | Hybrid wire-surface wave architecture for one-to-many communication in networks-on-chipabstractNetwork-on-chip (NoC) is a communication paradigm that has emerged to tackle different on-chip challenges and has satisfied different demands in terms of high performance and economical interconnect implementation. However, merely metal based NoC pursuit offers limited scalability with the relentless technology scaling, especially in one-to-many (1-to-M) communication. To meet the scalability demand, this paper proposes a new hybrid architecture empowered by both metal interconnects and Zenneck surface wave interconnects (SWI). This architecture, in conjunction with newly proposed routing and global arbitration schemes, avoids overloading the NoC and alleviates traffic hotspots compared to the trend of handling 1-to-M traffic as unicast. This work addresses the system level challenges for intra chip multicasting. Evaluation results, based on a cycle-accurate simulation and hardware description, demonstrate the effectiveness of the proposed architecture in terms of power reduction ratio of 4 to 12X and average delay reduction of 25X or more, compared to a regular NoC. These results are achieved with negligible hardware overheads. Ammar Karkar, Nizar Dahir, Ra'ed Al-Dujaily, Kenneth Tong, Terrence S. T. Mak, Alexandre Yakovlev |
DATE | 6 |
| 2014 | Asynchronous design for new on-chip wide dynamic range power electronicsabstractAsynchronous circuits will play an important role in microelectronic systems in the future, especially in energy harvesting and autonomous (EHA) systems where such circuits will be able to offer robustness and deliver high efficiency in a wide range of power-energy conditions. The concept of Capacitor Bank Block (CBB) mechanisms was proposed to form the basis of electronics for powering asynchronous loads. These mechanisms will benefit EHA systems by enabling effective co-scheduling of computational tasks and energy supply. This paper demonstrates how the CBB mechanisms can themselves be controlled by asynchronous circuits, thereby forming a new type of power delivery units (PDU) that will be able to deliver power to intelligent digital logic in future EHA systems. These PDUs are superior to traditional power converters largely because the latter can only regulate sufficiently high power and energy levels (regular and periodic) as well as their controllers require stable power levels themselves. This makes them unsuitable for intermittent and sporadic conditions inherent to EHA systems. In this paper, a novel asynchronous control for the CBB is described. Experiments and analysis of the new PDUs, comprising CBBs and asynchronous control, are presented and discussed in detail. Delong Shang, Xuefu Zhang, Fei Xia 0001, Alexandre Yakovlev |
DATE | 4 |
| 2014 | Asynchronously assisted FPGA for variabilityabstractThe effect of variability has become increasingly significant as a result of technology geometry scaling. This paper describes Asynchronous Assisting Logic (AAL) blocks and the method of introducing them into modern FPGA architecture, in order to increase tolerance of the wide range latency variations caused by parametric variation, and temperature and supply voltage fluctuations. The proposed method leverages the availability of variation maps and suggests deploying configurable AAL blocks only into the variation critical paths - reinforcing rather rerouting/remapping. This method reduces the size overhead significantly which normally will be incurred by fully asynchronous designs. The proposed technique maintains the existing FPGA architecture allowing potential reuse of design flow. Simulations show correct functionality given regularly variable, randomly variable and capacitor switching energy harvester voltage supplies. Hock Soon Low, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
FPL | 4 |
| 2014 | Design and Implementation of Dynamic Thermal-Adaptive Routing Strategy for Networks-on-ChipabstractTechnology scaling is leading to extreme thermal challenges that make worst-case cooling system design unfavourable. On the other hand, on-chip communication, in terms of Network-on-Chip (NoC) workload, is expected to dominate Systems-on-Chip as a major heat source. In this paper a Runtime Thermal Management (RTM) design implementation for NoCs is proposed. Dynamic Programming Network (DPN) is introduced to implement the adaptive routing control logic and Ring Oscillators (ROs) are used for temperature sensing. Various challenges associated with DPN convergence and sensor accuracy and precision, such as isolating the IR drops and intra-chip process variations, are addressed. An FPGA implementation of the proposed strategy demonstrates promising results in terms of both thermal regulation and functionality with a variety of traffics. In terms of functionality, the proposed scheme is shown to be highly flexible in manoeuvring the packets away from hot regions. This results in up to 16% reduction in the maximum chip temperature and lowers chip thermal gradient by up to 51% compared with performance-driven routing. Moreover, the proposed scheme results in significantly slower chip heating which is reflected as up to 100% higher performance when the chip works under a thermal limit. These results imply that the proposed technique would improve thermal reliability and performance for future many-core VLSI systems. Nizar Dahir, Ghaith Tarawneh, Terrence S. T. Mak, Ra'ed Al-Dujaily, Alexandre Yakovlev |
PDP | 5 |
| 2014 | Modeling and Tools for Power Supply Variations Analysis in Networks-on-ChipabstractPower supply integrity has become a critical concern with the rapid shrinking feature size and the ever increasing power consumption in nanometre scale integration. In particular, on-chip communication in platforms such as networks-on-chip (NoC) dictates the power dissipation and overall system performance in multicore systems and embedded computing architectures. These architectures require a dedicated tool for analyzing the power supply noise which must embed distinctive communication characteristics and spatial parameters. In this paper, we present a tool dedicated to determining the on-chip VDDdrops due to communication workload in NoCs. This tool integrates a fast power grid model, an NoC simulator, an on-chip link model, and a microarchitectural power model for router. The model has been rigorously verified using SPICE simulations. The proposed model and tools are further exemplified through analyzing the impact of power supply noise for NoC links. Statistical timing analysis of NoC links in the presence of power supply noise was performed to evaluate the bit error rates (BERs). This work would enable better understanding of the tradeoffs existing in the design of NoCs, and the induced power supply noise due to on-chip communication. This understanding is crucial for the analysis of the quality of service (QoS) of communication fabrics in NoCs at the early design stages. Nizar Dahir, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev |
IEEE Trans. Computers | 4 |
| 2014 | Synthesis of Processor Instruction Sets from High-Level ISA SpecificationsabstractAs processors continue to get exponentially cheaper for end users following Moore’s law, the costs involved in their design keep growing, also at an exponential rate. The reason is ever increasing complexity of processors, which modern EDA tools struggle to keep up with. This paper focuses on the design of Instruction Set Architecture (ISA), a significant part of the whole processor design flow. Optimal design of an instruction set for a particular combination of available hardware resources and software requirements is crucial for building processors with high performance and energy efficiency, and is a challenging task involving a lot of heuristics and high-level design decisions. This paper presents a new compositional approach to formal specification and synthesis of ISAs. The approach is based on a new formalism, called Conditional Partial Order Graphs, capable of capturing common behavioural patterns shared by processor instructions, and therefore providing a very compact and efficient way to represent and manipulate ISAs. The Event-B modelling framework is used as a formal specification and verification back-end to guarantee correctness of ISA specifications. We demonstrate benefits of the presented methodology on several examples, including Intel 8051 microcontroller. Andrey Mokhov, Alexei Iliasov, Danil Sokolov, Maxim Rykunov, Alexandre Yakovlev, Alexander B. Romanovsky |
IEEE Trans. Computers | 5 |
| 2014 | Thermal Optimization in Network-on-Chip-Based 3D Chip Multiprocessors Using Dynamic Programming NetworksabstractThe substantial silicon density in 3D VLSI, albeit its numerous advantages, introduces serious thermal threats that would lead to faults and system failures. This article introduces a new strategy to effectively diffuse heat from NoC-based 3D CMPs. Runtime Dynamic Programming Network (DPN) is proposed to optimize routing directions and provide silicon temperature moderation. Both on-chip reliability and computational performance have been improved by 63% and 27%, respectively, with the DPN approach. This work enables a new avenue to explore the adaptability for future large-scale 3D integration. Nizar Dahir, Ra'ed Al-Dujaily, Terrence S. T. Mak, Alexandre Yakovlev |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2014 | Eliminating Synchronization Latency Using Sequenced LatchingabstractModern multicore systems have a large number of components operating in different clock domains and communicating through asynchronous interfaces. These interfaces use synchronizer circuits, which guard against metastability failures but introduce latency in processing the asynchronous input. We propose a speculative method that hides synchronization latency by overlapping it with computation cycles. We verify the correctness of our approach through a field programmable gate array implementation and apply it to a number of synthesized benchmarks. Synthesis results reveal that our approach achieves average savings of 135% and 204% in area costs and nearly 100% in power costs compared to two similar speculative techniques Ghaith Tarawneh, Alexandre Yakovlev, Terrence S. T. Mak |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Step Persistence in the Design of GALS Systems
Johnson Fernandes, Maciej Koutny, Marta Pietkiewicz-Koutny, Danil Sokolov, Alexandre Yakovlev |
Petri Nets | 5 |
| 2013 | Novel Multi-Layer Network Decomposition boosting acceleration of multi-core algorithmsabstractComplex networks are a technique for the modeling and analysis of large data sets in many scientific and engineering disciplines. Due to their excessive size conventional algorithms and single core processors struggle with the efficient processing of such networks. Employing multi-core graphic processing units (GPUs) could provide sufficient processing power for the analysis of such networks. However, commonly designed algorithms cannot exploit these massively parallel processing power for the analysis of such networks. In this paper, we present the Multi Layer Network Decomposition (MLND) approach which provides a general approach for parallel network analysis using multi-core processors via efficient partitioning and mapping of networks onto GPU architectures. Evaluation using a 336 core GPU graphic card demonstrated a 16x speed-up in complex network analysis relative to a CPU based approach. Athanasios K. Grivas, Terrence S. T. Mak, Alexandre Yakovlev, Jonny Wray |
ASAP | 3 |
| 2013 | Design-for-adaptivity of microarchitecturesabstractIn the last decade we have witnessed a steady trend towards functional diversification of hardware, because an application specific hardware component is a lot easier to design and optimise than a general-purpose one. Therefore, a modern microelectronics system often contains several application specific cores, each targeted for a particular function. As operating conditions issues are becoming more important, we start to see non-functional diversification in terms of performance and energy consumption; it is expected that a system can operate in a wide spectrum of environmental conditions and it should support a hierarchy of energy-saving modes. As a result, "mode-specific" processing cores are gaining popularity. The number of possible combinations of functional and nonfunctional variations of hardware components is becoming unmanageable and is leading to inefficient silicon utilisation. In this paper we explore a novel approach to hardware design which allows building computation systems capable of adjusting to operating conditions through dynamic reconfiguration. We demonstrate the approach by designing an asynchronous microprocessor core that can operate in a wide range of supply voltages and can adjust its functionality towards a specific application and operating mode. Our methodology is based on a novel model of hardware description and on self-timed design techniques. Maxim Rykunov, Andrey Mokhov, Danil Sokolov, Alexandre Yakovlev, Albert Koelmans |
ASAP | 4 |
| 2013 | Advances in asynchronous logic: from principles to GALS & NoC, recent industry applications, and commercial CAD toolsabstractThe growing variability and complexity of advanced CMOS technologies makes the physical design of clocked logic in large Systems-on-Chip more and more challenging. Asynchronous logic has been studied for many years and become an attractive solution for a broad range of applications, from massively parallel multi-media systems to systems with ultra-low power & low-noise constraints, like cryptography, energy autonomous systems, and sensor-network nodes. The objective of this embedded tutorial is to give a comprehensive and recent overview of asynchronous logic. The tutorial will cover the basic principles and advantages of asynchronous logic, some insights on new research challenges, and will present the GALS scheme as an intermediate design style with recent results in asynchronous Network-on-Chip for future Many Core architectures. Regarding industrial acceptance, recent asynchronous logic applications within the microelectronics industry will be presented, with a main focus on the commercial CAD tools available today. Alexandre Yakovlev, Pascal Vivet, Marc Renaudin |
DATE | 1 |
| 2013 | Towards reliable hybrid bio-silicon integration using novel adaptive control systemabstractHybrid bio-silicon networks are difficult to implement in practice due to variations of biological neuron bursting frequency. This causes the hybrid network to have inaccuracies and unreliability. The network may produce irregular bursts or incorrect spiking phase relationships if the electrical neuron bursting frequency is not suitable for biological neurons. To solve this potentially vital problem, a novel adaptive control system based on dynamic clamp is proposed. Biological measurement is combined with an adaptive controller to control to silicon neuron bursting periods in real time. We use a hybrid pyloric network which contains three real neurons and one electronic neuron as a case study. Simulation results indicate that the silicon neuron can follow the biological neuron bursting frequency in real time to achieve hybrid network functionalities. System settling time can be achieved in 303 milliseconds and percentage overshoot kept to 1%. We believe that our methodology is scalable to various larger bio-silicon hybrid neural networks. Patrick Degenaar, Graeme Coapes, Alexandre Yakovlev, Terrence S. T. Mak, Peter Andras 0001 |
ISCAS | 4 |
| 2013 | Wide-range, reference free, on-chip voltage sensor for variable Vdd operationsabstractIn future systems with relatively unreliable and unpredictable energy sources such as harvesters, the system Vdd may become non-deterministic. Reliable and accurate on-chip voltage sensors are therefore indispensible for the power and computation management of such systems. Stable and known references are also difficult to obtain in this environment. This paper describes a reference-free voltage sensor implemented using a speed independent (SI) SRAM cell and an inverter chain. It can work under a wide range of Vdd, and provides accurate measurements of Vdd over this operating range with a precision range from 50mV to 10mV. Unlike existing methods, the voltage information is directly generated as a digital code without any analog circuits. This is realized by exploiting the inherently different latency behaviors of different types of circuits under different Vdd. Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 3 |
| 2013 | Dynamic On-Chip Thermal Optimization for Three-Dimensional Networks-On-ChipabstractThe complex thermal behaviour prohibits the advancement of three-dimensional (3D) very-large-scale integration system. Particularly, the high-density through-silicon via based integration could lead to ultra-high temperature hotspots and permanent silicon device damage. In this paper, we introduce an adaptive strategy to effectively diffuse heat throughout the 3D geometry. This strategy employs a dynamic programming network to select and optimize the direction of data manoeuvre in a network-on-chip (NoC). We also developed a tool, which is based on the accurate HotSpot thermal model and SystemC cycle accurate model, to simulate the thermal system and evaluate our approach. We found that the proposed approach can significantly diffuse the hotspots from a 3D geometry and overall temperature can be significantly reduced. Given the same thermal constraints, the throughput performance of an adaptive NoC can also be improved. This work enables a new avenue to explore the on-chip adaptability for the future large-scale 3D integration. Ra'ed Al-Dujaily, Terrence S. T. Mak, Kai-Pui Lam, Fei Xia 0001, Alexandre Yakovlev, Chi-Sang Poon |
Comput. J. | 5 |
| 2013 | Concurrent Multiresource Arbiter: Design and ApplicationsabstractThis paper presents a novel type of asynchronous arbiter that allocates M interchangeable resources among N clients. This arbiter enables the concurrent utilization of multiple resources and is a useful device in various load-balancing circuits. Dedicated request signals from the resources and the clients are used in pairs to form each new grant. The 2 × 2 arbiter is examined as an accessible special case of the N × M arbiter. A concurrent implementation is compared to fully sequential design. It is shown that the sequential design can be more practical when the time between a grant and the withdrawal of the initial request is small. The concurrent design provides higher performance in a system with a longer resource utilization time. A scalable tiled structure is developed to extend the arbiter structure beyond 2 × 2 to support N clients and M resources. Models and subsequent implementations of the tiles are presented. The tiles can be repeated without the use of additional connecting logic, enabling the construction of arbiters of larger sizes. Several examples demonstrate the usage of the arbiter. Stanislavs Golubcovs, Delong Shang, Fei Xia 0001, Andrey Mokhov, Alexandre Yakovlev |
IEEE Trans. Computers | 5 |
| 2013 | Dynamic programming-based runtime thermal management (DPRTM): An online thermal control strategy for 3D-NoC systemsabstractComplex thermal behavior inhibits the advancement of three-dimensional (3D) very-large-scale-integration (VLSI) system designs, as it could lead to ultra-high temperature hotspots and permanent silicon device damage. This article introduces a new runtime thermal management strategy to effectively diffuse and manage heat throughout 3D chip geometry for a better throughput performance in networks on chip (NoC). This strategy employs a dynamic programming-based runtime thermal management (DPRTM) policy to provide online thermal regulation. Reactive and proactive adaptive schemes are integrated to optimize the routing pathways depending on the critical temperature thresholds and traffic developments. Also, when the critical system thermal limit is violated, an urgent throttling will take place. The proposed DPRTM is rigorously evaluated through cycle-accurate simulations, and results show that the proposed approach outperforms conventional approaches in terms of computational efficiency and thermal stability. For example, the system throughput using the DPRTM approach can be improved by 33% when compared to other adaptive routing strategies for a given thermal constraint. Moreover, the DPRTM implementation presented in this article demonstrates that the hardware overhead is insignificant. This work opens a new avenue for exploring the on-chip adaptability and thermal regulation for future large-scale and 3D many-core integrations. Ra'ed Al-Dujaily, Nizar Dahir, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2012 | Reconfigurable time interval measurement circuit incorporating a programmable gain time difference amplifierabstractTime interval measurement (TIM) is used in a wide range of applications, for example, physics experiments, dynamic testing of integrated circuits (IC), telecommunications, laser distance measurement, X-ray and UV imagers etc., requiring a range of measurement accuracy and resolution. In this work, a reconfigurable TIM is designed with an adjustable resolution range of 15 down to 0.5 ps and a measurement dynamic range of 480 to 16 ps to perform a variety of time related measurements which require different test specifications; such as set-up and hold time and jitter measurements. It is considered that a reconfigurable measurement system will occupy less chip area than a range of measurement circuits designed for one specific test. The reconfigurable TIM consists of two parts, a programmable time difference amplifier and 32 cells tapped delay line. The proposed programmable time difference amplifier is designed to have a variable gain ranging from 4 to 117 with a very wide dynamic input range. Ahmed Naif M. Alahmadi, Gordon Russell 0002, Alexandre Yakovlev |
DDECS | 3 |
| 2012 | VARMA - VARiability modelling and analysis toolabstractProcess parameter variability in IC manufacturing has become an increasingly important issue as feature scaling descends further into the deep submicron region. Within industry the development of EDA tools associated with “process-aware-design” has a high priority as the impact on circuit performance due to process variations is having increasingly adverse effects on yield and performance. VARMA is a variability analysis tool which enables optimisation of both manufacturing process and nano-electronic circuit design in order to avoid `manufacturing surprises' resulting in costly chip respins, delays in reaching the market place and the subsequent loss of profitability. Gordon Russell 0002, Frank P. Burns, Alexandre Yakovlev |
DDECS | 3 |
| 2012 | A scalable FPGA-based design for field programmable large-scale ion channel simulationsabstractThe design of systems to replicate complex neural functionality is a requirement for the development of next-generation prosthetic devices. The demands of such neural models are growing exponentially as we discover more about how brain systems function. It is therefore important for the electronic architectures involved to scale effectively in terms of latency, area and power usage in order to be able to process more advanced neural models. Within this paper a design is proposed that utilises the parallel nature and the resources available upon modern FPGAs to achieve a scalable and efficient method for the implementation of complex neural models, allowing for the simulation of 150000 ion channels concurrently. Graeme Coapes, Terrence S. T. Mak, Alexandre Yakovlev, Chi-Sang Poon |
FPL | 4 |
| 2012 | Intra-chip physical parameter sensor for FPGAS using flip-flop metastabilityabstractWe present a novel intra-chip physical parameter sensor that exploits the clock-to-q delay response of flip-flops. The proposed design relies on deliberately violating the setup and hold time conditions of a flip-flop to bring it into metastable states and increase its clock-to-q delay. Traditionally, this is an undesired effect because it can result in unpredictable system failures. In this work, this phenomenon is exploited to quantify variations in intra-chip physical parameters. Our design has three benefits over conventional ring-oscillator-based sensors; it consumes less device resources, has a higher precision and does not require a high clock frequency. We present a small-signal model of the proposed sensor and compare its performance with ring oscillators by conducting voltage and temperature-controlled experiments on an Altera Cyclone II FPGA device. Ghaith Tarawneh, Terrence S. T. Mak, Alexandre Yakovlev |
FPL | 3 |
| 2012 | Error detection and correction of single event upset (SEU) tolerant latchabstractSoft errors are a serious concern in state holder circuits at they can cause temporarily malfunctions. C-elements are one of the state holders that are widely used in asynchronous circuits. In this paper, our investigations focus on the vulnerability of different latch types based on C-elements with respect to soft errors. Our aim is to design single event upset (SEU) tolerant latch that has the capability of both detecting and correcting soft errors based on converting single rail to dual-rail configuration and Razor flip flop implementation. In the event of an SEU hitting sensitive nodes and causing the state to temporarily change, an error is generated and a shadow latch restores the correct data. We have demonstrated the functionality of our proposed latch by simulating the design using UMC90nm technology. We also measured error rate of our proposed latch by using an Altera Cyclone II FPGA board. We have obtained the voltage dependence of the error rates. The results show that our proposed latch has less than 1.5% faults propagated at the output. Norhuzaimin Julai, Alexandre Yakovlev, Alexandre V. Bystrov |
IOLTS | 2 |
| 2012 | Ultra-low power transmitterabstractThis paper presents a design of an ultra-low power UWB transmitter based on 4thand 5thderivative Gaussian pulse shapes implemented in UMC 90nm CMOS technology. The simulations show 119mV peak to peak pulse amplitude and the pulse width of 240 ps for the 5thderivative Gaussian pulse and 99.71mV pulse amplitude and 190 ps pulse width for the 4thderivative Gaussian pulse. Power consumption of the pulse generators are calculated 30.11 uW and 21.5 uW for the 5thand 4thderivative Gaussian pulse respectively at a 100MHz pulse repeating frequency (PRF). Ultra-low power radio transmission is important in such application contexts as wireless network nodes and sensors powered by energy harvesters. Mohsen Ghasempour, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 4 |
| 2012 | Mixed Radix Reed-Muller ExpansionsabstractThe choice of radix is crucial for multivalued logic synthesis. Practical examples, however, reveal that it is not always possible to find the optimal radix when taking into consideration actual physical parameters of multivalued operations. In other words, each radix has its advantages and disadvantages. Our proposal is to synthesize logic in different radices, so it may benefit from their combination. The theory presented in this paper is based on Reed-Muller expansions over Galois field arithmetic. The work aims to first estimate the potential of the new approach and to second analyze its impact on circuit parameters down to the level of physical gates. The presented theory has been applied to real-life examples focusing on cryptographic circuits where Galois Fields find frequent application. The benchmark results show that the approach creates a new dimension for the trade-off between circuit parameters and provides information on how the implemented functions are related to different radices. Ashur Rafiev, Andrey Mokhov, Frank P. Burns, Julian P. Murphy, Albert Koelmans, Alexandre Yakovlev |
IEEE Trans. Computers | 6 |
| 2012 | Embedded Transitive Closure Network for Runtime Deadlock Detection in Networks-on-ChipabstractInterconnection networks with adaptive routing are susceptible to deadlock, which could lead to performance degradation or system failure. Detecting deadlocks at runtime is challenging because of their highly distributed characteristics. In this paper, we present a deadlock detection method that utilizes runtime transitive closure (TC) computation to discover the existence of deadlock-equivalence sets, which imply loops of requests in networks-on-chip (NoCs). This detection scheme guarantees the discovery of all true deadlocks without false alarms in contrast with state-of-the-art approximation and heuristic approaches. A distributed TC-network architecture, which couples with the NoC infrastructure, is also presented to realize the detection mechanism efficiently. Detailed hardware realization architectures and schematics are also discussed. Our results based on a cycle-accurate simulator demonstrate the effectiveness of the proposed method. It drastically outperforms timing-based deadlock detection mechanisms by eliminating false detections and, thus, reducing energy wastage in retransmission for various traffic scenarios including real-world application. We found that timing-based methods may produce two orders of magnitude more deadlock alarms than the TC-network method. Moreover, the implementations presented in this paper demonstrate that the hardware overhead of TC-networks is insignificant. Ra'ed Al-Dujaily, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev, Maurizio Palesi |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2011 | Run-time deadlock detection in networks-on-chip using coupled transitive closure networksabstractInterconnection networks with adaptive routing are susceptible to deadlock, which could lead to performance degradation or system failure. Detecting deadlocks at run-time is challenging because of their highly distributed characteristics. In this paper, we present a deadlock detection method that utilizes run-time Transitive Closure (TC) computation to discover the existence of deadlock-equivalence sets, which imply loops of requests in networks-on-chip (NoC). This detection scheme guarantees the discovery of all true deadlocks without false alarms unlike state-of-the-art approximation and heuristic approaches. A distributed TC-network architecture which couples with the NoC architecture is also presented to realize the detection mechanism efficiently. Our results based on a cycle-accurate simulator demonstrate the effectiveness of the TC-network method. It drastically outperforms timing-based deadlock detection mechanisms by eliminating false detections and thus reducing energy dissipation in various traffic scenarios. For example, timing based methods may produce two orders of magnitude more deadlock alarms than the TC-network method. Moreover, the implementations presented in this paper demonstrate that the hardware overhead of TC-networks is insignificant. Ra'ed Al-Dujaily, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev, Maurizio Palesi |
DATE | 4 |
| 2011 | Redressing timing issues for speed-independent circuits in deep submicron ageabstractThe class of speed independent (SI) circuits opens a promising way towards tolerating process variations. However, the fundamental assumption of speed independent circuit is that forks in some wires (usually, large percentage of wires) in such circuits are isochronic; this assumption is more and more challenged by the shrinking technology. This paper suggests a method to generate the weakest timing constraints for a SI circuit to work correctly under bounded delays in wires. The method works for all SI circuits and the generated timing constraints are significantly weaker than those suggested in the current literature claiming the weakest formally proved conditions. Terrence S. T. Mak, Alexandre Yakovlev |
DATE | 3 |
| 2011 | Energy-modulated computingabstractFor years people have been designing electronic and computing systems focusing on improving performance but "keeping power and energy consumption in mind". This is a way to design energy-aware or power-efficient systems, where energy is considered as a resource whose utilization must be optimized in the realm of performance constraints. Increasingly, energy and power turn from optimization criteria into constraints, sometimes as critical as, for example, reliability and timing. Furthermore, quanta of energy or specific levels of power can shape the system's action. In other words, the system's behavior, i.e. the way how computation and communication is carried out, can be determined or modulated by the flow of energy into the system. This view becomes dominant when energy is harvested from the environment. In this paper, we attempt to pave the way to a systematic approach to designing computing systems that are energy-modulated. To this end, several design examples are considered where power comes from energy harvesting sources with limited power density and unstable levels of power. Our design examples include voltage sensors based on self-timed logic and speed-independent SRAM operating in the dynamic range of Vdd 0.2-1V. Overall, this work advocates the vision of designing systems in which a certain quality of service is delivered in return for a certain amount of energy. Alexandre Yakovlev |
DATE | 1 |
| 2011 | Variation tolerant asynchronous FPGA (abstract only)abstractThis paper describes the realization of an interconnect Delay Insensitive (DI) FPGA architecture with distributed asynchronous control. This architecture maintains the basic block structure of traditional FPGAs allowing the potential use of existing FPGA design tools in block design. This asynchronous FPGA architecture is mainly aimed at tolerating the unpredictable delay variations caused by process and environment variations in current and future VLSI technology nodes and also targets low power operations, including modes such as dynamic voltage scaling and variable Vdd, as in applications featuring energy harvesting. This is achieved by making the longer inter-block interconnects DI, keeping the computational logic single-rail, and removing global clocks. Hock Soon Low, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
FPGA | 4 |
| 2011 | Reconfigurable controllers for synchronization via waggingabstractSynchronization via wagging is a method by which a high bandwidth data signal can be partitioned into several lower bandwidth data signals in order to increase the synchronization time of a master-slave latch configuration, and by consequence the mean time between failure for each of the latches in the lower bandwidth array of devices. Furthermore, reconfigurable controller hardware grants the circuit designer direct control over the synchronization time of the array of master-slave latches via the use of one-hot control codes.This work assesses the benefits of unified reconfigurable controller designs over brute force methods when accounting for effects such as process variations. The reconfiguration protocol is discussed, and three separate controller implementations for a wagging synchronizer are compared in a UMC 90 nm technology with operational frequencies of 37 GHz, 19 GHz, and 12 GHz and average power consumption between 1.3 ¼W and 7.5 ¼W per cell in the typical case. Furthermore, the area cost of a unified reconfigurable control device is shown to have a linear growth in complexity as compared the exponential growth present when utilizing selective hardware replication of configurable modes. Conclusions are then drawn outlining future directions for research. James Sebastian Guido, Alexandre Yakovlev |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | Formal modelling and transformations of processor instruction setsabstractInstruction sets of modern processors contain hundreds of instructions defined on a relatively small set of datapath components and distinguished by their codes and the order in which they activate these components. Optimal design of an instruction set for a particular combination of available hardware components and software requirements is crucial for system performance and is a challenging task involving a lot of heuristics and high-level design decisions. The overall design process is significantly complicated by inefficient representation of instructions, which are usually described individually despite the fact that they share a lot of common behavioural patterns. This paper presents a new methodology for compact graph representation of processor instruction sets, which gives the designer a new high-level perspective for reasoning on large sets of instructions without having to look at each of them individually. This opens the way for various transformation and optimisation procedures, which are formally defined and explained on several examples, as well as practically evaluated on an FPGA platform. Andrey Mokhov, Danil Sokolov, Maxim Rykunov, Alexandre Yakovlev |
MEMOCODE | 4 |
| 2011 | Communication centric on-chip power grid models for networks-on-chipabstractAdverse effects of unreliable on-chip power supply delivery are exacerbated due to the rapid shrinking of device dimensions and the ever increasing power consumptions in nanometre-scale integration. Power supply integrity becomes a critical concern. Particularly, on-chip communication networks, such as networks-on-chip (NoC), dictates power dissipations and overall system performance in multi-core systems and emerging embedded computing architectures. These new communication centric architectures require dedicated power grid model that embeds distinctive communication characteristics and spatial parameters for analysing impacts of power supply voltage drop and noise. In this paper, we present a new on-chip power delivery model that captures the on-chip communication patterns and power grid dynamics. This model integrates cycle-accurate simulation of networks-on-chip to analyze the impact of different design entities on power supply noise. The model has been rigorously evaluated. Novel observations of power delivery integrity due to communication network design are presented. This model provides a unique and communication-centric perspective to analyse power supply integrity that leads to future robust and reliable multi-core system design. Nizar Dahir, Terrence S. T. Mak, Alexandre Yakovlev |
VLSI-SoC | 3 |
| 2011 | Flat ArbitersabstractA new way of constructing N-way arbiters is proposed. The main idea is to perform arbitrations between all pairs of requests, and then make decision on what grant to issue based on their outcomes. Crucially, all the mutual exclusion elements in such Andrey Mokhov, Victor Khomenko, Alexandre Yakovlev |
Fundam. Informaticae | 3 |
| 2011 | A Novel Power Delivery Method for Asynchronous Loads in Energy Harvesting SystemsabstractFor systems depending on power harvesting, a fundamental contradiction in the power delivery chain has existed between conventional synchronous computational loads requiring relatively stable Vdd and power harvesters unable to supply it. DC/DC conversion has therefore been an integral part of such systems to resolve this contradiction. On the other hand, asynchronous computational loads, in addition to their potential power-saving capabilities, can be made tolerant to a much wider range of Vdd variance. This may open up opportunities for much more energy efficient methods of power delivery. This article presents in-depth investigations into the behavior and performance of different on-chip power delivery methods driving both asynchronous and synchronous loads directly from a harvester source. A novel power delivery method, which employs a capacitor bank for adaptively storing the energy from power harvesters depending on load and source conditions, is developed. Its advantages, especially when driving asynchronous loads, are demonstrated through comprehensive comparative analysis. Xuefu Zhang, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2011 | Security Evaluation of Balanced 1-of- n CircuitsabstractA new balanced library is presented which consists of novel mixed 1-of-2 and 1-of-4 components based on N-nary logic. Cryptographic circuit specifications are refined and passed to optimization and mapping tools for mapping to a library of power-balanced components. Logic optimization tools are then applied to generate secure synchronous circuits for layout generation. This paper presents a new technique for evaluating the security of such circuits in particular those which offer a higher level of protection. A security metric is introduced which is based on the common selection function that is widely used in differential power analysis attacks and a correlation measure similar to the one used in correlation power analysis attacks. This is used to compare the security level for these kinds of balanced circuits that are more difficult to attack. This paper shows that the circuits generated are more efficient and can offer more security than alternative solutions. Frank P. Burns, Alexandre V. Bystrov, Albert Koelmans, Alexandre Yakovlev |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2010 | Asynchronous design, Quo Vadis?abstractThis talk will briefly discuss the history of the asynchronous design methodology, as its principles and approaches have evolved through the various notions of circuits and systems operating without a global clock, i.e., asynchronous FSMs, speed-independent and delay-insensitive circuits, self-timed systems, multi-synchronous and GALS SOCs. Along the way, and focusing more on the present day hiatus, opportunities, success stories and problems will be discussed as they are being faced in academic research and industrial exploitation. The speaker will be happy to share his experiences of near thirty years in the research he has been involved in the area of self-timed systems and circuits. These will cover designing systems with multiple clock domains, various signalling schemes and protocols, self-timed circuit synthesis, verification, Petri nets, timing-elastic and power-adaptive systems. Should he be particularly adventurous on the day, he might even try to make some predictions into future developments. Alexandre Yakovlev |
DDECS | 1 |
| 2010 | A Reconfigurable Hebbian Eigenfilter for Neurophysiological Spike Train AnalysisabstractThe emergence of multi-electrode array enables the study of real-time neurophysiological activities across multiple regions of the brain. However, the real-time extracellular action potentials recorded on any electrode represent the simultaneous electrical activity of an unknown number of neurons which present a critical challenge to the accuracy of interpretation and identification of the neural circuitry in the subsequent analysis. In this paper, we present a principal component analysis approach utilizing Hebbian eigenfilter to identify the corresponding electrical activities of each neuron, namely spike sorting. The Hebbian eigenfilter greatly simplifies the computational complexity of eigen-projection. An efficient FPGA-based Hebbian eigenfilter is proposed. The performance, accuracy and power consumption of our Hebbian eigenfilter are thoroughly evaluated through synthetic spike trains. The proposal enables real-time spike sorting and analysis, and leads the way towards future motor and cognitive neuroprosthetics. Bo Yu 0014, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev, Yihe Sun, Chi-Sang Poon |
FPL | 5 |
| 2010 | Stochastic analysis of power, latency and the degree of concurrencyabstractConcurrent processing has become the default mode of operation in on-chip systems. Silicon has become cheap enough for having hardware facilities to support very large scale concurrent processing on chip. As a result the availability and applicability of power is becoming more of a limiting factor than logic. However, the advantage of parallelism in reducing power consumption will soon become unrealistic because of the limited scope of reducing Vdd beyond threshold voltage, leaving the reduction of concurrency (through the partial shut-down of system blocks) as a realistic means of reducing power consumption when needed. A stochastic modelling approach is presented in this paper which can integrate the degree of concurrency as a parameter into power and latency analysis. This will facilitate a system design and management regime where the degree of concurrency is used as a means of control to achieve power and performance goals. Yuan Chen 0002, Isi Mitrani, Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 5 |
| 2010 | Asynchronous FPGA architecture with distributed controlabstractAsynchronous techniques have become more significant with continued scaling of VLSI technologies. This paper proposes an asynchronous FPGA architecture. Different from previous methods of introducing asynchrony into FPGAs, our method seeks to preserve the current FPGA cell structure as much as possible, whilst achieving delay insensitivity in the inter-cell interconnects. By using David Cells as the central technique in the delay insensitive clock replacement, this method is conducive to the establishment of an automatic design and synthesis flow. It also particularly caters for low power designs, where current FPGA solutions are not effective yet. Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 3 |
| 2010 | Highly parallel multi-resource arbitersabstractMulti-resource multi-client arbiters are becoming more important in on-chip systems because of the increasing significance of dynamic, run-time, allocation of various system performance resources such as power and computation and communication facilities. Arbiters, for example, can be used to limit the amount of concurrency for regulating voltage droops, and for balancing load and traffic. This paper describes the design of multi-resource arbiters with high degrees of concurrency. By using freezing logic, this design method guarantees correct computation whilst simplifies the implementation. Quick release mechanisms and the implementation of the multi-token concept through the duplication of the client requests help improve the efficiency. Delong Shang, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 3 |
| 2010 | Conditional Partial Order Graphs: Model, Synthesis, and ApplicationabstractThe paper introduces a new formal model for specification and synthesis of control paths in the context of asynchronous system design. The model, called Conditional Partial Order Graph (CPOG), captures concurrency and choice in a system's behavior in a compact and efficient way. It has advantages over widely used interpreted Petri Nets and Finite State Machines for a class of systems which have many behavioral scenarios defined on the same set of actions, e.g., CPU microcontrollers. The CPOG model has potential applications in the area of microcontrol synthesis and brings new methods for modeling concurrency into the application domain of modern and future processor architectures. The paper gives the formal definition of the CPOG model, formulates and solves the problem of CPOG synthesis, and introduces various optimization techniques. The presented ideas can be applied for CPU control synthesis as well as for synthesis of different kinds of event-coordination circuits often used in data coding and communication in digital systems, as demonstrated with several application examples. Andrey Mokhov, Alexandre Yakovlev |
IEEE Trans. Computers | 2 |
| 2010 | Throughput Optimization for Area-Constrained Links With Crosstalk Avoidance MethodsabstractThe effect of crosstalk avoidance codes on the throughput of fixed width communication channels is studied. Closed form expressions of the throughput which incorporate the dimensions of the interconnects and the wiring overheads incurred by such techniques are derived for lines under different buffering conditions. These formulae are utilized to optimize the bandwidth of constrained-area parallel buses under different latency and power constraints. Our results are confirmed by the simulations we have performed in Spectre for a UMC CMOS 90-nm technology. Basel Halak, Alexandre Yakovlev |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Workcraft - A Framework for Interpreted Graph Models
Ivan Poliakov, Victor Khomenko, Alexandre Yakovlev |
Petri Nets | 3 |
| 2009 | Connection-centric network for spiking neural networksabstractA reconfigurable network architecture applied to spiking neural networks is presented. For hardware platforms for neural networks that implement some degree of realism of interest to neuroscientists, connectivity between neurons can be a major limitation. Recent data indicates that neurons in the brain form clusters of connections. Through the combination of this data and a routing scheme that uses a hybrid of short-range direct connectivity and an AER (address event representation) network, the presented architecture aims to provide a useful amount of inter-neuron connectivity. A connection-centric design can provide opportunities for NoCs such as optimising power, bandwidth or introducing redundancy. A method of mapping a network to the architecture is discussed, along with results of optimal hardware specifications for a given set of network parameters. Robin Emery, Alexandre Yakovlev, E. Graeme Chester |
NOCS | 2 |
| 2009 | Synthesis of Nets with Step Firing PoliciesabstractThe unconstrained step semantics of Petri nets is impractical for simulating and modelling applications. In the past, this inadequacy has been alleviated by introducing various flavours of maximally concurrent semantics, as well as priority orders. In this paper, we introduce a general way of controlling step semantics of Petri nets through step firing policies that restrict the concurrent behaviour of Petri nets and so improve their execution and modelling features. In a nutshell, a step firing policy disables at each marking a subset of enabled steps which could otherwise be executed. We discuss various examples of step firing policies and then investigate the synthesis problem for Petri nets controlled by such policies. Using generalised regions of step transition systems, we provide an axiomatic characterisation of those transition systems which can be realised as reachability graphs of Petri nets controlled by a given step firing policy. We also provide two different decision and synthesis algorithms for PT-nets and step firing policies based on linear rewards of steps, where the reward for firing a single transition is either fixed or it depends on the current net marking. The simplicity of the algorithms supports our claim that the proposed approach is practical. Philippe Darondeau, Maciej Koutny, Marta Pietkiewicz-Koutny, Alexandre Yakovlev |
Fundam. Informaticae | 4 |
| 2008 | A Symbolic Algorithm for the Synthesis of Bounded Petri Nets
Josep Carmona 0001, Jordi Cortadella, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Alexandre Yakovlev |
Petri Nets | 6 |
| 2008 | Synthesis of Nets with Step Firing Policies
Philippe Darondeau, Maciej Koutny, Marta Pietkiewicz-Koutny, Alexandre Yakovlev |
Petri Nets | 4 |
| 2008 | Bandwidth-Centric Optimisation for Area-Constrained Links with Crosstalk Avoidance MethodsabstractThe effect of crosstalk avoidance codes on the throughput of fixed width communication channels is studied. Closed form expressions of the throughput which incorporate the dimensions of the interconnects and the wires overheads by such techniques are derived for lines under different buffering conditions. These formulae are utilised to optimise the bandwidth of fixed width parallel buses under different latency and reliability constraints. Our results are confirmed by the simulations we have performed in Spectre for a UMC CMOS 90 nm technology. Basel Halak, Alexandre Yakovlev |
DATE | 2 |
| 2008 | Conditional Partial Order Graphs and Dynamically Reconfigurable Control SynthesisabstractThe paper introduces a new formal model for specifying control paths in the context of asynchronous system design. The model, called conditional partial order graph (CPOG), is capable of capturing concurrency and choice in a system's behaviour in a compact and efficient way. A problem of CPOG synthesis is formulated and solved; various CPOG optimisation techniques are presented. The introduced model can be used for the specification of system behaviour and for synthesis of area-efficient dynamically reconfigurable controllers. The synthesis of a controller is based on a novel generic architecture, called transition sequence encoder (TSE). The synthesized controllers are speed independent and thus very robust to parametric variations. The ideas presented in the paper can be applied for CPU control synthesis as well as for synthesis of different kinds of event-coordination circuits often used in data coding and communication in digital systems. Andrey Mokhov, Alexandre Yakovlev |
DATE | 2 |
| 2008 | Serialized Asynchronous Links for NoCabstractThis paper proposes an asynchronous serialized link for NoC that can achieve the same levels of performance in terms of flits per second as a synchronous link but with a reduced number of wires in the point to point switch links and reduced power consumption. This is achieved by employing serialization in the asynchronous domain as opposed to synchronous to facilitate the removal of global clocking on the serial links. Based on transistor level simulations using 0.12 μm foundry models it has been shown that it is possible to achieve the same level of performance as synchronous but with 75% reduction in wires and 65% reduction in power for a 300 MFlit/s link with 8 buffers with a switch clock speed of 300 MHz. Furthermore the paper presents the design requirements arising from interfacing switches of synchronous NoC and asynchronous serial links. Simon Ogg, Enrico Valli, Bashir M. Al-Hashimi, Alexandre Yakovlev, Crescenzo D'Alessandro, Luca Benini |
DATE | 4 |
| 2008 | Conversion driven design of binary to mixed radix circuitsabstractA conversion driven design approach is described. It takes the outputs of mature and time-proven EDA synthesis tools to generate mixed radix datapath circuits in an endeavour to investigate the added relative advantages or disadvantages. An algorithm underpinning the approach is presented and formally described together with m-of-n encoded gate-level implementations. The application is found in a wide variety and overlapping areas of circuit design, here a subset are analysed where the method finds the strongest application: arithmetic circuits and hardware security. The obtained results are reported showing an increase in power consumption but with considerable improvement in resistance to differential power analysis (DPA). Ashur Rafiev, Julian P. Murphy, Danil Sokolov, Alexandre Yakovlev |
ICCD | 4 |
| 2008 | Implementation of Wave-Pipelined Interconnects in FPGAs
Terrence S. T. Mak, Crescenzo D'Alessandro, N. Pete Sedcole, Peter Y. K. Cheung, Alexandre Yakovlev, Wayne Luk |
NOCS | 5 |
| 2008 | Comments on the BCS Lecture "The Future of Computer Technology and its Implications for the Computer Industry" by Professor Steve FurberabstractDepartment of Electrical and Electronic Engineering, Imperial College London, UK Email: [email protected] Professor Furber has, in a clear and succinct manner, provided us with a pragmatic and honest overview of the challenges facing the computer industry in the future. We have had a brief history of its development. Most of us in the audience, I am sure, are encouraged by achievements made in the last 60 years. Some may even feel proud knowing that they have made their personal contributions. We have been warned of the imminent danger facing the industry, on reliability (or more accurately, unreliability), on escalating costs and on challenging business models that the industry operates under. We are illuminated with some light at the end of the tunnel. In particular, we learn about the UK's efforts in setting up the Microelectronics design Grand Challenges, and the interesting paradigm in computing by learning from nature through the working of the brain. Peter Y. K. Cheung, Alexandre Yakovlev |
Comput. J. | 2 |
| 2008 | Resolution of Encoding Conflicts by Signal Insertion and Concurrency Reduction Based on STG Unfoldings
Victor Khomenko, Agnes Madalinski, Alexandre Yakovlev |
Fundam. Informaticae | 3 |
| 2008 | Analysis of Static Data Flow Structures
Danil Sokolov, Ivan Poliakov, Alexandre Yakovlev |
Fundam. Informaticae | 3 |
| 2008 | Fault-Tolerant Techniques to Minimize the Impact of Crosstalk on Phase Encoded Communication ChannelsabstractAn on-chip intermodule self-timed communication system is considered in which symbols are encoded by means of phase difference between transitions of signals on parallel wires. The reliability of such a channel is governed and significantly lowered by capacitive crosstalk effects between adjacent wires. A more robust high-speed phase-encoded channel can be designed by minimizing its vulnerability to crosstalk noise. This paper investigates the impact of crosstalk on phase-encoded transmission channels. A functional fault model is presented to characterize the problem. Two fault-tolerant schemes are introduced which are based on information redundancy techniques and a partial-order coding concept. The area overheads, performance, and fault-tolerant capability of those methods are compared. It is shown that a substantial improvement in the performance can be obtained for four-wire channels when using the fault-tolerant design approach, at the expense of 25 percent of information capacity per symbol. Basel Halak, Alexandre Yakovlev |
IEEE Trans. Computers | 2 |
| 2007 | A C-element Latch Scheme with Increased Transient Fault Tolerance for Asynchronous CircuitsabstractA technique of constructing dual-rail muller pipelines tolerant to transient faults is proposed. The pipeline datapath is either an NCL-D or NCL-X circuit. A dedicated controller implements the error recovery protocol, which significantly improves fault tolerance with respect to the earlier rail synchronization method. A case study featuring transient fault simulation is presented. K. T. Gardiner, Alexandre Yakovlev, Alexandre V. Bystrov |
IOLTS | 2 |
| 2007 | Impact of strain on the design of low-power high-speed circuitsabstractIn this article, we explore the impact of strain on circuit performance when strained silicon (s-Si) devices are used for designing low-power high-speed circuits. Emphasis has been given on the evaluation of noise characteristics and low-power performance along with the delay characteristics under different channel straining conditions. An inverter circuit has been used for performance evaluation through simulation where the device simulator is calibrated with experimental device data. The result shows a great promise for s-Si technology in digital applications which require high throughput and low power. Hiran Ramakrishnan, Koushik Maharatna, Sanatan Chattopadhyay, Alexandre Yakovlev |
ISCAS | 4 |
| 2007 | NoC Communication Strategies Using Time-to-Digital ConversionabstractA radical approach to high-speed on-chip communication between computational modules is proposed. Data communication is performed over multiple serial buses, where the time difference between events is used to encode and decode data on a number of wires. We present results obtained through a proof-of-concept implementation on FPGA and simulations on a 0.18mum technology Crescenzo D'Alessandro, Nikolaos Minas, Keith Heron, David Kinniment, Alexandre Yakovlev |
NOCS | 5 |
| 2007 | Reducing Interconnect Cost in NoC through Serialized Asynchronous LinksabstractThis work investigates the application of serialization as a means of reducing the number of wires in NoC combined with asynchronous links in order to simplify the clocking of the link. Throughput is reduced but savings in routing area and reduction in power could make this attractive Simon Ogg, Enrico Valli, Crescenzo D'Alessandro, Alexandre Yakovlev, Bashir M. Al-Hashimi, Luca Benini |
NOCS | 4 |
| 2007 | Automating Synthesis of Asynchronous Communication Mechanisms
Kyller Costa Gorgônio, Jordi Cortadella, Fei Xia 0001, Alexandre Yakovlev |
Fundam. Informaticae | 4 |
| 2007 | Direct Mapping of Low-Latency Asynchronous Controllers From STGsabstractA method for an automated synthesis of low-latency asynchronous controllers is presented. It is based on a direct mapping approach and starts from an initial specification in the form of a signal transition graph (STG). This STG is split into a device and an environment, which synchronize via a communication net that models wires. The device is represented as a tracker and a bouncer. The tracker follows the state of the environment and provides reference points to the device outputs. The bouncer communicates with the environment and generates output events in response to the input events according to the state of the tracker. This two-level architecture provides an efficient interface to the environment and is convenient for subsequent mapping into a circuit netlist. A set of optimization heuristics is developed to reduce the latency and size of the circuit. As a result of this paper, a software tool called OptiMist has been developed. Its low algorithmic complexity allows large specifications to be synthesized, which is not possible for the tools based on state-space exploration. OptiMist successfully interfaces conventional EDA design flow for simulation, timing analysis, and place-and-route Danil Sokolov, Alexandre V. Bystrov, Alexandre Yakovlev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2007 | Measuring Deep Metastability and Its Effect on Synchronizer PerformanceabstractPresent measurement techniques do not allow synchronizer reliability to be measured in the region of most interest, that is, beyond the first half cycle of the synchronizer clock. We describe methods of extending the measurement range, in which the number of metastable events generated is increased by four orders of magnitude and events with long metastable times are selected from the large number of more normal events. The relationship found between input times and the resulting output times is dependent on accurate measurement of input time distributions with deviations of less than 10 ps. We show how the distribution of to clock times at the input can be characterized in the presence of noise and how predictions of failure rates for long synchronizer times can be made. Anomalies such as the increased failure rates in a master-slave synchronizer produced by the back edge of the clock are explained and demonstrated. David Kinniment, Charles E. Dike, Keith Heron, Gordon Russell 0002, Alexandre Yakovlev |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2007 | Registers for Phase Difference Based LogicabstractA logic design style known as phase difference-based logic (PDBL) has several benefits with respect to security and testing. An existing design method for PDBL circuits has so far been lacking an important component, a register. In this paper, we present the design of a speed independent PDBL register and a timed PDBL register, which can be used in asynchronous or synchronous circuits. Comparisons are presented in terms of speed, size, and power consumption. Delong Shang, Alexandre Yakovlev, Albert Koelmans, Danil Sokolov, Alexandre V. Bystrov |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Low-Cost Online Testing of Asynchronous HandshakesabstractA new low-cost low-complexity checker for online testing of asynchronous interfaces in globally-asynchronous locally-synchronous circuits is proposed. The solution is fully based upon the standard gate libraries. The checker itself is fully offline testable. It also provides a fault-locating functionality, which is achieved by combining the online mode with scan techniques Delong Shang, Alexandre Yakovlev, Frank P. Burns, Fei Xia 0001, Alexandre V. Bystrov |
ETS | 2 |
| 2006 | Cost-aware synthesis of asynchronous circuits based on partial acknowledgementabstractDesigning asynchronous circuits by reusing existing synchronous tools has become a promising solution to the problem of poor CAD support in asynchronous world. A straightforward way is to structurally map the gates in a synchronous netlist to their functionally equivalent modules which use delay-insensitive codes. Different trade-offs exist in previous methods between the overheads of the implementations and their robustness. The aim of this paper is to optimise the area of asynchronous circuits using partial acknowledgement concept. We employ this concept in two design flows, which are implemented in a software tool to evaluate the efficiency of the method. The benchmark results show the average reduction in area by 28% and in the number of inter-functional module wires that require timing verification by 67%, compared to NCL-X. Yu Zhou 0006, Danil Sokolov, Alexandre Yakovlev |
ICCAD | 3 |
| 2006 | Online Testing by Protocol DecompositionabstractComparison between synchronous and asynchronous models leads to a protocol-based fault model for asynchronous circuits. Protocol monitoring of the control path is separated from data comparison in the data path. A novel protocol decomposition technique is used to extract simple protocols from behaviour of a complex circuit. This technique is implemented as a software tool. An asynchronous checker model, implementation and simulation results are presented. Coverage of internal faults of the checker is calculated Deepali Koppad, Danil Sokolov, Alexandre V. Bystrov, Alexandre Yakovlev |
IOLTS | 4 |
| 2006 | Virtual self-timed blocks for systems-on-chipabstractIntellectual properties (IP cores) are widely used as pre-designed and reusable units in various system-on-chip (SOC) designs, but their integration has presented difficulties for system designers. In this paper, we propose an approach to better reuse IP cores while maintain energy efficiency for SOC systems. Here we employ a so called self-timed event processor (STEP) to make each IP core into a virtual self-timed block. Much of the IP cores' pre-designed properties can be preserved and the new SOC systems that use virtual self-timed blocks can be more energy efficient. A MATLAB based investigation is carried out on an example STEP processor. Yuan Chen 0002, Fei Xia 0001, Alexandre Yakovlev |
ISCAS | 3 |
| 2006 | Logic Synthesis for Asynchronous Circuits Based on STG Unfoldings and Incremental SAT
Victor Khomenko, Maciej Koutny, Alexandre Yakovlev |
Fundam. Informaticae | 3 |
| 2006 | Buffered Asynchronous Communication Mechanisms
Fei Xia 0001, Ian G. Clark, Alexandre Yakovlev, E. Graeme Chester |
Fundam. Informaticae | 4 |
| 2005 | A Multi-version Data Model and Semantic-Based Transaction Processing Protocol
Alexandre Yakovlev |
ADBIS | 1 |
| 2005 | Modeling and Verification of Globally Asynchronous and Locally Synchronous Ring ArchitecturesabstractThe paper demonstrates a prevalent global deadlock situation resulting from a local deadlock in a GALS (globally asynchronous and locally synchronous) ring architecture. We present a novel design for building systems which are tolerant to such deadlocks arising in the local modules. The paper concentrates on the modeling of the proposed design methodology and its correctness is proved with the help of a public domain verification tool. Sohini Dasgupta, Alexandre Yakovlev |
DATE | 2 |
| 2005 | Power-Balanced Self Checking Circuits for Cryptographic ChipsabstractCryptographic chips are highly susceptible to fault injection and power analysis attacks, which easily lets an attacker gain secret keys intended to be secure. In addition to these issues, testability circuitry is frequently manipulated to induce faults and undesired behavior. When power-balanced dual-rail (1-of-2) logic, a return-to-spacer protocol and power-balanced totally self checking checkers with redundant transistors are used together, they significantly improve and enforce security; simultaneously facilitating secure on-line testability. In this paper, we propose and show how to implement such circuits in cryptographic chips, the fruits of which are high reliability, testability and security. Julian P. Murphy, Alexandre V. Bystrov, Alexandre Yakovlev |
IOLTS | 3 |
| 2005 | On-Line Testing of Globally Asynchronous CircuitsabstractThe problem of on-line testing of asynchronous circuits is analyzed, several infrastructures are proposed and a self-checking tree checker is designed. The checker uses on-demand self-test, which reduces power consumption and guarantees bounded self-test period. Objects under test are tested by observing protocols at their primary inputs and outputs. The checker does not slow down the functional system, as it only samples the signals. The protocols are checked by identifying enabled and refused signal transitions in each state of the system. The fault coverage of internal faults of the checker is calculated. Simulation results are included. Delong Shang, Alexandre V. Bystrov, Alexandre Yakovlev, Deepali Koppad |
IOLTS | 3 |
| 2005 | Design and Analysis of Dual-Rail Circuits for Security ApplicationsabstractDual-rail encoding, return-to-spacer protocol, and hazard-free logic can be used to resist power analysis attacks by making energy consumed per clock cycle independent of processed data. Standard dual-rail logic uses a protocol with a single spacer, e.g., all-zeros, which gives rise to energy balancing problems. We address these problems by incorporating two spacers; the spacers alternate between adjacent clock cycles. This guarantees that all gates switch in every clock cycle regardless of the transmitted data values. To generate these dual-rail circuits, an automated tool has been developed. It is capable of converting synchronous netlists into dual-rail circuits and it is interfaced to industry CAD tools. Dual-rail and single-rail benchmarks based upon the advanced encryption standard (AES) have been simulated and compared in order to evaluate the method and the tool. Danil Sokolov, Julian P. Murphy, Alexandre V. Bystrov, Alexandre Yakovlev |
IEEE Trans. Computers | 4 |
| 2004 | Improving the Security of Dual-Rail Circuits
Danil Sokolov, Julian P. Murphy, Alexandre V. Bystrov, Alexandre Yakovlev |
CHES | 4 |
| 2004 | An Asynchronous Synthesis Toolset Using VerilogabstractWe present a new CAD tool set for generating asynchronous circuits from high-level Verilog level-sensitive specifications. Initially, high-level Verilog descriptions are compiled and converted into a novel intermediate Petri net format. The intermediate format is subsequently passed to optimization tools and mapping tools where it is directly mapped into asynchronous datapath and control circuits using David cells (DCs). Finally, logic optimization tools are applied to generate speed-independent (SI) circuits. The speed independent circuits generated perform well compared to circuits generated by existing asynchronous tools. Frank P. Burns, Delong Shang, Albert Koelmans, Alexandre Yakovlev |
DATE | 4 |
| 2004 | MATLAB Models of ACMS in Control Systems
Fei Xia 0001, E. Graeme Chester, Alexandre Yakovlev, Ian G. Clark |
ICINCO (3) | 4 |
| 2004 | Detecting State Encoding Conflicts in STG Unfoldings Using SAT
Victor Khomenko, Maciej Koutny, Alexandre Yakovlev |
Fundam. Informaticae | 3 |
| 2004 | Design and Analysis of a Self-Timed Duplex Communication SystemabstractCommunication-centric design is a key paradigm for systems-on-chips (SoCs), where most computing blocks are predesigned IP cores. Due to the problems with distributing a clock across a large die, future system designs are more asynchronous or self-timed. For portable, battery-run applications, power and pin efficiency is an important property of a communication system where the cost of a signal transition on a global interconnect is much greater than for internal wires in logic blocks. We address this issue by designing an asynchronous communication system aimed at power and pin efficiency. Another important issue of SoC design is design productivity. It demands new methods and tools, particularly for designing communication protocols and interconnects. The design of a self-timed communication system is approached employing formal techniques supported by verification and synthesis tools. The protocol is formally specified and verified with respect to deadlock-freedom and delay-insensitivity using a Petri-net-based model-checking tool. A protocol controller has been synthesized by a direct mapping of the Petri net model derived from the protocol specification. The logic implementation was analyzed using the Cadence toolkit. The results of SPICE simulation show the advantages of the direct mapping method compared to logic synthesis. Alexandre Yakovlev, Steve Furber, René Krenz, Alexandre V. Bystrov |
IEEE Trans. Computers | 1 |
| 2003 | Visualization and Resolution of Coding Conflicts in Asynchronous Circuit Design
Agnes Madalinski, Alexandre V. Bystrov, Victor Khomenko, Alexandre Yakovlev |
DATE | 4 |
| 2003 | STG Optimisation in the Direct Mapping of Asynchronous Circuits
Danil Sokolov, Alexandre V. Bystrov, Alexandre Yakovlev |
DATE | 3 |
| 2002 | Visualization of Partial Order Models in VLSI Design FlowabstractSummary form only given. A new method, algorithms and tool for the visualisation of a finite complete prefix (FCP) of a Petri net (PN) or a signal transition graph are presented. A transformation is defined that converts such a prefix into a two-level model. At the top level, it has a finite state machine (FSM), describing modes of operation and transitions between them. At the low level, there are marked graphs, which can be drawn as waveforms, embedded into the top level nodes. The models of both levels are abstractions traditionally used by electronics engineers. The resultant model is completed trace equivalent to the original prefix. Moreover, the branching structure of the latter is preserved as much as possible. Alexandre V. Bystrov, Maciej Koutny, Alexandre Yakovlev |
DATE | 3 |
| 2002 | Detecting State Coding Conflicts in STGs Using Integer ProgrammingabstractThe paper presents a new method for checking unique and complete state coding, the crucial conditions in the synthesis of asynchronous control circuits from signal transition graphs (STGs). The method detects state coding conflicts in an STG using its partial order semantics (unfolding prefix) and an integer programming technique. This leads to huge memory savings compared to methods based on reachability graphs, and also to significant speedups in many cases. In addition, the method produces execution paths leading to an encoding conflict. Finally, the approach is extended to checking the normalcy property of STGs, which is a necessary condition for their implementability using gales whose characteristic functions, are monotonic. Victor Khomenko, Maciej Koutny, Alexandre Yakovlev |
DATE | 3 |
| 2002 | Lazy transition systems and asynchronous circuit synthesis withrelative timing assumptionsabstractThis paper presents a design flow for timed asynchronous circuits. It introduces lazy transitions systems as a new computational model to represent the timing information required for synthesis. The notion of laziness explicitly distinguishes between the enabling and the firing of an event in a transition system. Lazy transition systems can be effectively used to model the behavior of asynchronous circuits in which relative timing assumptions can be made on the occurrence of events. These assumptions can be derived from the information known a priori about the delay of the environment and the timing characteristics of the gates that will implement the circuit. The paper presents the necessary conditions to generate circuits and a synthesis algorithm that exploits the timing assumptions for optimization. It also proposes a method for back-annotation that derives a set of sufficient timing constraints that guarantee the correctness of the circuit. Jordi Cortadella, Michael Kishinevsky, Steven M. Burns, Alex Kondratyev, Luciano Lavagno, Kenneth S. Stevens, Alexander Taubin, Alexandre Yakovlev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2000 | WCET Analysis of Superscalar Processors Using Simulation With Coloured Petri Nets
Frank P. Burns, Albert Koelmans, Alexandre Yakovlev |
Real Time Syst. | 3 |
| 2000 | Synchronous and asynchronous A-D conversionabstractAnalog-digital (A-D) converters with a fixed conversion time are subject to errors due to metastability. It is shown that an asynchronous converter in which the conversion time is not bounded is faster, on average, than the synchronous design. Real-time applications require the data to be produced within a fixed time, and failures may occur with the long conversion times that can arise with fully asynchronous converters. For these applications, we show that an internally asynchronous bounded time converter is both faster and more reliable than a synchronous converter. David Kinniment, Alexandre Yakovlev |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | Automatic Synthesis and Optimization of Partially Specified Asynchronous SystemsabstractA method for automating the synthesis of asynchronous control circuits from high level (CSP-like) and/or partial STG (involving only functionally critical events) specifications is presented.The method solves two key subtasks in this new, more flexible, design flow: handshake expansion, i.e. inserting reset events with maximum concurrency, and event reshuffling under interface and concurrency constraints, by means of concurrency reduction.In doing so, the algorithm optimizes the circuit both for size and performance.Experimental results show a significant increase in the solution space explored when compared to existing CSP-based or STG-based synthesis tools. Alex Kondratyev, Jordi Cortadella, Michael Kishinevsky, Luciano Lavagno, Alexandre Yakovlev |
DAC | 5 |
| 1999 | What is the cost of delay insensitivity?abstractDeep submicron technology calls for new design techniques, in which wire and gate delays are accounted to have equal or nearly equal effect on circuit behaviour. Asynchronous speed-independent (SI) circuits, whose behaviour is only robust to gate delay variations, may be too optimistic. On the other hand, building circuits totally delay-insensitive (DI), for both gates and wires, is impractical. The paper presents an approach for automated synthesis of globally DI and locally SI circuits. It is based on order relaxation, a simple graphical transformation of a circuit's behavioural specification, for which the Signal Transition Graph, an interpreted Petri net, is used. The method is successfully tested on a set of benchmarks and a realistic design example. It proves effective showing average cost of DI interfacing at about 40% for area and 20% for speed. Hiroshi Saito, Alex Kondratyev, Jordi Cortadella, Luciano Lavagno, Alexandre Yakovlev |
ICCAD | 5 |
| 1999 | Asynchronous microprocessors: From high level model to FPGA implementation
Lee Lloyd, Keith Heron, Albert Koelmans, Alexandre Yakovlev |
J. Syst. Archit. | 4 |
| 1999 | Logic decomposition of speed-independent circuitsabstractLogic decomposition is a well-known problem in logic synthesis, but it poses new challenges when targeted to speed-independent circuits. The decomposition of a gate into smaller gates must preserve not only the functional correctness of a circuit but also speed independence, i.e., hazard freedom under unbounded gate delays. This paper presents a new method for logic decomposition of speed-independent circuits that solves the problem in two major steps: (1) logic decomposition of complex gates and (2) insertion of new signals that preserve hazard freedom. The method is shown to be more general than previous approaches and its effectiveness is evaluated by experiments on a set of benchmarks. Alex Kondratyev, Jordi Cortadella, Michael Kishinevsky, Luciano Lavagno, Alexandre Yakovlev |
Proc. IEEE | 5 |
| 1999 | Decomposition and technology mapping of speed-independent circuits using Boolean relationsabstractThis paper presents a new technique for decomposition and technology mapping of speed-independent circuits. An initial circuit implementation is obtained in the form of a netlist of complex gates, which may not be available in the design library. The proposed method iteratively performs Boolean decomposition of each such gate F into a two-input combinational or sequential gate G available in the library and two gates H/sub 1/ and H/sub 2/ simpler than F, while preserving the original behavior and speed-independence of the circuit. To extract functions for H/sub 1/ and H/sub 2/ the method uses Boolean relations as opposed to the less powerful algebraic factorization approach used in previous methods. After logic decomposition, the overall library matching and optimization is carried out. Logic resynthesis, performed after speed-independent signal insertion for H/sub 1/ and H/sub 2/, allows for sharing of decomposed logic. Overall, this method is more general than the existing techniques based on restricted decomposition architectures, and thereby leads to better results in technology mapping. Jordi Cortadella, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Enric Pastor, Alexandre Yakovlev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 1998 | Unfolding and Finite Prefix for Nets with Read Arcs
Walter Vogler, Alexei L. Semenov, Alexandre Yakovlev |
CONCUR | 3 |
| 1998 | Lazy transition systems: application to timing optimization of asynchronous circuitsabstractThis paper introduces bzy Transitions Systems &zTSs).The notion of laziness exDlicitlv distinguishes behveen the enabling and the firing of an e;ent in"a transition system.LzTSS can be effectively used to model the behavior of asynchronous circuits in whicfi relative timing assumptions cm-be made on the occurrence of events.These assumptions can be derived from the information known a priori about fie de]ay of fie environment and the timing characteristics of the gates that will implement the circuit.The paper presents necessary conditions to synthesize circuits with a correct behavior under the given timing assumr3tions.Preliminary results show that significant area and performance improvements can be obtained by exploiting the extra "don't care" space implicitly provided by the Iazmess of the events. Jordi Cortadella, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Alexander Taubin, Alexandre Yakovlev |
ICCAD | 6 |
| 1998 | Designing Control Logic for Counterflow Pipeline Processor Using Petri Nets
Alexandre Yakovlev |
Formal Methods Syst. Des. | 1 |
| 1998 | Analysing Superscalar Processor Architectures with Coloured Petri Nets
Frank P. Burns, Albert Koelmans, Alexandre Yakovlev |
Int. J. Softw. Tools Technol. Transf. | 3 |
| 1998 | Deriving Petri Nets for Finite Transition SystemsabstractThis paper presents a novel method to derive a Petri net from any specification model that can be mapped into a state-based representation with arcs labeled with symbols from an alphabet of events (a Transition System, TS). The method is based on the theory of regions for Elementary Transition Systems (ETS). Previous work has shown that, for any ETS, there exists a Petri Net with minimum transition count (one transition for each label) with a reachability graph isomorphic to the original Transition System. Our method extends and implements that theory by using the following three mechanisms that provide a framework for synthesis of safe Petri nets from arbitrary TSs. First, the requirement of isomorphism is relaxed to bisimulation of TSs, thus extending the class of synthesizable TSs to a new class called Excitation-Closed Transition Systems (ECTS). Second, for the first time, we propose a method of PN synthesis for an arbitrary TS based on mapping a TS event into a set of transition labels in a PN. Third, the notion of irredundant region set is exploited, to minimize the number of places in the net without affecting its behavior. The synthesis method can derive different classes of place-irredundant Petri Nets (e.g., pure, free choice, unique choice) from the same TS, depending on the constraints imposed on the synthesis algorithm. This method has been implemented and applied in different frameworks. The results obtained from the experiments have demonstrated the wide applicability of the method. Jordi Cortadella, Michael Kishinevsky, Luciano Lavagno, Alexandre Yakovlev |
IEEE Trans. Computers | 4 |
| 1998 | Hazard-free implementation of speed-independent circuitsabstractThis paper develops a theoretical framework for the hazard-free gate-level implementation of speed-independent circuits specified by event-based models, such as signal transition graphs (for processes with AND causality and input choice) or their extension, called change diagrams (which allow OR-causality). It presents sufficient conditions, called the generalized monotonous cover requirements, for a hazard-free circuit to be built within a standard implementation structure. This structure consists of two-level simple-gate combinational logic and a row of latches, either a C-element or an RS-latch. A set of semantic-preserving transformations is defined that can be applied to an original behavioral description of the circuit so as to produce its specification in the form that satisfies the monotonous cover requirement. The transformations are applied at the event-based representation level (to avoid state explosion) and proved to be effective. The main result of the paper is therefore twofold: 1) the proof that any speed-independent behavior can be implemented at the gate level without hazards and 2) an efficient method for constructing such an implementation. Experimental results show that the proposed method compares very favorably, in area and performance, to the previously known techniques. Alex Kondratyev, Michael Kishinevsky, Alexandre Yakovlev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1997 | Synthesis of Speed-Independent Circuits from STG-Unfolding SegmentabstractThis paper presents a novel technique for synthesis of speed-independentcircuits. It is based on partial order representation ofthe state graph called STG-unfolding segment. The new methoduses approximation technique to speed up the synthesis process.The method is illustrated on the basic implementation architecture.Experimental results demonstrating its efficiency are presented anddiscussed. Alexei L. Semenov, Alexandre Yakovlev, Enric Pastor, Marco A. Peña, Jordi Cortadella |
DAC | 2 |
| 1997 | Decomposition and technology mapping of speed-independent circuits using Boolean relationsabstractPresents a new technique for the decomposition and technology mapping of speed-independent circuits. An initial circuit implementation is obtained in the form of a netlist of complex gates, which may not be available in the design library. The proposed method iteratively performs Boolean decomposition of each such gate F into a two-input combinational or sequential gate G, which is available in the library, and two gates H/sub 1/ and H/sub 2/, which are simpler than F, while preserving the original behavior and speed-independence of the circuit. To extract functions for H/sub 1/ and H/sub 2/, the method uses Boolean relations, as opposed to the less powerful algebraic factorization approach used in previous methods. After logic decomposition, overall library matching and optimization is carried out. Logic resynthesis, performed after speed-independent signal insertion for H/sub 1/ and H/sub 2/, allows for the sharing of decomposed logic. Overall, this method is more general than existing techniques based on restricted decomposition architectures, and thereby leads to better results in technology mapping. Jordi Cortadella, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Enric Pastor, Alexandre Yakovlev |
ICCAD | 6 |
| 1997 | A region-based theory for state assignment in speed-independent circuitsabstractState assignment problems still need satisfactory solutions to make asynchronous circuit synthesis more practical. A well-known example of such a problem is that of complete state coding (CSC), which happens when a pair of different states in a specification has the same binary encoding. A standard way to approach state coding conflicts is to insert new state signals into the original specification in such a way that the original behavior remains intact. This paper proposes a method which improves over existing approaches by coupling generality, optimality, and efficiency. The method is based on the use of a class of "ground objects", called regions, that play the role of a bridge between state-based specifications (transition systems, TS's) and event-based specifications (signal transition graphs, STG's), We need to deal with both types of specification because designers usually prefer a timing diagram-like notation, such as STG, while optimization and cost analysis work better at the state level. A region in a transition system is a set of states that corresponds to a place in an STG (or the underlying Petri net). Regions are tightly connected with a set of properties that are to be preserved across the state encoding process, namely, 1) trace equivalence between the original and the encoded specification, and 2) implementability as a speed-independent circuit. We will build on a theoretical body of work that has shown the significance of regions for such property-preserving transformations, and describe a set of algorithms aimed at efficiently solving the encoding problem. The algorithms have been implemented in a software tool called petrify. Unlike many existing tools, petrify represents the encoded specification as an STG. This significantly improves the readability of the result (compared to a state-based description in which concurrency is represented implicitly by interleaving), and allows the designer to be more closely involved in the synthesis process. The efficiency of the method is demonstrated on a number of "difficult" examples. Jordi Cortadella, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Alexandre Yakovlev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 1996 | Methodology and Tools for State Encoding in Asynchronous Circuit SynthesisabstractThis paper proposes a state encoding method for asynchronous circuits based on the theory of regions. A region in a Transition System is a set of states that “behave uniformly” with respect to a given transition (value change of an observable signal), and is analogue to a place in a Petri net. Regions are tightly connected with a set of properties that must be preserved across the state encoding process, namely: (1) trace equivalence between the original and the encoded specification, and (2) implementability as a speed-independent circuit. We build on a theoretical body of work that has shown the significance of regions for such property-preserving transformations, and describe a set of algorithms aimed at efficiently solving the encoding problem. The algorithms have been implemented in a software tool called petrify. Unlike many existing tools, petrify represents the encoded specification as an STG, and thus allows the designer to be more closely involved in the synthesis process. The efficiency of the method is demonstrated on a number of “difficult” examples. Jordi Cortadella, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Alexandre Yakovlev |
DAC | 5 |
| 1996 | Verification of asynchronous circuits using Time Petri Net unfoldingabstractArticle Free Access Share on Verification of asynchronous circuits using time Petri net unfolding Authors: Alexei Semenov Department of Computing Science, University of Newcastle, Newcastle upon Tyne NE1 7RU, England Department of Computing Science, University of Newcastle, Newcastle upon Tyne NE1 7RU, EnglandView Profile , Alexandre Yakovlev Department of Computing Science, University of Newcastle, Newcastle upon Tyne NE1 7RU, England Department of Computing Science, University of Newcastle, Newcastle upon Tyne NE1 7RU, EnglandView Profile Authors Info & Claims DAC '96: Proceedings of the 33rd annual Design Automation ConferenceJune 1996 Pages 59–62https://doi.org/10.1145/240518.240530Published:01 June 1996Publication History 24citation370DownloadsMetricsTotal Citations24Total Downloads370Last 12 Months8Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Alexei L. Semenov, Alexandre Yakovlev |
DAC | 2 |
| 1996 | On the Models for Asynchronous Circuit Behaviour with OR Causality
Alexandre Yakovlev, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Marta Pietkiewicz-Koutny |
Formal Methods Syst. Des. | 1 |
| 1996 | A Unified Signal Transition Graph Model for Asynchronous Control Circuit Synthesis
Alexandre Yakovlev, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
Formal Methods Syst. Des. | 1 |
| 1996 | Modelling, analysis and synthesis of asynchronous control circuits using Petri nets
Alexandre Yakovlev, Albert Koelmans, Alexei L. Semenov, David Kinniment |
Integr. | 1 |
| 1995 | On hazard-free implementation of speed-independent circuitsabstractNo abstract available. Alex Kondratyev, Michael Kishinevsky, Alexandre Yakovlev |
ASP-DAC | 3 |
| 1995 | Synthesizing Petri nets from state-based modelsabstractThis paper presents a method to synthesize labeled Petri nets from state-based models. Although state-based models (such as finite state machines) are a powerful formalism to describe the behavior of sequential systems, they cannot explicitly express the notions of concurrency, causality and conflict Petri nets can naturally capture these notions. The proposed method in based on deriving an elementary transition system (ETS) from a specification model. Previous work has shown that for any ETS there exists a Petri net with minimum transition count (one transition for each label) with a reachability graph isomorphic to the original ETS. This paper presents the first known approach to obtain an ETS from a non-elementary TS and derive a place-irredundant Petri net. Furthermore, by imposing constraints on the synthesis method, different classes of Petri nets can be derived from the same reachability graph (pure, free choice, unique choice). This method has been implemented and efficiently applied in different frameworks: Petri net composition, synthesis of Petri nets from asynchronous circuits, and resynthesis of Petri nets. Jordi Cortadella, Michael Kishinevsky, Luciano Lavagno, Alexandre Yakovlev |
ICCAD | 4 |
| 1994 | Basic Gate Implementation of Speed-Independent CircuitsabstractExisting methods for synthesis of speedindependent circuits under unbounded delay model have difficulties in combining the generality of formal approach with the practicality of the implementation architectures used at the logic level.This paper presents a characteristic property of the state graph specification, called Monotonous Cover requirement, implying its hazard-free implementation within the standard structure of a two-level SOP logic and a row of latches.The overall synthesis procedure ensures satisfiability of this condition by applying the generalised state assignment approach. Alex Kondratyev, Michael Kishinevsky, Bill Lin 0001, Peter Vanbekbergen, Alexandre Yakovlev |
DAC | 5 |
| 1994 | A low latency asynchronous arbitration circuitabstractWe present an asynchronous circuit for an arbiter cell that can be used to construct cascaded multiway arbitration circuits. The circuit is completely speed-independent. It has a short response delay at the input request-grant handshake link due to both a) the propagation of requests in parallel with starting arbitration and b) the concurrent resetting of request-grant handshakes in different cascades of a request-grant propagation chain.> Alexandre Yakovlev, A. Petrov, Luciano Lavagno |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1992 | A unified signal transition graph model for asynchronous control circuit synthesisabstractBoth low-level (analysis-oriented) and high-level (specification-oriented) models for asynchronous circuits and the environment where they operate, together with strong equivalence results between the properties at the low levels, are described. One interesting side result is the precise characterization of classical static and dynamic hazards in terms of the model. Consequently the designer can check the specification and directly decide if the behavior of any implementation will depend, e.g., on the delays of the signals described by such specification.> Alexandre Yakovlev, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
ICCAD | 1 |
| 1992 | On Limitations and Extensions of STG Model for Designing Asynchronous Control CircuitsabstractA number of limitations of the current status of signal transition graphs (STGs), a model which has recently become popular for designing asynchronous interface circuits, are discussed. The major syntactic and semantic restrictions that can be lifted are safety, free-choice net structure and binary signal labeling. A number of instructive examples of interface control circuit specifications, which are semantically correct yet free from such restrictions, are presented. Adequate techniques for analysis and implementation of the extended STG model are discussed.> Alexandre Yakovlev |
ICCD | 1 |
| 1989 | Analyzing Semantics of Concurrent Hardware Specifications
Leonid Ya. Rosenblum, Alexandre Yakovlev |
ICPP (3) | 2 |
| 1988 | Signal Graphs: A Model for Designing Concurrent Logic
Alex Kondratyev, Leonid Ya. Rosenblum, Alexandre Yakovlev |
ICPP (1) | 3 |