Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ashur Rafiev

dblp:52/350 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
2since 2021 · last 2024
0000-0002-7387-5970ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 35% Hardware accelerators and domain-specific architectures · 27% Cloud and datacenter computing · 18%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Human-computer interaction and pervasive computing
1 paper
Health and well-being technologies · 77% Design research and methods · 23%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
0.712023
REDRESS: Generating Compressed Models for Edge Inference Using Tsetlin Machines · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.712023
REDRESS: Generating Compressed Models for Edge Inference Using Tsetlin Machines · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Cloud and datacenter computing › resource management
dynamic resource management
0.412020
PARMA: Parallelization-Aware Run-Time Management for Energy-Efficient Many-Core Systems · IEEE Trans. Computers 2020
Parallel and multicore computing
many-core systems
0.412020
PARMA: Parallelization-Aware Run-Time Management for Energy-Efficient Many-Core Systems · IEEE Trans. Computers 2020
Parallel and multicore computing
processor allocation
0.412020
PARMA: Parallelization-Aware Run-Time Management for Energy-Efficient Many-Core Systems · IEEE Trans. Computers 2020
Health and well-being technologies
mental health technology
0.212016
Challenges for Designing new Technology for Health and Wellbeing in a Complex Mental Healthcare Context · CHI 2016
Electronic design automation
logic synthesis
0.112012
Mixed Radix Reed-Muller Expansions · IEEE Trans. Computers 2012
Electronic design automation › logic synthesis
multivalued logic synthesis
0.112012
Mixed Radix Reed-Muller Expansions · IEEE Trans. Computers 2012
Electronic design automation › logic synthesis
reed-muller expansion
0.112012
Mixed Radix Reed-Muller Expansions · IEEE Trans. Computers 2012
Design research and methods › experience design
experience-centered design
0.112016
Challenges for Designing new Technology for Health and Wellbeing in a Complex Mental Healthcare Context · CHI 2016
Integrated circuit design
digital circuit design
0.012012
Mixed Radix Reed-Muller Expansions · IEEE Trans. Computers 2012
Integrated circuit design
finite field arithmetic
0.012012
Mixed Radix Reed-Muller Expansions · IEEE Trans. Computers 2012

Methods — techniques the papers use, named apart from their topics

include-encoding · 1.3bit-parallel inference · 1.3algorithm-hardware co-design · 1.3performance counter modeling · 0.4energy-delay product · 0.4energy per instruction · 0.4experience-centered design · 0.2mixed radix expansion · 0.1galois field arithmetic · 0.1
YearPublicationVenuePosition
2024 An Event-Driven Approach to Genotype Imputation on a Custom RISC-V Cluster
abstract
This article proposes an event-driven solution to genotype imputation, a technique used to statistically infer missing genetic markers in DNA. The work implements the widely accepted Li and Stephens model, primary contributor to the computational complexity of modern x86 solutions, in an attempt to determine whether further investigation of the application is warranted in the event-driven domain. The model is implemented using graph-based Hidden Markov Modeling and executed as a customized forward/backward dynamic programming algorithm. The solution uses an event-driven paradigm to map the algorithm to thousands of concurrent cores, where events are small messages that carry both control and data within the algorithm. The design of a single processing element is discussed. This is then extended across multiple cores and executed on a custom RISC-V NoC cluster called POETS. Results demonstrate how the algorithm scales over increasing hardware resources and a multi-core run demonstrates a 270X reduction in wall-clock processing time when compared to a single-threaded x86 solution. Optimisation of the algorithm via linear interpolation is then introduced and tested, with results demonstrating a wall-clock reduction time of ∼ 5 orders of magnitude when compared to a similarly optimised x86 solution.
Jordan Morris, Ashur Rafiev, Graeme M. Bragg, Mark Vousden, David B. Thomas, Alexandre Yakovlev, Andrew D. Brown
IEEE ACM Trans. Comput. Biol. Bioinform.2
2023 REDRESS: Generating Compressed Models for Edge Inference Using Tsetlin Machines
abstract
Inference at-the-edge using embedded machine learning models is associated with challenging trade-offs between resource metrics, such as energy and memory footprint, and the performance metrics, such as computation time and accuracy. In this work, we go beyond the conventional Neural Network based approaches to explore Tsetlin Machine (TM), an emerging machine learning algorithm, that uses learning automata to create propositional logic for classification. We use algorithm-hardware co-design to propose a novel methodology for training and inference of TM. The methodology, called REDRESS, comprises independent TM training and inference techniques to reduce the memory footprint of the resulting automata to target low and ultra-low power applications. The array of Tsetlin Automata (TA) holds learned information in the binary form as bits: {0,1}, called excludes and includes, respectively. REDRESS proposes a lossless TA compression method, called the include-encoding, that stores only the information associated with includes to achieve over 99% compression. This is enabled by a novel computationally minimal training procedure, called the Tsetlin Automata Re-profiling, to improve the accuracy and increase the sparsity of TA to reduce the number of includes, hence, the memory footprint. Finally, REDRESS includes an inherently bit-parallel inference algorithm that operates on the optimally trained TA in the compressed domain, that does not require decompression during runtime, to obtain high speedups when compared with the state-of-the-art Binary Neural Network (BNN) models. In this work, we demonstrate that using REDRESS approach, TM outperforms BNN models on all design metrics for five benchmark datasets viz. MNIST, CIFAR2, KWS6, Fashion-MNIST and Kuzushiji-MNIST. When implemented on an STM32F746G-DISCO microcontroller, REDRESS obtained speedups and energy savings ranging 5-5700× compared with different BNN models.
Sidharth Maheshwari, Tousif Rahman, Rishad A. Shafik, Alexandre Yakovlev, Ashur Rafiev, Lei Jiao 0001, Ole-Christoffer Granmo
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 PARMA: Parallelization-Aware Run-Time Management for Energy-Efficient Many-Core Systems
abstract
Performance and energy efficiency considerations have shifted computing paradigms from single-core to many-core architectures. At the same time, traditional speedup models such as Amdahl's Law face challenges in the run-time reasoning for system performance and energy efficiency, because these models typically assume limited variations of the parallel fraction. Moreover, the parallel fraction, which varies dynamically in workloads, is generally unknown at run-time without application-level instrumentation. This article describes novel performance/energy trade-off models based on realistic architectural considerations, which describe the parallel fraction and speedup as functions of performance counter values available in modern processors, removing the need for application-level instrumentation. These are then used to develop a Parallelization-Aware Run-time Management (PARMA) approach. PARMA aims at controlling core allocations and operating voltage/frequency points for energy efficiency, according to the varying workload parallel fractions. The efficacy of our models and the PARMA approach is extensively validated using a number of PARSEC benchmark applications, involving two performance/energy trade-off metrics: energy-delay-product (EDP), typically used in high-performance applications and energy per instruction (EPI), suitable for energy-aware applications. Up to 48 and 68 percent improvements in EDP and EPI have been observed using the PARMA approach compared with parallelization-agnostic methods.
Mohammed A. Noaman Al-Hayanni, Ashur Rafiev, Fei Xia 0001, Rishad A. Shafik, Alexander B. Romanovsky, Alexandre Yakovlev
IEEE Trans. Computers2
2016 Challenges for Designing new Technology for Health and Wellbeing in a Complex Mental Healthcare Context
abstract
This paper describes the challenges and lessons learned in the experience-centered design (ECD) of the Spheres of Wellbeing, a technology to promote the mental health and wellbeing of a group of women, suffering from significant mental health problems and living in a medium secure hospital unit. First, we describe how our relationship with mental health professionals at the hospital and the aspirations for person-centric care that we shared with them enabled us, in the design of the Spheres, to innovate outside traditional healthcare procedures. We then provide insights into the challenges presented by the particular care culture and existing services and practices in the secure hospital unit that were revealed through our technology deployment. In discussing these challenges, our design enquiry opens up a space to make sense of experience living with complex mental health conditions in highly constrained contexts within which the deployment of the Spheres becomes an opportunity to think about wellbeing in similar contexts.
Anja Thieme, John C. McCarthy 0002, Paula Johnson, Stephanie Phillips, Jayne Wallace, Siân E. Lindley, Karim Ladha, Daniel Jackson 0002, Diana Nowacka, Ashur Rafiev, Cassim Ladha, Thomas Nappey, Mathew Kipling, Peter C. Wright, Thomas D. Meyer, Patrick Olivier
CHI10
2016 Selective abstraction and stochastic methods for scalable power modelling of heterogeneous systems
abstract
With the increase of system complexity in both platforms and applications, power modelling of heterogeneous systems is facing grand challenges from the model scalability issue. To address these challenges, this paper studies two systematic methods: selective abstraction and stochastic techniques. The concept of selective abstraction via black-boxing is realised using hierarchical modelling and cross-layer cuts, respecting the concepts of boxability and error contamination. The stochastic aspect is formally underpinned by Stochastic Activity Networks (SANs). The proposed method is validated with experimental results from Odroid XU3 heterogeneous 8-core platform and is demonstrated to maintain high accuracy while improving scalability.
Ashur Rafiev, Fei Xia 0001, Alexei Iliasov, Rem Gensh, Ali Aalsaud, Alexander B. Romanovsky, Alexandre Yakovlev
FDL1
2016 Power-Aware Performance Adaptation of Concurrent Applications in Heterogeneous Many-Core Systems
abstract
Modern embedded systems execute multiple applications, both sequentially and concurrently. These applications are exercised on heterogeneous platforms generating varying power consumption and system workloads (CPU or memory intensive or both). As a result, determining the most energy-efficient system configuration (i.e. the number of parallel threads, their core allocations and operating frequencies) tailored for each kind of workload and application scenario is extremely challenging. In this paper, we propose a novel runtime optimization approach with the aim of achieving maximized power normalized performance considering dynamic variation of workload and application scenarios. Fundamental to this approach is a comprehensive study to investigate the tradeoffs between inter-application concurrency with performance and power consumption under different system configurations. Using real experimental measurements on an Odroid XU-3 heterogeneous platform with a number of PARSEC benchmark applications, we model power normalized performance (in terms of IPS/Watt) underpinning analytical power and performance models, derived through multivariate linear regression (MLR). Using these models, we show that with increasing number of concurrent CPU intensive applications show variable gains in IPS/Watt compared to the memory intensive applications in both sequential and concurrent application scenarios. Furthermore, we demonstrate that it is possible to continuously adapt system configuration through a low-cost and linear-complexity runtime algorithm, which can improve the IPS/Watt by up to 125% compared to the existing approach.
Ali Aalsaud, Rishad A. Shafik, Ashur Rafiev, Fei Xia 0001, Sheng Yang 0003, Alexandre Yakovlev
ISLPED3
2015 A Formal Specification and Prototyping Language for Multi-core System Management
abstract
We relate the experience of a defining a formal domain specific language (DSL) for the construction and reasoning about OS-level management logic of multi-core systems. The approach is based on a novel, iterative development principle where results of prototyping studies feed back into the next language revision. We illustrate the DSL with several examples of executable scripts.
Alexei Iliasov, Ashur Rafiev, Fei Xia 0001, Rem Gensh, Alexander B. Romanovsky, Alexandre Yakovlev
PDP2
2013 BinCam: Designing for Engagement with Facebook for Behavior Change
Rob Comber, Anja Thieme, Ashur Rafiev, Nick Taylor 0002, Nicole C. Krämer, Patrick Olivier
INTERACT (2)3
2012 Mixed Radix Reed-Muller Expansions
abstract
The choice of radix is crucial for multivalued logic synthesis. Practical examples, however, reveal that it is not always possible to find the optimal radix when taking into consideration actual physical parameters of multivalued operations. In other words, each radix has its advantages and disadvantages. Our proposal is to synthesize logic in different radices, so it may benefit from their combination. The theory presented in this paper is based on Reed-Muller expansions over Galois field arithmetic. The work aims to first estimate the potential of the new approach and to second analyze its impact on circuit parameters down to the level of physical gates. The presented theory has been applied to real-life examples focusing on cryptographic circuits where Galois Fields find frequent application. The benchmark results show that the approach creates a new dimension for the trade-off between circuit parameters and provides information on how the implemented functions are related to different radices.
Ashur Rafiev, Andrey Mokhov, Frank P. Burns, Julian P. Murphy, Albert Koelmans, Alexandre Yakovlev
IEEE Trans. Computers1
2008 Conversion driven design of binary to mixed radix circuits
abstract
A conversion driven design approach is described. It takes the outputs of mature and time-proven EDA synthesis tools to generate mixed radix datapath circuits in an endeavour to investigate the added relative advantages or disadvantages. An algorithm underpinning the approach is presented and formally described together with m-of-n encoded gate-level implementations. The application is found in a wide variety and overlapping areas of circuit design, here a subset are analysed where the method finds the strongest application: arithmetic circuits and hardware security. The obtained results are reported showing an increase in power consumption but with considerable improvement in resistance to differential power analysis (DPA).
Ashur Rafiev, Julian P. Murphy, Danil Sokolov, Alexandre Yakovlev
ICCD1