Tughrul Arslan

dblp:32/362 · DBLP profile ↗
← Back
150ranked-venue papers
0as first author
29since 2021 · last 2026
0000-0001-8176-5803ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 84 · 10 since 2021Artificial intelligence and machine learning · 32 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 since 2021Software engineering, systems software and programming languages · 7 · 1 since 2021Computer networks · 4 · 3 since 2021
YearPublicationVenuePosition
2026 Dynamic multi-kernel convolutional network with noise injected features for audio-only speech enhancement
Nasir Saleem, Sami Bourouis, Fazal-E. Wahab, Kia Dashtipour, Tughrul Arslan, Amir Hussain 0001
Neurocomputing5
2026 FVOR-YOLO: A Real-Time Model for Fruits and Vegetables Detection in Complex Supermarket Self-Checkout Environments
abstract
Recent advances in Computer Vision (CV) and Artificial Intelligence (AI) have rendered automated Object Detection (OD) an indispensable component of modern supermarket self-checkout (SCO) systems, particularly for barcode-less loose items such as fruits and vegetables. However, the effective deployment of these systems presents a significant technical challenge, which is to simultaneously achieve high accuracy and real-time inference speed on resource-constrained edge hardware. To address this challenge, this paper introduces FVOR-YOLO, a novel detection architecture derived from the YOLOv12 framework. The proposed model incorporates a lightweight MobileViTv2 module to replace key components of the original YOLOv12 backbone. This approach not only enhances global modelling capabilities but also preserves the local inductive biases of Convolutional Neural Networks (CNNs), resulting in a powerful yet lightweight structure. Furthermore, FVOR-YOLO integrates a parallel attention module, PCS (Parallel Channel and Spatial Group-wise Enhance), to selectively amplify salient CA and SGE features, thereby enhancing feature discriminability for OD. To address the problem of limited datasets, a bespoke fruit and vegetables image dataset is built to emulate complex supermarket SCO environments. The proposed model’s efficacy is confirmed through empirical validation. On a high-performance NVIDIA RTX 5080 GPU, the model achieves a mAP@[0.5:0.95] of 66.1%. This represents a significant 4.2% performance gain over the YOLOv12 baseline. The high accuracy is delivered alongside a rapid inference speed of 123.1 FPS. Furthermore, when deployed on the NVIDIA Jetson AGX Orin edge device, the model maintains its efficacy. The model achieves a mAP@[0.5:0.95] of 63.5%, representing a 4.3% improvement over YOLOv12, and delivers a real-time inference speed of 74.1 FPS. Collectively, these results validate FVOR-YOLO as an efficient and practical solution, particularly for edge AI applications in retail. The model is well-suited for advanced SCO systems that operate with limited computational resources.
Jiabin Jia, Tughrul Arslan
IEEE Internet Things J.3
2026 Multimodal Cognitive Load Estimation With Radio Frequency Sensing and Pupillometry in Complex Auditory Environments
abstract
The detection of listening effort or cognitive load (CL) has been a major research challenge in recent years. Most conventional techniques utilise physiological or audio-visual sensors and are privacy-invasive and computationally complex. The challenges of synchronization, data alignment and accessibility limitations potentially increase the noise and error probability, compromising the accuracy of CL estimates. This innovative work presents a multi-modal, non-invasive and privacy-preserving approach that combines Radio Frequency (RF) and pupillometry sensing to address these challenges. Custom RF sensors are first designed and developed to capture blood flow changes in specific brain regions with high spatial resolution. Next, multi-modal fusion with pupillometry sensing is proposed and shown to offer a robust assessment of cognitive and listening effort through pupil size and pupil dilation. Our novel approach evaluates RF sensing to estimate CL from cerebral blood flow variations utilizing pupillometry as a baseline. A first-of-its-kind, multi-modal dataset is collected as a new benchmark resource in a controlled environment with participants to comprehend target speech with varying background noise levels. The framework is statistically evaluated using intraclass correlation for pupillometry data (average ICC> 0.95). The correlation between pupillometry and RF data is established through Pearson's correlation (average PCC> 0.79). Further, CL is classified into high and low categories based on RF data using K-means clustering. Future work involves integrating RF sensors with glasses to estimate listening effort for hearing-aid users and utilising RF measurements to optimize speech enhancement based on individual's listening effort and complexity of acoustic environment.
Usman Anwar, Adeel Hussain, Mandar Gogate, Kia Dashtipour, Tughrul Arslan, Amir Hussain 0001, Peter Lomax
IEEE J. Biomed. Health Informatics5
2025 Investigating Gender Bias in Text-to-Audio Generation Models
Aarish Shah Mohsin, Mohammad Nadeem, Shahab Saquib Sohail, Tughrul Arslan, Mandar Gogate, Nasir Saleem, Amir Hussain 0001
INTERSPEECH4
2025 Anchored by Sound: Indoor Trajectory Mapping with Activity-Driven Audio Anchors
abstract
Indoor trajectory mapping is essential for applications such as health status monitoring, context-aware computing, and tracking systems. These applications often require accurate location information in environments where the Global Navigation Satellite System (GNSS) is unavailable or unreliable. Infrastructure-free approaches, such as inertial sensor-based Pedestrian Dead Reckoning (PDR), have gained popularity due to their ease of deployment and low cost. However, a key limitation of PDR is not only its susceptibility to cumulative drift over time but also its lack of absolute position references, which prevents reliable alignment of the trajectory with real-world coordinates. To address this challenge, we propose an activity-driven audio-IMU fusion solution that uses Factor Graph Optimization (FGO) to accurately reconstruct indoor walking trajectories on a room-scale floor plan. Specifically, we use two fixed microphone arrays to estimate absolute coordinates from activity-driven audio events. These events act as sparse, high-confidence anchors to map PDR trajectories. We validate our system through experiments in a 6 m × 6 m room. The proposed fusion solution achieves a mean positioning error of 0.31 m, with 68% and 95% of errors below 0.38 m and 0.53 m, respectively. The results demonstrate the potential of sparse, activity-driven audio anchoring to improve infrastructure-free indoor trajectory mapping.
Bingnan Duan, Yinhuan Dong, Ilari Vallivaara, Tughrul Arslan
IPIN4
2025 Ancestry Tree Clustering for Particle Filter Diversity Maintenance
abstract
We propose a method for linear-time diversity maintenance in particle filtering. It clusters particles based on ancestry tree topology: closely related particles in sufficiently large subtrees are grouped together. The main idea is that the tree structure implicitly encodes similarity without the need for spatial or other domain-specific metrics. This approach, when combined with intra-cluster fitness sharing and the protection of particles not included in a cluster, effectively prevents premature convergence in multimodal environments while maintaining estimate compactness. We validate our approach in a multimodal robotics simulation and a real-world multimodal indoor environment. We compare the performance to several diversity maintenance algorithms from the literature, including Deterministic Resampling and Particle Gaussian Mixtures. Our algorithm achieves high success rates with little to no negative effect on compactness, showing particular robustness to different domains and challenging initial conditions.
Ilari Vallivaara, Bingnan Duan, Yinhuan Dong, Tughrul Arslan
IPIN4
2025 Wearable RF Sensing System with Edge AI Inference for In-vivo Cognitive Load Classification
abstract
Cognitive load (CL) refers to the mental effort required to process information. Monitoring increased CL is important as it can indicate the onset of cognitive decline and neurodegeneration, allowing for early intervention and management. Traditional techniques using bio-physiological and audio-visual sensors are privacy-invasive and computationally complex. The synchronization, data alignment, and accessibility problems with these techniques can lead to increased noise and errors, reducing the accuracy of CL estimates. This paper presents a first-of-its-kind Radio Frequency (RF) based sensing system that effectively monitors and classifies cognitive load states by detecting cerebral blood flow variations through backscattered RF signal strength. The RF sensors are designed and miniaturized using microwave computational software and fabricated sensors are integrated with glasses to estimate in-vivo CL variations with on-edge processing and classification. The system is validated through user-oriented audio-only (AO) and audio-visual (AV) trials. Participants are tested on their ability to comprehend target speech with different levels of background noise, assessing the impact on CL. The statistical features from the RF reflection data are processed using machine learning (ML) and deep learning (DL) algorithms implemented on a Raspberry Pi. The participants’ CL is classified into high, medium and low categories separately for AO and AV trials. The Multilayer Perceptron (MLP) achieves an overall accuracy of 85% for AO trials and 66.2% for AV trials, with average training and testing times of 6 seconds and 0.001 seconds, respectively. The promising results suggest that the system is an effective, portable, and low-cost alternative for on-edge CL estimation.
Usman Anwar, Yinhuan Dong, Tughrul Arslan, Amir Hussain 0001, Peter Lomax
ISCAS3
2025 Comparative Analysis of AI-driven Feature Selection and Machine Learning Framework for Efficient Radar Fall Detection
abstract
Radar sensors are increasingly viable for fall detection, yet optimizing feature selection and machine learning (ML) model efficiency remains essential for practical deployment. While the literature extensively explores complex AI models for feature extraction and fall detection, a notable gap exists in optimizing the feature selection process for radar data to enhance computational efficiency. This paper addresses this gap by presenting a comprehensive comparative evaluation of AI-driven feature selection techniques -Variance Threshold (VT), Spearman Correlation, Recursive Feature Elimination (RFE) and embedded methods - applied across various ML classifiers to balance accuracy and processing speed for fall detection. Using FMCW radar data from 400 fall and 500 non-fall samples collected from volunteers aged 26 to 78, results demonstrate that tree-based classifiers, particularly Extra Trees with RFE, provide an optimal balance, achieving a 0.08 ms runtime with 98% accuracy, ideally suited for real-time precision-critical fall detection. Additionally, Random Forest and Gradient Boosting with VT method achieved comparable accuracy of 97.22% with runtimes of 0.12 and 0.17 ms, respectively. This work presents a scalable framework for fall detection with insights into the strengths and limitations of various feature selection strategies, advancing real-time radar-based fall detection and embedded system design in healthcare applications.
Nazia Gillani, Tughrul Arslan
ISCAS2
2025 A Robust Multifloor Wi-Fi Fingerprinting Approach Based on Geospatial Cells for Indoor Localization
abstract
Wi-Fi fingerprinting uses wireless signals to locate smartphone users indoors effectively. This method is popular because many buildings have numerous wireless access points (APs) available. However, due to the complex multifloor indoor environment with its inherent signal dynamics, it suffers from: 1) computationally inefficient and complex algorithms designed to overcome the spatial unevenness and feature sparsity of the dataset; 2) high sensitivity to environmental dynamics; and 3) low accuracies due to signal fluctuations and inconsistencies between the physical and signal domains. To address these issues, we formulated a robust Wi-Fi fingerprinting approach for multifloor indoor localization that accurately finds the user’s residing floor and location coordinates. The computational complexities are reduced by organizing the reference points (RPs) into geospatial cells (G-cells), extracting useful statistical features, and finding the localized cells (LCs) leveraging the knowledge of common APs with the test point (CATP) and the nearest RP (N-RP). The algorithm removes outlier N-RPs, and the localized KNN algorithm (L-KNN) finds the real-time user’s coordinates accurately and efficiently. The floor identification algorithm, comprising two modules, segregates LCs into a layered format and utilizes information from neighboring floors’ APs common to the test point (TP) to overcome spatial unevenness between training and test samples. Extensive experiments are conducted on three real-world datasets. The results demonstrate that the proposed approach has superior localization accuracy and robustness compared to the best methods in the literature.
Sahibzada Muhammad Ahmad Umair, Ayesha Kanwal, Haseeb Hussain, Tughrul Arslan
IEEE Internet Things J.4
2025 A deep learning approach for non-invasive Alzheimer's monitoring using microwave radar data
abstract
Over 50 million people globally suffer from Alzheimer's disease (AD), emphasizing the need for efficient, early diagnostic tools. Traditional methods like Magnetic Resonance Imaging (MRI) and Computed Tomography (CT) scans are expensive, bulky, and slow. Microwave-based techniques offer a cost-effective, non-invasive, and portable solution, diverging from conventional neuroimaging practices. This article introduces a deep learning approach for monitoring AD , using realistic numerical brain phantoms to simulate scattered signals via the CST Studio Suite. The obtained data is preprocessed using normalization, standardization, and outlier removal to ensure data integrity. Furthermore, we propose a novel data augmentation technique to enrich the dataset across various AD stages. Our deep learning approach combines Recursive Feature Elimination (RFE) with Principal Component Analysis (PCA) and Autoencoders (AE) for optimal feature selection. Convolution Neural Network (CNN) is combined with Gated Recurrent Unit (GRU), Bidirectional Long Short Term Memory (Bidirectional-LSTM), and Long Short-Term Memory (LSTM) to improve classification performance. The integration of RFE-PCA-AE significantly elevates performance, with the CNN+GRU model achieving an 87% accuracy rate, thus outperforming existing studies.
Farhatullah, Xin Chen 0012, Deze Zeng, Rahmat Ullah, Rab Nawaz, Jiafeng Xu, Tughrul Arslan
Neural Networks7
2024 Edged based audio-visual speech enhancement demonstrator
Song Chen 0005, Mandar Gogate, Kia Dashtipour, Jasper Kirton-Wingate, Adeel Hussain, Faiyaz Doctor, Tughrul Arslan, Amir Hussain 0001
INTERSPEECH7
2024 Enhanced Pedestrian Trajectory Reconstruction Using Bidirectional Extended Kalman Filter and Automatic Refinement
abstract
Pedestrian trajectory reconstruction is essential for applications such as human behaviour analysis and smart transportation systems. Existing methods often struggle with data gaps and inaccuracies due to the limitations of location-acquisition technologies like GPS and WiFi. While some algorithms have been developed to estimate trajectories using IMU data, they only provide relative information and lack absolute location context, necessitating additional efforts and technologies to compensate for this limitation. This paper proposes an enhanced pedestrian trajectory reconstruction method that leverages GPS and IMU data to provide accurate location information across outdoor and indoor areas. The method integrates a bi-directional Extended Kalman Filter (EKF) to fuse GPS and IMU data, iterating forward and backward trajectories. An automatic refinement approach is applied to merge the forward and backward trajectories against multiple residuals, resulting in the final pedestrian trajectory. Evaluations demonstrate the method’s effectiveness using all available GPS data and under challenging conditions with only 5% GPS data. When using all GPS data, the method achieves a median pointwise distance of 4.76 meters and a $\mathbf{6 8 \%}$ CDF error of 7.10 meters. With 5% GPS data, the median pointwise distance is 10.95 meters, with a 68% CDF error of 13.82 meters. These results highlight the method’s robustness and potential for practical applications where GPS data may be sparse or unreliable.
Yinhuan Dong, Kiros Kwan, Tughrul Arslan
IPIN3
2024 Saying goodbyes to rotating your phone: Magnetometer calibration during SLAM
abstract
While Wi-Fi positioning is still more common indoors, using magnetic field features has become widely known and utilized as an alternative or supporting source of information. Magnetometer bias presents significant challenge in magnetic field navigation and SLAM. Traditionally, magnetometers have been calibrated using standard sphere or ellipsoid fitting methods and by requiring manual user procedures, such as rotating a smartphone in a figure-eight shape. This is not always feasible, particularly when the magnetometer is attached to heavy or fast-moving platforms, or when user behavior cannot be reliably controlled. Recent research has proposed using map data for calibration during positioning. This paper takes a step further and verifies that a pre-collected map is not needed; instead, calibration can be done as part of a SLAM process. The presented solution uses a factorized particle filter that factors out calibration in addition to the magnetic field map. The method is validated using smartphone data from a shopping mall and mobile robotics data from an office environment. Results support the claim that magnetometer calibration can be achieved during SLAM with comparable accuracy to manual calibration. Furthermore, the method seems to slightly improve manual calibration when used on top of it, suggesting potential for integrating various calibration approaches.
Ilari Vallivaara, Yinhuan Dong, Tughrul Arslan
IPIN3
2024 Context-Aware Audio-Visual Speech Enhancement Based on Neuro-Fuzzy Modeling and User Preference Learning
abstract
It is estimated that by 2050 approximately one in ten individuals globally will experience disabling hearing impairment. In the presence of everyday reverberant noise, a substantial proportion of individual users encounter challenges in speech comprehension. This study introduces a novel application of neuro-fuzzy modeling that synergizes and fuses audio-visual speech enhancement (AV SE) with an initial user preference learning based framework. Specifically, our approach uniquely integrates multimodal AV speech data with innovative SE methods and fuzzy inferencing techniques. This integration is further enriched by incorporating a user-preference learning model that adapts to environmental and user-specific contexts, including signal-to-noise ratios, sound power, and the quality of visual information. The proposed framework facilitates the incorporation of clinical measures such as user cognitive load (or listening effort) with real-world uncertainty to steer the system outputs. We employ an adaptive fuzzy neural network to derive the most effective Sugeno fuzzy inference model, employing particle swarm optimization to ensure optimal SE by considering sound power, ambient noise levels, and visual quality. Experimental results utilize our new benchmark AV multitalker challenge dataset to demonstrate the superiority of our user preference-informed, context-aware AV SE approach in enhancing speech intelligibility and quality in challenging noisy conditions, marking a significant advancement over conventional methods while reducing energy consumption. The conclusion supports the ecological scalability of our approach and its potential for real-world applications, setting a new benchmark in AV SE research, paving the way for future assistive hearing and communication technologies.
Song Chen 0005, Jasper Kirton-Wingate, Faiyaz Doctor, Usama Arshad, Kia Dashtipour, Mandar Gogate, Zahid Halim, Ahmed Yassin Al-Dubai, Tughrul Arslan, Amir Hussain 0001
IEEE Trans. Fuzzy Syst.9
2023 5G-IoT Cloud based Demonstration of Real-Time Audio-Visual Speech Enhancement for Multimodal Hearing-aids
Ankit Gupta 0008, Abhijeet Bishnu, Mandar Gogate, Kia Dashtipour, Tughrul Arslan, Ahsan Adeel, Amir Hussain 0001, Tharmalingam Ratnarajah, Mathini Sellathurai
INTERSPEECH5
2023 Towards Two-point Neuron-inspired Energy-efficient Multimodal Open Master Hearing Aid
Adewale Adetomi, Khubaib Ahmed, Amir Hussain 0001, Tughrul Arslan, Ahsan Adeel
INTERSPEECH5
2023 A Multimodal Graph Fingerprinting Method for Indoor Positioning Systems
abstract
WiFi fingerprinting has been extensively studied for years to provide location estimation in indoor scenarios. Researchers have used various machine learning algorithms to match the online fingerprint to the pre-collected offline fingerprints, which have location labels, for location estimation. However, neither conventional machine learning algorithms nor modern deep neural networks explore the geometric relations between WiFi access points and the location where the fingerprint was taken. Therefore, they cannot capture the non-Euclidean nature of the WiFi fingerprint. Furthermore, prior research has indicated that fusing multiple modalities can improve positioning performance compared to using only WiFi. Therefore, this study proposes a novel Multimodal Graph Fingerprinting method for indoor positioning systems. The proposed method constructs a multimodal graph at the location of the user’s smart terminal by fusing radio frequency signals, electromagnetic field (EMF) strength, and inertial sensor measurements. A hierarchical deep graph neural network is developed to learn the relations between the multimodal graphs and their locations by capturing the features of the identities (such as MAC addresses, WiFi Received Signal Strength (RSS), and EMF data) and the topology information. Experiments on a real dataset built on a university campus demonstrate that the proposed model can achieve a median positioning error of 2.1m by fusing different modalities.
Yinhuan Dong, Tughrul Arslan, Yunjie Yang 0001
IPIN2
2023 Live Demonstration: Unlocking the Potential of Two-Point Neuronal Cells for Energy-Efficient Training of Deep Networks
abstract
This paper is the first live demonstration of the transformative computational potential of context-sensitive two-point layer 5 pyramidal cells (L5PCs). We will showcase a Multi-Processor System on Chip (MPSoC)-based implementation of a biologically plausible L5PC-driven deep neural network (DNN), termed multisensory cooperative computing (MCC). This will be shown to effectively process heterogeneous real-world audio-visual data consuming far less energy compared to state-of-the-art ‘point’ neuron-driven DNNs. Our approach opens new cross-disciplinary avenues for future on-chip DNN training implementations and posits a radical shift in current neuromorphic computing paradigms.
Ahsan Adeel, Adewale Adetomi, William A. Phillips, Khubaib Ahmed, Amir Hussain 0001, Tughrul Arslan
ISCAS7
2023 Wearable RF Sensing and Imaging System for Non-invasive Vascular Dementia Detection
abstract
Vascular dementia is the second most common form of dementia prevalent in old age groups, and is also one of the leading causes of mortality. Timely diagnosis and detection of vascular dementia is critical to avoid brain damage. Brain imaging is an essential tool for diagnosis and determines future treatment options available to the patient. Currently, Magnetic Resonance Imaging (MRI), Positron Emission Tomography (PET), Computerized Tomography (CT) scan and carotid ultrasound are mainly used for brain imaging and vascular dementia diagnosis. However, these technologies are expensive, require extensive medical supervision and are not easily accessible. This can cause substantial delay in diagnosis and potentially lead to irreversible damage. This paper presents a first-of-its-kind, miniaturized octagonal monopole-patch antenna (OMPA) sensor for vascular dementia detection. It is designed to operate as part of a portable device and can effectively diagnose underlying causes like brain infarction, stroke and blood clots at the initial stage. The developed sensor designs are validated using microwave computational software and fabricated models are experimentally verified using artificial stroke and blood clot targets inside an artificial brain model. Simulated and measured reflection coefficient results are consistent, and target objects are detected successfully. The findings show that the prototype device is viable as an efficient, portable and low-cost alternate for vascular dementia detection.
Usman Anwar, Tughrul Arslan, Amir Hussain 0001, Peter Lomax
ISCAS2
2023 Live Demonstration: Cloud-based Audio-Visual Speech Enhancement in Multimodal Hearing-aids
abstract
Hearing loss is among the most serious public health problems, affecting as much as 20% of the worldwide population. Even cutting-edge multi-channel audio-only speech enhancement (SE) algorithms used in modern hearing aids face significant hurdles since they typically magnify noises while failing to boost speech understanding in crowded social environments. Recently, for the first time we proposed a novel integration of 5G cloud-radio access network, internet of things (IoT), and strong privacy algorithms to develop 5G IoT enabled hearing aid (HA) [1]. In this demonstration, we show the first-ever transceiver (PHY layer) model for cloud-based audio-visual (AV) SE, which meets the requirements for high data rate and low latency of forthcoming multi-modal HAs (such as Google glasses with integrated HAs). Even in highly noisy conditions like cafés, clubs, conferences, meetings, etc., the transceiver [2] transmits raw AV information from a hearing aid system to a cloud-based platform and obtains a clear signal. In Fig. 1, we illustrate an example of our cloud-based AV SE hearing aid demonstration. Herein, the left-side computer and Universal Software Radio Peripheral (USRP) x310 function as IoT systems (hearing aids), the right-side USRP serves as an access point or base station, and the right-side computer serves as a cloud-server for operating NN-based SE models. Please take note that the channel between the HA device and the cloud is defined as an uplink channel, whereas the channel between the access point (cloud) and the HA device is defined as a downlink channel. Given the time-varying sensitivity of the data received at HA devices, the uplink channel can therefore handle a variety of data rates. As a result, a customized long-term evolution (LTE)-based frame structure is developed for uplink transmission of data. It provides error-correction codes in the 1.4 MHz and 3 MHz bandwidths with a variety of modulations and code rates. Furthermore, the cloud access point simply supports a limited transmission rate because it only transmits audio data to the HA equipment. In order to support real-time AV SE, a modified frame structure for LTE with 1.4 MHz of bandwidth is developed. The AV SE algorithm receives cropped lip images of the target speaker and a noisy speech spectrogram, and it produces an ideal binary mask that lessens the noise-dominant regions while improving the speech-dominant areas. We use the depth-wise separable convolutions, reduced STFT window size of 32 ms, smaller STFT window shift of 8 ms, and 64 convolutions in the audio feature extraction layers of our Cochlea-Net [3] multi-modal AV SE neural network architecture to reduce processing latency. Furthermore, the visual feature extraction framework is employed. Our proposed architecture can handle streaming data frame-by-frame. Thus, the users will experience for the first time the real-world development of a physical layer transceiver that can perform AV SE in real-time under strict latency and data rate requirements. For this demonstration, we will bring two computers and two USRP x310 devices.
Abhijeet Bishnu, Ankit Gupta 0008, Mandar Gogate, Kia Dashtipour, Tughrul Arslan, Ahsan Adeel, Amir Hussain 0001, Mathini Sellathurai, Tharmalingam Ratnarajah
ISCAS5
2023 Wearable RF Device for Monitoring Brain Activities in the Ageing Population
abstract
With the increasing rate of ageing population increase, the prevalence of age-related brain diseases is also increasing at a fast rate. The deterioration of brain-related tasks and cognitive abilities is one of the major challenges being faced by the ageing population today. It is noted that electrical signals that are transmitted by neurons in the brain, are a good way to indicate the health and performance of the brain for an individual. However, current systems, such as EEGs and f-MRIs are bulky, time-consuming, and overall uncomfortable or inconvenient for older people to use. In this paper, a wearable RF device is presented that utilises integrated textile-based RF sensors and a flexible switching circuit to monitor the electrical activity in the brain. The device was validated with a brain phantom model that was fabricated with materials that mimicked the dielectric properties of an actual brain. In addition, a function generator was used in order to generate voltage signals at 4 Hz, 12 Hz, and 35 Hz in the phantom, which emulates the presence of electrical activity in the brain. Results showed the reflected signals captured by the sensors were able to differentiate between the voltage signals that were generated at different frequency levels. In addition, the flexible switching circuit was able to switch between each active sensor efficiently to capture all the sensors' data.
Imran M. Saied, Tughrul Arslan
ISCAS2
2023 LAMANet: A Real-Time, Machine Learning-Enhanced Approximate Message Passing Detector for Massive MIMO
abstract
Model-driven machine learning for signal detection in the physical layer of mobile communication systems combines well-known detector structures with learned parameters. Recent work has shown high detection performance in massive multiple-input–multiple-output (MIMO) detection; however, thorough complexity analysis and real-time processing hardware are lacking. This work proposes a novel machine learning enhanced approximate message passing (AMP) algorithm named LAMANet and its hardware implementation. The algorithm solves some major challenges of previous proposals such as the complete loss of performance in untrained detectors and the still high computational complexity compared to traditional massive MIMO detection methods. We provide a comprehensive complexity comparison, simulations of the symbol error rate (SER) performance over realistic channel models, and a field-programmable gate array (FPGA) implementation capable of processing LAMANet in real time. The results show a similar detection performance of LAMANet to previous machine learning-enhanced algorithms, while the computational effort is reduced to a level where real-time computation in hardware becomes comparable to traditional detection methods.
Stefan Brennsteiner, Tughrul Arslan, John S. Thompson, Andrew C. McCormick
IEEE Trans. Very Large Scale Integr. Syst.2
2022 An Encoded LSTM Network Model for WiFi-based Indoor Positioning
abstract
WiFi received signal strength (RSS)-based finger-printing has been widely adopted in many indoor positioning systems due to its implementation simplicity and low computational complexity. Recently, some studies have explored the potential temporal features among WiFi data to provide better positioning accuracy using Long Short-Term Memory (LSTM). However, the large volume of invalid RSS signals/values in the radio map degrades the training performance and the positioning accuracy. Some recent research has shown effort to reduce the impact of the problem mentioned above using deep learning. However, these either lack efficiency (repeated training needed or training resource wasted) or cannot provide precision position estimations. This paper proposes a novel Encoded LSTM network model to efficiently extract main features from the WiFi RSS data and provide high positioning accuracy. Experimental results based on an open-source crowdsourced dataset show that the encoded LSTM can reduce the dimension of the WiFi fingerprints from 992 to 64 and achieve a mean positioning error of 7.37m. Compared to the benchmark results based on the same dataset, the encoded LSTM shows the lowest mean error, which outperforms 13 popular positioning algorithms. The proposed encoded LSTM can also provide a 10% improvement in positioning accuracy in comparison to other conventional LSTM models.
Yinhuan Dong, Tughrul Arslan, Yunjie Yang 0001
IPIN2
2022 A WiFi Fingerprint Augmentation Method for 3-D Crowdsourced Indoor Positioning Systems
abstract
WiFi received signal strength (RSS)-based finger-printing has attracted much attention in indoor positioning in the past decade. One WiFi fingerprint comprises multiple RSS values annotated with the location (reference point) where the signals are obtained. The positioning accuracy of WiFi RSS fingerprinting-based indoor positioning systems is highly reliant on the data volume of the observed signals. However, taking such data in a large complex indoor area is usually time-consuming and labor-intensive. In recent years, crowdsourcing approaches have been proposed to collect WiFi data and record the location by utilizing the trajectories of common users to reduce the burden of constructing the database. Nevertheless, crowdsourced data is usually sensitive to crowd density. The data coverage is usually not enough to cover the entire targeted environment to provide good positioning accuracy, particularly at the beginning stage of constructing a database. Besides, it is also expected that some regions do not have enough fingerprints to provide good positioning performance since they do not have as many visitors as others. Therefore, this paper proposes a WiFi fingerprint augmentation method to generate more fingerprints by predicting RSS values on unsurveyed locations through a multivariate Gaussian process regression (MGPR) model. Evaluations are conducted on an open-source crowdsourced WiFi fingerprint dataset collected in an actual multi-floor university building. The experiment results show that the proposed WiFi fingerprint augmentation method can enhance the global data coverage (considering the entire building) to reduce the positioning error by 5% to 20%. Also, the proposed method can sharply reduce the positioning error in some indoor regions by improving local data density (considering a 2D region on a certain floor).
Yinhuan Dong, Tughrul Arslan, Yunjie Yang 0001, Yingda Ma
IPIN2
2022 Deep Unfolding-based Detection for Quantized Massive MU-MIMO-OFDM Systems
abstract
This paper explores the difficulties of massive multi-user (MU) multiple-input multiple-output (MIMO) orthogonal frequency-division multiplexing (OFDM) detection with low-precision quantization. To solve these problems, we propose QMMO-Net, a novel deep unfolding (DU)-based detection scheme that fuses the architecture specialized for quantized MIMO-OFDM detection with data-driven techniques. To handle the severe distortions from coarse quantization, we add multiple trainable parameters to increase the model flexibility. With the help of the proposed differentiable proximal operator and DU tools, these parameters including a vector can be jointly optimized. Simulation results demonstrate that QMMO-Net outperforms traditional and DU-based detection algorithms in coarsely quantized MU-MIMO-OFDM systems. By combining the power of domain knowledge with data, our QMMO-Net has strong robustness to the non-linear effects of coarse quantization and the co-channel interference caused in high user load scenarios.
John S. Thompson, Tughrul Arslan
VTC Spring3
2022 A Deep Unfolding Network for Massive Multi-user MIMO-OFDM Detection
abstract
This paper proposes a novel deep unfolding (DU)- based iterative detection algorithm for massive multi-user (MU) multiple-input multiple-output (MIMO) orthogonal frequency-division multiplexing (OFDM) systems. Our model-driven algorithm fuses over-relaxed alternating direction method of multipliers (OR-ADMM) with DU tools, which combines the power of domain knowledge and data. By unfolding the iterations into neural network layers and designing a differentiable multi-level projection function, learnable parameters in the algorithm can be optimized via deep-learning techniques. For i.i.d. Gaussian channels, only one set of parameters is sufficient for all channel realizations. Different sets of parameters are learned to handle severe frequency selectivity and channel correlation under realistic 3GPP-3D channels. Our simulations demonstrate that this scheme outperforms traditional and DU-based detection algorithms in MU-MIMO-OFDM systems with similar or lower complexity. According to comparison results, our network achieves faster convergence and performs robustly on different channels, especially when the number of Tx and Rx antennas are equal.
John S. Thompson, Tughrul Arslan
WCNC3
2022 A Real-Time Deep Learning OFDM Receiver
abstract
Machine learning in the physical layer of communication systems holds the potential to improve performance and simplify design methodology. Many algorithms have been proposed; however, the model complexity is often unfeasible for real-time deployment. The real-time processing capability of these systems has not been proven yet. In this work, we propose a novel, less complex, fully connected neural network to perform channel estimation and signal detection in an orthogonal frequency division multiplexing system. The memory requirement, which is often the bottleneck for fully connected neural networks, is reduced by ≈ 27 times by applying known compression techniques in a three-step training process. Extensive experiments were performed for pruning and quantizing the weights of the neural network detector. Additionally, Huffman encoding was used on the weights to further reduce memory requirements. Based on this approach, we propose the first field-programmable gate array based, real-time capable neural network accelerator, specifically designed to accelerate the orthogonal frequency division multiplexing detector workload. The accelerator is synthesized for a Xilinx RFSoC field-programmable gate array, uses small-batch processing to increase throughput, efficiently supports branching neural networks, and implements superscalar Huffman decoders.
Stefan Brennsteiner, Tughrul Arslan, John S. Thompson, Andrew C. McCormick
ACM Trans. Reconfigurable Technol. Syst.2
2021 EnSuRe: Energy & Accuracy Aware Fault-tolerant Scheduling on Real-time Heterogeneous Systems
abstract
This paper proposes an energy efficient real-time scheduling strategy called EnSuRe, which (i) executes real-time tasks on low power consuming primary processors to enhance the system accuracy by maintaining the deadline and (ii) provides reliability against a fixed number of transient faults by selectively executing backup tasks on high power consuming backup processor. Simulation results reveal that EnSuRe consumes nearly 25% less energy, compared to existing techniques, while satisfying the fault tolerance requirements. EnSuRe is also able to achieve 75% system accuracy with 50% system utilisation. Further, the obtained simulation outcomes are validated on benchmark tasks via a fault injection framework on Xilinx ZYNQ APSoC heterogeneous dual core platform.
Sangeet Saha, Adewale Adetomi, Xiaojun Zhai, Server Kasap, Shoaib Ehsan, Tughrul Arslan, Klaus D. McDonald-Maier
IOLTS6
2021 Hardware Accelerator for Wearable and Portable Radar-Based Microwave Breast Imaging Systems
abstract
This paper proposes a novel hardware accelerator for wearable and portable radar-based microwave breast imaging applications. The hardware accelerator is implemented on a low-cost and compact FPGA that is much cheaper than other accelerators developed in previous studies. The proposed hardware accelerator is demonstrated by using experimental tests with different configurations. Data was obtained from a wearable antenna array system and processed through microwave imaging in space-time (MIST) algorithm. The results from the proposed hardware accelerator show an improvement in experimental execution time that is 4 times faster than those obtained from computer-based methods. The fast execution times make the use of the proposed hardware accelerator suitable for real-time detection. In addition, the proposed hardware accelerator had the lowest power consumption among other similar accelerators developed in previous work, thus making it a safe and efficient system to be implemented in wearable microwave imaging systems. The use of a low-cost FPGA provides a promising method for hardware acceleration with existing microwave imaging systems. This research provides a foundation for future work to be developed in building a complete system that utilises low-cost and low requirement-based FPGAs that can be customized for portable and wearable breast cancer imaging applications.
Imran M. Saied, Tughrul Arslan, Rahmat Ullah, Fengzhou Wang
ISCAS2
2020 Proxy Circuits for Fault-Tolerant Primitive Interfacing in Reconfigurable Devices Targeting Extreme Environments
abstract
Continuous interface access to device-level primitives in reconfigurable devices in extreme environments is key to reliable operation. However, it is possible for a primitive's interface controller, which is static to be rendered non-operational by a permanent damage in the controller's circuitry. In order to mitigate this, this paper proposes the use of relocatable proxy circuits to provide remote interfacing capability to primitives from anywhere on a reconfigurable device. A demonstration with device register read controller shows that an improvement in fault-tolerance can be achieved.
Adewale Adetomi, Sangeet Saha, Klaus D. McDonald-Maier, Tughrul Arslan
ISCAS4
2020 RecNet: Deep Learning-Based OFDM Receiver with Semi-Blind Channel Estimation
abstract
This paper proposes a novel deep learning-based system for channel estimation and signal detection in an orthogonal frequency-division multiplexing (OFDM) system. Different from viewing the whole OFDM system as a black box in the literature, the proposed receiving system is divided into two low-complexity neural networks (NNs). The first NN is designed for semi-blind channel estimation and the second NN recovers the original signals based on the channel state information (CSI) obtained from the first network. Simulation results show that this system offers a competitive accuracy for channel estimation. Specifically, our simulations show better robustness when a very small number of pilots are used compared with traditional channel estimation methods. With the help of the estimated CSI, the NN for signal detection converges within much fewer epochs than data-driven solutions. Due to fast convergence and small size of the NNs, this system can be trained within a short time to adapt to different channel models. Compared with the traditional schemes, our whole system has a better performance under low signal-to-noise ratio (SNR).
Tughrul Arslan
ISCAS2
2020 Non-Invasive RF Technique for Detecting Different Stages of Alzheimer's Disease and Imaging Beta-Amyloid Plaques and Tau Tangles in the Brain
abstract
This paper describes a novel approach of detecting different stages of Alzheimer's disease (AD) and imaging beta-amyloid plaques and tau tangles in the brain using RF sensors. Dielectric measurements were obtained from grey matter and white matter regions of brain tissues with severe AD pathology at a frequency range of 200 MHz to 3 GHz using a vector network analyzer and dielectric probe. Computational models were created on CST Microwave Suite using a realistic head model and the measured dielectric properties to represent affected brain regions at different stages of AD. Simulations were carried out to test the performance of the RF sensors. Experiments were performed using textile-based RF sensors on fabricated phantoms, representing a human brain with different volumes of AD-affected brain tissues. Experimental data was collected from the sensors and processed in an imaging algorithm to reconstruct images of the affected areas in the brain. Measured dielectric properties in brain tissues with AD pathology were found to be different from healthy human brain tissues. Simulation and experimental results indicated a correlated shift in the captured reflection coefficient data from RF sensors as the amount of affected brain regions increased. Finally, images reconstructed from the imaging algorithm successfully highlighted areas of the brain affected by plaques and tangles as a result of AD. The results from this study show that RF sensing can be used to identify areas of the brain affected by AD pathology. This provides a promising new non-invasive technique for monitoring the progression of AD.
Imran M. Saied, Tughrul Arslan, Siddharthan Chandran, Tara Spires-Jones, Suvankar Pal
IEEE Trans. Medical Imaging2
2020 Enabling Dynamic Communication for Runtime Circuit Relocation
abstract
Runtime circuit relocation has been proposed for mitigating the effect of permanent damages in reconfigurable hardware, such as field-programmable gate arrays (FPGAs), with potentials to improve reliability and reduce or eliminate system downtime. However, a major obstacle to the adoption of circuit relocation is the presence of static communication links between the circuits. Existing solutions to this are either computationally expensive or counterintuitive to system reliability. This article proposes a dynamic communication mechanism that is able to circumvent the static links. The clock buffers in a typical FPGA use independent wires and, thus, do not constitute static routing. These are repurposed as network links to provide dynamic communication for relocatable circuits, with a demonstrator based on a four-node star network showing a bandwidth of 428.58 Mb/s for a 32-bit payload at an overhead of only 144 slices on a seven-series FPGA.
Adewale Adetomi, Godwin Enemali, Tughrul Arslan
IEEE Trans. Very Large Scale Integr. Syst.3
2019 A Dynamic Feature Fusion Strategy for Magnetic Field and Wi-Fi Based Indoor Positioning
abstract
The paper presents a dynamic feature fusion strategy to provide more accurate and reliable location estimates in areas of weak discernibility for indoor magnetic signals. Experimental evaluations were carried out in two different scenarios to confirm improved performance of the proposed strategy. The results achieved is more than a 25% and 45% improvement in the average error distance compared to conventional magnetic-field/Wi-Fi localisation system and single magnetic localisation system respectively.
Yichen Du, Tughrul Arslan, Qianqian Shen
IPIN2
2019 Editorial TVLSI Positioning - Continuing and Accelerating an Upward Trajectory
abstract
I. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5].
Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber
IEEE Trans. Very Large Scale Integr. Syst.3
2018 Characterization of Clock Buffers for On-Chip Inter-Circuit Communication in Xilinx FPGAs
abstract
Resource underutilization can occur in FPGAs if there is not enough routing resource to connect circuit elements in a region of the chip area. To alleviate this, we have proposed the use of clock buffers for on-chip data routing. Moreover, dynamism in communication for reliability is facilitated by using clock buffers for communication. This is because the clock routing network is independent of the general interconnect. In this paper, we present different configurations of the clock buffers in a Xilinx 7 series FPGA and characterize them based on the achievable speed of communication.
Adewale Adetomi, Godwin Enemali, Tughrul Arslan
ISCAS3
2018 Wideband Textile Antenna for Monitoring Neurodegenerative Diseases
abstract
Wearable microwave diagnostic systems could pave the way to providing an effective and continuous patient monitoring system in the future. To realise wearable systems for head diagnostics, a novel compact flexible antenna made from flexible textile materials is presented in this paper. The sensing antenna is fed with a microstrip transmission line incorporating a stepped monopole structure. Three patches of different sizes are excited to obtain a wideband performance. The conductive part and the substrate of the antenna is made from a flexible conductive textile and a 6 mm thick felt material respectively. The antenna operates from 1 GHz to 4.2 GHz when in direct contact with an artificial head phantom. Moreover, the antenna also exhibits directional radiation patterns with a simulated 6 dB front-to-back ratio at the center frequency of 2.5 GHz. Simulation results show that the antenna is suitable for wearable microwave head diagnostic systems. In addition, experimental results were performed using artificial head phantoms that mimicked the pathological effects of a human brain affected with neurodegenerative diseases, specifically, brain atrophy and ventricle enlargement. The results show that the textile-based antenna is capable of monitoring the pathological effects of the disease successfully.
Imran M. Saied, Tughrul Arslan
PIMRC2
2017 Relocating Encrypted Partial Bitstreams by Advance Task Address Loading
abstract
The ability to relocate hardware tasks in FPGAs is an attractive task management technique, especially in reconfigurable operating systems. A method of relocation involves the modification of the location address of the task while it is being configured. However, the use of encryption to protect bitstreams requires that decryption is done on-chip before relocation. This usually results in a significant resource overhead, arising from the introduced decryption circuit. This paper presents Advance Task Address Loading (ATAL), a unique solution that involves loading the unencrypted task address ahead of the encrypted task's configuration frame data. We have developed a software named Splixbit, which processes the bitstream offline, and a corresponding hardware configuration controller that configures the bitstream on the FPGA. Our results confirmed the possibility of avoiding on-chip dedicated decryption circuit in relocating encrypted partial bitstreams.
Adewale Adetomi, Godwin Enemali, Tughrul Arslan
FCCM3
2017 Relocation-aware communication network for circuits on Xilinx FPGAs
abstract
The parallelism of hardware and the dynamic reconfigurability of FPGAs enable multiple hardware tasks to run concurrently, and also time-share resources by being swapped in and out of the device during runtime. More than ever before, these capabilities are being employed in systems with high-reliability requirements. To improve reliability, a method often used is circuit relocation. However, the static nature of conventional FPGA communication interconnects is a bane to flexible runtime relocation. This paper employs a novel network architecture to enable dynamic communication and thus improve the flexibility of circuit relocation. By using the clock infrastructure of the FPGA as the physical network links for tasks in a 4-node star network, we have shown that dynamic communication between relocatable circuits can be achieved without incurring any overheads of time and resources, save for only 32 slices used for the Network Interface.
Adewale Adetomi, Godwin Enemali, Tughrul Arslan
FPL3
2017 Magnetic field indoor positioning system based on automatic spatial-segmentation strategy
abstract
We present an automatic spatial-segmentation method for a novel magnetic-field hybrid indoor positioning system. Unlike conventional approaches that implement magnetic field alone for the whole experimental space, our approach employs an efficient automatic segmentation method to make up the constraints of magnetic-field indoor positioning after fully evaluating local magnetic-field characteristics. By generating new partitioned databases, this method takes full advantage of characteristic changes in the geomagnetic field according to the distribution of disturbance objects inside buildings. Moreover, it reduces the processing time, as new databases contribute only to the region partition, and no further location calculation is involved. Experimental evaluations carried out with two different sets of hybrid techniques confirm satisfactory performance of the indoor localisation based on automatic spatial segmentation achieving more than a 1 meters' improvement in the average error distance compared to conventional magnetic-field localisation systems.
Yichen Du, Tughrul Arslan
IPIN2
2017 Isolated beacon identification using statistical approach
abstract
Location-based services (LBS) have an increasing demand for high accuracy but accurate positioning is available only for outdoor areas. Multiple indoor positioning algorithms have been researched and improved to provide better accuracy. For the purpose of the internet of things Bluetooth 4 (BLE) was introduced. However, a Bluetooth 4 signal is highly affected by obstruction attenuation and multipath. Deploying the BLE beacon could impose a significant cost if a large number of beacons need to be employed to avoid the obstacle attenuation and multipath effects. This paper proposes an algorithm to identify a beacon placed in a different room (isolated beacon). The proposed algorithm has a 68% probability of estimating the correct condition with an accuracy of up to 78% for a high attenuation obstruction. By implementing the algorithm, the results show an average accuracy improvement of 48% with RMSE error decreasing from 7.8 metres to 4.64 metres.
Arief Affendi Juri, Tughrul Arslan, Yichen Du
IPIN2
2017 A placement management circuit for efficient realtime hardware reuse on FPGAs targeting reliable autonomous systems
abstract
Reconfigurable hardware such as FPGAs offer promising platform for the development of embedded autonomous systems. This is due to their unique combination of high performance and flexibility. However, state-of-the-art FPGAs have large reconfiguration time, which often leads to missed deadlines in real-time systems. They also suffer from considerable fragmentation during runtime placement, leading to poor chip area utilization. In this paper, we present a novel hardware-placement management circuit to address these limitations by offering circuit reuse and a low-cost defragmentation. Its implementation occupies only 1852 FPGA slices. Our results showed that over 70% of configuration time was circumvented compared to state-of-the-art techniques. In addition, up to 56% improvement in reuse efficiency was observed.
Godwin Enemali, Adewale Adetomi, Tughrul Arslan
ISCAS3
2016 Camera-aided region-based magnetic field indoor positioning
abstract
It has been shown that local magnetic field (MF) anomalies can be used in accurate global self-localisation with fingerprinting. However, MF anomalies can only affect limited areas, and the low discernibility of received local MF signals may result in many positions having the same MF-Location information in areas far away from disturbances. This is mainly due to the sensitivity limitations of sensors embedded in mobile phones. This makes it challenging to distinguish between different positions with the same MF value. To address this problem, this paper proposes a new camera-assisted region-based MF fingerprinting technique, which takes maximum advantage of MF-based indoor positioning. Unlike using MF positioning alone for the whole training space, the basic idea of this multipronged system is to use camera-based positioning in areas with fewer disturbances to assist MF positioning targeting attaining a more location estimates. We have calculated the accuracy rate of the proposed system and compared it with MF-only and camera-only localisation systems. The results suggest that the proposed system performs significantly better than these two systems.
Yichen Du, Tughrul Arslan, Arief Affendi Juri
IPIN2
2016 Dual scaling and sub-model based PnP algorithm for indoor positioning based on optical sensing using smartphones
abstract
Traditionally, optical positioning has been an area that attracted significant interest for specialised applications such as robotic navigation. With the recent development and integration of cameras in mobile devices, optical positioning is gaining further interest in indoor positioning applications targeting human navigation. Furthermore, the release of the Google Tango platform has attracted the interest of several researchers aiming to improve optical positioning using 3D mapping. However, estimating a 3D pose based on PnP problem has been a challenge which affects positioning accuracy. This paper proposes two novel strategies to improve PnP algorithms, and hence accuracy. The first “scaling” strategy, is based on minimising model size, with the second, “sub-model” strategy, involving the selection of only the related area of the model to be used. The proposed strategies also have the advantage of limiting the error to the set scale size. The scale and sub-model strategies show an average improvement of 1.61 and 3.33 m, respectively.
Arief Affendi Juri, Tughrul Arslan, Yichen Du
IPIN2
2015 Microkernel Architecture and Hardware Abstraction Layer of a Reliable Reconfigurable Real-Time Operating System (R3TOS)
abstract
This article presents a new solution for easing the development of reconfigurable applications using Field-Programable Gate Arrays (FPGAs). Namely, our Reliable Reconfigurable Real-Time Operating System (R3TOS) provides OS-like support for partially reconfigurable FPGAs. Unlike related works, R3TOS is founded on the basis of resource reusability and computation ephemerality. It makes intensive use of reconfiguration at very fine FPGA granularity, keeping the logic resources used only while performing computation and releasing them as soon as it is completed. To achieve this goal, R3TOS goes beyond the traditional approach of using reconfigurable slots with fixed boundaries interconnected by means of a static communication infrastructure. Instead, R3TOS approaches a static route-free system where nearly everything is reconfigurable. The tasks are concatenated to form a computation chain through which partial results naturally flow, and data are exchanged among remotely located tasks using FPGA’s reconfiguration mechanism or by means of “removable” routing circuits. In this article, we describe the R3TOS microkernel architecture as well as its hardware abstraction services and programming interface. Notably, the article presents a set of novel circuits and mechanisms to overcome the limitations and exploit the opportunities of Xilinx reconfigurable technology in the scope of hardware multitasking and dependability.
Xabier Iturbe, Khaled Benkrid, Chuan Hong, Ali Ebrahim, Raul Torrego, Tughrul Arslan
ACM Trans. Reconfigurable Technol. Syst.6
2014 A fast and scalable FPGA damage diagnostic service for R3TOS using BIST cloning technique
abstract
This paper presents a new technique to be used in the context of reconfigurable computing to accelerate the online diagnosis of permanent damage on Xilinx FPGAs using Built-In Self Tests (BISTs). Detecting and locating permanently damaged resources with precision is central to keep the system implemented on the FPGA flawless at all times; i.e. upcoming hardware tasks are mapped to available functional resources, circumventing the use of the damaged ones. The proposed diagnostic technique exploits the Multiple Frame Write (MFW) feature available in Xilinx FPGAs to “clone” (i.e. replicate) a single basic BIST circuit along arbitrarily sized and shaped areas on the FPGA without incurring large time overheads. Hence, the proposed technique allows for creating at runtime on-demand tailored BIST circuits to satisfy any diagnosis requirements that may rise up. Moreover, the proposed solution allows for saving memory in the system as it only requires storing basic BIST circuits. Finally, the paper presents a diagnostic service for a Reliable Reconfigurable Real-Time Operating System (R3TOS) that is based on the BIST cloning technique and works in cooperation with the R3TOS fault-handling and recovery mechanisms.
Ali Ebrahim, Tughrul Arslan, Xabier Iturbe
FPL2
2013 Reconfigurable feeding network for GSM/GPS/3G/WiFi and global LTE applications
abstract
This paper presents a novel miniaturized reconfigurable and switchable feeding network to cover GSM, GPS, 3G, WiFi and global LET standards. The feeding network consists of four conventional Wilkinson power dividers which can be individually reconfigured in length using PIN diodes switches. By controlling the bias voltages of these PIN diodes, the operating frequency of the proposed design can be converted between four different bands: 600MHz-900MHz, 1.2GHz-1.6GHz, 1.8GHz-2.2GHz and 2.4GHz-2.6GHz. The first frequency band (600MHz-900MHz) is applied to satisfy the applications of LTE US (700MHz), LTE UK (800MHz) and GSM (850MHz, 900MHz). The second band (1.2GHz-1.6GHz) targets GPS L1 (1.575GHz) and GPS L2 (1.227GHz). Different GSM (1800MHz, 1900MHz) and 3G standards (UMTS, W-CDMA, TD-SCDMA and CDMA2000) are located in the third frequency band (1.8GHz-2.2GHz). The last band (2.4GHz-2.6GHz) is used to cover WiFi (2.45GHz) and LTE Europe (2.6GHz). The miniaturized and optimized feeding network exhibits good performance for S-Parameters in each band, which includes low return loss, equal power splitting and suitable insertion loss. Within the simulation environment, three types of PIN diode models were constructed and investigated in order to improve accuracy. The feeding network is implemented on an FR4 substrate. Fabrication and measurement results closely correlate with those obtained during design simulations. The reconfigurable feeding network can be particularly applied to commercial multiband communication systems.
Tughrul Arslan, Khaled Benkrid, Ahmed O. El-Rayis, Nakul Haridas
ISCAS2
2013 Faceted array antennas for adaptive beamforming applications
abstract
In this paper, the performance of adaptive array antennas conformed to different faceted structures, namely, 2-faceted, 3-faceted, 4-faceted and 8-faceted structures, are studied and compared with each other in the context of adaptive beamforming. Faceted array antennas are chosen because they can achieve a wider scanning range compared to planar arrays. Each of the faceted arrays consists of eight circularly polarised antennas operating at 2.4 GHz. Beamforming is achieved according to two different optimisation criteria, which are the minimised mean square error (MMSE) and the maximised signal-to-interference-plus-noise ratio (MSINR). The adaptive algorithm, Least Mean Square (LMS), and the bio-inspired algorithm, Particle Swarm Optimisation (PSO), are used to calculate the complex excitations of the array elements in order to satisfy the optimisation criteria. The directivity of the 3-faceted array is high, 12.59 dBi, even when the desired signal is away from boresight. The results obtained suggest that the 3-faceted array is the most suitable for adaptive array applications compared to the other faceted structures.
Nurul Hazlina Noordin, Tughrul Arslan, Brian W. Flynn, Ahmet T. Erdogan
PIMRC2
2013 A reconfigurable feed network for a dual circularly polarised antenna array
abstract
A novel miniaturised reconfigurable, switchable feed network for a four elements dual circularly polarised antenna array is proposed. The four feeds are in phase quadrature to give a phase shift of 90obetween each neighbouring antenna element, producing a circularly polarised pattern. The feed network consists of three conventional Wilkinson power dividers which can be individually reconfigured in length using PIN diodes switches. By controlling the bias voltages of these PIN diodes, the 90ophase increment can be allocated in either clockwise or counter-clockwise directions. Two orthogonal patterns with left-hand circular polarisation (LHCP) and right-hand circular polarisation (RHCP) are generated. The operating frequency is 2.45GHz, targeting WiFi applications. After miniaturisation and optimisation, the feed network achieves a low return loss, equal power splitting, low insertion loss and accurate phase difference. Fabrication and measurement results closely correlate with those obtained from simulations. The reconfigurable feed network is particularly useful for diversity reception to combat channel fading in indoor wireless LAN applications.
Tughrul Arslan, Brian W. Flynn
PIMRC2
2013 Dynamic Fault-Tolerant three-dimensional cellular genetic algorithms
Asmaa Al-Naqi, Ahmet T. Erdogan, Tughrul Arslan
J. Parallel Distributed Comput.3
2013 Adaptive three-dimensional cellular genetic algorithm for balancing exploration and exploitation processes
Asmaa Al-Naqi, Ahmet T. Erdogan, Tughrul Arslan
Soft Comput.3
2013 R3TOS: A Novel Reliable Reconfigurable Real-Time Operating System for Highly Adaptive, Efficient, and Dependable Computing on FPGAs
abstract
Despite the clear potential of FPGAs to push the current power wall beyond what is possible with general-purpose processors, as well as to meet ever more exigent reliability requirements, the lack of standard tools and interfaces to develop reconfigurable applications limits FPGAs' user base and makes their programming not productive. R3TOS is our contribution to tackle this problem. It provides systematic OS support for FPGAs, allowing the exploitation of some of the most advanced capabilities of FPGA technology by inexperienced users. What makes R3TOS special is its nonconventional way of exploiting on-chip resources: These are used indistinguishably for carrying out either computation or communication tasks at different times. Indeed, R3TOS does not rely on any static infrastructure apart from its own core circuitry, which is constrained to a specific region within the FPGA where it is implemented. Thus, the rest of the device is kept free of obstacles, with the spare resources ready to be used as and whenever needed. At runtime, the hardware tasks are scheduled and allocated with the dual objective of improving computation density and circumventing damaged resources on the FPGA.
Xabier Iturbe, Khaled Benkrid, Chuan Hong, Ali Ebrahim, Raul Torrego, Imanol Martinez, Tughrul Arslan, Jon Pérez 0001
IEEE Trans. Computers7
2013 High-Efficiency Customized Coarse-Grained Dynamically Reconfigurable Architecture for JPEG2000
abstract
This brief presents an efficient implementation of JPEG2000 encoding algorithm based on an architecture consisting of a coarse-grained dynamically reconfigurable instruction cell array and an embedded advanced RISC machine core. In this implementation, different tasks within the JPEG2000 encoding algorithm are allocated with proper computational resources to achieve high throughput. The proposed architecture is dynamically reconfigured for different tasks during the encoding process. Simulation results demonstrate that the proposed architecture provides a throughput of up to 52.1 f/s (or 19.18 ms/frame) to encode a 256 × 256 standard Lena test image. Compared with various digital signal processor & very long instruction word-based JPEG2000 solutions, the proposed architecture provides significant advantages in terms of throughput and energy consumption.
Ahmet T. Erdogan, Tughrul Arslan
IEEE Trans. Very Large Scale Integr. Syst.3
2011 Methods and Mechanisms for Hardware Multitasking: Executing and Synchronizing Fully Relocatable Hardware Tasks in Xilinx FPGAs
abstract
This paper presents the details of a novel technique which allows for the implementation and execution of completely relocatable hardware tasks onto dynamically reconfigurable FPGAs. Our novel technique harnesses the internal configuration access port (ICAP) for inter-task communication and synchronization, leading to very little logic overheads. The advantages of this technique include fault-tolerance, as tasks could be relocated freely on the fabric to circumvent damaged resources, and high performance, due to better exploitation of the logic fabric. The work is part of a larger effort in our group which aims to build a fully operational dynamically reconfigurable computer which would satisfy the often conflicting requirements of high performance, fault-tolerance and high level programming.
Xabier Iturbe, Khaled Benkrid, Tughrul Arslan, Raul Torrego, Imanol Martinez
FPL3
2011 Indoor positioning with floor determination in multi story buildings
abstract
Floor determination is one of the challenges in indoor positioning research. To date indoor positioning solutions for floor determination have been mainly based on either Fingerprinting [2] or RF-ID [3]. While these solutions have been able to locate persons or equipments accurately even in multi story buildings, these can not be considered as time and cost efficient solutions. Therefore, in this research we present new Wi-Fi based indoor positioning algorithms in order to accommodate the need for limited resources solution in terms of deployment time and cost. While finger-printing and RF-ID based solutions have been targeting sub-meter position accuracy, this research focuses on floor determination only. In addition, in this research we have highlighted two possible approaches by using available Wi-Fi signals for floor determination.
Firas Alsehly, Tughrul Arslan, Zankar Sevak
IPIN2
2011 Multi-objective evolutionary optimizations of a space-based reconfigurable sensor network under hard constraints
Erfu Yang, Ahmet T. Erdogan, Tughrul Arslan, Nick Barton
Soft Comput.3
2010 Fault tolerance through automatic cell isolation using three-dimensional cellular genetic algorithms
abstract
In this paper we propose a new algorithmic approach to achieve fault tolerance based on three dimensional cellular genetic algorithms (3D-cGAs). Herein, 3D architecture is targeted due to its amenability to implementation with current advanced custom silicon chip technology. The proposed approach is designed to exploit the inherent features of a cGA in which the genetic diversity is used as the key factor in identifying and isolating faulty individuals. A new migration schema is proposed as a mitigation technique. Several configurations concerning migration and selection intensity are considered. The approach is tested using four benchmark test functions and two real world problems which present different levels of difficulty. The overall results show that the proposed approach is able to cope with up to 40% soft errors (SEUs).
Asmaa Al-Naqi, Ahmet T. Erdogan, Tughrul Arslan
IEEE Congress on Evolutionary Computation3
2010 Lattice reconfiguration vs. local selection criteria for diversity tuning in cellular GAs
abstract
This paper aims to compare the effect of dynamically controlling the exploration-exploitation trade-off in cellular Genetic Algorithms (cGAs) from two perspectives: first, through lattice reconfiguration while dynamically changing the grid-neighbourhood ratio and thus taking advantage of their inherent structural properties; second, through local selection using a recently developed method known as anisotropic selection which allows to modify the overall population selection pressure at a local level. For both perspectives, the dynamic control of selection pressure is implemented constantly (every n generations) or adaptively based on the loss of diversity at the phenotype or the genotype space. Benchmark problems ranging from academic to real and combinatorial problems have been tackled in order to fairly compare both approaches. Statistical significance tests have also been carried out to support the results herein presented.
Alicia Morales-Reyes, Ahmet T. Erdogan, Tughrul Arslan
IEEE Congress on Evolutionary Computation3
2010 Optimising Self-Timed FPGA Circuits
abstract
This paper introduces a novel synchronous to asynchronous logic conversion tool targeted specifically for a synchronous field programmable gate array (FPGA). This tool augments the synchronous FPGA design flow and removes the clock network to implement an asynchronous control network in its place. We evaluate the timing performance benefits of the methods used to implement the asynchronous control network on synchronous FPGA fabric. Industrial video processing circuits are used to demonstrate the iterative timing improvements the tool makes to asynchronous control networks in each circuit. The targeted design constraints used in the tool are intended to improve the robustness and predictability of the placed circuits. This allows the timing benefits of asynchronous bundled data circuits easier to achieve, making asynchronous circuits a viable design option on modern FPGAs.
Phillip David Ferguson, Aristides Efthymiou, Tughrul Arslan, Danny Hume
DSD3
2010 ATB: Area-Time response Balancing algorithm for scheduling real-time hardware tasks
abstract
This paper describes a novel scheduling algorithm for the execution of hardware tasks with real-time constraints onto partially and dynamically reconfigurable FPGAs. The Area-Time response Balancing scheduling algorithm (ATB) is inspired by the well-known Earliest Deadline First (EDF) algorithm, which is extended with a technique for reducing the fragmentation on FPGA's reconfigurable area. This technique promotes the reuse of the resources that are released when great area tasks finish their execution by smaller area tasks as long as the real-time constraints permit to do so. Providing an exclusively time-based algorithm, such as EDF, with support for dealing with area-related issues ensures the best results. Simulation results reported in this paper show that ATB misses 23% less deadlines than EDF. Moreover, since FPGA's damaged resources provoke unpredictable fragmentation on the device, ATB is currently the best scheduling option to be used in a Reliable Reconfigurable Real-Time Operating System (R3TOS).
Xabier Iturbe, Khaled Benkrid, Tughrul Arslan, Imanol Martinez, Mikel Azkarate-askatsua
FPT3
2010 A parallel hybrid merge-select sorting scheme for K-best LSD MIMO decoder on a dynamically reconfigurable processor
abstract
In this paper, we propose a parallel hybrid merge-select sorting approach for the implementation of K-best list sphere detection (LSD) multi-input multi-output (MIMO) decoder based on a recently developed novel Reconfigurable Instruction Cell Array (RICA). Several popular sorting algorithms adopted in MIMO decoding are analyzed and mapped onto our proposed platform. We discuss the targeted K-best LSD algorithm as well as the sorting scheme variations which have been tailored for our RICA architecture. Simulation results prove that our proposed hybrid sorting approach can significantly reduce the number of comparison and swap operations when selecting the K-best candidates. Our results show that a 50% speedup can be achieved compared to traditional single bubble sorting based K-best LSD MIMO decoder.
Zong Wang, Ahmet T. Erdogan, Tughrul Arslan
PIMRC3
2010 High level modeling and automated generation of heterogeneous SoC architectures with optimized custom reconfigurable cores and on-chip communication media
Balal Ahmad, Ali Ahmadinia, Tughrul Arslan
J. Syst. Archit.3
2010 Interframe Bus Encoding Technique and Architecture for MPEG-4 AVC/H.264 Video Compression
abstract
In this paper, we propose an implementation of a data encoder to reduce the switched capacitance on a system bus. Our technique focuses on transferring raw video data for multiple reference frames between off- and on-chip memories in an MPEG-4 AVC/H.264 encoder. This technique is based on entropy coding to minimize bus transition. Existing techniques exploit the correlation between neighboring pixels. In our proposed technique, we exploit pixel correlation between two consecutive frames. Our method achieves a 58% power saving compared to an unencoded bus when transferring pixels on a 32-b off-chip bus with a 15-pF capacitance per wire.
Asral Bahari, Tughrul Arslan, Ahmet T. Erdogan
IEEE Trans. Very Large Scale Integr. Syst.2
2009 A distributed cellular GA based architecture for real time GPS attitude determination
abstract
This paper investigates a distributed cellular Genetic Algorithm (dcGA) for the implementation of a GPS attitude determination system. Previously, a cellular GA architecture has been proposed considering several implementations; however, comparison among these reveals that accuracy is compromised when the population size is increased. In this paper, a distributed configuration approach is proposed and compared with previous implementations in the literature; a significant improvement in terms of accuracy is reported without increasing computational cost.
Alicia Morales-Reyes, Ahmet T. Erdogan, Tughrul Arslan
IEEE Congress on Evolutionary Computation3
2009 An ILP formulation for task mapping and scheduling on multi-core architectures
abstract
Multi-core architectures are increasingly being adopted in the design of emerging complex embedded systems. Key issues of designing such systems are on-chip interconnects, memory architecture, and task mapping and scheduling. This paper presents an integer linear programming formulation for the task mapping and scheduling problem. The technique incorporates profiling-driven loop level task partitioning, task transformations, functional pipelining, and memory architecture aware data mapping to reduce system execution time. Experiments are conducted to evaluate the technique by implementing a series of DSP applications on several multi-core architectures based on dynamically reconfigurable processor cores. The results demonstrate that the proposed technique is able to generate high-quality mappings of realistic applications on the target multi-core architecture, achieving up to 1.3times parallel efficiency by employing only two dynamically reconfigurable processor cores.
Wei Han 0001, Ahmet T. Erdogan, Tughrul Arslan
DATE5
2009 Carbon Nanotube Interconnects for Low-power High-speed Applications
abstract
This paper investigates the prospects of mixed bundle of carbon nanotubes (CNT) as low-power high-speed interconnects for future VLSI applications. The power dissipation and delay of CNT bundle interconnects are examined and compared with that of the Cu interconnects at the 32-nm technology node. We evaluated and compared various performance metrics of interconnects with both CMOS and carbon nanotube field effect transistor (CNFET) driver and FO4 load using transmission line model. The results show that CNT bundle consumes 1.5 to 4 folds smaller power than Cu for intermediate and global interconnects. The CNT bundle interconnects are also faster than Cu except for local interconnects. It is concluded that the mixed bundle of CNTs is a promising candidate for intermediate and global interconnects in future technologies.
Naushad Alam, Abdul Kadir Kureshi, Mohd. Hasan, Tughrul Arslan
ISCAS4
2009 Leakage Reduction in FPGA Routing Multiplexers
abstract
There is a pressing need to reduce the static power consumption in FPGAs implemented in deep submicron so that it can also be used in portable battery operated devices. The multiplexer based interconnect matrix of an FPGA consumes most of the static power consumption. This paper investigates reducing leakage power in unused FPGA routing multiplexers by controlling their inputs at the deep submicron 22 nm technology node. HSPICE simulation using Berkeley Predictive technology models (BPTM) on different sizes and topologies of routing multiplexers show that the minimum leakage vector at the 22 nm technology node significantly varies from that at 65 nm node. This is due to higher gate leakage and output stage loading effects. The application of this vector results in 20% more leakage power saving as compared to the existing approaches. This technique saves significant leakage power because most of the routing multiplexers are unused in an FPGA. Moreover, there is no area overhead contrary to the existing approaches.
Mohd. Hasan, Abdul Kadir Kureshi, Tughrul Arslan
ISCAS3
2009 Subthreshold Deep Submicron Performance Investigation of CMOS and DTCMOS Biasing Schemes for Reconfigurable Computing
abstract
This paper investigates subthreshold CMOS logic for ultra low power applications on next generation reconfigurable devices. The performance characteristics of key digital building blocks such as arithmetic units, multiplexers and look-up-tables have been analyzed in terms of speed, power dissipation and power delay product using Berkeley predictive technology models at 22 nm technology node for both the conventional subthreshold CMOS (CMOS) and the dynamic threshold subthreshold CMOS (DTCMOS). Simulation results show that DTCMOS has lower PDP and sensitivities to process variations compared to CMOS for the digital blocks. Moreover, the PDP of blocks can be further improved by using longer channel lengths.
Abdul Kadir Kureshi, Naushad Alam, Mohd. Hasan, Tughrul Arslan
ISCAS4
2009 A Low Power Reconfigurable Heterogeneous Architecture for a Mobile SDR System
abstract
The main challenge in designing a mobile wireless software defined radio (SDR) system is to provide a solution that has high flexibility, hardware-like throughput, low power consumption, in addition to ease of programmability. In this paper, the authors propose a new architecture for SDR that is based on a reconfigurable instruction cell array (RICA). The architecture targets the IEEE 802.11g standard that includes Viterbi decoding, which is a key performance bottleneck. One of the salient novel features in this architecture, compared to existing solutions, is adopting a multi processor frame segmentation scheme when implement the 802.11 physical layer of the above standard. The paper describes the architecture, the associated software design flow, and performance efficiency. We demonstrate that the architecture can achieve a raw data throughput of 30.6 Mbps for an 802.11g receiver at a core power of 8.1 mW.
Zong Wang, Tughrul Arslan
ISCAS2
2009 Optimization of Reconfigurable Multi-core SOCs for Multi-standard Applications
Ali Ahmadinia, Tughrul Arslan, Hernando Fernandez-Canque
KES (2)2
2009 Multicore Architectures With Dynamically Reconfigurable Array Processors for Wireless Broadband Technologies
abstract
Wireless Internet-access technologies have significant market potential, particularly the Worldwide Interoperability for Microwave Access (WiMAX) protocol which can offer data rates of tens of megabits per second. A significant demand for embedded high-performance WiMAX solutions is forcing designers to seek single-chip multicore systems that offer competitive advantages in terms of all performance metrics, such as speed, power, and area. Through the provision of a degree of flexibility similar to that of a DSP and performance and power consumption advantages approaching that of an application-specific integrated circuit, emerging dynamically reconfigurable (DR) processors are proving to be strong candidates for processing cores in future high-performance multicore-processor systems. This paper presents several new single-chip multicore architectures for the WiMAX application based on recently emerging coarse-grained DR processor cores. A simulation platform is proposed in order to explore and implement various multicore solutions combining different memory architectures and task-partitioning schemes. This paper describes the different architectures, the simulation environment, and several task-partitioning methods and demonstrates that up to 7.3 and 12 times speedup can be achieved by employing eight and ten DR processor cores for both the WiMAX transmitter and receiver sections, respectively. A comparison with other WiMAX multicore solutions is given in order to demonstrate that our best solution delivers a high throughput at relatively low area cost.
Wei Han 0001, Mark Muir, Ioannis Nousias, Tughrul Arslan, Ahmet T. Erdogan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2009 Low-Power H.264 Video Compression Architectures for Mobile Communication
abstract
This paper presents a method to reduce the computation and memory access for variable block size motion estimation (ME) using pixel truncation. Previous work has focused on implementing pixel truncation using a fixed-block-size (16 times 16 pixels) ME. However, pixel truncation fails to give satisfactory results for smaller block partitions. In this paper, we analyze the effect of truncating pixels for smaller block partitions and propose a method to improve the frame prediction. Our method is able to reduce the total computation and memory access compared to conventional full-search method without significantly degrading picture quality. With unique data arrangement, the proposed architectures are able to save up to 53% energy compared to the conventional full-search architecture. This makes such architectures attractive for H.264 application in future mobile devices.
Asral Bahari, Tughrul Arslan, Ahmet T. Erdogan
IEEE Trans. Circuits Syst. Video Technol.2
2008 Fault tolerant cellular Genetic Algorithm
abstract
This paper presents a cellular Genetic Algorithm (cGA) which aims at realizing a fault tolerant platform based on the inherent ability of cGAs to deal with Single Hard Errors (SHE) that could permanently affect the operation of a system. To attain this objective it is indispensable to control the parameters of the cGA which directly affect the efficiency and accuracy of its search process. Among the overall set of parameters, the migration rate and frequency, the grid size, and the shape and size of local neighbourhoods have a remarkable effect on the cGA performance. By appropriately controlling these parameters, the complex search space (presenting multi-peak fitness-function) associated with the practical case study of the investigation herein presented, is conveniently explored in terms of efficiency and efficacy. Initially, fitness score registers have been identified as critical for proper systempsilas operation. In case, SHEs occur at these registers, the algorithm will ignore possible good solutions and rapidly spread bad individuals. Experiments results show the faults effects regarding convergence time, search rate and results accuracy, as well as the cGA improvement on faulty scenarios when migration is applied following different selection and replacement criteria or increasing selection intensity through different local neighbourhoods configurations.
Alicia Morales-Reyes, Evangelos F. Stefatos, Ahmet T. Erdogan, Tughrul Arslan
IEEE Congress on Evolutionary Computation4
2008 Evolutionary techniques for precise and real-time implementation of low-power FIR filters
abstract
This paper presents an evolutionary based reconfigurable framework that aims at implementing and reconfiguring precise and low-power FIR filters within short amount of time. Five evolutionary techniques are evaluated for their efficiency to drive the evolution of FIR filters upon the same custom reconfigurable hardware substrate. From a hardware perspective, our architecture composes a novel topology that achieves hardware economy and does not introduce hardware dependencies between different coefficients within the targeted coefficient-set. Three novel evolutionary techniques are proposed that guarantee accurate, prompt and low-power implementation of FIR filters. Each evolutionary technique mainly emphasizes on one or two out of the three investigated parameters (accuracy, power-consumption and real-time adaptation) and hence the designer can select one of these techniques, based on the nature and the needs of the targeted application.
Evangelos F. Stefatos, Tughrul Arslan, Alister Hamilton
IEEE Congress on Evolutionary Computation2
2008 A novel shifting balance theory-based approach to optimization of an energy-constrained modulation scheme for wireless sensor networks
abstract
This paper presents a new approach to optimization of an energy-constrained modulation scheme for wireless sensor networks by taking advantage of a novel bio-inspired optimization algorithm. The algorithm is inspired by Wrightpsilas shifting balance theory (SBT) of evolution in population genetics. The total energy consumption of an energy-constrained modulation scheme is minimized by using the new SBT-based optimization algorithm. The results obtained by this new algorithm are compared with other popular optimization algorithms. Numerical experiments are performed to demonstrate that the SBT-based algorithm could be used as an efficient optimizer for solving the optimization problems arising from currently emerging energy-efficient wireless sensor networks.
Erfu Yang, Nick Barton, Tughrul Arslan, Ahmet T. Erdogan
IEEE Congress on Evolutionary Computation3
2008 Efficient Implementation of Wireless Applications on Multi-core Platforms Based on Dynamically Reconfigurable Processors
abstract
Wireless internet access technologies such as WiMAX have significant market potential. The high demand for embedded high performance WiMAX solutions is forcing designers to seek multi-core systems which offer competitive advantages in terms of all performance metrics. By providing the flexibility of a DSP with performance and power consumption approaching that of an ASIC, emerging dynamically reconfigurable processors are proving to be stronger candidates than conventional general-purpose processors or DSPs for future multi-core systems. This paper presents several multi-core solutions, based on newly emerging dynamically reconfigurable processor cores targeting WiMAX based applications. Coming with a SystemC trace-driven multi-core simulator, a simulation platform has been proposed in order to explore and implement various multi-core solutions combining different task partitioning strategies and inter-process communication methods.
Wei Han 0001, Mark Muir, Ioannis Nousias, Tughrul Arslan, Ahmet T. Erdogan
CISIS5
2008 Automated Dynamic Throughput-constrained Structural-level Pipelining in Streaming Applications
abstract
Stream processing applications such as image signal processing demand high throughput. However, customers increasingly demand runtime flexibility in their designs, which cannot be provided by custom ASIC solutions. Currently, reconfigurable processors tend to offer insufficient throughput for widespread use in streaming applications. This paper demonstrates how structural-level pipelining techniques can be applied to rapidly dynamically reconfigurable computing architectures, in order to increase throughput. This is done by automatically inserting registers into the datapath of performance critical code sections that have already been optimised into a single configuration context. A new algorithm is presented to choose the insertion point of pipeline stage registers in order to meet a specified throughput whilst minimising register resource usage. The paper then demonstrates a new approach where properties of dynamic reconfiguration can be utilised to perform the tasks of pipeline stage initialisation and flushing. The technique is demonstrated on a real-life application: the demosaic filter in a standard image signal processing pipe used in modern digital cameras, and can be seen to boost the throughput from 16MPixels/s to 51MPixels/s on an example reconfigurable processor.
Mark Muir, Tughrul Arslan, Iain Lindsay
DATE2
2008 Dynamically programmable Reed Solomon processor with embedded Galois Field multiplier
abstract
This work presents a novel reconfigurable Galois field multiplier embedded in a dynamically reconfigurable processor for real time programmable Reed Solomon (RS) encoder and decoder targeting various communication standards. The fundamental operation in Reed-Solomon encoding and decoding is the multiplication over Galois field (GF). The reconfigurable GF multiplier with single instruction multiple data (SIMD) support is presented here, as an instruction set extension to the processor. The processor supports the RS coding to be programmable for Galois Field (28) with its sixteen primitive polynomials and for all supported data block sizes. Various optimization techniques have been applied in order to enhance the processor throughput. The throughput achieved for RS (204,188) is up to 202 Mbps for the encoder demonstrating a future proof flexible design.
Ahmed O. El-Rayis, Tughrul Arslan, Ahmet T. Erdogan
FPT3
2008 A low power reconfigurable heterogeneous architecture for a mobile SDR system
abstract
The main challenge in designing a mobile wireless Software Defined Radio (SDR) system is to provide a solution that has high flexibility, hardware-like throughput, low power consumption, in addition to ease of programmability. In this paper, the authors propose a new architecture for SDR that is based on a reconfigurable instruction cell array (RICA). The architecture targets the IEEE 802.11g standard that includes Viterbi decoding, which is a key performance bottleneck. One of the salient novel features in this architecture, compared to existing solutions, is adopting a multi processor frame segmentation scheme when implement the 802.11 physical layer of the above standard. The paper describes the architecture, the associated software design flow, and performance efficiency. We demonstrate that the architecture can achieve a raw data throughput of 30.6Mbps for an 802.11g receiver at a core power of 8.1mW.
Zong Wang, Tughrul Arslan
FPT2
2008 Nyquist-rate analog-to-digital converter specification for Zero-IF UMTS receiver
abstract
This paper presents a mathematical analysis method to derive the required specification for an analog-to-digital converter (ADC) to be used in a zero-IF UMTS receiver architecture targeting mobile station application. Wireless standard specifications are used to derive the design plan for a Nyquist converter type. Our research demonstrates that the minimum resolution should be 10 bits for a practical realization that includes design imperfections with the minimum required sampling frequency being 10 MHz.
Zulhakimi Razak, Tughrul Arslan
ISCAS2
2008 Code Compression and Decompression for Coarse-Grain Reconfigurable Architectures
abstract
This paper presents a code compression and on-the-fly decompression scheme suitable for coarse-grain reconfigurable technologies. These systems pose further challenges by having an order of magnitude higher memory requirement due to much wider instruction words than typical VLIW/TTA architectures. Current compression schemes are evaluated. A highly efficient and novel dictionary-based lossless compression technique is implemented and compared against a previous implementation for a reconfigurable system. This paper looks at several conflicting design parameters, such as the compression ratio, silicon area, latency, and power consumption. Compression ratios in the range of 0.32 to 0.44 are recorded with the proposed scheme for a given set of test programs. With these test programs, a 60% overall silicon area saving is achieved, even after the decompressor hardware overhead is taken into account. The proposed technique may be applied to any architecture which exhibits common characteristics to the example reconfigurable architecture targeted in this paper.
Nazish Aslam, Mark Milward, Ahmet T. Erdogan, Tughrul Arslan
IEEE Trans. Very Large Scale Integr. Syst.4
2008 The Reconfigurable Instruction Cell Array
abstract
This paper presents a novel instruction cell-based reconfigurable computing architecture for low-power applications, thereafter referred to as the reconfigurable instruction cell array (RICA). For the development of the RICA, a top-down software driven approach was taken and revealed as one of the key design decisions for a flexible, easy to program, low-power architecture. These features make RICA an architecture that inherently solves the main design requirements of modern low-power devices. Results show that it delivers considerably less power consumption when compared to leading VLIW and low-power digital signal processors, but still maintaining their throughput performance.
Sami Khawam, Ioannis Nousias, Mark Milward, Mark Muir, Tughrul Arslan
IEEE Trans. Very Large Scale Integr. Syst.6
2007 VLSI Design of Multi Standard Turbo Decoder for 3G and Beyond
abstract
Turbo decoding architectures have greater error correcting capability than any other known code. Due to their excellent performance turbo codes have been employed in several transmission systems such as CDMA2000, WCDMA (UMTS), ADSL, IEEE 802.16 metropolitan networks etc. The computation kernel of the algorithm is very similar and we have exploited this commonality for a turbo decoder VLSI design suitable for deployment using platform based system on chip methodologies. Turbo and Viterbi components of the unified array are also individually reconfigurable for different standards. This supports the 4G concept that user can be simultaneously connected to several access technologies (for example Wi-Fi, 3G, GSM etc.) and can seamlessly move between them. A new normalization scheme for turbo decoding is presented to suit reconfigurable mappings. We have also shown dynamic reconfiguration methodology for a context switch between turbo and Viterbi decoders which does not waste any clock cycles. The reconfigurable turbo decoder fabric is implemented reusing components of Viterbi decoder on a 180 nm UMC process technology.
Imran Ahmed 0001, Tughrul Arslan
ASP-DAC2
2007 Implementation of a Real Time Programmable Encoder for Low Density Parity Check Code on a Reconfigurable Instruction Cell Architecture
abstract
This paper presents a real time programmable irregular low density parity check (LDPC) encoder as specified in the IEEE P802.16E/D7 standard. The encoder is programmable for frame sizes from 576 to 2304 and for five different code rates. H matrix is efficiently generated and stored for a particular frame size and code rate. The encoder is implemented on reconfigurable instruction cell architecture (RA) which has recently emerged as an ultra low power, high performance, ANSI-C programmable embedded core. Different general and architecture specific optimization techniques are applied to enhance the throughput. With RA, a throughput from 10 to 19 Mbps has been achieved.
Zahid Khan, Tughrul Arslan
ASP-DAC2
2007 A Novel Reconfigurable Low Power Distributed Arithmetic Architecture for Multimedia Applications
abstract
The use of reconfigurable cores in system on chip (SoC) designs is increasingly becoming a trend. Such cores are being used for their flexibility, powerful functionality and low power consumption. Distributed arithmetic (DA) is a powerful algorithm widely used in many fields of multimedia for its efficiency. This paper presents a novel reconfigurable adder-based architecture for DA to realize the inner product which is the key computation in many digital signal processing applications. 1D DCT is mapped onto the architecture. Compared with some existing ASIC designs, the new architecture achieves good performance in area, speed and power.
Tughrul Arslan, Ahmet T. Erdogan
ASP-DAC2
2007 A multi-objective algorithm for the design of high performance reconfigurable architectures with embedded decoding
abstract
The increasing demand for FPGAs and reconfigurable hardware targeting high performance low power applications has lead to an increasing requirement for new computer aided design methodology for rapid design of such high performance reconfigurable embedded FPGA cores. In this paper, a power aware genetic algorithm is presented for automatic generation of target specific embedded reconfigurable cores from a library of logic blocks. The paper compares four decoding techniques within the algorithm, and explores the impact they have on the optimality of the architecture.
Wing On Fung, Tughrul Arslan
IEEE Congress on Evolutionary Computation2
2007 Pipelined implementation of a real time programmable encoder for low density parity check code on a reconfigurable instruction cell architecture
abstract
This paper presents pipelined implementation of a real time programmable irregular low density parity check (LDPC) encoder as specified in the IEEE P802.16E/D7 standard. The encoder is programmable for frame sizes from 576 to 2304 and for five different code rates. H matrix is efficiently generated and stored for a particular frame size and code rate. The encoder is implemented on reconfigurable instruction cell architecture which has recently emerged as an ultra low power, high performance, ANSI-C programmable embedded core. Different general and architecture specific optimization techniques are applied to enhance the throughput. With the architecture, a throughput from 10 to 19 Mbps has been achieved. The maximum throughput achieved with pipelining/multi-core is 78 Mbps
Zahid Khan, Tughrul Arslan
DATE2
2007 A new pipelined implementation for minimum norm sorting used in square root algorithm for MIMO-VBLAST systems
Zahid Khan, Tughrul Arslan, John S. Thompson, Ahmet T. Erdogan
DATE2
2007 Code Compressor and Decompressor for Ultra Large Instruction Width Coarse-Grain Reconfigurable Systems
abstract
This paper presents a code compression and on-the-fly decompression scheme suitable for coarse-grain reconfigurable technologies. A novel unit-grouping dictionary based compression technique utilizing special control bits to increase the effective storage capacity of the dictionaries is implemented and compared against an existing suitable technique for an example reconfigurable system. Compressions ratios in the range of 40%-59% are recorded with new scheme.
Nazish Aslam, Mark Milward, Ioannis Nousias, Tughrul Arslan, Ahmet T. Erdogan
FCCM4
2007 Mapping Real Time Operating System on Reconfigurable Instruction Cell Based Architectures
abstract
This paper presents the porting of an RTOS Micro C/OS-II on a novel reconfigurable instruction cell based architecture which fills the gap between DSP, FPGA and ASIC with high performance, high flexibility and ANSI-C support. WiMAX physical layer program has been implemented on the target architecture with the RTOS support. A semaphore based synchronization scheme is used to improve the task independence. The research lays a foundation for further exploration of multithreading on multiple target architectures.
Mark Muir, Ioannis Nousias, Tughrul Arslan, Ahmet T. Erdogan
FCCM4
2007 System-level Modelling and Analysis of Embedded Reconfigurable Cores for Wireless Systems
abstract
A new system-level approach is needed to incorporate reconfigurability in IP-inegration design flow, in order to speed up the designer's productivity. To incorporate reconfiguration aspects of IPs, a multiple-context representation of the different functionalities is used that will be mapped on the re-configurable block during different run-time periods. Co-simulation scenario is proposed as a part of a system-on-chip (SoC) design and modelling. SystemC-HDL co-simulation scenario provides a way of checking interoperability of a single designed HW module with the SystemC model. As a case study, novel reconfigurable FFT and Viterbi architectures are modelled in SystemC, and co-simulated in a C-based WiMAX system. Area and power consumption of main blocks in WiMAX are analysed.
Ali Ahmadinia, Balal Ahmad, Ahmet T. Erdogan, Tughrul Arslan
FPL4
2007 The Design of Multitasking Based Applications on Reconfigurable Instruction Cell Bsed Architectures
abstract
This paper presents a new direct implementation of a popular RTOS with an associated application - the WiMAX physical layer - on reconfigurable computing architectures. A novel coarse-grained reconfigurable instruction cell based architecture is chosen as the target architecture. Firstly an RTOS - Micro C/OS-II - was ported to the target architecture, and then the WiMAX physical layer program was partitioned into multiple OS tasks which communicate with each other through the synchronization approaches provided by this RTOS. The WiMAX physical layer program has been also implemented on the ARM7TDMI processor. The results show that the performance of the target architecture is much better than the ARM7TDMI, and not limited by the bottleneck of memory latency.
Wei Han 0001, Ioannis Nousias, Mark Muir, Tughrul Arslan, Ahmet T. Erdogan
FPL4
2007 H.264/AVC In-Loop De-Blocking Filter Targeting a Dynamically Reconfigurable Instruction Cell Based Architecture
abstract
We present a new De-Blocking Filter module fully optimised for use on a recently introduced dynamically reconfigurable, instruction cell based architecture. The module consists of a novel combination of standard software transforms alongside architecture specific techniques and aims to reduce reconfiguration overheads and increase utilisation of resources. Our proposed filter outperforms the standard FFMpeg based filter code on the target architecture by 4.5 times.
Adam Major, Ioannis Nousias, Sami Khawam, Mark Milward, Mark Muir, Tughrul Arslan
FPL7
2007 A Multi Objective GA based Physical Placement Algorithm for Heterogeneous Dynamically Reconfigurable Arrays
abstract
This paper presents the preliminary results of a physical placement algorithm for heterogeneous Dynamically Reconfigurable Arrays (DRA), based on a multi-objective, multi-threaded GA. The algorithm deals with the spatial and temporal nature of the configurations used in DRAs, in an attempt to find a suitable layout for a wide range of applications, since general applicability is a key criteria for DRAs.
Ioannis Nousias, Sami Khawam, Mark Milward, Mark Muir, Tughrul Arslan
FPL5
2007 Algorithmic Level Design Space Exploration Tool for Creation of Highly Optimized Synthesizable Circuits
abstract
This paper presents some of the results obtained by using a prototype algorithmic level design space exploration tool currently under development. The tool is based upon multi-objective evolutionary algorithms. The paper highlights the tool's benefits and discusses its current abilities in terms of its experimental applications.
Nazish Aslam, Tughrul Arslan, Ahmet T. Erdogan
ICASSP (2)2
2007 Code Compression and Decompression for Instruction Cell Based Reconfigurable Systems
abstract
Code compression has been applied to embedded systems to minimize the silicon area utilized for program memories, and lower the power consumption. More recently, it has become a necessity for multiple-issue architectures, such as VLIW and TTA, to permit a viable realization of these designs. In this paper, a code compression and decompression scheme suitable for newly emerging reconfigurable technologies is presented, which pose further challenges by having an order of magnitude higher memory requirement due to much wider instruction words than typical VLIW/TTA architectures. Two dictionary-based lossless compression schemes are implemented and compared for an example reconfigurable system. This paper looks at several conflicting design parameters, such as the compression ratio, silicon area and speed. Test programs for a 2D DCT, minimum error, wimax and H.264 have been evaluated with compression ratios in the range of 41% to 62% recorded with the best scheme.
Nazish Aslam, Mark Milward, Ioannis Nousias, Tughrul Arslan, Ahmet T. Erdogan
IPDPS4
2007 Radiation Hardened Coarse-Grain Reconfigurable Architecture for Space Applications
abstract
Technology trends are such that single event effects (SEE) are likely to become even more of a concern for the future. Decreasing feature sizes, lower operating voltage, and higher speeds, all conspire to increase susceptibility to single event upsets (SEU). Upset in avionics is an established concern. Upset at the ground level is becoming a concern for manufacturers of microelectronics for terrestrial applications. The use of flip-chip packaging and multiple levels of metals further exacerbate the problem. Typical methods of mitigation that either increase the transistor count or reduce IC performance are not acceptable to commercial manufacturers. SOI technology may help in this regard, but is not a magic bullet to end all SEE concerns. We present unique schemes to model and rectify single event disruption in combinatorial and synchronous parts of a reconfigurable architecture. We compare our scheme with different schemes already introduced and results are reported to prove the efficacy of the proposed radiation hardened reconfigurable architecture.
Sajid Baloch, Tughrul Arslan, Adrian Stoica
IPDPS2
2007 Low power variable block size motion estimation using pixel truncation
abstract
This paper presents a method of low-power variable-block-size motion estimation using pixel truncation. Previous work focused on implementing pixel truncation using fixed-block-size motion estimation. However, pixel truncation fails to give satisfactory results for smaller block partitions. In this paper, we analyse the effect of truncating pixels for smaller block partitions and propose a method to improve the frame prediction. To further reduce power consumption, we adopt low-complexity matching criteria for the highly truncated bit. The low-complexity matching criteria can work together with pixel truncation to reduce computational complexity without significantly degrading picture quality.
Asral Bahari, Tughrul Arslan, Ahmet T. Erdogan
ISCAS2
2007 Integrated Heterogenous Modelling for Power Estimation of Single Processor based Reconfigurable SoC Platform
abstract
Various instruction and transaction based power estimation techniques for processor and on-chip buses have been proposed in the past. In this paper, the authors propose a heterogeneous power model to estimate the power utilized by complete processor based reconfigurable system-on-chip (SoC) platform. The proposed model estimates the power consumed by the SoC platform using instruction-based model as well as transaction-based model. In addition the authors estimate the power consumed by various bus arbitration policies used in the on-chip communication
Prakash Srinivasan, Ali Ahmadinia, Ahmet T. Erdogan, Tughrul Arslan
ISCAS4
2006 Image Registration of Printed Circuit Boards using Hybrid Genetic Algorithm
abstract
In this paper, hybridization of hill-climbing (HC) and elitism (E) with a specially tailored genetic algorithm (GA) for image registration of printed circuit boards (PCBs) placed arbitrarily on a conveyor belt during inspection is proposed to maximize the robustness of the existing framework. These hybrid methods are investigated individually and in combination for accuracy, reliability and performance. Experimental results highlight the potential of the hybrid GA (HGA) that consists of all methods in combination because of the most accurate findings and significantly more reliable than GA alone. However, there is a compensation on performance, though it converges efficiently in terms of number of generations.
Syamsiah Mashohor, Jonathan R. Evans, Tughrul Arslan
IEEE Congress on Evolutionary Computation3
2006 An Incremental Evolutionary Strategy for the Design of FIR Filters Targeting Real-Time Applications
abstract
This paper introduces a new methodology for the design of finite impulse response filters within an evolvable hardware platform targeting real-time adaptation. The methodology incorporates a technique, which directs the search by focusing it into regions responsible for the evolution of individual coefficients. The methodology has been evaluated with three different paradigms that include two low-pass filters (29-order and 34-order, respectively) and a 41-order pass-band finite impulse response filter. The proposed method is compared with a conventional evolutionary strategy that has been commonly used for the implementation of digital filters resulting in significant improvements in terms of quality, convergence speed and computational efficiency.
Evangelos F. Stefatos, Tughrul Arslan
IEEE Congress on Evolutionary Computation2
2006 Non-Uniform search domain based Genetic algorithm for the optimization of real time FFT Processor architectures
abstract
This paper presents a GA for optimization of word length coefficients in a pipelined FFT processor. The algorithm optimizes memory and buses both at the I/O interfaces within the processor datapath. This provides a complex search space in which the algorithm needs balance optimization parameters against error. A special feature of the GA is the use of nonuniform operators which allow tuning the search to provide an optimal optimization with minimum number of generations. The paper describes the algorithm, the concept of non uniform operators through the mutation operation. The results show the effect of both uniform and non uniform sampling on the quality of the optimization, turbulence towards convergence, and the speed of convergence.
Nasri Sulaiman, Tughrul Arslan
IEEE Congress on Evolutionary Computation2
2006 System-level scheduling on instruction cell based reconfigurable systems
abstract
This paper presents a new operation chaining reconfigurable scheduling algorithm (CRS) based on list scheduling that maximizes instruction level parallelism available in distributed high performance instruction cell based reconfigurable systems. Unlike other typical scheduling methods, it considers the placement and routing effect, register assignment and advanced operation chaining compilation technique to generate higher performance scheduled code. The effectiveness of this approach is demonstrated here using a recently developed industrial distributed reconfigurable instruction cell based architecture [ 11]. The results show that schedules using this approach achieve equivalent throughput to VLIW architectures but at much lower power consumption.
Ioannis Nousias, Mark Milward, Sami Khawam, Tughrul Arslan, Iain Lindsay
DATE5
2006 A Reconfigurable Viterbi Decoder for a Communication Platform
abstract
A new large constraint length, soft decision Viterbi decoder fabric is presented for deployment using platform based system on chip methodologies. The decoder can be reconfigured for standards such as CDMA2000, WCDMA (UMTS), ADSL, IEEE 802.11, and GSM. Maximum resource allocation and performance is achieved by reusing components within turbo decoder base array. This cross platform Viterbi decoder is reconfigurable between different trellis types, constraint lengths and rates making it ideal for a unified multi-standard telecommunication platform. In addition, the authors also propose a novel technique for dynamic reconfiguration in order to achieve faster context switching between different mappings. The reconfigurable fabric is implemented as a subset of turbo decoder array on a 180 nm UMC process technology
Imran Ahmed 0001, Tughrul Arslan
FPL2
2006 An Efficient Fault Tolerance Scheme for Preventing Single Event Disruptions in Reconfigurable Architectures
abstract
Reconfigurable architectures are becoming increasingly popular with space related design engineers as they are inherently flexible to meet multiple requirements and offer significant performance and cost savings for critical applications. As the microelectronics industry has advanced, integrated circuit (IC) design and reconfigurable architectures (FPGAs, reconfigurable SoC and etc) have experienced dramatic increase in density and speed. These advancements have serious implications for the reconfigurable architectures when used in space environment where IC is subject to total ionization dose (TID) and single event effects as well. Due to transient nature of single event upsets (SEUs), these are most difficult to avoid in space-borne reconfigurable architectures. We present a unique SEU fault tolerance technique based upon double redundancy with comparison to overcome the overheads associated with the conventional schemes
Sajid Baloch, Tughrul Arslan, Adrian Stoica
FPL2
2006 Integrating the Electronics of the Control-Loops of the JPL/Boeing Gyroscope Within an Evolvable Hardware Architecture
abstract
This paper presents an autonomous custom reconfigurable architecture that employs evolvable hardware technology to accomplish the electronics of the control-loops of the JPL/Boeing gyroscope. The proposed adaptive hardware presents two functional modes. The switch between these two is controlled by an evolutionary strategy with only objective the most efficient adaptation of the system's functionality. The choice of operational mode mainly depends on the density of faults that occur on the hardware. Simulation results have shown that our architecture is able to adapt the functionality of the sensor's electronics within the presence of permanent stuck-at faults that occur on the user and configuration memory of the system. Moreover, the analysis of the power consumption reveals that the electronics that are accomplished within our reconfigurable architecture consume significantly less power compared with equivalent circuits, which are designed within industrial FPGAs
Evangelos F. Stefatos, Tughrul Arslan, Didier Keymeulen, Ian Ferguson
FPL2
2006 A real time programmable encoder for low density parity check code targeting a reconfigurable instruction cell architecture
abstract
This paper presents a new real time programmable irregular low density parity check (LDPC) encoder as specified in the IEEE P802.16E/D7 standard. The encoder is programmable for frame sizes from 576 to 2304 and for five different code rates. H matrix is efficiently generated and stored for a particular frame size and code rate. The encoder is implemented on reconfigurable instruction cell based architecture which has recently emerged as an ultra low power, high performance, ANSI-C programmable embedded core. Different general and technology specific optimization techniques are applied in order to achieve a throughput, ranging from 10 to 19 Mbps
Zahid Khan, Tughrul Arslan
FPT2
2006 An Efficient Decoder Scheme for Double Binary Circular Turbo Codes
abstract
Recently, double binary circular turbo code has received tremendous attention. Due to its better error-correcting capability than classical turbo code, it has commenced practical applications in current communication standards, such as DVB-RSC and IEEE 802.16 (Wimax). However current decoding schemes will incur a huge computation complexity. In this paper, authors present a novel decoding scheme for double binary circular turbo codes, which will not only reduce the computation complexity, but also give at least 0.5 dB performance gain compared with current decoding schemes
Cheng Zhan, Tughrul Arslan, Ahmet T. Erdogan, Scott MacDougall
ICASSP (4)2
2006 Low Power Cordic IP Core Implementation
abstract
There is a high demand for low power and efficient implementation of complex arithmetic operations in many Digital Signal Processing (DSP) algorithms. The CORDIC algorithm is suitable to be implemented in DSP systems since its calculation for complex arithmetic is simple and elegant. However, the large number of iterations involved in CORDIC operation limits its speed performance seriously and also consumes large power. This paper presents three CORDIC IP cores which were implemented using a new CORDIC algorithm. Each of them has one of more distinctive performance in terms of power, area, speed and flexibility due to their different architectures.
Jong Hun Han, Ahmet T. Erdogan, Tughrul Arslan
ICASSP (3)4
2006 A stochastic multi-objective algorithm for the design of high performance reconfigurable architectures
abstract
The increasing demand for FPGAs and reconfigurable hardware targeting high performance low power applications has lead to an increasing requirement for new high performance reconfigurable embedded FPGA cores. This paper presents a multi-objective population based algorithm which given a library of basic blocks and a list of constraints, identifies an optimum reconfigurable embedded reconfigurable core suitable for the target application.
Wing On Fung, Tughrul Arslan
IPDPS2
2006 A low energy VLSI design of random block interleaver for 3GPP turbo decoding
abstract
In this paper hardware architecture for internal random block interleaver compliant with the 3rd Generation Partnership Project (3GPP) turbo decoding is described. The complexity of this algorithm results in other implementations using large memories as address tables. In this implementation real time address computation avoids the use of pre-computed address storage. This greatly reduces the load on the processor and gives significant improvements in area and power. ASIC synthesis results on 0.18 mum CMOS UMC technology demonstrate the efficiency of the proposed VLSI interleaver architecture
Indrajit Ahmed, Tughrul Arslan
ISCAS2
2006 An embedded low power reconfigurable fabric for finite state machine operations
abstract
Generic reconfigurable finite state machine (FSM) array architecture is presented in this paper. The architecture has been customized for reconfiguration targeting applications requiring large number of states with the added advantage of low power consumption. Examples of up to 256 states have been targeted in this paper. Compared with commercial FPGA devices, the architecture provides the following reductions: up to 89.3% in power consumption, up to 55.2% in area and around 8% in delay time
Tughrul Arslan, Ahmet T. Erdogan
ISCAS2
2006 Low-power implementation of FIR filters within an adaptive reconfigurable architecture
abstract
This paper presents a custom very-large-scale-integration architecture, which consists of a reconfigurable hardware substrate and a hybrid-genetic algorithm responsible for resolving the optimal configuration for the reconfigurable components of the substrate. The reconfigurable hardware is specifically tailored for the implementation of multiplier-less symmetrical finite-impulse-response filters based on the primitive operator technique, while the architecture of the hybrid-genetic algorithm aims to improving the quality of the realized filters and speeding-up the time required for their realization. Power analysis demonstrates that the filters, which are implemented by our architecture, consume considerably less power than industrial field-programmable-gate-arrays, targeting similar applications.
Evangelos F. Stefatos, I. Bravos, Tughrul Arslan
ISCAS3
2006 A novel equaliser architecture with dynamic length optimisation
abstract
This paper presents a novel architecture for tap-length optimisation of the linear LMS equaliser. No analysis has previously been carried out to determine any tradeoff that exists in circuit area against power saving achieved. A low-complexity length update algorithm is employed to dynamically adjust and optimise the number of taps in the linear equaliser according to channel conditions. The results show that the chosen algorithm presents minimal overhead and reduces power consumed due to optimisation of the equaliser length. This paper presents the first complete architectural VLSI implementation of the length optimised equaliser and includes a performance study in terms of area and power
Mark P. Tennant, Ahmet T. Erdogan, Tughrul Arslan, John S. Thompson
ISCAS3
2006 A sensor system on chip for wireless microsystems
abstract
Recent years have seen the rapid development of microsensor technology, system on chip design, wireless technology and ubiquitous computing. When assembled into a complex microsystem the technologies become powerful tools in medical diagnostics, environmental monitoring and personal connectivity. In this paper we describe the demonstration of a silicon chip that has all the attributes required of a microsystem for use in these applications. The design methodology we have employed is a variant of the system on chip approach whereby many intellectual property blocks are integrated at a high level in the design flow. Our intellectual property blocks include the analogue sensor instrumentation for temperature and pH, a data multiplexing and conversion module, a digital platform based around an 8-bit microcontroller, data encoding for spread-spectrum wireless transmission and a RF section requiring very few off-chip components. The chip has been fully evaluated and tested by connection to external sensors. Each block has well defined interfaces so that they can be easily reused in future designs targeted to different applications
Lei Wang 0029, Nizamettin Aydin, A. Astaras, Mansour Ahmadian, Paul A. Hammond, T. B. Tang, Erik A. Johannessen, Tughrul Arslan, Steve P. Beaumont, Brian W. Flynn, Alan F. Murray, Jonathan M. Cooper, David R. S. Cumming
ISCAS8
2006 Analysis and Implementation of Multiple-Input, Multiple-Output VBLAST Receiver From Area and Power Efficiency Perspective
abstract
This paper presents an analysis of the vertical Bell Laboratories layered space time (VBLAST) receiver used in a multiple-input multiple-output (MIMO) wireless system from the hardware implementation perspective and identifies those processing elements that consume more area and power due to complex signal processing. This paper models a scalable VBLAST receiver based on minimum mean square error (MMSE) nulling criteria assuming a block flat fading channel. After identifying the major area and power consuming blocks, this paper proposes two area and power efficient VLSI architectures for the block that computes pseudoinverse of the channel matrix. This paper discusses different tradeoff issues in both architectures and compares them with the architectures in the literature
Zahid Khan, Tughrul Arslan, John S. Thompson, Ahmet T. Erdogan
IEEE Trans. Very Large Scale Integr. Syst.2
2005 The development of high performance FFT IP cores through hybrid low power algorithmic methodology
abstract
This paper presents a solution based on parallel-pipelined architectures for high throughput and power efficient FFT IP cores. Low power consumption can be gained through the combination of hybrid low power algorithms and architectures. A number of IP cores have been implemented for the comparison of the impact of parameterization on power/area/speed performance. The results show that up to 55% and 52% power saving can be achieved by the combination of the above techniques for 64-point 4-parallel-pipelined FFT and 16-point 2-parallel-pipelined FFT respectively, as compared to R4SDC pipelined FFTs.
Wei Han 0001, Ahmet T. Erdogan, Tughrul Arslan, Mohd. Hasan
ASP-DAC3
2005 A high performance synthesisable unsymmetrical reconfigurable fabric for heterogeneous finite state machines
abstract
Abstract- The use of synthesizable reconfigurable cores in system on chip (SoC) designs is increasingly becoming a trend. Such domain-special cores are being used for their flexibility, powerful function and low power consumption. A reconfigurable Finite State Machine (FSM) is constantly required for the purpose of control in any reconfigurable SoC. This paper presents a novel unbalanced unsymmetrical reconfigurable architecture for generic FSM; Compared with commercial FPGA devices, the new architecture results in area reduction of 43 % and power consumption decrease of 82%. I
Tughrul Arslan, Sami Khawam, Iain Lindsay
ASP-DAC2
2005 An AMBA AHB-based reconfigurable SOC architecture using multiplicity of dedicated flyby DMA blocks
abstract
We propose a System-on-Chip (SoC) architecture for reconfigurable applications based on the AMBA High-Speed Bus (AHB). The architecture features multiple low-area flyby DMA blocks for transferring configuration data. Furthermore, the architecture eliminates the use of energy-consuming instructions used in comparable commercial reconfigurable SoCs. The flyby DMA blocks achieve a reduction of up to 98% in the number of gates found in general-purpose DMA controllers. The DMA blocks also achieve the flyby throughput which halves the number of clock cycles used in conventional DMA for data transfer. We also demonstrate the presence of parallel processing which contributes to improved system performance of the proposed architecture over commercial comparatives.
Adeoye Olugbon, Sami Khawam, Tughrul Arslan, Ioannis Nousias, Iain Lindsay
ASP-DAC3
2005 Automatic synthesis and scheduling of multirate DSP algorithms
abstract
Abstract- To date, most high-level synthesis systems do not automatically solve present design problems, such as those related to timing associated with the physical implementation of multirate DSP architectures. Whilst others do not trade off area/speed of algorithm efficiently for such architectures. An automatic synthesis methodology based on both retiming techniques together with folding transformations is presented in this paper in order to solve timing problems associated with the implementation of multirate DSP algorithms. We demonstrate that techniques for modeling computational unit latencies, which can influence parameterisations of a multirate DSP IP core, can lead to highly efficient solutions. This is illustrated using a polyphase IIR IDCT example. Using the folding transformation, the control circuit for a hardware sharing multirate DSP is also presented. I
Mark Milward, Sami Khawam, Ioannis Nousias, Tughrul Arslan
ASP-DAC5
2005 A domain specific reconfigurable Viterbi fabric for system-on-chip applications
abstract
A novel embedded dynamically reconfigurable fabric for implementing the Viterbi algorithm in a System-on-Chip device is presented in this paper. The proposed reconfigurable fabric can support Viterbi implementations for different standards, such as GSM, IS-95, CDMA and Wireless LAN. Our results illustrate that the proposed architecture has superior power consumption and throughput characteristics and it is demonstrated a 80% reduction in power consumption over generic field programmable gate array (FPGA) and 40 times improvement in throughput over digital signal processor (DSP), respectively. Thus, the reconfigurable system-on-chip platform based on this kind of domain specific reconfigurable fabrics is an efficient solution for the high-performance portable communication systems.
Cheng Zhan, Tughrul Arslan, Sami Khawam, Iain Lindsay
ASP-DAC2
2005 Elitist selection schemes for genetic algorithm based printed circuit board inspection system
abstract
This paper presents the implementation of a number of elitist schemes for a low cost printed circuit board (PCB) inspection system. This strategy also aims to explore the role of tournament and roulette-wheel in improving the existing system when using a deterministic selection scheme. In this system, GA is used to detect rotation angle and displacement of PCB placed arbitrarily on a conveyor belt passing under the camera. Deterministic, tournament and roulette-wheel selection scheme have been compared in terms of maximum fitness, rate of accuracy and computation time. The finding shows that deterministic outperformed the other two schemes in all categories and still proves to be an ideal candidate for GA-based PCB inspection system. The modifications on population size and implementation of center block image matching technique also contributed to the improvement of computational time of the system.
Syamsiah Mashohor, Jonathan R. Evans, Tughrul Arslan
Congress on Evolutionary Computation3
2005 Techniques for the evolution of pipelined linear transforms
abstract
Linear transforms are used in an enormous variety of signal processing tasks. When implemented in hardware, multiplierless transforms have advantages in terms of power, silicon area and longest-path delay. The design space for multiplierless transforms can be large and highly multimodal, so powerful search techniques such as evolutionary algorithms are necessary. This paper investigates two ways in which the performance of such an evolutionary hardware design system might be improved. The first technique involves a reduction in the search space, while the second technique is a novel graph crossover operator.
Robert Thomson 0003, Tughrul Arslan
Congress on Evolutionary Computation2
2005 Low Power Domain-Specific Reconfigurable Array for Discrete Wavelet Transforms Targeting Multimedia Applications
abstract
Domain-specific heterogeneous reconfigurable arrays provide high performance over generic field programmable gate arrays (FPGAs) while maintaining the flexibility for that particular domain. This paper introduces an embedded domain-specific reconfigurable array that targets discrete wavelet transforms (DWT) and also presents different configurations of the array to prove its suitability for complex algorithms which are part of changing standards like JPEG and MPEG etc. The proposed array is flexible in order to accommodate various 5/3 and 9/7 discrete wavelet transforms which makes it quite suitable for multimedia applications. Experimental results demonstrate that the proposed architecture is over 31% more efficient in terms of power consumption over standard FPGAs.
Sajid Baloch, Imran Ahmed 0001, Tughrul Arslan, Adrian Stoica
FPL3
2005 Implementation of an efficient two-step SOVA turbo decoder for wireless communication systems
abstract
The authors present a soft-input soft-output (SISO) turbo decoder based on the two-step soft-output Viterbi algorithm (SOVA). The turbo decoder has been implemented with the trace-back algorithm (TBA). It achieves power and area reduction by employing an additional transition metric unit (TMU) for the generation of reliability values, instead of using a FIFO block for retrieving the previously stored reliability values in the survivor memory unit (SMU). Simulation results are provided for bit error rate (BER) performance using constraint lengths of K=3, 4, and 5. Implementation results for K=5 are also provided showing 46% area and 15% power savings compared to a conventional SOVA decoder
Jong Hun Han, Ahmet T. Erdogan, Tughrul Arslan
GLOBECOM3
2005 Multiplier-less based parallel-pipelined FFT architectures for wireless communication applications
abstract
This paper proposes two novel parallel-pipelined FFT architectures, based on multiplier-less implementation, targeting wireless communication applications, such as IEEE 802.11 wireless baseband chip and MC-CDMA receiver. The proposed parallel-pipelined architectures have the advantages of high throughput and high power efficiency. The multiplier-less architecture uses shift and addition operations to realize complex multiplications. By combining a new commutator architecture, and a low power butterfly with this approach, the resulting power and area savings are up to 31% and 20% respectively, for 64-point and 16-point FFTs, as compared to parallel-pipelined FFTs based on Booth coded Wallace tree multipliers.
Wei Han 0001, Tughrul Arslan, Ahmet T. Erdogan, Mohd. Hasan
ICASSP (5)2
2005 A direct-sequence spread-spectrum communication system for integrated sensor microsystems
abstract
Some of the most important challenges in health-care technologies have been identified to be development of noninvasive systems and miniaturization. In developing the core technologies, progress is required in pushing the limits of miniaturization, minimizing the costs and power consumption of microsystems components, developing mobile/wireless communication infrastructures and computing technologies that are reliable. The implementation of such miniaturized systems has become feasible by the advent of system-on-chip technology, which enables us to integrate most of the components of a system on to a single chip. One of the most important tasks in such a system is to convey information reliably on a multiple-access-based environment. When considering the design of telecommunication system for such a network, the receiver is the key performance critical block. The paper describes the application environment, the choice of the communication protocol, the implementation of the transmitter and receiver circuitry, and research work carried out on studying the impact of input data characteristics and internal data path complexity on area and power performance of the receiver. We provide results using a test data recorded from a pH sensor. The results demonstrate satisfying functionality, area, and power constraints even when a degree of programmability is incorporated in the system.
Nizamettin Aydin, Tughrul Arslan, David R. S. Cumming
IEEE Trans. Inf. Technol. Biomed.2
2004 Evolutionary recovery of electronic circuits from radiation induced faults
abstract
Radiation hard technologies for electronics are the conventional approach for survivability in high radiation environments. This work presents a novel approach based on evolvable hardware. The key idea is to reconfigure a programmable device, in-situ, to compensate, or bypass its degraded or damaged components. The paper demonstrates the approach using a JPL-developed reconfigurable device, a field programmable transistor array (FPTA), which shows recovery from radiation damage when reconfigured under the control of evolutionary algorithms. Experiments with total radiation dose up to 350kRad show that while the functionality of a variety of circuits, including a rectifier and a digital to analog converter implemented on a FPTA-2 chip, is degraded/lost at levels before l00KRad, the correct functionality can be recovered through the proposed evolutionary approach. The evolutionary algorithm controls the state of about 1,500 switches that determine configurations on the FPTA-2 programmable device. Evolution is able to use the resources of the reconfigurable cells, even radiation damaged components, to synthesize a new solution.
Adrian Stoica, Didier Keymeulen, Vu Duong, Ricardo Salem Zebulum, Ian Ferguson, Taher Daud, Tughrul Arslan, Xin Guo 0002
IEEE Congress on Evolutionary Computation7
2004 Efficient Implementations of Mobile Video Computations on Domain-Specific Reconfigurable Arrays
abstract
Mobile video processing as defined in standards like MPEG-4 and H.263 contains a number of timeconsuming computations that cannot be efficiently executed on current hardware architectures. The authors recently introduced a reconfigurable SoC platform that permits a low-power, high-throughput and flexible implementation of the motion estimation and DCT algorithms. The computations are done using domainspecific reconfigurable arrays that have demonstrated up to 75% reduction in power consumption when compared to generic FPGA architecture, which makes them suitable for portable devices. This paper presents and compares different configurations of the arrays to efficiently implementing DCT and motion estimation algorithms. A number of algorithms are mapped into the various reconfigurable fabrics demonstrating the flexibility of the new reconfigurable SoC architecture and its ability to support a number of implementations having different performance characteristics.
Sami Khawam, Sajid Baloch, Arjun Pai, Imran Ahmed 0001, Nizamettin Aydin, Tughrul Arslan, Fred Westall
DATE6
2004 Unidirectional Switch-Boxes for Synthesizable Reconfigurable Arrays
abstract
The new trend in designing reconfigurable system-on-chip (SoC) by using embedded FPGAs provides numerous improvements to ASIC designs due to added flexibility and improvements in functionality. Further advantages are introduced if FPGA is provided as a synthesis to core that fits well in the SoC architecture and design flow. This paper proposes and investigates different design of switch-boxes suitable for synthesis in reconfigurable coarse-grain architectures. The various designs are evaluated and compared in terms of power consumption, area, delays and routability.
Sami Khawam, Tughrul Arslan, Fred Westall
FCCM2
2004 Switch-box design for synthesizable coarse-grain arrays for system-on-chip applications
abstract
The use of synthesizable embedded FPGAs in a reconfigurable systems-on-chip (SoC) provides numerous improvements to ASIC designs due to the added flexibility and improvements in functionality. Such arrays have been proposed earlier by the authors and it was found that, as with all reconfigurable architectures, most of the power and area consumed is in the programmable interconnects mesh. This paper focuses on the design of optimized synthesizable switch-boxes that can be used in reconfigurable coarse-grain architectures. The paper compares different switch-boxes in terms of area, power, delay and mutability and proposes new designs optimized for directional data-flow which are found to provide up to 47% less area and 22% less power with only an increase of 10% in routed wirelength and delays when compared to existing designs.
Sami Khawam, Tughrul Arslan
FPT2
2004 Self-recovery experiments in extreme environments using a field programmable transistor array
abstract
Temperature and radiation tolerant electronics, as well as long life survivability are key capabilities required for future NASA missions. Current approaches to electronics for extreme environments focus on component level robustness and hardening. However, current technology can only ensure very limited lifetime in extreme environments. This paper describes novel experiments that allow adaptive in-situ circuit redesign/reconfiguration during operation in extreme temperature and radiation environments. This technology would complement material/device advancements and increase the mission capability to survive harsh environments. The approach is demonstrated on a mixed-signal programmable chip (FPTA-2), which recovers functionality for temperatures until 28/spl deg/C and with total radiation dose up to 250kRad.
Adrian Stoica, Didier Keymeulen, Tughrul Arslan, Vu Duong, Ricardo Salem Zebulum, Ian Ferguson, Xin Guo 0002
FPT3
2004 Domain specific reconfigurable fabric targeting Viterbi algorithm
abstract
This work presents a novel embedded reconfigurable fabric targeting efficient implementation of the Viterbi decoder within a system-on-chip device. The proposed reconfigurable fabric can support constraint lengths ranging from 3 to 9, and code rates in the range 1/2-1/3.Our results demonstrate that this novel architecture has superior throughput and power consumption characteristics when compared to generic DSPs and FPGAs respectively.
Cheng Zhan, Sami Khawam, Tughrul Arslan
FPT3
2004 Synthesizable Reconfigurable Array Targeting Distributed Arithmetic for System-on-Chip Applications
abstract
Summary form only given. Domain-specific reconfigurable arrays are embedded arrays optimized for one domain of applications providing performance improvements over generic embedded field programmable gate arrays (FPGAs). An embedded reconfigurable array that targets distributed arithmetic (DA) implementations is presented. DA includes calculations that are commonly found in multimedia applications, such as filtering and discrete cosine transform (DCT). Two benchmark DCT circuits are implemented on the array, on conventional FPGAs and on hardwired cores. The performance measured shows considerable improvements in area, power consumption and timing when comparing the presented array with FPGAs. Experimental results are provided which demonstrate the suitability of our architecture in low-power system-on-chip platforms targeting portable mobile devices.
Sami Khawam, Tughrul Arslan, Fred Westall
IPDPS2
2004 Evolutionary design and adaptation of high performance digital filters within an embedded reconfigurable fault tolerant hardware platform
Ben I. Hounsell, Tughrul Arslan, Robert Thomson 0003
Soft Comput.2
2003 On the impact of modelling, robustness and diversity to the performance of a multi-objective evolutionary algorithm for digital VLSI system design
abstract
This paper describes the operation of an evolutionary algorithm (EA) for the creation of linear digital VLSI circuit designs. The EA can produce hardware designs from a behavioural description of a problem. The designs are based upon a library of high-level components. The EA performs a multi-objective search, using models of the longest-path delay and the silicon area of a design. These models are based upon the properties of real-world components, implementable in a 0.18 micron technology. The accuracy of these models is investigated. Two important aspects of multi-objective evolution are the population diversity, and the variability of the results. Both of these areas are examined. The population diversity is assessed in terms of conflict between the objectives, and the robustness of the EA is experimentally investigated.
Robert Thomson 0003, Tughrul Arslan
IEEE Congress on Evolutionary Computation2
2003 A genetic algorithm for energy efficient device scheduling in real-time systems
abstract
Most embedded systems have tight constraints on power consumption because the amount of power available to these systems is limited due to the limitation of battery life. For this reason energy consumption is an important parameter in evaluating performance of embedded systems. DPM (dynamic power management) has gained considerable attention over the last few years as a way to save energy in device that can be turned on and off by operating system control. Scheduling is very important for DPM since it directly affects the efficiency of DPM. We have implemented a customised genetic algorithm which generates a near-optimal device schedule for a set of real-time tasks, with the goal of minimising the power consumed. When compared with other schedulers, the genetic algorithm based system is shown to have less memory and time requirements and scale far better as the problem complexity is increased.
Lirong Tian, Tughrul Arslan
IEEE Congress on Evolutionary Computation2
2003 Power/Area Analysis and Optimization of a DS-SS receiver for an Integrated Sensor Microsystem
abstract
Communication systems targeting miniaturized sensor microsystem networks are characterized by their restricted power and area constraints. When considering the design of telecommunication system for such a network, the receiver is the key performance critical block. This paper describes research work carried out on studying the impact of input data characteristics and resolution and internal data path complexity on area and power performance of the receiver. We have constructed a number of transmitter/receiver architectures and analyzed their power/area. We demonstrate that up to 59% and 11% savings in area and power respectively could be achieved by optimizing input data size and internal register width for a particular application while maintaining signal quality.
Nizamettin Aydin, Tughrul Arslan, David R. S. Cumming
DSD2
2003 Domain-Specific Reconfigurable Array for Distributed Arithmetic
Sami Khawam, Tughrul Arslan, Fred Westall
FPL2
2003 A Genetic Algorithm for Energy Efficient Device Scheduling in Real-Time Systems
Lirong Tian, Tughrul Arslan
GECCO2
2003 Crosstalk Immune Coding from Area and Power Perspective for high performance AMBA based SoC systems
Zahid Khan, Tughrul Arslan, Ahmet T. Erdogan
VLSI-SOC2
2002 An evolutionary algorithm for the multi-objective optimisation of VLSI primitive operator filters
abstract
This paper introduces an evolvable hardware system for the generation of optimised FIR filter designs. This system converts the frequency domain specification of a filter directly into a circuit netlist. The filter designs are optimised with respect to an accurate model of the silicon area and latency.
Robert Thomson 0003, Tughrul Arslan
IEEE Congress on Evolutionary Computation2
2002 GPS attitude determination using a genetic algorithm
abstract
In this paper, a new technique that uses a specially tailored genetic algorithm is proposed for attitude determination via GPS carrier phase observables. The technique overcomes restrictions due to computational overheads incurred by existing techniques such as the ambiguity function method. We present experimental results which show that the algorithm is able to efficiently search the complex search space imposed by the problem in addition to being immune to cycle slips compared to other conventional methods.
Jiangning Xu, Tughrul Arslan, Dejun Wan, Qing Wang 0026
IEEE Congress on Evolutionary Computation2
2002 FFT coefficient memory reduction technique for OFDM applications
abstract
There is a strong need to implement long FFT's in applications like orthogonal frequency division multiplexing (OFDM), radars and sonars etc. It is highly desirable to reduce the size and power requirements of the FFT so as to realize single chip long FFT based systems targeting portable applications. This paper presents a novel technique to reduce the coefficient memory almost by a factor of four by exploiting the relationships among the coefficient values thereby significantly reducing the area and power requirements of the hardware.
Mohd. Hasan, Tughrul Arslan
ICASSP2
2002 An embedded extension algorithm for the lifting based Discrete Wavelet Transform in JPEG2000
abstract
This paper presents a novel algorithmic technique for reducing power and area consumption of the Discrete Wavelet Transform (DWT) when implemented in VLSI hardware. The technique reduces power by combining the data-extension procedure into the lifting-based DWT core resulting in a significant reduction in the amount of memory requirements and associated read/write operation together with a reduction in the number of arithmetic operations. This in turn has the effect of reducing the amount of switched capacitance within the hardware unit. We demonstrate that the new algorithm can be used to obtain more than 50% reduction in the amount of memory required leading to significant reduction in area and power.
Kay-Chuan Benny Tan, Tughrul Arslan
ICASSP2
2001 Synthesis of low-power DSP systems using a genetic algorithm
abstract
This paper presents a new tool for the synthesis of low-power VLSI designs, specifically, those designs targeting digital signal processing applications. The synthesis tool genetic algorithm for low-power synthesis (GALOPS) uses a genetic algorithm to apply power-reducing transformations to high-level signal-processing designs, producing designs that satisfy power requirements as well as timing and area constraints. GALOPS uses problem-specific genetic operators that are specifically tailored to incorporate VLSI-based digital signal processing design knowledge. A number of signal-processing benchmarks are used to facilitate the analysis of low-power design tools, and to aid in the comparison of results. Results demonstrate that GALOPS achieves significant power reductions in the presented benchmark designs. In addition, GALOPS produces a family of unique solutions for each design, all of which satisfy the multiple design objectives, providing flexibility to the VLSI designer.
Mark S. Bright, Tughrul Arslan
IEEE Trans. Evol. Comput.2
2000 A novel genetic algorithm for the automated design of performance driven digital circuits
abstract
Presents a genetic algorithm for the design of high-performance arithmetic circuits for evolvable hardware applications. A distinct feature of the algorithm is its ability to directly evolve and evaluate circuits in a hardware description language (HDL), within a novel environment termed the Virtual Chip. Because the Virtual Chip evolves circuit structures within a HDL, detailed simulation and analysis of each circuit is possible with any technology-specific component library. This feature allows accurate analysis of performance issues such as timing and area. The paper describes the genetic algorithm and the hardware evaluation environment, and provides results with a number of benchmark arithmetic circuits evolved under different performance-driven timing and area constraints. Our results reveal that the genetic algorithm is able to exploit the flexibility provided by a novel chromosome architecture, and utilise a combination of primitive gates and macro components from a component library in order to produce circuits which operate well within timing restrictions. The validity of our results are further supported by comparing the performance of functionally equivalent circuits generated using standard high-level design methodologies.
Ben I. Hounsell, Tughrul Arslan
CEC2
2000 A Novel Evolvable Hardware Framework for the Evolution of High Performance Digital Circuits
Ben I. Hounsell, Tughrul Arslan
GECCO2
2000 An order based segmentation algorithm for low power implementation of digital filters
abstract
The paper presents a new algorithm for low power implementation of digital filters. The algorithm reduces power consumption through a two phased strategy, which targets reducing the switched capacitance within the multiplier circuit. The first phase involves the segmentation of coefficients into more primitive components which could in turn be processed through a single shift and a more primitive multiplication operations. The second phase exploits the correlation among the new set of coefficients at the coefficient input of the multiplier for more reduction in switched capacitance. The paper describes the algorithm and its evaluation environment and provides results with a number of filter examples demonstrating up to 65% reduction in power compared to conventional filtering.
Ahmet T. Erdogan, Tughrul Arslan
ICASSP2
2000 A hybrid segmentation and block processing algorithm for low power implementation of digital filters
abstract
The paper proposes a hierarchical algorithm for low power implementation of digital filters. The algorithm processes data and coefficients in blocks of fixed sizes. During the manipulation of each block, coefficients are segmented into two primitive components. The accumulative effect of processing a sequence of blocks and segmentation results in up to 80% reduction in power consumption compared to conventional filtering. The paper describes the implementation of the hierarchical algorithm, its constituent components and the power evaluation environment developed. Results are provided which demonstrate the effectiveness of the algorithm.
Ahmet T. Erdogan, Tughrul Arslan
ISCAS2