EDBT 2026 Demo / reviewers in the wild / expert
Theocharis Theocharides
dblp:59/3933 · also Theo Theocharides
· DBLP profile ↗
63ranked-venue papers
2as first author
21since 2021 · last 2026
0000-0001-7222-9152ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 2 first-author · 11 since 2021Software engineering, systems software and programming languages · 11 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Computer networks · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: Enhancing Resilience, Efficiency, and Trustworthiness of Edge AI in Safety-Critical Systems (GuardAI)abstractAI at the network edge promises real-time perception and decision-making in safety-critical domains such as aerial robotics, autonomous vehicles, and 5G-enabled infrastructures. Yet, operating under resource constraints, dynamic, and adversarial conditions exposes edge AI systems to fragility, inefficiency, and security risks that threaten their safe operation. GuardAI, a Horizon Europe project, introduces a framework for resilient and trustworthy edge AI that unites three pillars: adversarial robustness, context-enhanced inference, and security-by-design. Initial project results include a diffusion-based adversarial purification framework optimized for real-time operation, lightweight deep unrolling architectures for LiDAR super-resolution with built-in outlier removal, and robust uncertainty quantification modules to improve confidence calibration. It further develops a context-enhanced inference engine that integrates visual, spatial, and operational context across multi-agent systems, and a risk-aware defense recommender that autonomously selects mitigation strategies based on evolving threat landscapes. Through representative Use Cases, covering monitoring with Unmanned Aerial Vehicle, decentralized 5G network analytics, and secure perception in connected autonomous vehicles, GuardAI demonstrates how robust and adaptive AI can be achieved within stringent edge constraints. Together, these technologies lay the groundwork for a new generation of secure, context-aware, and certifiable AI systems that can be trusted to operate autonomously in the physical world. Antonis D. Savva, Mehmet Demirel, Yeshwanth Kumar Adimoolam, Rafaella Elia, Alexandros Gkillas, Erion-Vasilis M. Pikoulis, Amalia Damianou, Charmaine Barker, Daniel Bethell, Ahmed Salah Tawfik Ibrahim, Filippo Cugini, Francesco Paolucci, Kyriakos Vlachos, Simos Gerasimou, Antonios Lalas, Konstantinos Votis, Aris S. Lalos, Christos Kyrkou, Theocharis Theocharides |
DATE | 20 |
| 2026 | ARMOR: Architectural Reliability via Multi-Objective Optimization for Early-Exit Neural Networks
Michalis Kontos, Georgios Konstantinidis 0003, Theocharis Theocharides, Maria K. Michael |
ETS | 3 |
| 2026 | Loss Landscape Topology Reveals Why Simple Baselines are Competitive at 3D Point Cloud Segmentation Under Class Imbalance
Antonis D. Savva, Christos Kyrkou, Theocharis Theocharides |
ICPR (6) | 3 |
| 2026 | A reliability- and latency-driven task allocation framework for workflow applications in the edge-hub-cloud continuum
Andreas Kouloumpris, Georgios L. Stavrinides, Maria K. Michael, Theocharis Theocharides |
Future Gener. Comput. Syst. | 4 |
| 2026 | Exact, Efficient, and Reliable Multiobjective and Multiconstrained IoT Workflow Scheduling in Edge-Hub-Cloud Cyber-Physical SystemsabstractEmerging Internet of Things (IoT)-enabled cyber-physical applications, such as autonomous critical infrastructure inspection, demand low-latency, energy-efficient, and reliable execution across resource-constrained edge devices with heterogeneous multicore processors and diverse sensing and actuating capabilities, in collaboration with a hub device and a cloud server. These workflow-based applications comprise interdependent tasks that must be executed under stringent deadline, reliability, capability, memory, storage, and energy constraints. Given their critical nature, exact optimization is necessary to obtain optimal schedules that ensure dependable operation. Existing scheduling approaches, both exact and heuristic, fail to jointly address all these objectives and constraints. To this end, we propose an exact multi-objective and multi-constrained workflow scheduling approach for edge-hub-cloud cyber-physical systems, based on continuous-time mixed integer linear programming. The proposed formulation jointly optimizes latency, energy, and reliability, while holistically addressing timing and resource constraints. To enhance reliability while avoiding the overhead of unnecessary task replicas, it selectively employs task duplication. We evaluate our approach against a widely used heuristic, which we extend to ensure a fair and meaningful comparison, using a real-world IoT workflow and synthetic task graphs of varying sizes, across different system configurations and objective trade-offs. The proposed method consistently outperforms the heuristic, achieving up to 29.83%, 33.96%, and 28.49% average improvements in latency, energy, and reliability, respectively, while attaining practical runtimes. Overall, the experimental results demonstrate the effectiveness of our approach under various system configurations and objective trade-offs, and show its practical scalability to task graphs of sizes relevant to the targeted applications and system architecture. Andreas Kouloumpris, Georgios L. Stavrinides, Maria K. Michael, Theocharis Theocharides |
IEEE Internet Things J. | 4 |
| 2025 | Reliability Assessment of Early Exit Deep Convolutional Neural NetworksabstractMachine learning inference deployed on edge devices is subject to limited resources, which pushes for better energy and resource efficiency, while ensuring high performance. At the algorithmic level, dynamic Deep Neural Networks (dDNNs) have been proposed in an effort to improve performance and energy efficiency which is particularly useful in time-critical systems such as autonomous navigation and search and rescue operations. These systems are often deployed in harsh environments, making the hardware vulnerable to soft errors. This work investigates the reliability of such dDNNs, specifically convolutional, focusing on early exit approaches, in an effort to better comprehend the impact of the early exits on the overall reliability of dDNNs. We consider transient faults and use hardware-aware software fault models, combined with a fine-grained fault taxonomy, to derive failures-in-time (FIT) rate and assess reliability. We observe a new category of faults specific to dDNNs and we refer to them as non-materialized faults. In our assessment, we use a new fault taxonomy that is more fine-grained than existing ones, also including the new category of non-materialized faults. We explore various early exit scenarios, having as baseline the Google Inception V1 network deployed on an NVDLA accelerator to generate FIT rates and analyze the faults per network architecture block and operation type. Furthermore, we investigate the impact of faults on the overall performance (accuracy) and execution time (average inference time) of the neural network models. Based on our analysis we observe that the overall FIT rate remains similar among different early exit implementations (8.40-8.89), while the distribution of the critical faults is influenced by the network’s design-time decisions on early exit block placement and architecture. Georgios Konstantinidis 0003, Maria K. Michael, Theocharis Theocharides |
ATS | 3 |
| 2025 | Jointly-Optimized Trajectory Generation and Camera Control for 3D Coverage PlanningabstractThis work proposes a jointly optimized trajectory generation and camera control approach, enabling an autonomous agent, such as an unmanned aerial vehicle (UAV) operating in 3D environments, to plan and execute coverage trajectories that maximally cover the surface area of a 3D object of interest. Specifically, the UAV's kinematic and camera control inputs are jointly optimized over a rolling planning horizon to achieve complete 3D coverage of the object. The proposed controller incorporates ray-tracing into the planning process to simulate the propagation of light rays, thereby determining the visible parts of the object through the UAV's camera. This integration enables the generation of precise look-ahead coverage trajectories. The coverage planning problem is formulated as a rolling finite-horizon optimal control problem and solved using mixed-integer programming techniques. Extensive real-world and synthetic experiments validate the performance of the proposed approach. Savvas Papaioannou, Panayiotis Kolios, Theocharis Theocharides, Christoforos Panayiotou, Marios M. Polycarpou |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Convolutional Channel-Wise Competitive Learning for the Forward-Forward AlgorithmabstractThe Forward-Forward (FF) Algorithm has been recently proposed to alleviate the issues of backpropagation (BP) commonly used to train deep neural networks. However, its current formulation exhibits limitations such as the generation of negative data, slower convergence, and inadequate performance on complex tasks. In this paper we take the main ideas of FF and improve them by leveraging channel-wise competitive learning in the context of convolutional neural networks for image classification tasks. A layer-wise loss function is introduced that promotes competitive learning and eliminates the need for negative data construction. To enhance both the learning of compositional features and feature space partitioning, a channel-wise feature separator and extractor block is proposed that complements the competitive learning process. Our method outperforms recent FF-based models on image classification tasks, achieving testing errors of 0.58%, 7.69%, 21.89%, and 48.77% on MNIST, Fashion-MNIST, CIFAR-10 and CIFAR-100 respectively. Our approach bridges the performance gap between FF learning and BP methods, indicating the potential of our proposed approach to learn useful representations in a layer-wise modular fashion, enabling more efficient and flexible learning. Our source code and supplementary material are available at https://github.com/andreaspapac/CwComp. Andreas Papachristodoulou, Christos Kyrkou, Stelios Timotheou, Theocharis Theocharides |
AAAI | 4 |
| 2024 | Optimal Multi-Constrained Workflow Scheduling for Cyber-Physical Systems in the Edge-Cloud ContinuumabstractThe emerging edge-hub-cloud paradigm has enabled the development of innovative latency-critical cyber-physical applications in the edge-cloud continuum. However, this paradigm poses multiple challenges due to the heterogeneity of the devices at the edge of the network, their limited computational, communication, and energy capacities, as well as their different sensing and actuating capabilities. To address these issues, we propose an optimal scheduling approach to minimize the overall latency of a workflow application in an edge-hub-cloud cyber-physical system. We consider multiple edge devices cooperating with a hub device and a cloud server. All devices feature heterogeneous multicore processors and various sensing, actuating, or other specialized capabilities. We present a comprehensive formulation based on continuous-time mixed integer linear programming, encapsulating multiple constraints often overlooked by existing approaches. We conduct a comparative experimental evaluation between our method and a well-established and effective scheduling heuristic, which we enhanced to consider the constraints of the specific problem. The results reveal that our technique outperforms the heuristic, achieving an average latency improvement of 13.54% in a relevant real-world use case, under varied system configurations. In addition, the results demonstrate the scalability of our method under synthetic workflows of varying sizes, attaining a 33.03% average latency decrease compared to the heuristic. Andreas Kouloumpris, Georgios L. Stavrinides, Maria K. Michael, Theocharis Theocharides |
COMPSAC | 4 |
| 2024 | Embedded and Real-Time Anomalous Command Classification in Unmanned Ground Vehicle OperationsabstractUnmanned Ground Vehicles (UGVs) are increasingly used in safety-critical applications, typically controlled under challenging conditions that may cause stress and fatigue on the operator. This can potentially compromise the safety of the mission due to involuntary movements by the operator, resulting in abnormal commands issued on the remote controller of the vehicle. Such movements can be detected by evaluating the mental state of the operator, the operational context of the UGV, and of course real-time movement detection by the operator. To detect such anomalous commands, we propose a three-stage Machine Learning (ML) based approach, which is suitable for embedded implementation providing real-time classification. Firstly, we detect whether there is any movement on the controller and classify the type of the given movement. Next, we classify whether the operator is under incremental stress and finally discern, in relationship to the UGV's operational context extracted by the state of the UGV is considered normal or not. We use a dataset collected through real-world scenarios, to evaluate the proposed approach, and we evaluate the performance of the proposed approach on an embedded platform (Jetson Xavier NX), extracting relevant metrics such as processing time, energy consumption, and classification accuracy. Our findings demonstrate successful high classification rates for the mental state of the operator (96%) and the ground vehicle movements (74%), as well as the associated involuntary command recognition, with low energy$(0.43mJ)$and time$(0.39s)$requirements. Rafaella Elia, Theocharis Theocharides |
VLSI-SoC | 2 |
| 2024 | An optimization framework for task allocation in the edge/hub/cloud paradigm
Andreas Kouloumpris, Georgios L. Stavrinides, Maria K. Michael, Theocharis Theocharides |
Future Gener. Comput. Syst. | 4 |
| 2024 | Introduction to the Special Issue on tinyMLabstractSpecial Issue Part 1 (Issue 3) and Part 2 (Issue 4) of AIEDAM are based on a workshop on Learning and Creativity held at the 2002 conference on Artificial Intelligence in Design, AID '02 (www.cad.strath.ac.uk/AID02_workshop/Workshop_webpage.html; Gero, ... Theocharis Theocharides, Charlotte Frenkel, Lukas Cavigelli |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2023 | True Rank Guided Efficient Neural Architecture Search for End to End Low-Complexity Network Discovery
Shahid Siddiqui, Christos Kyrkou, Theocharis Theocharides |
CAIP (1) | 3 |
| 2023 | Introducing Convolutional Channel-wise Goodness in Forward-Forward LearningabstractThis paper introduces a Channel-wise Goodness Function (CWG) that enhances the Forward-Forward through the use of Convolutional Neural Networks.The CWG function facilitates simultaneous feature extraction and separation, eliminating the requirement for constructing negative data and leading to faster convergence rates.The approach employs a two-component loss function that maximizes positive goodness and minimizes negative goodness.This enables the model to learn class-specific features to outperform recent non-backpropagation approaches on basic image classification datasets and shorten the gap with the well-established backpropagation methods. Andreas Papachristodoulou, Christos Kyrkou, Stelios Timotheou, Theocharis Theocharides |
ESANN | 4 |
| 2023 | Machine Learning for Emergency Management: A Survey and Future OutlookabstractEmergency situations encompassing natural and human-made disasters, as well as their cascading effects, pose serious threats to society at large. Machine learning (ML) algorithms are highly suitable for handling the large volumes of spatiotemporal data that are generated during such situations. Hence, over the years, they have been utilized in emergency management to aid first responders and decision-makers in such situations and ultimately improve disaster prevention, preparedness, response, and recovery. In this survey article, we highlight relevant work in this area by first focusing on the commonalities of emergency management applications and key challenges that ML algorithms need to address. Then, we present a categorization of relevant works across all the emergency management phases and operations, highlighting the main algorithms used. Based on our review, we conclude that ML algorithms can provide the basis for tackling different activities across the emergency management phases with a unified algorithmic framework that can solve a large set of problems. Finally, through the systematic literature review, we provide promising future directions for utilizing ML algorithms more effectively in emergency management applications. More importantly, we identify the need for better generalization of algorithms, improved explainability, and trustworthiness of ML algorithms with respect to the emergency management personnel, as well as more efficient ways of addressing the challenges associated with building appropriate datasets. Christos Kyrkou, Panayiotis Kolios, Theocharis Theocharides, Marios M. Polycarpou |
Proc. IEEE | 3 |
| 2023 | Distributed Search Planning in 3-D Environments With a Dynamically Varying Number of AgentsabstractIn this work, a novel distributed search-planning framework is proposed, where a dynamically varying team of autonomous agents cooperate in order to search multiple objects of interest in three-dimension (3-D). It is assumed that the agents can enter and exit the mission space at any point in time, and as a result the number of agents that actively participate in the mission varies over time. The proposed distributed search-planning framework takes into account the agent dynamical and sensing model, and the dynamically varying number of agents, and utilizes model predictive control (MPC) to generate cooperative search trajectories over a finite rolling planning horizon. This enables the agents to adapt their decisions on-line while considering the plans of their peers, maximizing their search planning performance, and reducing the duplication of work. Savvas Papaioannou, Panayiotis Kolios, Theocharis Theocharides, Christoforos Panayiotou, Marios M. Polycarpou |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | A Comprehensive Solution for Securing Connected and Autonomous VehiclesabstractWith the advent of Connected and Autonomous Vehicles (CAVs) comes the very real risk that these vehicles will be exposed to cyber-attacks by exploiting various vulnerabilities. This paper gives a technical overview of the H2020 CARAMEL project (currently in the intermediate stage) in which Artificial Intelligent (AI)-based cybersecurity for CAVs is the main goal. Most of the possible scenarios are considered, by which an adversary can generate attacks on CAVs, such as attacks on camera sensors, GPS location, Vehicle to Everything (V2X) message transmission, the vehicle's On-Board Unit (OBU), etc. The counter-measures to these attacks and vulnerabilities are presented via the current results in the CARAMEL project achieved by implementing the designed security algorithms. Mohsin Kamal, Christos Kyrkou, Nikos Piperigkos, Andreas Papandreou, Andreas Kloukiniotis, Jordi Casademont, Natlia Porras Mateu, Daniel Baos Castillo, Rodrigo Diaz Rodriguez, Nicola Gregorio Durante, Petros Kapsalas, Aris S. Lalos, Konstantinos Moustakas, Christos Laoudias, Theocharis Theocharides, Georgios Ellinas |
DATE | 16 |
| 2022 | AirCamRTM: Enhancing Vehicle Detection for Efficient Aerial Camera-based Road Traffic MonitoringabstractEfficient road traffic monitoring is playing a fundamental role in successfully resolving traffic congestion in cities. Unmanned Aerial Vehicles (UAVs) or drones equipped with cameras are an attractive proposition to provide flexible and infrastructure-free traffic monitoring. However, real-time traffic monitoring from UAV imagery poses several challenges, due to the large image sizes and presence of non-relevant targets. In this paper, we propose the AirCam-RTM framework that combines road segmentation and vehicle detection to focus only on relevant vehicles, which as a result, improves the monitoring performance by ~2 × and provides ~ 18% accuracy improvement. Furthermore, through a real experimental setup we qualitatively evaluate the performance of the proposed approach, and also demonstrate how it can be used for real-time traffic monitoring using UAVs. Rafael Makrigiorgis, Nicolas Hadjittoouli, Christos Kyrkou, Theocharis Theocharides |
WACV | 4 |
| 2022 | Introduction to the Special Issue on Accelerating AI on the Edge - Part 1abstractIntroduction to the Special Issue on Accelerating AI on the Edge -Part 1Machine Learning (ML) is nowadays embedded in several computing devices, consumer electronics, and cyber-physical systems.Smart sensors are deployed everywhere, in applications such as autonomous vehicles, robotics, wearables and perceptual computing devices, and intelligent algorithms power our connected world.These devices collect and aggregate volumes of data, and in doing so, they augment our society in multiple ways: from healthcare, to social networks, to consumer electronics, and many more.To process these immense volumes of data, machine learning is emerging as the de facto analysis tool, that powers several aspects of our Big Data society.Applications spanning from infrastructure (smart agriculture, smart cities, intelligent transportation systems, smart grids, to name a few), to social networks and content delivery, to e-commerce and smart factories, and emerging concepts such as self-driving cars and autonomous/assistive robots, are powered by advanced machine learning technologies.These emerging systems require real-time inference and decision support; such scenarios therefore may use customized hardware accelerators, are typically bound by limited compute, memory and energy resources, and are restricted to limited connectivity and bandwidth.Thus, near-sensor computation and near-sensor intelligence are starting to emerge as necessities, in order to continue supporting the paradigm shift of our connected world.The need for real-time intelligent data analytics (especially in the era of Big Data) for decision support near the data acquisition points, emphasizes the need of revolutionizing the way we design, build, test and verify processors, accelerators and systems that facilitate machine learning (and deep learning in particular) implemented in resource-constrained environments for use at the edge and the fog.As such, traditional Von Neumann architectures may no longer be sufficient and suitable, primarily because of limitations in both performance and energy efficiency caused especially by large amounts of data movement.Furthermore, due to the connected and critical nature of such systems, security and reliability are also critically important.To facilitate AI at the edge, we need to re-focus on problems such as design, verification, architecture, scheduling and allocation policies, optimization, and many more, for determining the most efficient way to implement these novel applications within a resource-constrained system, which may or may not be connected.Acceleration of AI at the edge therefore, is a fast-growing field of machine learning technologies and applications including algorithms, hardware, and software capable of performing on-device sensor (vision, audio, IMU, biomedical, etc.) data analytics at extremely low power, typically in the mW range and below, and hence enabling a variety of always-on use-cases and targeting battery-operated devices.This special issue therefore targets research Muhammad Shafique 0001, Theocharis Theocharides, Hai Li 0001, Chun Jason Xue |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | Introduction to the Special Issue on Accelerating AI on the Edge - Part 2
Muhammad Shafique 0001, Theocharis Theocharides, Hai Li 0001, Chun Jason Xue |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2021 | TinyML: Current Progress, Research Challenges, and Future RoadmapabstractTinyML: tiny in size, BIG in impact!This paper highlights the current progress, challenges and open research opportunities in the domain of tinyML, benchmarking, and emerging applications for Edge-AI. Muhammad Shafique 0001, Theocharis Theocharides, Vijay Janapa Reddi, Boris Murmann |
DAC | 2 |
| 2020 | Operation and Topology Aware Fast Differentiable Architecture SearchabstractDifferentiable architecture search (DARTS) has gained significant attention amongst neural architecture search approaches due to its effectiveness in finding competitive network architectures with affordable computational complexity. However, DARTS' search space is designed such that even a randomly sampled architecture performs reasonably well. Moreover, due to the complexity of search architectural building block or cell, it is unclear whether these are certain operations or the cell topology that contributes most to achieving higher final accuracy. In this work, we dissect the DARTS's search space to understand which components are most effective in producing better architectures. Our experiments show that: (1) Good architectures can be discovered regardless of the search network depth; (2) Seperable convolution with 3x3 kernel is the most effective operation in this search space; and (3) The cell topology also has substantial effect on the accuracy. Based on these insights, we propose an efficient search approach referred to as eDARTS, which searches on a pre-specified cell having good topology with increased attention to important operations, using a shallow search supernet. Moreover, we propose some optimizations for eDARTS that significantly speed up the search as well as alleviate the well known skip connection aggregation problem of DARTS. eDARTS achieves an error rate of 2.53% on CIFAR-10 using a 3.1M parameters model whereas the search cost is less than 30 minutes. Shahid Siddiqui, Christos Kyrkou, Theocharis Theocharides |
ICPR | 3 |
| 2020 | Extracting the fundamental diagram from aerial footageabstractEfficient traffic monitoring is playing a fundamental role in successfully tackling congestion in transportation networks. Congestion is strongly correlated with two measurable characteristics, the demand and the network density that impact the overall system behavior. At large, this system behavior is characterized through the fundamental diagram of a road segment, a region or the network. In this paper we devise an innovative way to obtain the fundamental diagram through aerial footage obtained from drone platforms. The derived methodology consists of 3 phases: vehicle detection, vehicle tracking and traffic state estimation. We elaborate on the algorithms developed for each of the 3 phases and demonstrate the applicability of the results in a real-world setting. Rafael Makrigiorgis, Panayiotis Kolios, Stelios Timotheou, Theocharis Theocharides, Christoforos Panayiotou |
VTC Spring | 4 |
| 2020 | Jointly-Optimized Searching and Tracking with Random Finite SetsabstractIn this paper, we investigate the problem of joint searching and tracking of multiple mobile targets by a group of mobile agents. The targets appear and disappear at random times inside a surveillance region and their positions are random and unknown. The agents have limited sensing range and receive noisy measurements from the targets. A decision and control problem arises, where the mode of operation (i.e., search or track) as well as the mobility control action for each agent, at each time instance, must be determined so that the collective goal of searching and tracking is achieved. We build our approach upon the theory of random finite sets (RFS) and we use Bayesian multi-object stochastic filtering to simultaneously estimate the time-varying number of targets and their states from a sequence of noisy measurements. We formulate the above problem as a non-linear binary program (NLBP) and show that it can be approximated by a genetic algorithm. Finally, to study the effectiveness and performance of the proposed approach we have conducted extensive simulation experiments. Savvas Papaioannou, Panayiotis Kolios, Theocharis Theocharides, Christoforos Panayiotou, Marios M. Polycarpou |
IEEE Trans. Mob. Comput. | 3 |
| 2019 | Building Robust Machine Learning Systems: Current Progress, Research Challenges, and OpportunitiesabstractMachine learning, in particular deep learning, is being used in almost all the aspects of life to facilitate humans, specifically in mobile and Internet of Things (IoT)-based applications. Due to its state-of-the-art performance, deep learning is also being employed in safety-critical applications, for instance, autonomous vehicles. Reliability and security are two of the key required characteristics for these applications because of the impact they can have on human's life. Towards this, in this paper, we highlight the current progress, challenges and research opportunities in the domain of robust systems for machine learning-based applications. Jeff Zhang 0001, Kang Liu 0017, Faiq Khalid, Muhammad Abdullah Hanif, Semeen Rehman, Theocharis Theocharides, Alessandro Artussi, Muhammad Shafique 0001, Siddharth Garg |
DAC | 6 |
| 2019 | Reliability-Aware Task Allocation Latency Optimization in Edge ComputingabstractNowadays notable computing is shifted away from the cloud and performed onto the Internet of Things (IoT) devices. This necessity emerges due to the growing needs not only for real-time decision support but also for real-time data processing. When used in critical applications such as search and rescue missions or monitoring and control of critical infrastructure, the overall reliable operation of the application running on these devices becomes a major challenge, especially as system reliability is an application - and h/w - dependent measure. Moreover, performance and energy are typically constrained and vary depending on where the computation takes place, as well as, on the communication channels between the devices. Hence, the problem of task allocation under reliability performance-energy constraints becomes even more complex in such cloud/hub/edge computing paradigms. In this work, we use a mathematical programming based framework to derive an optimal task allocation based on multiple operational constraints (latency and energy in both computation and communication), while taking into consideration the reliability demands of the application. We consider an architecture consisting of an edge node, an intermediate node (hub), and the cloud infrastructure, and evaluate our approach using a real-life use-case where the proposed framework minimizes the overall latency of the application while considering the reliability demands of each executed task. Andreas Kouloumpris, Maria K. Michael, Theocharis Theocharides |
IOLTS | 3 |
| 2019 | Towards an Embedded and Real-Time Joint Human-Machine Monitoring Framework: Dataset optimization Techniques for Anomaly DetectionabstractUnmanned Remotely Operated Vehicles (ROVs) are widely used across many civil application domains including real-time monitoring, security and surveillance, and search and rescue missions. Most of these applications require the human operator to control the ROV under stressful conditions and harsh environments. As such, the remote-control operator is prone to sometimes issuing anomalous commands, because of either unwanted hand or finger motion or even irrational decisions, results of fatigue, stress, etc. To enable detection of such anomalies, we propose the use of a joint human-ROV monitoring framework, by monitoring the human operator's bio-signals and the ROV's sensory data. The framework is anticipated to run on the ROV, enabling it to recognize and possibly ignore anomalies, potentially paving the way for a shared control algorithm. In this paper therefore, we present a first step towards achieving this goal, focusing on optimizing the fused dataset consisting of the aforementioned signals and investigated different techniques such as feature extraction and statistical component analysis, in an effort to reduce the dimensionality of the dataset. To this end, we present a dataset constructed by surface Electromyography (sEMG) signals from various operators, fused with the ROV's inertial sensors. Through our proposed optimizations, we are able to reduce both the data size as well as the necessary features and signal components, while maintaining the ability to detect anomalies with at least 85% accuracy depending on the dimensionality reduction technique (over raw data). We evaluated our dataset over a variety of classifier configurations and embedded platforms with noteworthy energy and performance benefits. Rafaella Elia, George Plastiras, Theocharis Theocharides |
VLSI-SoC | 3 |
| 2018 | Edge Intelligence: Challenges and Opportunities of Near-Sensor Machine Learning ApplicationsabstractThe number of connected IoT devices is expected to reach over 20 billion by 2020. These range from basic sensor nodes that log and report the data for cloud processing, to the ones on the edge, that are capable of processing and analyzing the incoming information and taking an action accordingly. Machine learning, and in particular deep learning, is the defacto processing paradigm for intelligently processing these immense volumes of data. However, the resource inhibited environment of edge devices, owing to their limited energy budget, and low compute capabilities, render them a challenging platform for deployment of desired data analytics, particularly in realtime applications. In this paper therefore, we argue that for a wide range of emerging applications edge intelligence is a necessary evolutionary need, and thus we provide a summary of the challenges and opportunities that arise from this need. We showcase through a case study regarding computer vision for commercial drones, how these opportunities can be taken advantage, and how some of the challenges can be potentially addressed. George Plastiras, Maria Terzi, Christos Kyrkou, Theocharis Theocharides |
ASAP | 4 |
| 2018 | An overview of next-generation architectures for machine learning: Roadmap, opportunities and challenges in the IoT eraabstractThe number of connected Internet of Things (IoT) devices are expected to reach over 20 billion by 2020. These range from basic sensor nodes that log and report the data to the ones that are capable of processing the incoming information and taking an action accordingly. Machine learning, and in particular deep learning, is the de facto processing paradigm for intelligently processing these immense volumes of data. However, the resource inhibited environment of IoT devices, owing to their limited energy budget and low compute capabilities, render them a challenging platform for deployment of desired data analytics. This paper provides an overview of the current and emerging trends in designing highly efficient, reliable, secure and scalable machine learning architectures for such devices. The paper highlights the focal challenges and obstacles being faced by the community in achieving its desired goals. The paper further presents a roadmap that can help in addressing the highlighted challenges and thereby designing scalable, high-performance, and energy efficient architectures for performing machine learning on the edge. Muhammad Shafique 0001, Theocharis Theocharides, Christos-Savvas Bouganis, Muhammad Abdullah Hanif, Faiq Khalid, Rehan Hafiz, Semeen Rehman |
DATE | 2 |
| 2018 | DroNet: Efficient convolutional neural network detector for real-time UAV applicationsabstractUnmanned Aerial Vehicles (drones) are emerging as a promising technology for both environmental and infrastructure monitoring, with broad use in a plethora of applications. Many such applications require the use of computer vision algorithms in order to analyse the information captured from an on-board camera. Such applications include detecting vehicles for emergency response and traffic monitoring. This paper therefore, explores the trade-offs involved in the development of a single-shot object detector based on deep convolutional neural networks (CNNs) that can enable UAVs to perform vehicle detection under a resource constrained environment such as in a UAV. The paper presents a holistic approach for designing such systems; the data collection and training stages, the CNN architecture, and the optimizations necessary to efficiently map such a CNN on a lightweight embedded processing platform suitable for deployment on UAVs. Through the analysis we propose a CNN architecture that is capable of detecting vehicles from aerial UAV images and can operate between 5-18 frames-per-second for a variety of platforms with an overall accuracy of ~ 95%. Overall, the proposed architecture is suitable for UAV applications, utilizing low-power embedded processors that can be deployed on commercial UAVs. Christos Kyrkou, George Plastiras, Theocharis Theocharides, Stylianos I. Venieris, Christos-Savvas Bouganis |
DATE | 3 |
| 2018 | Optimizing the Detection Performance of Smart Camera Networks Through a Probabilistic Image-Based ModelabstractNetworks of smart cameras, equipped with on-board processing and communication infrastructure, are increasingly being deployed in a variety of different application fields, such as security and surveillance, traffic monitoring, industrial monitoring, and critical infrastructure protection. The task(s) that a network of smart cameras executes in these applications, e.g., activity monitoring and object identification, can be severely degraded due to errors in the detection module. However, in most cases, higher level tasks and decision making processes in smart camera networks (SCNs) assume ideal detection capabilities for the cameras, which is often not the case due to the probabilistic nature of the detection process, especially for low-cost cameras with limited capabilities. Realizing that it is necessary to introduce robustness in the decision process, this paper presents results toward uncertainty-aware SCNs. Specifically, we introduce a flexible uncertainty model that can be used to characterize the detection behavior in a camera network. We also show how to utilize the model to formulate detection-aware optimization algorithms that can be used to reconfigure the network in order to improve the overall detection efficiency and thus increase the effective number of detected targets. We evaluate our proposed model and algorithms using a network of Raspberry-Pi-based smart cameras that reconfigure in order to improve the detection performance based on the position of targets in the area. The experimental results in the laboratory as well as in a human monitoring application and extensive simulation results indicate that the proposed solutions are able to improve the robustness and reliability of SCNs. Christos Kyrkou, Eftychios G. Christoforou, Stelios Timotheou, Theocharis Theocharides, Christoforos Panayiotou, Marios M. Polycarpou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Intelligent embedded and real-time ANN-based motor control for multi-rotor unmanned aircraft systemsabstractConstant technological advancements in commercial multirotor unmanned aerial vehicles (drones) resulted in their deployment in more and more applications, ranging from entertainment to disaster management and many more domains. However, in contrast to their powerful and diverse entrance into our lifestyle and society, they do not yet provide sufficient intrinsic fail-safe mechanisms to prevent accidents that may occur due to technical problems or unforeseen flight incidents such as turbulent winds, inexperienced pilots, and so on. Therefore, in the current study, we propose the use of an integrated intelligent motor controller, which is trained to recognize incidents directly from the on-board sensors (barometer, gyroscope, compass and accelerometer) and react in real-time, adjusting the drone's motors. The goal is to provide a small, intelligent, low-power, real-time, built-in controller for multirotor UAVs that will be able to understand a dangerous scenario right before it happens, start taking counter measures to keep the drone safe, and provide the pilot with a bigger reaction-time window. We propose the use of an artificial neural network, implemented in a lightweight embedded processing board, that is able to recognize and react in real-time to various turbulent situations. Experimental results suggest that our controller is able to respond properly and timely to wind changes (turbulence) allowing the drone to maintain its expected state and path. George Michael, Nectarios Efstathiou, Kyriacos Mantis, Theocharis Theocharides, Danilo Pau |
VLSI-SoC | 4 |
| 2016 | Emulation-based hierarchical fault-injection framework for coarse-to-fine vulnerability analysis of hardware-accelerated approximate algorithms
Ioannis Chadjiminas, Ioannis Savva, Christos Kyrkou, Maria K. Michael, Theocharis Theocharides |
DATE | 5 |
| 2016 | A Holistic Approach Towards Intelligent Hotspot Prevention in Network-on-Chip-Based MulticoresabstractTraffic hotspots, a severe form of network congestion, can be caused unexpectedly in a network-on-chip (NoC) due to the immanent spatio-temporal unevenness of application traffic. Hotspots reduce the NoC's effective throughput, where in the worst-case scenario, network traffic flows can be frozen indefinitely. To alleviate this problematic phenomenon several adaptive routing algorithms employ online load-balancing schemes, aiming to reduce the possibility of hotspots arising. Since most are not explicitly hotspotagnostic, they cannot completely prevent hotspot formation(s) as their reactive capability to hotspots is merely passive. This paper presents a pro-active Hotspot-Preventive Routing Algorithm (HPRA) which uses the advance knowledge gained from network embedded artificial neural network-based (ANN) hotspot predictors to guide packet routing in mitigating any unforeseen near-future hotspot occurrences. First, these ANN-based predictors are trained offline and during multicore operation they gather online statistical data to predict about-to-be-formed hotspots, promptly informing HPRA to take appropriate hotspot-preventive action(s). Next, in a holistic approach, additional ANN training is performed with data acquired after HPRA interferes, so as to further improve hotspot prediction accuracy; hence, the ANN mechanism does not only predict hotspots, but is also aware of changes that HPRA imposes upon the interconnect infrastructure. Evaluation results, including utilizing real application traffic traces gathered from parallelized workload executions onto a chip multiprocessor architecture, show that HPRA can improve network throughput up to 81 percent when compared with prior-art. Hardware synthesis results affirm the HPRA mechanism's moderate overhead requisites. Vassos Soteriou, Theocharis Theocharides, Elena Kakoulli |
IEEE Trans. Computers | 2 |
| 2016 | A Low-Cost Real-Time Embedded Stereo Vision System for Accurate Disparity Estimation Based on Guided Image FilteringabstractStereo matching, a key element towards extracting depth information from stereo images, is widely used in several embedded consumer electronic and multimedia systems. Such systems demand high processing performance and accurate depth perception, while their deployment in embedded and mobile environments implies that cost, energy and memory overheads need to be minimized. Hardware acceleration has been demonstrated in efficient embedded stereo vision systems. To this end, this paper presents the design and implementation of a hardware-based stereo matching system able to provide high accuracy and concurrently high performance for embedded vision devices, which are associated with limited hardware and power budget. We first implemented a compact and efficient design of the guided image filter, an edge-preserving filter, which reduces the hardware complexity of the implemented stereo algorithm, while at the same time maintains high-quality results. The guided filter design is used in two parts of the stereo matching pipeline, showing that it can simplify the hardware complexity of the Adaptive Support Weight aggregation step, and efficiently enable a powerful disparity refinement unit, which improves matching accuracy, even though cost aggregation is based on simple, fixed support strategies. We implemented several variants of our design on a Kintex-7 FPGA board, which was able to process HD video (1,280 × 720) in real-time (60 fps), using ~57.5k and ~71k of the FPGA's logic (CLB) and register resources, respectively. Additionally, the proposed stereo matching design delivers leading accuracy when compared to state-of-the-art hardware implementations based on the Middlebury evaluation metrics (at least 1.5 percent less bad matching pixels). Christos Ttofis, Christos Kyrkou, Theocharis Theocharides |
IEEE Trans. Computers | 3 |
| 2016 | Embedded Hardware-Efficient Real-Time Classification With Cascade Support Vector MachinesabstractCascade support vector machines (SVMs) are optimized to efficiently handle problems, where the majority of the data belong to one of the two classes, such as image object classification, and hence can provide speedups over monolithic (single) SVM classifiers. However, SVM classification is a computationally demanding task and existing hardware architectures for SVMs only consider monolithic classifiers. This paper proposes the acceleration of cascade SVMs through a hybrid processing hardware architecture optimized for the cascade SVM classification flow, accompanied by a method to reduce the required hardware resources for its implementation, and a method to improve the classification speed utilizing cascade information to further discard data samples. The proposed SVM cascade architecture is implemented on a Spartan-6 field-programmable gate array (FPGA) platform and evaluated for object detection on 800×600 (Super Video Graphics Array) resolution images. The proposed architecture, boosted by a neural network that processes cascade information, achieves a real-time processing rate of 40 frames/s for the benchmark face detection application. Furthermore, the hardware-reduction method results in the utilization of 25% less FPGA custom-logic resources and 20% peak power reduction compared with a baseline implementation. Christos Kyrkou, Christos-Savvas Bouganis, Theocharis Theocharides, Marios M. Polycarpou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Real-Time Obstacle Avoidance for Mobile Robots via Stereoscopic Vision Using Reconfigurable Hardware (Abstract Only)abstractAn embedded, real-time, and low power obstacle avoidance system is a critical component towards fully autonomous robots that can be used in safety missions, space exploration, and transportation systems among others. In this paper a complete prototyping platform for the evaluation of obstacle avoidance systems and autonomous robots is realized on reconfigurable hardware. An efficient stereo vision algorithm for producing the necessary 3D and an obstacle avoidance subsystem were both implemented on an ATLYS Spartan-6 FPGA board equipped with a VmodCam stereo camera module. A modified FDX Vantage 1/10 electric car platform was used for testing the proposed architecture in indoor and outdoor real-world scenes. The system receives stereo image data from the VmodCam module and a decision-making algorithm is applied on a specified Region of Interest (RoI) on the produced disparity map. The algorithm outputs the direction that the robot should move to in order to avoid any obstacles present. Experimental evaluation results indicate that the FPGA-based robotic platform can avoid obstacles in real-time (i.e. can process and identify obstacles within a 1/30th of a second that a stereo image takes to be processed) in both indoor and outdoor environments, with 91.7% accuracy, equivalent to software implementations. The overall power consumption of the proposed architecture, excluding the electronic car platform, is 6 W, making it ideal for use on mobile robots, without becoming a significant drain on its battery life. Martinianos Papadopoulos, Christos Ttofis, Christos Kyrkou, Theocharis Theocharides |
FPGA | 4 |
| 2015 | In-field vulnerability analysis of hardware-accelerated computer vision applicationsabstractIn this paper, we propose an FPGA-based emulation framework that can provide dynamic vulnerability analysis for hardware-accelerated computer vision applications. The framework can be integrated alongside the targeted application, to allow for run-time, in-field, dynamically adjusted vulnerability analysis in real-world conditions, taking into consideration the non-deterministic parameters of the computer vision algorithm computations. We evaluate the proposed framework in real-time using an FPGA platform, for an obstacle avoidance (OA) computer vision application and its disparity estimation kernel to study the impact of Single-Event Upsets (SEUs). Ioannis Chadjiminas, Christos Kyrkou, Theocharis Theocharides, Maria K. Michael, Christos Ttofis |
FPL | 3 |
| 2015 | Cooperative fault-tolerant target tracking in Camera Sensor NetworksabstractCamera Sensor Networks (CSN) are becoming increasingly popular in a variety of security and safety-critical applications including public space surveillance, monitoring of attack-sensitive facilities, and critical infrastructure protection. Cameras in such networks are equipped with high-resolution visual sensors and on-board processors, while featuring wireless communication capabilities. These features enable the execution of various tasks, such as area coverage, activity recognition and target tracking, in a cooperative fashion. However, the performance of CSN may be compromised when faults occur, either due to unintentional software and hardware faults or as the result of a malicious attack. Paving the way for fault tolerance in CSN-based target tracking, we introduce a flexible fault model that can be used to generate different types of erroneous behaviour, thus simulating realistic faults in CSN. We also propose a fault-tolerant decentralized solution for tracking a target that passes through the area monitored by the CSN. Our simulation results indicate that the proposed solution is able to track the target reliably despite the presence of faults. Christos Laoudias, P. Tsangaridis, Marios M. Polycarpou, Christoforos Panayiotou, Christos Kyrkou, Theocharis Theocharides |
ICC | 6 |
| 2015 | A Hardware-Efficient Architecture for Accurate Real-Time Disparity Map EstimationabstractEmerging embedded vision systems utilize disparity estimation as a means to perceive depth information to intelligently interact with their host environment and take appropriate actions. Such systems demand high processing performance and accurate depth perception while requiring low energy consumption, especially when dealing with mobile and embedded applications, such as robotics, navigation, and security. The majority of real-time dedicated hardware implementations of disparity estimation systems have adopted local algorithms relying on simple cost aggregation strategies with fixed and rectangular correlation windows. However, such algorithms generally suffer from significant ambiguity along depth borders and areas with low texture. To this end, this article presents the hardware architecture of a disparity estimation system that enables good performance in both accuracy and speed. The architecture implements an adaptive support weight stereo correspondence algorithm that integrates image segmentation information in an attempt to increase the robustness of the matching process. The article also presents hardware-oriented algorithmic modifications/optimization techniques that make the algorithm hardware-friendly and suitable for efficient dedicated hardware implementation. A comparison to the literature asserts that an FPGA implementation of the proposed architecture is among the fastest implementations in terms of million disparity estimations per second (MDE/s), and with an overall accuracy of 90.21%, it presents an effective processing speed/disparity map accuracy trade-off. Christos Ttofis, Christos Kyrkou, Theocharis Theocharides |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2014 | High-quality real-time hardware stereo matching based on guided image filteringabstractStereo matching is a vital task in several emerging embedded vision applications requiring high-quality depth computation and real-time frame-rate. Although several stereo matching dedicated-hardware systems have been proposed in recent years, only few of them focus on balancing accuracy and speed. This paper proposes a hardware-based stereo matching architecture that aims to provide high accuracy and concurrently high performance in embedded vision applications. The proposed architecture integrates a compact and efficient design of the recently proposed guided image filter; an edge-preserving filter that reduces the hardware complexity of the implemented stereo algorithm, while at the same time maintains high-quality results. A prototype of the architecture has been implemented on a Kintex-7 FPGA board, achieving 60 fps for 720p resolution images. Moreover, the proposed design delivers leading accuracy when compared to state-of-the-art hardware implementations. Christos Ttofis, Theocharis Theocharides |
DATE | 2 |
| 2014 | A high performance hardware architecture for portable, low-power retinal vessel segmentation
Dimitris Koukounis, Christos Ttofis, Agathoklis Papadopoulos, Theocharis Theocharides |
Integr. | 4 |
| 2013 | FPGA-based acceleration of cascaded support vector machines for embedded applications (abstract only)abstractSupport Vector Machines (SVMs) are considered one of the most popular classification algorithms yielding high accuracy rates. However, SVMs often require processing a large number of support vectors, making the classification process computationally demanding, and hence it is challenging to meet real-time processing constraints imposed by many embedded applications. In order to improve SVM classification times the cascade classification scheme has been proposed. However, even in this case real-time performance is still challenging to achieve without exploiting the throughput and processing requirements of each cascade stage. Hence the design of an FPGA-based accelerator for cascaded SVM processing is proposed; in addition to a hardware reduction method in order to reduce the implementation requirements of the cascade SVM leading to significant resource savings. The accelerator was implemented on a Virtex 5 FPGA platform and evaluated using face detection as the target application on 640×480 resolution images. It was compared against FPGA implementations of the same cascade processing architecture but without using the reduction method, and a single parallel SVM classifier. The accelerator is capable an average performance of 70 frames-per-second, achieving a speed-up of 5× over the single parallel SVM classifier. Furthermore, the hardware reduction method results in the utilization of 43% less FPGA LUT resources, with only 0.7% reduction in classification accuracy. Christos Kyrkou, Christos-Savvas Bouganis, Theocharis Theocharides |
FPGA | 3 |
| 2013 | Hardware acceleration of retinal blood vasculature segmentationabstractRetinal vessel tree extraction is a complex and computationally intensive task used in several medical and biometric applications. The emergence of portable biometric authentication applications, as well as on-site biomedical diagnostics, raises the need for hardware-accelerated, power-efficient architectures that can satisfy the performance and accuracy requirements of retinal vessel tree extraction. As such, this paper presents a VLSI implementation of a retina vessel segmentation system, in an attempt to illustrate the advantages and performance benefits that result from a dedicated VLSI solution. The proposed design implements an unsupervised, vessel segmentation algorithm, which utilizes match filtering with signed integers to enhance the difference between the blood vessels and the rest of the retina. The design simplifies the process of obtaining a binary map of the vessel tree by using parallel processing and efficient resource sharing, thus offering real-time performance. FPGA-based simulation results indicate significant performance improvements (up to 90x) when compared to existing hardware and software implementations. Dimitris Koukounis, Christos Ttofis, Theocharis Theocharides |
ACM Great Lakes Symposium on VLSI | 3 |
| 2013 | A hardware-efficient architecture for embedded real-time cascaded support vector machines classificationabstractThis work presents an optimized architecture for cascaded SVM processing, along with a hardware reduction method for the implementation of the additional stages in the cascade, leading to significant improvements. The architecture was implemented on a Virtex 5 FPGA platform and evaluated using face detection as the target application on 640×480 resolution images. Additionally, it was compared against implementations of the same cascade processing architecture but without using the reduction method, and a single parallel SVM classifier. The proposed architecture achieves an average performance of 70 frames-per-second, demonstrating a speed-up of 5× over the single parallel SVM classifier. Furthermore, the hardware reduction method results in the utilization of 43% less hardware resources, with only 0.7% reduction in classification accuracy. Christos Kyrkou, Theocharis Theocharides, Christos-Savvas Bouganis |
ACM Great Lakes Symposium on VLSI | 2 |
| 2013 | FPGA-based hardware acceleration for local complexity analysis of massive genomic data
Agathoklis Papadopoulos, Ioannis Kirmitzoglou, Vasilis J. Promponas, Theocharis Theocharides |
Integr. | 4 |
| 2013 | Edge-Directed Hardware Architecture for Real-Time Disparity Map ComputationabstractStereo Vision, a technique aimed at inferring depth information from stereo images, has been used in a wide range of computer vision applications, with real-time requirements in emerging embedded vision systems. Computation of the disparity map, a vital step in extracting depth information from stereo images, requires a significant amount of computational resources. As such, existing software implementations require high-end hardware platforms to achieve real-time frame rates, suggesting that dedicated hardware mechanisms might be more suitable for embedded applications. In this paper, we present a disparity map computation architecture targeting embedded stereo vision applications with hard real-time requirements. The architecture integrates a hardware edge detection mechanism that reduces the search space, improving the overall performance, and is configurable in terms of various application parameters, making it suitable for a number of application environments. The paper also presents a study on the impact of the various parameters in terms of the performance and hardware/power overheads. An experimental prototype of the architecture was implemented on the Xilinx ML505 FPGA Evaluation Platform, achieving 50 Frames Per Second (fps) for 1,280 × 1,024 image sizes. Moreover, the quality of the disparity maps generated by the proposed system is comparable to other existing hardware implementations featuring local stereo correspondence methods. Christos Ttofis, Stavros Hadjitheophanous, Athinodoros S. Georghiades, Theocharis Theocharides |
IEEE Trans. Computers | 4 |
| 2013 | A hardware architecture for real-time object detection using depth and edge informationabstractEmerging embedded 3D vision systems for robotics and security applications utilize object detection to perform video analysis in order to intelligently interact with their host environment and take appropriate actions. Such systems have high performance and high detection-accuracy demands, while requiring low energy consumption, especially when dealing with embedded mobile systems. However, there is a large image search space involved in object detection, primarily because of the different sizes in which an object may appear, which makes it difficult to meet these demands. Hence, it is possible to meet such constraints by reducing the search space involved in object detection. To this end, this article proposes a depth and edge accelerated search method and a dedicated hardware architecture that implements it to provide an efficient platform for generic real-time object detection. The hardware integration of depth and edge processing mechanisms, with a support vector machine classification core onto an FPGA platform, results in significant speed-ups and improved detection accuracy. The proposed architecture was evaluated using images of various sizes, with results indicating that the proposed architecture is capable of achieving real-time frame rates for a variety of image sizes (271 fps for 320 × 240, 42 fps for 640 × 480, and 23 fps for 800 × 600) compared to existing works, while reducing the false-positive rate by 52%. Christos Kyrkou, Christos Ttofis, Theocharis Theocharides |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2012 | Opportunities from the use of FPGAs as platforms for bioinformatics algorithmsabstractThis paper presents an in-depth look of how FPGA computing can offer substantial speedups in the execution of bioinformatics algorithms, with specific results achieved to date for a broad range of algorithms. Examples and case studies are presented for sequence comparison (BLAST, CAST), multiple sequence alignment (MAFFT, T-Coffee), RNA and protein secondary structure prediction (Zuker, Predator), gene prediction (Glimmer/GlimmerHMM) and phylogenetic tree computation (RAxML), running on mainstream FPGA technologies as well as high-end FPGA-based systems (Convey HC1, BeeCube). This work also presents technological and other obstacles that need to be overcome in order for FPGA computing to become a mainstream technology in Bioinformatics. Grigorios Chrysos 0001, Euripides Sotiriades, Christos Rousopoulos, Apostolos Dollas, Agathoklis Papadopoulos, Ioannis Kirmitzoglou, Vasilis J. Promponas, Theocharis Theocharides, George Petihakis 0001, Jacques Lagnel, Panagiotis Vavylis, George Kotoulas |
BIBE | 8 |
| 2012 | Towards accurate hardware stereo correspondence: A real-time FPGA implementation of a segmentation-based adaptive support weight algorithmabstractDisparity estimation in stereoscopic vision is a vital step for the extraction of depth information from stereo images. This paper presents the hardware implementation of a disparity estimation system that enables good performance in both accuracy and speed. The architecture implements an adaptive support weight stereo correspondence algorithm, which integrates information obtained from image segmentation, in an attempt to increase the robustness of the matching process. The proposed system integrates optimization techniques that make the algorithm hardware-friendly and suitable for embedded vision systems. A prototype of the architecture was implemented on an FPGA, achieving 30 fps for 640×480 image sizes. The quality of the disparity maps generated by the proposed system is also better than other existing hardware implementations featuring fixed support local correspondence methods. Christos Ttofis, Theocharis Theocharides |
DATE | 2 |
| 2012 | Towards systolic hardware acceleration for local complexity analysis of massive genomic dataabstractModern biological research has greatly benefited from genomics. Such research however requires extensive computational power, traditionally employed on large-scale cluster machines as well as multi-core systems. Recent research in reconfigurable architectures however suggests that FPGA-based acceleration of genomic algorithms greatly improves the performance and energy efficiency when compared to multi-core systems and clusters. In this work, we present an initial attempt for massive systolic acceleration of the popular CAST algorithm employed by biologists for complexity analysis of genomic data. CAST is used for detecting (and subsequently masking) low-complexity regions (LCRs) in protein sequences. We designed and implemented a high-performance hardware-accelerated version of CAST for which we built an FPGA prototype, and benchmarked its performance against serial and multithreaded versions of the CAST algorithm in software. The proposed architecture achieves remarkable speedup compared to both serial and multithreaded CAST implementations ranging from approx. 100x-9500x, depending on the dataset features, such as low-complexity content and sequence length distribution. Such performance may enable complex analyses of voluminous sequence datasets, and has the potential to interoperate with other hardware architectures for protein sequence analysis. Agathoklis Papadopoulos, Vasilis J. Promponas, Theocharis Theocharides |
ACM Great Lakes Symposium on VLSI | 3 |
| 2012 | HPRA: A pro-active Hotspot-Preventive high-performance routing algorithm for Networks-on-ChipsabstractThe inherent spatio-temporal unevenness of traffic flows in Networks-on-Chips (NoCs) can cause unforeseen, and in cases, severe forms of congestion, known as hotspots. Hotspots reduce the NoC's effective throughput, where in the worst case scenario, the entire network can be brought to an unrecoverable halt as a hotspot(s) spreads across the topology. To alleviate this problematic phenomenon several adaptive routing algorithms employ online load-balancing functions, aiming to reduce the possibility of hotspots arising. Most, however, work passively, merely distributing traffic as evenly as possible among alternative network paths, and they cannot guarantee the absence of network congestion as their reactive capability in reducing hotspot formation(s) is limited. In this paper we present a new pro-active Hotspot-Preventive Routing Algorithm (HPRA) which uses the advance knowledge gained from network-embedded Artificial Neural Network-based (ANN) hotspot predictors to guide packet routing across the network in an effort to mitigate any unforeseen near-future occurrences of hotspots. These ANNs are trained offline and during multicore operation they gather online buffer utilization data to predict about-to-be-formed hotspots, promptly informing the HPRA routing algorithm to take appropriate action in preventing hotspot formation(s). Evaluation results across two synthetic traffic patterns, and traffic benchmarks gathered from a chip multiprocessor architecture, show that HPRA can reduce network latency and improve network throughput up to 81% when compared against several existing state-of-the-art congestion-aware routing functions. Hardware synthesis results demonstrate the efficacy of the HPRA mechanism. Elena Kakoulli, Vassos Soteriou, Theocharis Theocharides |
ICCD | 3 |
| 2012 | A Parallel Hardware Architecture for Real-Time Object Detection with Support Vector MachinesabstractObject detection applications are often associated with real-time performance constraints that stem from the embedded environment that they are often deployed in. Consequently, researchers have proposed dedicated hardware architectures, utilizing a variety of classification algorithms targeting object detection. Support Vector Machines (SVMs) is among the most popular classification algorithms used in object detection yielding high accuracy rates. However, existing SVM hardware implementations attempting to speed up SVM classification, have either targeted only simple applications, or SVM training. As such, there are limited proposed hardware architectures that are generic enough to be used in a variety of object detection applications. Hence, this paper presents a parallel array architecture for SVM-based object detection, in an attempt to show the advantages, and performance benefits that stem from a dedicated hardware solution. The proposed hardware architecture provides parallel processing, resource sharing among the processing units, and efficient memory management. Furthermore, the size of the array is scalable to the hardware demands, and can also handle a variety of applications such as multiclass classification problems. A prototype of the proposed architecture was implemented on an FPGA platform and evaluated using three popular detection applications, demonstrating real-time performance (40-122 fps for a variety of applications). Christos Kyrkou, Theocharis Theocharides |
IEEE Trans. Computers | 2 |
| 2012 | Intelligent Hotspot Prediction for Network-on-Chip-Based Multicore SystemsabstractHotspots are network-on-chip (NoC) routers or modules in multicore systems which occasionally receive packetized data from other networked element producers at a rate higher than they can consume it. This adverse phenomenon may greatly reduce the performance of NoCs, especially when wormhole flow-control is employed, as backpressure can cause the buffers of neighboring routers to quickly fill-up leading to a spatial spread in congestion. This can cause the network to saturate prematurely where in the worst scenario the NoC may be rendered unrecoverable. Thus, a hotspot prevention mechanism can be greatly beneficial, as it can potentially enable the interconnection system to adjust its behavior and prevent the rise of potential hotspots, subsequently sustaining NoC performance. The inherent unevenness of traffic patterns in an NoC-based general-purpose multicore system such as a chip multiprocessor, due to the diverse and unpredictable access patterns of applications, produces unexpected hotspots whose appearance cannot be known a priori, as application demands are not predetermined, making hotspot prediction and subsequently prevention difficult. In this paper, we present an artificial neural network-based (ANN) hotspot prediction mechanism that can be potentially used in tandem with a hotspot avoidance or congestion-control mechanism to handle unforeseen hotspot formations efficiently. The ANN uses online statistical data to dynamically monitor the interconnect fabric, and reactively predicts the location of an about to-be-formed hotspot(s), allowing enough time for the multicore system to react to these potential hotspots. Evaluation results indicate that a relatively lightweight ANN-based predictor can forecast hotspot formation(s) with an accuracy ranging from 65% to 92%. Elena Kakoulli, Vassos Soteriou, Theocharis Theocharides |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Depth-directed hardware object detectionabstractObject detection is a vital task in several emerging applications, requiring real-time detection frame-rate and low energy consumption for use in embedded and mobile devices. This paper proposes a hardware-based, depth-directed search method for reducing the search space involved in object detection, resulting in significant speed-ups and energy savings. The proposed architecture utilizes the disparity values computed from a stereoscopic camera setup, in an attempt to direct the detection classifier to regions that contain objects of interest. By eliminating large amounts of search data, the proposed system achieves both performance gains and reduced energy consumption. FPGA simulation results indicate performance speedups up to 4.7 times and high energy savings ranging from 41-48%, when compared to the traditional sliding window approach. Christos Kyrkou, Christos Ttofis, Theocharis Theocharides |
DATE | 3 |
| 2011 | FPGA-Accelerated Object Detection Using Edge InformationabstractObject detection is a vital task in several existing as well as emerging applications, requiring real-time processing and low energy consumption, and often with limited available hardware budget in the case of embedded and mobile devices. This paper proposes an FPGA-based object detection system that utilizes edge information to reduce the search space involved in object detection. By eliminating large amounts of search data, the proposed system achieves both performance gains, and reduced energy consumption, while requiring minimal additional hardware, making it suitable for resource-constrained FPGAs. Implementation results on an FPGA indicate performance speedups up to 4.9 times, and high energy savings ranging from 73-78%, when compared to the traditional sliding window approach for FPGA implementations. Christos Kyrkou, Christos Ttofis, Theocharis Theocharides |
FPL | 3 |
| 2011 | Towards optimal CMOS lifetime via unified reliability modeling and multi-objective optimizationabstractReliability of CMOS devices emerges as a vital design constraint, evidenced by several CMOS failure mechanisms. Such mechanisms have traditionally been modeled independently, using statistical approximation techniques to estimate Mean-Time-to-Failure (MTTF) rates. This paper proposes a unified framework that integrates the existing failure models into a multi-objective optimization engine, in an attempt to provide a pareto-optimal solution indicating the suggested operating conditions of a system for a given technology and size (in transistors), in an effort to maximize its lifetime reliability. In addition to the existing failure mechanisms, the framework also considers a proposed system-level leakage power estimation model, as leakage is interdependent on temperature, and as such impacts system reliability. The framework can be used in several design scenarios, such as thermal-aware task scheduling. Agathoklis Papadopoulos, Theocharis Theocharides, Maria K. Michael |
ISCAS | 2 |
| 2011 | A Flexible Parallel Hardware Architecture for AdaBoost-Based Real-Time Object DetectionabstractReal-time object detection is becoming necessary for a wide number of applications related to computer vision and image processing, security, bioinformatics, and several other areas. Existing software implementations of object detection algorithms are constrained in small-sized images and rely on favorable conditions in the image frame to achieve real-time detection frame rates. Efforts to design hardware architectures have yielded encouraging results, yet are mostly directed towards a single application, targeting specific operating environments. Consequently, there is a need for hardware architectures capable of detecting several objects in large image frames, and which can be used under several object detection scenarios. In this work, we present a generic, flexible parallel architecture, which is suitable for all ranges of object detection applications and image sizes. The architecture implements the AdaBoost-based detection algorithm, which is considered one of the most efficient object detection algorithms. Through both field-programmable gate array emulation and large-scale implementation, and register transfer level synthesis and simulation, we illustrate that the architecture can detect objects in large images (up to 1024 × 768 pixels) with frame rates that can vary between 64-139 fps for various applications and input image frame sizes. Christos Kyrkou, Theocharis Theocharides |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Towards hardware stereoscopic 3D reconstruction a real-time FPGA computation of the disparity mapabstractStereoscopic 3D reconstruction is an important algorithm in the field of Computer Vision, with a variety of applications in embedded and real-time systems. Existing software-based implementations cannot satisfy the performance requirements for such constrained systems; hence an embedded hardware mechanism might be more suitable. In this paper, we present an architecture of a 3D reconstruction system for stereoscopic images, which we implement on Virtex2 Pro FPGA. The architecture uses a Sobel edge detector to achieve real-time (75 fps) performance, and is configurable in terms of various application parameters, making it suitable for a number of application environments. The paper also presents a design exploration on algorithmic parameters such as disparity range, correlation window size, and input image size, illustrating the impact on the performance for each parameter. Stavros Hadjitheophanous, Christos Ttofis, Athinodoros S. Georghiades, Theocharis Theocharides |
DATE | 4 |
| 2010 | A reconfigurable MPSoC-based QAM modulation architectureabstractQAM is a widely used multi-level modulation technique, with a variety of applications in data radio communication systems. Most existing implementations of QAM-based systems use high levels of modulation in order to meet the high data rate constraint of emerging applications. This work presents the architecture of a highly-parallel MPSoC-based QAM modulator that offers multi-rate modulation. The proposed MPSoC architecture is modular and provides flexibility via dynamic reconfiguration of the QAM, offering high data rates (more than 1 Gbps), even at low modulation levels (16-QAM). Furthermore, the proposed QAM implementation integrates a hardware-based resource allocation algorithm for dynamic load balancing. Christos Ttofis, Agathoklis Papadopoulos, Theocharis Theocharides, Maria K. Michael, Demosthenes Doumenis |
VLSI-SoC | 3 |
| 2009 | Towards embedded runtime system level optimization for MPSoCs: on-chip task allocationabstractNext generation multiprocessor systems-on-chip (MPSoCs) are expected to contain numerous processing elements, interconnected via on-chip networks, executing real-time applications. It is anticipated that runtime optimization algorithms which dynamically adjust system parameters with the purpose of optimizing the system's operation, will be embedded in the system software and/or hardware. In this paper, we present a methodology for simulating and evaluating system-level optimization algorithms, demonstrated by the case of on-chip dynamic task allocation applied to generic MPSoC architectures. Through this methodology, we are able to show that dynamic, system-level bidding-based task allocation can improve system performance, when compared to a round robin allocation, in popular MPSoC applications. Theocharis Theocharides, Maria K. Michael, Marios M. Polycarpou, Ajit Dingankar |
ACM Great Lakes Symposium on VLSI | 1 |
| 2005 | A low latency router supporting adaptivity for on-chip interconnectsabstractThe increased deployment of System-on-Chip designs has drawn attention to the limitations of on-chip interconnects. As a potential solution to these limitations, Networks-on -Chip (NoC) have been proposed. The NoC routing algorithm significantly influences the performance and energy consumption of the chip. We propose a router architecture which utilizes adaptive routing while maintaining low latency. The two-stage pipelined architecture uses look ahead routing, speculative allocation, and optimal output path selection concurrently. The routing algorithm benefits fromcongestionaware flow control, making better routing decisions. We simulate and evaluate the proposed architecture in terms of network latency and energy consumption. Our results indicate that the architecture is effective in balancing the performance and energy of NoC designs. Jongman Kim, Dongkook Park, Theocharis Theocharides, Narayanan Vijaykrishnan, Chita R. Das |
DAC | 3 |
| 2004 | Thermal-Aware IP Virtualization and Placement for Networks-on-Chip ArchitectureabstractNetworks-on-chip (NoC), a new SoC paradigm, has been proposed as a solution to mitigate complex on-chip interconnect problems. NoC architecture consists of a collection of IP cores or processing elements (PEs) interconnected by on-chip switching fabrics or routers. Hardware virtualization, which maps logic processing units onto PEs, affects the power consumption of each PE and the communications among PEs. The communication among PEs affects the overall performance and router power consumption, and it depends on the placement of PEs. Therefore, the temperature distribution profile of the chip depends on the IP core virtualization and placement. In this paper, we present an IP virtualization and placement algorithm for generic regular network on chip (NoC) architecture. The algorithm attempts to achieve a thermal balanced design while minimizing the communication cost via placement. Our framework can also realize hardware virtualization which can further accomplish better performance. A case study on low density parity checks (LDPC) decoder is presented to evaluate our algorithm. Wei-Lun Hung, Charles Addo-Quaye, Theocharis Theocharides, Yuan Xie 0001, Narayanan Vijaykrishnan, Mary Jane Irwin |
ICCD | 3 |