Panagiotis Mousouliotis

dblp:167/3542 · also Panagiotis G. Mousouliotis · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
4since 2021 · last 2023
0000-0001-9621-924XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2023 A Novel Integrated Simulation Framework for Cyber-Physical Systems Modelling
abstract
The growing use of Cyber-Physical Systems (CPS) in a plethora of domains (e.g. healthcare, industry, smart homes, transportation, etc.) triggers an urgent demand for simulation frameworks that can simulate in an integrated manner all the components (i.e. CPUs, Memories, Networks, Physical Environment) of a system-under-design(SuD). By utilizing such a simulator, software design can proceed in parallel with physical development which results in the reduction of the so important time-to-market. The main problem, however, is that currently there is a shortage of such simulation frameworks; most simulators used for modelling the digital aspects of CPS applications (i.e. full-system CPU/Mem/Peripheral simulators) lack any support of the CPS physical aspects and vice versa. The presented fully-distributed simulation framework (APOLLON) is the first known open-source, high-performance simulator that can handle holistically complex CPSs including processors, peripherals, networks and physical aspects of them. APOLLON is an extension of the COSSIM simulation framework and it integrates, in a novel and efficient way, a combined processing and network simulator with the widely-used Ptolemy physical simulator, in a transparent way. Our highly integrated approach is further augmented with Machine Learning capabilities by implementing Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) recurrent neural networks in both the Cyber and Physical domains, enabling users to develop their complex recurrent neural networks significantly fast and accurately. APOLLON has been evaluated when executing a number of benchmarks and real-world use cases; the end results demonstrate that the presented approach has up to 99% accuracy in the reported SuD aspects.
Nikolaos Tampouratzis, Panagiotis Mousouliotis, Ioannis Papaefstathiou
IEEE Trans. Parallel Distributed Syst.2
2021 High Speed Implementation of the Deformable Shape Tracking Face Alignment Algorithm
abstract
The 2D facial landmark alignment method, implemented in C++ in the open source libraries DLIB and Deformable Shape Tracking (DEST), is used in several applications such as driver drowsiness detection. The most challenging of these applications require fast video frame processing. Therefore, the alignment of the facial landmarks in a single video frame has to be performed with the minimum possible latency without precision loss. In this paper, the DEST implementation of the face alignment method that is based on regression trees is heavily restructured to reduce latency. The resulting face alignment predictor is implemented in C. The elimination of multiple nested routine calls, excessive argument copying, type conversions and integrity checks lead to a software implementation that is 240 times faster than the one provided in the DEST library. Moreover, the structure of the new face alignment predictor is appropriate for hardware implementation on a Field Programmable Gate Array (FPGA) for further acceleration1.
Nikos Petrellis, Stavros Zogas, Panagiotis Christakos, Georgios Keramidas, Panagiotis Mousouliotis, Nikos S. Voros, Christos P. Antonopoulos
DSD5
2021 An Open-source Implementation of LSTM and GRU in the Ptolemy Simulation Framework
abstract
Ptolemy II [1] is an open-source software framework for modelling, simulation and design of concurrent, heterogeneous, real-time systems, including distributed/parallel systems [2]. These systems can be described using the combination of different formal as well as computation models. Ptolemy also includes machine learning libraries providing support for particle filtering, model-predictive control, hidden Markov models, and various statistical analysis tools. However, one of the main problems Ptolemy users face is the lack of fundamental recurrent neural network structures. In this paper, an LSTM and a GRU recurrent neural network are implemented in Ptolemy II framework, in order to extend its machine learning library, enabling users to develop their complex recurrent neural networks in significantly less time. The presented work has been verified through a real-world weather forecasting use case; the results demonstrate that our approach has identical accuracy with one of the most widely used machine learning library (i.e. Keras) in all cases. To further increase the impact of our approach, the complete source code is freely distributed to the community.
Vasilis Daoulas, Nikolaos Tampouratzis, Panagiotis Mousouliotis, Ioannis Papaefstathiou
DS-RT3
2021 Challenges Towards Hardware Acceleration of the Deformable Shape Tracking Application
abstract
In the context of this paper, a shape tracking application based on landmark alignment is transformed to support implementation in Field Programmable Gate Arrays (FPGAs). Towards this direction, several challenges are posed since a) computational intensive operations have to be replaced by faster ones, b) specific loops have to be modified (e.g., unrolled) to support the implementation of operations in parallel with different hardware resources, c) multiple pretrained models have to be compared in terms of speed and accuracy, d) partial loading of the pre-trained models has to be examined in order to fit their parameters in the Block Random Access Memories (BRAMs) of the FPGA for faster access, and e) alternative arithmetic representations have to be evaluated for higher speed and reduced resources.The C++ Deformable Shape Tracking (DEST) implementation of face alignment that is based on an Ensemble of Regression Trees is employed in our approach. The DEST application uses Eigen library routines to implement algebraic operations which are proved to be quite slow. The achievements of this paper, concern the replacement of appropriate Eigen calls in time critical paths with fast C code that can be directly used to synthesize reconfigurable hardware implementations. The elimination of the computational intensive Eigen calls has already improved the speed of the face alignment application by more than 240 times. In this paper we examine how the modified source code structure of the DEST application can be used to address the challenges described above.
Nikos Petrellis, Panagiotis Christakos, Stavros Zogas, Panagiotis Mousouliotis, Georgios Keramidas, Nikos S. Voros, Christos P. Antonopoulos
VLSI-SoC4
2020 SqueezeJet-3: An Accelerator Utilizing FPGA MPSoCs for Edge CNN Applications
abstract
Most FPGA-based Convolutional Neural Network (CNN) hardware accelerators target the datacenter rather than edge processing units. To further fill this gap, this work presents SqueezeJet-3 - a novel FPGA-based embedded system, consisting of software and hardware, for accelerating edge CNN inference. Even though SqueezeJet-3 is optimized for accelerating small ImageNet class CNNs, such as SqueezeNet v1.1, on low-end lowcost FPGA SoC devices, it can also be used for the acceleration of larger CNNs, such as the VGG16. Evaluation of our accelerator reveals better or comparable performance with that triggered by the current state-of-the-art similar tools and systems.
Panagiotis Mousouliotis, Ioannis Papaefstathiou, Loukas Petrou
FCCM1
2015 Augmented reality for maintenance application on a mobile platform
abstract
Pose estimation is a major requirement for any augmented reality (AR) application. Cameras and inertial measurement units (IMUs) have been used for pose estimation not only in AR but also in many other fields. The level of accuracy and pose update required in an AR application is more demanding than in any other field. In certain AR applications, (maintenance for example) a small change in pose can cause a huge deviation in the rendering of the virtual content. This misleads the user in terms of an object location and can display incorrect information. Further, the huge amount of processing power required for the camera based pose estimation results in a bulky system. This reduces the mobility and ergonomics of the system. This demonstration shows a fast pose estimation using a camera and an IMU on a mobile platform for augmented reality in a maintenance application.
N. S. Lakshmprabha, Stathis Kasderidis, Panagiotis Mousouliotis, Loukas Petrou, Olga Beltramello
VR3