Sophia Karagiorgou

dblp:83/7569 · also Sofia Karagiorgou · DBLP profile ↗
← Back
11ranked-venue papers in the field
2as first author
6since 2021 · last 2024
0000-0002-1099-8463ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 7Database Systems & Data Management · 3 (2 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2024 Adversarial Explanations for Informed Civilian and Environmental Protection
abstract
Combating crime and conditions of high physical risk in cities, the environment, and critical infrastructures requires a multifaceted approach. For sensitive problems, such as advanced situational awareness in the fields of civilian applications and environmental protection, Artificial Intelligence (AI) and Neural Network (NN) adoption has been slow due to concerns about their reliability, leading to several algorithms for explaining their decisions. Despite the possibilities for AI in critical infrastructure protection and civilian applications, many challenges still exist. For instance: (i) there are complex and high risks meaning that AI systems need to be transparent and interpretable to gain decision-maker trust; (ii) AI models may be vulnerable to imperceptible manipulations of input data even without any knowledge about the AI technique that is used; (iii) the need to efficiently process distributed, multimodal and big data coming from different, but however cheap, Internet of Things (IoT) and sensory devices (e.g., drones, cameras, accelerometers, telemetry, geomagnetic field, and proximity sensors); and (iv) many AI methods based on Machine Learning (ML) require huge amounts of training data, resulting in a Big Data computation problem. We introduce, benchmark, and demonstrate an adversarial explanations approach that we can efficiently tackle both adversarial robustness and explanation complexity of AI systems. To achieve this, we train robustified NNs and transparent explainers on big imagery data and leverage the attacks’ knowledge as explanations to gain greater fidelity to the AI model. The merit of the proposed approach is that the new and robustified model has a great performance against new, unseen types of perturbations and attacks. This way, we pave the adoption of more informed and responsible AI integration in sensitive application domains.
Theodora Anastasiou, Ioannis Pastellas, Sophia Karagiorgou
IEEE Big Data3
2024 CEASEFIRE: An AI-Powered System for Combating Illicit Firearms Trafficking
abstract
Modern technologies have enabled illicit firearms trafficking to partially merge with cybercrime, while also allowing its off-line aspects to become increasingly complex. The online trade of firearms, their components, 3D blueprints and illicit substances carried out by criminals on both the surface Web and dark Web is increasingly difficult to address as a consequence of the exponential growth in the amount of information disseminated on the Internet. On the other hand, law enforcement agencies are confronted with significant challenges that require the development of sophisticated technological solutions capable of processing large volumes of data, identifying relevant information in a timely manner and creating networks of connections between potential criminal groups. This article presents a real-world practical system, namely the CEASEFIRE one, powered by advanced artificial intelligence technologies that can assist law enforcement personnel in addressing the above described challenges.
Jorgen Cani, Ioannis Mademlis, Marina Mancuso, Caterina Paternoster, Emmanouil Adamakis, George Margetis, Sylvie Chambon, Alain Crouzil, Loubna Lechelek, Georgia Dede, Spyridon Evangelatos, George Lalas, Franck Mignet, Pantelis Linardatos, Konstantinos Kentrotis, Henryk Gierszal, Piotr Tyczka, Sophia Karagiorgou, George Pantelis, Georgios Stavropoulos, Konstantinos Votis, Georgios Th. Papadopoulos
IEEE Big Data18
2024 On Energy-aware and Verifiable Benchmarking of Big Data Processing targeting AI Pipelines
abstract
As Artificial Intelligence (AI) is revolutionizing various industries and applications, understanding the hardware requirements and energy consumption of AI pipelines in Big Data (BD) applications has become increasingly essential. This paper presents a comprehensive, scalable framework, designed to systematically measure hardware resources, energy usage, and model performance across two prominent data modalities: tabular data and images. The framework is generalizable, facilitating replicability across the AI research community, and encourages the deployment of AI models with comprehensive metrics beyond traditional accuracy, promoting the optimization of pipelines for real-world scenarios. Through detailed benchmarking, we identify EfficientNet as a standout model for image classification, and XGBoost for tabular data, both excelling in their respective domains. Notably, our findings show that Graphics Processing Units (GPUs) account for approximately 90% of total energy consumption in image-based tasks, while Central Processing Units (CPUs) are responsible for around 50% of energy use in tabular data processing. The merit of our innovative proposed framework combines information theory and probability theory to enhance our understanding of AI model performance in Edge-to-Cloud (E2C) applications that demand efficient Big Data processing in distributed environments. By seamlessly integrating energy efficiency with hardware optimization, it enables realtime monitoring of energy consumption and computing resources in containerized environments, providing precise insights for optimizing AI workloads. This framework facilitates scalable AI deployment on resource-constrained edge devices, reducing energy consumption while enhancing AI model robustness and interpretability, thereby promoting greater trust and transparency in AI-powered decision-making for critical real-world applications. This emphasizes the importance of multi-objective optimization for more sustainable and efficient Big Data AI workflows.
Georgios Theodorou, Sophia Karagiorgou, Christos Kotronis
IEEE Big Data2
2024 Benchmarking of Different YOLO Models for CAPTCHAs Detection and Classification
abstract
This paper provides an analysis and comparison of the YOLOv5, YOLOv8 and YOLOv10 models for webpage CAPTCHAs detection using the datasets collected from the web and darknet as well as synthetized data of webpages. The study examines the nano (n), small (s), and medium (m) variants of YOLO architectures and use metrics such as Precision, Recall, F1 score, mAP@50 and inference speed to determine the real-life utility. Additionally, the possibility of tuning the trained model to detect new CAPTCHA patterns efficiently was examined as it is a crucial part of real-life applications. The image slicing method was proposed as a way to improve the metrics of detection on oversized input images which can be a common scenario in webpages analysis. Models in version nano achieved the best results in terms of speed, while more complexed architectures scored better in terms of other metrics.
Mikolaj Wysocki, Henryk Gierszal, Piotr Tyczka, George Pantelis, Sophia Karagiorgou
IEEE Big Data5
2023 Harvesting Large Textual and Multimedia Data to Detect Illegal Activities on Dark Web Marketplaces
abstract
During the last decades, the dark web has become the ground for criminal activities, enabling for illegal content sharing, as well as marketplaces selling drugs and firearms. The 2023 Internet Organised Crime Threat Assessment (IOCTA) of Europol’s European Cybercrime Centre (EC3) presented the dark web as the one of the top crime ecosystems. Analyzing the .onion sites hosting marketplaces is of interest to law enforcement, security researchers, and big data analysts. The capability of automatically harvesting web content from web servers enables Law Enforcement Agencies (LEAs) to collect and preserve data prone to serve as potential clues or evidence in investigations. The use of sophisticated protocols and the inherent complexity of the Dark Web makes it difficult for security agencies to identify and investigate these activities through conventional methods. The sheer size, unpredictable ecosystem, and anonymity provided by the Dark Web are the essential confrontations to trace the criminals. Therefore, it is a crucial step to discover the potential solutions towards cyber-crimes evaluating the sailing Dark Web crime threats. In this paper, we devise Artificial Intelligence and Big Data processing to extract insights and investigate how the Dark Web facilitates crime and dynamically maintains marketplaces with illegal goods exchange. The scientific contribution of this paper entails novel textual and multimedia analytics over large collections of data to collect evidence for further investigations and support visual reporting and alerting mechanisms. The conclusions include practical implications of Dark Web content retrieval and archival, such as investigation clues and evidence, and related future research topics.
Georgia Chatzimarkaki, Sophia Karagiorgou, Mariza Konidi, Dimitrios Alexandrou, Thanassis D. Bouras, Spyridon Evangelatos
IEEE Big Data2
2023 MobiSpaces: An Architecture for Energy-Efficient Data Spaces for Mobility Data
abstract
In this paper, we present an architecture for mobility data spaces enabling trustworthy and reliable data operations along with its main constituent parts. The architecture makes use of a data lake for scalable storage of diverse mobility data sets, on top of which separate computing and storage layers are implemented to allow independent scaling with a data operations toolbox providing all data operations. Furthermore, to cater for mobility analytics, machine learning and artificial intelligence support, an edge analytics suite is provided that encompasses distributed algorithms for mobility analytics and federated learning, thereby exploiting edge computing technologies. In turn, this is supported by a resource allocator that monitors the energy consumption of data-intensive operations and provides this information to the platform for intelligent task placement in edge devices, aiming at energy-efficient operations. As a result, an end-to-end platform is proposed that combines data services and infrastructure services towards supporting mobility application domains, such as urban and maritime.
Christos Doulkeridis, Georgios M. Santipantakis, Nikolaos Koutroumanis, George Makridis, Vasilis Koukos, George S. Theodoropoulos, Yannis Theodoridis, Dimosthenis Kyriazis, Pavlos Kranas, Diego Burgos, Ricardo Jiménez-Peris, Mariana M. G. Duarte, Mahmoud Attia Sakr, Esteban Zimányi, Anita Graser, Clemens Heistracher, Kristian Torp, Ioannis Chrysakis, Theofanis Orphanoudakis, Evgenia Kapassa, Marios Touloupou, Jürgen Neises, Petros Petrou, Sophia Karagiorgou, Rosario Catelli, Domenico Messina, Marcelo Corrales Compagnucci, Matteo Falsetta
IEEE Big Data24
2015 A MapReduce based k-NN joins probabilistic classifier
abstract
Water management field has concentrated great interest, with the potential to affect the long term well-being, the societal economy and security. In parallel, it imposes specific research challenges which have not been already met, due to the lack of fine-grained data. Knowledge extraction and decision making for efficient management in the energy field has attracted a lot of interest in Big Data research. However, the water domain is strikingly absent, with minimal focused work on data exploitation and useful information extraction. The goal of this work is to discover persistent and meaningful knowledge from water consumption data and provide efficient and scalable big data management and analysis services. We propose a novel methodology which exploits machine learning techniques and introduces a robust probabilistic classifier which is able to operate on data of arbitrary dimensionality and of huge volume. It also provides added value services and new operation models for the water management domain, inducing sustainable behavioural changes for consumers, which can further raise social awareness. It does so through a new k-Nearest Neighbour based algorithm, developed in a parallel and distributed environment, which operates over Big Data and discovers useful knowledge about consumption classes and other water related attitudinal properties. A detailed experimental evaluation assesses the effectiveness and efficiency of the algorithm on prediction precision along with the provision of analytics. The results show that this method is prosperous and provides accurate and interesting results that allow us to identify useful characteristics, not only for the households, but also for the water utilities.
Georgios Chatzigeorgakidis, Sophia Karagiorgou, Spiros Athanasiou, Spiros Skiadopoulos
IEEE BigData2
2015 A comparison and evaluation of map construction algorithms using vehicle tracking data
Mahmuda Ahmed, Sophia Karagiorgou, Dieter Pfoser, Carola Wenk
GeoInformatica2
2015 Crowdsourcing urban form and function
abstract
Urban form and function have been studied extensively in urban planning and geographical information science. However, gaining a greater understanding of how they merge to define the urban morphology remains a substantial scientific challenge. Toward this goal, this paper addresses the opportunities presented by the emergence of crowdsourced data to gain novel insights into form and function in urban spaces. We are focusing in particular on information harvested from social media and other open-source and volunteered datasets (e.g. trajectory and OpenStreetMap data). These data provide a first-hand account of form and function from the people who define urban space through their activities. This novel bottom-up approach to study these concepts complements traditional urban studies to provide a new lens for studying urban activity. By synthesizing recent advancements in the analysis of open-source data, we provide a new typology for characterizing the role of crowdsourcing in the study of urban morphology. We illustrate this new perspective by showing how social media, trajectory, and traffic data can be analyzed to capture the evolving nature of a city’s form and function. While these crowd contributions may be explicit or implicit in nature, they are giving rise to an emerging research agenda for monitoring, analyzing, and modeling form and function for urban design and analysis.
Andrew T. Crooks, Dieter Pfoser, Andrew Jenkins, Arie Croitoru, Anthony Stefanidis, Duncan Smith, Sophia Karagiorgou, Alexandros Efentakis, George Lamprianidis
Int. J. Geogr. Inf. Sci.7
2013 Segmentation-based road network construction
abstract
This work proposes a novel method that converts movement trajectories into a hierarchical transportation network. It utilizes an improved map construction algorithm on segmented input data based on types of movement. The produced hierarchical road network layers are then combined into a single network. This segmentation addresses the challenges imposed by noisy, low sampling rate trajectories and provides for a mechanism to accommodate automatic map maintenance on updates. An experimental evaluation is conducted using trajectories derived from GPS tracking taxi fleets and utility vehicles in Berlin, Vienna and Athens.
Sophia Karagiorgou, Dieter Pfoser, Dimitrios Skoutas 0001
SIGSPATIAL/GIS1
2012 On vehicle tracking data-based road network generation
abstract
Road networks are important datasets for an increasing number of applications. However, the creation and maintenance of such datasets pose interesting research challenges. This work proposes an automatic road network generation algorithm that takes vehicle tracking data in the form of trajectories as input and produces a road network graph. This effort addresses the challenges of evolving map data sets, specifically by focusing on (i) automatic map-attribute generation (weights), (ii) automatic road network generation, and (iii) by providing a quality assessment. An experimental study assesses the quality of the algorithms by generating a part of the road network of Athens, Greece, using trajectories derived from GPS tracking a school bus fleet.
Sophia Karagiorgou, Dieter Pfoser
SIGSPATIAL/GIS1