Panagiotis Liakos

dblp:120/9595 · DBLP profile ↗
← Back
19ranked-venue papers in the field
15as first author
7since 2021 · last 2026
0000-0003-4569-7801ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 9 (6 first)Information Retrieval & Web Search · 5 (5 first)Big Data, Cloud & Distributed Data Systems · 4 (3 first)Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
YearPublicationVenuePosition
2026 PLA-Mux: Multiplexing Piece-Wise Linear Approximations on Edge-Assisted Sensor Networks
Xenophon Kitsios, Panagiotis Liakos, Katia Papakonstantinopoulou, Yannis Kotidis
MDM2
2024 How to Make your Duck Fly: Advanced Floating Point Compression to the Rescue
Panagiotis Liakos, Katia Papakonstantinopoulou, Thijs Bruineman, Mark Raasveldt, Yannis Kotidis
EDBT1
2024 Flexible grouping of linear segments for highly accurate lossy compression of time series data
Xenophon Kitsios, Panagiotis Liakos, Katia Papakonstantinopoulou, Yannis Kotidis
VLDB J.2
2023 Sim-Piece: Highly Accurate Piecewise Linear Approximation through Similar Segment Merging
abstract
Approximating series of timestamped data points using a sequence of line segments with a maximum error guarantee is a fundamental data compression problem, termed as piecewise linear approximation (PLA). Due to the increasing need to analyze massive collections of time-series data in diverse domains, the problem has recently received significant attention, and recent PLA algorithms that have emerged do help us handle the overwhelming amount of information, at the cost of some precision loss. More specifically, these algorithms entail a trade-off between the maximum precision loss and the space savings achieved. However, advances in the area of lossless compression are undercutting the offerings of PLA techniques in real datasets. In this work, we propose Sim-Piece, a novel lossy compression algorithm for time-series data that optimizes the space requirements of representing PLA line segments, by finding the minimum number of groups we can organize these segments into, to represent them jointly. Our experimental evaluation demonstrates that our approach readily outperforms competing techniques, attaining compression ratios with more than twofold improvement on average over what PLA algorithms can offer. This allows for providing significantly higher accuracy with equivalent space requirements. Moreover, our algorithm, due to the simplicity of its merging phase, imposes little overhead while compacting the PLA description, offering a significantly improved trade-off between space and running time. The aforementioned benefits of our approach significantly improve the efficiency in which we can store time-series data, while allowing a tight maximum error in the representation of their values.
Xenophon Kitsios, Panagiotis Liakos, Katia Papakonstantinopoulou, Yannis Kotidis
Proc. VLDB Endow.2
2022 On Compressing Temporal Graphs
abstract
Contemporary data-systems empowering the daily human activity are routinely represented with graphs. During the last decade, the volume growth of such systems has been unprece-dented. This hinders the timely analysis of the formed networks due to existing physical memory limitations and significant I/O overheads. Graph compression techniques have managed to reduce memory requirements and allow for representing such networks using a few bits-per-edge. Respective approaches offer succinct mappings for social, biological, and information networks while allowing for the efficient access of sought graph elements. Despite their success, such methods mostly focus on static graphs, and predominantly offer access to either a snapshot or an aggregated view of a network. In reality however, networks change over time and, in many instances, we are interested in capturing and studying this evolution. In this paper we propose a framework for compressing emerging temporal graphs based on a dual-representation which articulates both network structure and corresponding temporal information. We empirically establish properties exhibited by community-networks regarding their time aspect(s) and harness these features in our proposed repre-sentation. Our experimental evaluation demonstrates that our approach for compressing temporal graphs readily outperforms competing techniques, attaining compression ratios that are on average around 60% of the space required by state-of-the-art techniques. Moreover, our memory-efficient representation yields more than 70 % faster graph compression and orders of magnitude quicker retrieval of graphs' elements, especially when it comes to large-scale networks. Finally, our framework is the first effort we are aware of, that considers actual time instead of time steps. This helps us attain better control for the size of our representation and reap further memory savings.
Panagiotis Liakos, Katia Papakonstantinopoulou, Theodore Stefou, Alex Delis
ICDE1
2022 Chimp: Efficient Lossless Floating Point Compression for Time Series Databases
abstract
Applications in diverse domains such as astronomy, economics and industrial monitoring, increasingly press the need for analyzing massive collections of time series data. The sheer size of the latter hinders our ability to efficiently store them and also yields significant storage costs. Applying general purpose compression algorithms would effectively reduce the size of the data, at the expense of introducing significant computational overhead. Time Series Management Systems that have emerged to address the challenge of handling this overwhelming amount of information, cannot suffer the ingestion rate restrictions that such compression algorithms would cause. Data points are usually encoded using faster, streaming compression approaches. However, the techniques that contemporary systems use do not fully utilize the compression potential of time series data, with implications in both storage requirements and access times. In this work, we propose a novel streaming compression algorithm, suitable for floating point time series data. We empirically establish properties exhibited by a diverse set of time series and harness these features in our proposed encodings. Our experimental evaluation demonstrates that our approach readily outperforms competing techniques, attaining compression ratios that are competitive with slower general purpose algorithms, and on average around 50% of the space required by state-of-the-art streaming approaches. Moreover, our algorithm outperforms all earlier techniques with regards to both compression and access time , offering a significantly improved trade-off between space and speed. The aforementioned benefits of our approach - in terms of all space requirements, compression time and read access - significantly improve the efficiency in which we can store and analyze time series data.
Panagiotis Liakos, Katia Papakonstantinopoulou, Yannis Kotidis
Proc. VLDB Endow.1
2022 Rapid Detection of Local Communities in Graph Streams
abstract
We examine the problem of uncovering communities in complex real-world networks whose elements and their respective associations manifest as streams of data. Community detection is applied in emerging computational environments and concerns critical applications in diverse areas including social computing, web analysis, IoT and biology. Despite the already expended related research efforts, the task of revealing the community structure of massive and rapidly-evolving networks remains very challenging. More specifically, there is an emerging need for online approaches that ingest graph data as a stream. In this paper, we propose a streaming-graph community-detection algorithm that expands seed-sets of nodes to communities. We consider an online setting and process a stream of edges while aiming to form communities on-the-fly using partial knowledge of the graph structure. We use space-efficient structures to maintain very limited information regarding the nodes of the graph and the sought communities, so as to effectively process large scale networks. In addition to our novel streaming approach, we develop a technique that increases the accuracy of our algorithm considerably and additionally propose a new clustering algorithm that allows for automatically deriving the size of the communities we seek to detect. Using ground-truth communities for a wide range of large real-word and synthetic networks, our experimental evaluation shows that our approach does achieve accuracy comparable, and oftentimes better, to the state-of-the-art non-streaming community detection algorithms. More importantly, we attain significant improvements in both execution time and memory requirements.
Panagiotis Liakos, Katia Papakonstantinopoulou, Alexandros Ntoulas, Alex Delis
IEEE Trans. Knowl. Data Eng.1
2020 A Sentiment Analysis Service Platform for Streamed Multilingual Tweets
abstract
Micro-blogging and social-media platforms are now prominent forums for disseminating information, opinions and commentaries. Among these, Twitter enjoys an in-excess of 330M base of users who continually produce and consume information snippets. Users collectively create a voluminous and multi-lingual corpus in a very broad range of topics on a daily basis. The discourse generated in the blogosphere is often of prime interest and importance to individuals, organizations, and companies. These actors would certainly like to periodically receive an overall assessment of demonstrated "sentiments" on specific issues by automatically classifying tweets expressed in different languages in conjunction with big-data analytics. In this paper, we propose a scalable service platform that employs multilingual sentiment analysis to classify streamed-tweets and yields analytics for selected topics in real-time. We discuss the main component of our Spark-enabled platform as we seek to offer an effective big-data service that can: 1) dynamically handle voluminous as well as high-rate tweettraffic through a multi-component application exploiting the latest software developments, 2) accurately identify messages originated by non-genuine user-accounts, and 3) utilize the Spark machine-learning library (MLib) to successfully classify streamed multi-lingual messages in real-time, using multiple potentially distributed executors. To empower our service platform, we have adopted training sets and developed sentiment analysis (SA) models for English, French, and Greek that help classify streamed tweetswith high accuracy. While experimenting with our distributed analytical platform, we establish both accurate and real-time classification for tweetsexpressed in the above European languages.
Ioanna Karageorgou, Panagiotis Liakos, Alex Delis
IEEE BigData2
2018 Realizing Memory-Optimized Distributed Graph Processing
abstract
A multitude of contemporary applications heavily involve graph data whose size appears to be ever-increasing. This trend shows no signs of subsiding and has caused the emergence of a number of distributed graph processing systems including Pregel, Apache Giraph, and GraphX. However, the unprecedented scale now reached by real-world graphs hardens the task of graph processing due to excessive memory demands even for distributed environments. By and large, such contemporary graph processing systems employ ineffective in-memory representations of adjacency lists. Therefore, memory usage patterns emerge as a primary concern in distributed graph processing. We seek to address this challenge by exploiting empirically-observed properties demonstrated by graphs generated by human activity. In this paper, we propose 1) three compressed adjacency list representations that can be applied to any distributed graph processing system, 2) a variable-byte encoded representation of out-edge weights for space-efficient support of weighted graphs, and 3) a tree-based compact out-edge representation that allows for efficient mutations on the graph elements. We experiment with publicly-available graphs whose size reaches two-billion edges and report our findings in terms of both space-efficiency and execution time. Our suggested compact representations do reduce respective memory requirements for accommodating the graph elements up-to 5 times if compared with state-of-the-art methods. At the same time, our memory-optimized methods retain the efficiency of uncompressed structures and enable the execution of algorithms for large scale graphs in settings where contemporary alternative structures fail due to memory errors.
Panagiotis Liakos, Katia Papakonstantinopoulou, Alex Delis
IEEE Trans. Knowl. Data Eng.1
2017 COEUS: Community detection via seed-set expansion on graph streams
abstract
We examine the problem of effective identification of community structure of a network whose elements and their respective relationships manifest through streams. The problem has recently garnered much interest as it appears in emerging computational environments and concerns critical applications in diverse areas including social computing, web analysis, IoT and biology. Despite the already expended research efforts in detecting communities in networks, the unprecedented volume that real-world networks now reach, renders the task of revealing community structures extremely burdensome. The sheer size of such networks oftentimes makes their representation in main memory impossible. Thus, processing the developing graphs to extract the underlying communities remains an open challenge. In this paper, we propose a graph-stream community detection algorithm that expands seed-sets of nodes to communities. We consider a stream of edges and aim at processing them to form communities without maintaining the entire graph structure. Instead, we maintain very limited information regarding the nodes of the graph and the communities we seek. In addition to our novel streaming approach, we both develop a technique that increases the accuracy of our algorithm considerably and propose a new clustering algorithm that allows for automatically deriving the size of the communities we seek to detect. Our experimental evaluation using ground-truth communities for a wide range of large real-word networks shows that our proposed approach does achieve accuracy comparable or even better to that of state-of-the-art non-streaming community detection algorithms. More importantly, the attained improvements in both execution time and memory space requirements are remarkable.
Panagiotis Liakos, Alexandros Ntoulas, Alex Delis
IEEE BigData1
2017 Rhea: Adaptively sampling authoritative content from social activity streams
abstract
Processing the full activity stream of a social network in real time is oftentimes prohibitive in terms of both storage and computational cost. One way to work around this problem is to take a sample of the social activity and use this sample to feed into applications such as content recommendation, opinion mining, or sentiment analysis. In this paper, we study the problem of extracting samples of authoritative content from a social activity stream. Specifically, we propose an adaptive stream sampling approach, termed Rhea, that processes a stream of social activity in real-time and samples the content of users that are more likely to provide influential information. To the best of our knowledge, Rhea is the first algorithm that dynamically adapts over time to account for evolving trends in the activity stream. Thus, we are able to capture high quality content from emerging users that contemporary white-list based methods ignore. We evaluate Rhea using two popular social networks reaching up to half a billion posts. Our results show that we significantly outperform previously proposed methods in terms of both recall and precision, while also offering remarkably more accurate ranking.
Panagiotis Liakos, Alexandros Ntoulas, Alex Delis
IEEE BigData1
2016 Scalable link community detection: A local dispersion-aware approach
abstract
Real-life systems involving interacting objects are typically modeled as graphs and can often grow very large in size. Revealing the community structure of such systems is crucial in helping us better understand their complex nature. However, the ever-increasing size of real-world graphs, and our evolving perception of what a community is, make the task of community detection very challenging. One such challenge, is the discovery of the possibly overlapping communities of a given node in a billion-node graph. This problem is very common in modern large social networks like Facebook and Linkedln. In this paper, we propose a scalable local community detection approach to efficiently unfold the communities of individual target nodes in a given network. Our goal is to reveal the groupings formed around nodes (e.g., users) by leveraging the relations of the different contexts the nodes participate in. Our algorithm, termed Local Dispersion-aware Link Communities or LDLC, measures the similarity of pairs of links in the graph as well as the extent of their participation in multiple contexts. Then, it determines the ordering that we should group the links in order to form communities. Our approach is not affected by constraints existent in previous techniques (e.g., the need for several seed nodes or the need to collapse multiple overlapping communities to one). Our experimental evaluation using ground-truth communities for a wide range of large real-world networks show that LDLC significantly outperforms state-of-the-art methods on both accuracy and efficiency.
Panagiotis Liakos, Alexandros Ntoulas, Alex Delis
IEEE BigData1
2016 Memory-Optimized Distributed Graph Processing through Novel Compression Techniques
abstract
A multitude of contemporary applications now involve graph data whose size continuously grows and this trend shows no signs of subsiding. This has caused the emergence of many distributed graph processing systems including Pregel and Apache Giraph. However, the unprecedented scale now reached by real-world graphs hardens the task of graph processing even in distributed environments and the current memory usage patterns rapidly become a primary concern for such contemporary graph processing systems. We seek to address this challenge by exploiting empirically-observed properties demonstrated by graphs that are generated by human activity. In this paper, we propose three space-efficient adjacency list representations that can be applied to any distributed graph processing system. Our suggested compact representations reduce respective memory requirements for accommodating the graph elements up to 5 times if compared with state-of-the-art methods. At the same time, our memory-optimized methods retain the efficiency of uncompressed structures and enable the execution of algorithms for large scale graphs in settings where contemporary alternative structures fail due to memory errors.
Panagiotis Liakos, Katia Papakonstantinopoulou, Alex Delis
CIKM1
2016 On the Impact of Social Cost in Opinion Dynamics
Panagiotis Liakos, Katia Papakonstantinopoulou
ICWSM1
2015 An Interactive Freight-Pooling Service for Efficient Last-Mile Delivery
abstract
The existing practices of the urban section of freight transport chain result in traffic congestion, air pollution and resources being wasted. We focus on the final stage of freight distribution and propose an interactive freight-pooling service, in an effort to reduce the undesirable effects and the cost of freight transport in urban areas. Our service empowers city and state authorities to orchestrate the distribution network through interactive interfaces. We break the problem into three distinct phases that collectively helps us set constraints related to the quality of service and find inexpensive routes. In this regard, our proposed freight-pooling approach becomes an attractive option for efficient distribution, that guarantees cost minimization without sacrificing the level of quality.
Panagiotis Liakos, Alex Delis
MDM (2)1
2015 A Distributed Infrastructure for Earth-Science Big Data Retrieval
abstract
Earth-Science data are composite, multi-dimensional and of significant size, and as such, continue to pose a number of ongoing problems regarding their management. With new and diverse information sources emerging as well as rates of generated data continuously increasing, a persistent challenge becomes more pressing: To make the information existing in multiple heterogeneous resources readily available. The widespread use of the XML data-exchange format has enabled the rapid accumulation of semi-structured metadata for Earth-Science data. In this paper, we exploit this popular use of XML and present the means for querying metadata emanating from multiple sources in a succinct and effective way. Thereby, we release the user from the very tedious and time consuming task of examining individual XML descriptions one by one. Our approach, termed Meta-Array Data Search (MAD Search), brings together diverse data sources while enhancing the user-friendliness of the underlying information sources. We gather metadata using different standards and construct an amalgamated service with the help of tools that discover and harvest such metadata; this service facilitates the end-user by offering easy and timely access to all metadata. The main contribution of our work is a novel query language termed xWCPS, that builds on top of two widely-adopted standards: XQuery and the Web Coverage Processing Service (WCPS). xWCPS furnishes a rich set of features regarding the way scientific data can be queried with. Our proposed unified language allows for requesting metadata while also giving processing directives. Consequently, the xWCPS-enabled MAD Search helps in both retrieval and processing of large data sets hosted in an heterogeneous infrastructure. We demonstrate the effectiveness of our approach through diverse use-cases that provide insights into the syntactic power and overall expressiveness of xWCPS. We evaluate MAD Search in a distributed environment that comprises five high-volume array-databases whose sizes range between 20 and 100 GB and so, we ascertain the applicability and potential of our proposal.
Panagiotis Liakos, Panagiota Koltsida, George Kakaletris, Peter Baumann 0001, Yannis E. Ioannidis, Alex Delis
Int. J. Cooperative Inf. Syst.1
2014 Pushing the Envelope in Graph Compression
abstract
We improve the state-of-the-art method for the compression of web and other similar graphs by introducing an elegant technique which further exploits the clustering properties observed in these graphs. The analysis and experimental evaluation of our method shows that it outperforms the currently best method of Boldi et al. by achieving a better compression ratio and retrieval time. Our method exhibits vast improvements on certain families of graphs, such as social networks, by taking advantage of their compressibility characteristics, and ensures that the compression ratio will not worsen for any graph, since it easily falls back to the state-of-the-art method.
Panagiotis Liakos, Katia Papakonstantinopoulou, Michael Sioutis
CIKM1
2014 On the Effect of Locality in Compressing Social Networks
Panagiotis Liakos, Katia Papakonstantinopoulou, Michael Sioutis
ECIR1
2012 Topic-Sensitive Hidden-Web Crawling
Panagiotis Liakos, Alexandros Ntoulas
WISE1