Katerina Doka

dblp:63/4123 · DBLP profile ↗
← Back
27ranked-venue papers
9as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 7 · 3 first-authorSystems, architecture and hardware · 7 · 3 first-authorSecurity and privacy · 5 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorComputer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2024 PROOF: Decentralized Platform for Verifiable Outsourced Computation
abstract
Decentralization is reshaping our digital landscape, promising improved control, security, and accessibility. This demo paper explores the potential of decentralization in Cloud Computing, presenting PROOF, a platform designed to enable task outsourcing to any - potentially untrusted - computational resource without compromising the credibility of the produced output. Implemented as an Ethereum dApp, PROOF handles task allocation, execution, and verification, leveraging Blockchain technology for transparent and secure transactions. It utilizes an auction mechanism to match outsourcing requests to resource providers, the IPFS peer-to-peer file system to serve as storage, and Docker containers to facilitate secure and standardized task execution across a network of individual providers. During the demonstration, the attendees will have the chance to interact with PROOF (a) as clients, requesting resources, delegating task processing to remote infrastructures and verifying the validity of results and (b) as resource providers engaging in auctions, performing computations either in an honest or in a malicious way and automatically receiving payments.
Spiros Grigoratos, Katerina Doka, Nectarios Koziris
ICBC2
2023 Graph-Centric Crypto Price Prediction
abstract
In this paper we propose, for the first time in the literature, a graph-centric approach for crypto price prediction through the use of Graph Neural Networks, which can naturally be applied over transaction graphs and thus exploit the connectivity and structural features of Blockchains' transaction network. We apply our approach to the prediction of the most popular cryptocurrency currently on the market, Bitcoin. We demonstrate the effectiveness of our methods for short and longer term predictions considering several variants of the Bitcoin transaction graphs, additional external information and various basic architectures, achieving results with mean absolute percentage error as low as 1.069 %. Compared to the state-of-the-art, our method exhibits more useful results, in terms of both accuracy and predictive power.
Charalampos Kleitsikas, Katerina Doka, Agis Politis, Nectarios Koziris
ICBC2
2022 One-off Disclosure Control by Heterogeneous Generalization
Olga Gkountouna, Katerina Doka, Mingqiang Xue, Jianneng Cao, Panagiotis Karras
USENIX Security Symposium2
2022 Enabling Transparent Acceleration of Big Data Frameworks using Heterogeneous Hardware
abstract
The ever-increasing demand for high performance Big Data analytics and data processing, has paved the way for heterogeneous hardware accelerators, such as Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), to be integrated into modern Big Data platforms. Currently, this integration comes at the cost of programmability since the end-user Application Programming Interface (APIs) must be altered to access the underlying heterogeneous hardware. For example, current Big Data frameworks, such as Apache Spark, provide a new API that combines the existing Spark programming model with GPUs. For other Big Data frameworks, such as Flink, the integration of GPUs and FPGAs is achieved via external API calls that bypass their execution models completely. In this paper, we rethink current Big Data frameworks from a systems and programming language perspective, and introduce a novel co-designed approach for integrating hardware acceleration into their execution models. The novelty of our approach is attributed to two key design decisions: a) support for arbitrary User Defined Functions (UDFs), and b) no modifications to the user level API. The proposed approach has been prototyped in the context of Apache Flink, and enables unmodified applications written in Java to run on heterogeneous hardware, such as GPU and FPGAs, transparently to the users. The performance evaluation of the proposed solution has shown performance speedups of up to 65x on GPUs and 184x on FPGAs for suitable workloads of standard benchmarks and industrial use cases against vanilla Flink running on traditional multi-core CPUs.
Maria Xekalaki, Juan José Fumero, Athanasios Stratikopoulos, Katerina Doka, Christos Katsakioris, Constantinos Bitsakos, Nectarios Koziris, Christos Kotselidis
Proc. VLDB Endow.4
2021 Clouseau: Blockchain-based Data Integrity for HDFS Clusters
abstract
As the volume of produced data is exponentially increasing, companies tend to rely on distributed systems to meet the surging demand for storage capacity. With the business workflows becoming more and more complex, such systems often consist of or are accessed by multiple independent, untrusted entities, which need to interact with shared data. In such scenarios, the potential conflicts of interest incentivize malicious parties to act in a dishonest way and tamper the data to their own benefit. The decentralized nature of the systems renders verifiable data integrity a strenuous but necessary task: The various parties should be able to audit changes and detect tampering when it happens.In this work, we focus on HDFS, the most common storage substrate for Big Data analytics. HDFS is vulnerable to malicious users and participating nodes and does not provide a trustful lineage mechanism, thus jeopardizing the integrity of stored data and the credibility of extracted insights. As a remedy, we present Clouseau, a blockchain-based system that provides verifiable integrity over HDFS, while it does not incur significant overhead at the critical path of read/write operations. During the demonstration, the attendees will have the chance to interact with Clouseau, corrupt data themselves, and witness how Clouseau detects malicious actions.
Alyzia Konsta, Ioannis Mytilinis, Katerina Doka, Sotirios Niarchos, Nectarios Koziris
ICDE3
2020 Fair Procedures for Fair Stable Marriage Outcomes
abstract
Given a two-sided market where each agent ranks those on the other side by preference, the stable marriage problem calls for finding a perfect matching such that no pair of agents prefer each other to their matches. Recent studies show that the number of stable solutions can be large in practice. Yet the classical solution to the problem, the Gale-Shapley (GS) algorithm, assigns an optimal match to each agent on one side, and a pessimal one to each on the other side; such a solution may fare well in terms of equity only in highly asymmetric markets. Finding a stable matching that minimizes the sex equality cost, an equity measure expressing the discrepancy of mean happiness among the two sides, is strongly NP-hard. Extant heuristics either (a) oblige some agents to involuntarily abandon their matches, or (b) bias the outcome in favor of some agents, or (c) need high-polynomial or unbounded time.We provide the first procedurally fair algorithms that output equitable stable marriages and are guaranteed to terminate in at most cubic time; the key to this breakthrough is the monitoring of a monotonic state function and the use of a selective criterion for accepting proposals. Our experiments with diverse simulated markets show that: (a) extant heuristics fail to yield high equity; (b) the best solution found by the GS algorithm can be very far from optimal equity; and (c) our procedures stand out in both efficiency and equity, even when compared to a non-procedurally fair approximation scheme.
Nikolaos Tziavelis, Ioannis Giannakopoulos, Rune Quist Johansen, Katerina Doka, Nectarios Koziris, Panagiotis Karras
AAAI4
2020 Efficient Compilation and Execution of JVM-Based Data Processing Frameworks on Heterogeneous Co-Processors
abstract
This paper addresses the fundamental question of how modern Big Data frameworks can dynamically and transparently exploit heterogeneous hardware accelerators. After presenting the major challenges that have to be addressed towards this goal, we describe our proposed architecture for automatic and transparent hardware acceleration of Big Data frameworks and applications. Our vision is to retain the uniform programming model of Big Data frameworks and enable automatic, dynamic Just-In-Time compilation of the candidate code segments that benefit from hardware acceleration to the corresponding format. In conjunction with machine learning-based device selection, that respect user-defined constraints (e.g., cost, time, etc.), we enable dynamic code execution on GPUs and FPGAs transparently to the user. In addition, we dynamically re-steer execution at runtime based on the availability of resources. Our preliminary results demonstrate that our approach can accelerate an existing Apache Flink application by up to 16.5x.
Christos Kotselidis, Sotirios Diamantopoulos, Orestis Akrivopoulos, Viktor Rosenfeld, Katerina Doka, Hazeef Mohammed, Georgios Mylonas, Vassilis Spitadakis, Will Morgan
DATE5
2020 Dynamic planar range skyline queries in log logarithmic expected time
Katerina Doka, Andreas Kosmatopoulos, Apostolos N. Papadopoulos, Spyros Sioutas, Kostas Tsichlas, Dimitrios Tsoumakos
Inf. Process. Lett.1
2019 Equitable Stable Matchings in Quadratic Time
abstract
Can a stable matching that achieves high equity among the two sides of a market be reached in quadratic time? The Deferred Acceptance (DA) algorithm finds a stable matching that is biased in favor of one side; optimizing apt equity measures is strongly NP-hard. A proposed approximation algorithm offers a guarantee only with respect to the DA solutions. Recent work introduced Deferred Acceptance with Compensation Chains (DACC), a class of algorithms that can reach any stable matching in O(n^4) time, but did not propose a way to achieve good equity. In this paper, we propose an alternative that is computationally simpler and achieves high equity too. We introduce Monotonic Deferred Acceptance (MDA), a class of algorithms that progresses monotonically towards a stable matching; we couple MDA with a mechanism we call Strongly Deferred Acceptance (SDA), to build an algorithm that reaches an equitable stable matching in quadratic time; we amend this algorithm with a few low-cost local search steps to what we call Deferred Local Search (DLS), and demonstrate experimentally that it outperforms previous solutions in terms of equity measures and matches the most efficient ones in runtime.
Nikolaos Tziavelis, Ioannis Giannakopoulos, Katerina Doka, Nectarios Koziris, Panagiotis Karras
NeurIPS3
2019 Building Ad-Hoc Clouds with CloudAgora
abstract
The public Cloud market has become a monopoly, where a handful of providers - which are by default considered as trusted entities - define the prices, accumulate knowledge from users' data and computations and strengthen their already privileged position. As a remedy we propose CloudAgora, a platform that democratizes the Cloud market by allowing individuals and companies alike to compete on equal terms as potential resource providers, while enabling users to access low-cost storage and computation without having to blindly trust any central authority. During the demo, the attendees will be able to interact with CloudAgora through an easy-to-use UI, which will allow them to act both as users and as providers. As users, the attendees will have the chance to request storage or compute resources, upload data and outsource task processing over remote infrastructures. As providers, they will be able to participate in auctions, serve requests and offer validity proofs upon request. Moreover, the audience will experience first hand how the underlying blockchain technology is used to record commitment policies, publicly verify off-chain services and trigger automatic micropayments.
Tasos Bakogiannis, Ioannis Mytilinis, Katerina Doka, Georgios I. Goumas
SRDS3
2018 The Vision of a HeterogeneRous Scheduler
abstract
Modern Big Data processing systems, scheduling platforms and cloud infrastructures employ specialized hardware accelerators such as GPUs, FPGAs, TPUs, ASICs, etc. to optimize the execution of resource intensive workloads such as Machine Learning, Artificial Intelligence or generic Data Analytics tasks. Nevertheless, this support is mostly a user-dependent, manual process that requires careful and educated decisions on both the amount and type of required resources to exploit the underlying hardware and achieve any user-defined higher level policies. In this work we present the initial design of the HeterogeneRous Scheduler (HRS), an intelligent scheduler that can make automated decisions on both how and where to map arbitrary data analytics tasks to underlying cloud hardware that may consist of a mix of hardware accelerators and clusters with general purpose CPUs. We experimentally evaluate the performance trade-offs between hardware accelerators and CPUs where we show that there are cases where one technology outperforms the other. We finally present an initial architecture of HRS where we depict its different components and their interactions with the Big Data framework and the cloud infrastructure.
Ioannis Mytilinis, Constantinos Bitsakos, Katerina Doka, Ioannis Konstantinou, Nectarios Koziris
CloudCom3
2018 Docker-Sec: A Fully Automated Container Security Enhancement Mechanism
abstract
The popularity of containers is constantly rising in the virtualization landscape, since they incur significantly less overhead than Virtual Machines, the traditional hypervisor-based counterparts, while enjoying better performance. However, containers pose significant security challenges due to their direct communication with the host kernel, allowing attackers to break into the host system and co-located containers more easily than Virtual Machines. Existing security hardening mechanisms are based on the enforcement of Mandatory Access Control rules, which exclusively allow specified, desired operations. However, these mechanisms entail explicit knowledge of the container functionality and behavior and require manual intervention and setup. To overcome these limitations, we present Docker-sec, a user-friendly mechanism for the protection of Docker containers throughout their lifetime via the enforcement of access policies that correspond to the anticipated (and legitimate) activity of the applications they enclose. Docker-sec employs two mechanisms: (a) Upon container creation, it constructs an initial, static set of access rules based on container configuration parameters; (b) During container runtime, the initial set is enhanced with additional rules that further restrict the container's capabilities, reflecting the actual application operations. Through a rich interaction with our system the audience will experience firsthand how Docker-sec can successfully protect containers from zero-day vulnerabilities in an automatic manner, with minimal overhead on the application performance.
Fotis Loukidis-Andreou, Ioannis Giannakopoulos, Katerina Doka, Nectarios Koziris
ICDCS3
2017 Isolation in Docker through Layer Encryption
abstract
Containers are constantly gaining ground in the virtualization landscape as a lightweight and efficient alternative to hypervisor-based Virtual Machines, with Docker being the most successful representative. Docker relies on union-capable file systems, where any action performed to a base image is captured as a new file system layer. This strategy allows developers to easily pack applications into Docker image layers and distribute them via public registries. However, this image creation and distribution strategy does not protect sensitive data from malicious privileged users (e.g., registry administrator, cloud provider), since encryption is not natively supported. We propose and demonstrate a mechanism for secure Docker image manipulation throughout its life cycle: The creation, storage and usage of a Docker image is backed by a data-at-rest mechanism, which maintains sensitive data encrypted on disk and encrypts/decrypts them on-the-fly in order to preserve their confidentiality at all times, while the distribution and migration of images is enhanced with a mechanism that encrypts only specific layers of the file system that need to remain confidential and ensures that only legitimate key holders can decrypt them and reconstruct the original image. Through a rich interaction with our system the audience will experience first-hand how sensitive image data can be safely distributed and remain encrypted at the storage device throughout the container's lifetime, bearing only a marginal performance overhead.
Ioannis Giannakopoulos, Konstantinos Papazafeiropoulos, Katerina Doka, Nectarios Koziris
ICDCS3
2016 Mix 'n' match multi-engine analytics
abstract
Current platforms fail to efficiently cope with the data and task heterogeneity of modern analytics workflows due to their adhesion to a single data and/or compute model. As a remedy, we present IReS, the Intelligent Resource Scheduler for complex analytics workflows executed over multi-engine environments. IReS is able to optimize a workflow with respect to a user-defined policy relying on cost and performance models of the required tasks over the available platforms. This optimization consists in allocating distinct workflow parts to the most advantageous execution and/or storage engine among the available ones and deciding on the exact amount of resources provisioned. Our current prototype supports 5 compute and 3 data engines, yet new ones can effortlessly be added to IReS by virtue of its engine-agnostic mechanisms. Our extensive experimental evaluation confirms that IReS speeds up diverse and realistic workflows by up to 30% compared to their optimal single-engine plan by automatically scattering parts of them to different execution engines and datastores. Its optimizer incurs only marginal overhead to the workflow execution performance, managing to discover the optimal execution plan within a few seconds, even for large-scale workflow instances.
Katerina Doka, Nikolaos Papailiou, Victor Giannakouris, Dimitrios Tsoumakos, Nectarios Koziris
IEEE BigData1
2015 Heterogeneous k-anonymization with high utility
abstract
Among the privacy-preserving approaches that are known in the literature, h-anonymity remains the basis of more advanced models while still being useful as a stand-alone solution. Applying h-anonymity in practice, though, incurs severe loss of data utility, thus limiting its effectiveness and reliability in real-life applications and systems. However, such loss in utility does not necessarily arise from an inherent drawback of the model itself, but rather from the deficiencies of the algorithms used to implement the model. Conventional approaches rely on a methodology that publishes data in homogeneous generalized groups. An alternative modern data publishing scheme focuses on publishing the data in heterogeneous groups and achieves higher utility, while ensuring the same privacy guarantees. As conventional approaches cannot anonymize data following this heterogeneous scheme, innovative solutions are required for this purpose. Following this approach, in this paper we provide a set of algorithms that ensure high-utility h-anonymity, via solving an equivalent graph processing problem.
Katerina Doka, Mingqiang Xue, Dimitrios Tsoumakos, Panagiotis Karras, Alfredo Cuzzocrea, Nectarios Koziris
IEEE BigData1
2015 k-Anonymization by Freeform Generalization
abstract
Syntactic data anonymization strives to (i) ensure that an adversary cannot identify an individual's record from published attributes with high probability, and (ii) provide high data utility. These mutually conflicting goals can be expressed as an optimization problem with privacy as the constraint and utility as the objective function. Conventional research using the k-anonymity model has resorted to publishing data in homogeneous generalized groups. A recently proposed alternative does not create such cliques; instead, it recasts data values in a heterogeneous manner, aiming for higher utility. Nevertheless, such works never defined the problem in the most general terms; thus, the utility gains they achieve are limited. In this paper, we propose a methodology that achieves the full potential of heterogeneity and gains higher utility while providing the same privacy guarantee. We formulate the problem of maximal-utility k-anonymization by freeform generalization as a network flow problem. We develop an optimal solution therefor using Mixed Integer Programming. Given the non-scalability of this solution, we develop an O(k n2) Greedy algorithm that has no time-complexity disadvantage vis-á-vis previous approaches, an O(k n2 log n) enhanced version thereof, and an O(k n3) adaptation of the Hungarian algorithm; these algorithms build a set of k perfect matchings from original to anonymized data, a novel approach to the problem. Moreover, our techniques can resist adversaries who may know the employed algorithms. Our experiments with real-world data verify that our schemes achieve near-optimal utility (with gains of up to 41%), while they can exploit parallelism and data partitioning, gaining an efficiency advantage over simpler methods.
Katerina Doka, Mingqiang Xue, Dimitrios Tsoumakos, Panagiotis Karras
AsiaCCS1
2015 An Equitable Solution to the Stable Marriage Problem
abstract
A stable marriage problem (SMP) of size n involves n men and n women, each of whom has ordered members of the opposite gender by descending preferability. A solution is a perfect matching among men and women, such that there exists no pair who prefer each other to their current spouses. The problem was formulated in 1962 by Gale and Shapley, who showed that any instance can be solved in polynomial time, and has attracted interest due to its application to any two-sided market. Still, the solution obtained by the Gale-Shapley algorithm is favorable to one side. Gusfield and Irving introduced the equitable stable marriage problem (ESMP), which calls for finding a stable matching that minimizes the distance between men's and women's sum-of-rankings of their spouses. Unfortunately, ESMP is strongly NP-hard, approximation algorithms therefor are impractical, while even proposed heuristics may run for an unpredictable number of iterations. We propose a novel, deterministic approach that treats both genders equally, while eschewing an exhaustive exploration of the space of all stable matchings. Our thorough experimental study shows that, in contrast to previous proposals, our method not only achieves high-quality solutions, but also terminates efficiently and repeatably on all tested large problem instances.
Ioannis Giannakopoulos, Panagiotis Karras, Dimitrios Tsoumakos, Katerina Doka, Nectarios Koziris
ICTAI4
2015 IReS: Intelligent, Multi-Engine Resource Scheduler for Big Data Analytics Workflows
abstract
Big data analytics tools are steadily gaining ground at becoming indispensable to businesses worldwide. The complexity of the tasks they execute is ever increasing due to the surge in data and task heterogeneity. Current analytics platforms, while successful in harnessing multiple aspects of this ``data deluge", bind their efficacy to a single data and compute model and often depend on proprietary systems. However, no single execution engine is suitable for all types of computation and no single data store is suitable for all types of data. To this end, we demonstrate IReS, the Intelligent Resource Scheduler for complex analytics workflows executed over multi-engine environments. Our system models the cost and performance of the required tasks over the available platforms. IReS is then able to match distinct workflow parts to the execution and/or storage engine among the available ones in order to optimize with respect to a user-defined policy. During the demo, the attendees will be able to execute workflows that match real use cases and parametrize the input datasets and optimization policy. The underlying platform supports multiple compute and data engines, allowing the user to choose any subset of them. Through the inspection of the produced plan, its execution and the collection and presentation of numerous cost and performance metrics, the audience will experience first-hand how IReS takes advantage of heterogeneous runtimes and data stores and effectively models operator cost and performance for actual and diverse workflows.
Katerina Doka, Nikolaos Papailiou, Dimitrios Tsoumakos, Christos Mantas, Nectarios Koziris
SIGMOD Conference1
2015 MoDisSENSE: A Distributed Spatio-Temporal and Textual Processing Platform for Social Networking Services
abstract
The amount of social networking data that is being produced and consumed daily is huge and it is constantly increasing. A user's digital footprint coming from social networks or mobile devices, such as comments and check-ins contains valuable information about his preferences. The collection and analysis of such footprints using also information about the users' friends and their footprints offers many opportunities in areas such as personalized search, recommendations, etc. When the size of the collected data or the complexity of the applied methods increases, traditional storage and processing systems are not enough and distributed approaches are employed. In this work, we present MoDisSENSE, an open-source distributed platform that provides personalized search for points of interest and trending events based on the user's social graph by combining spatio-textual user generated data. The system is designed with scalability in mind, it is built using a combination of latest state-of-the art big data frameworks and its functionality is offered through easy to use mobile and web clients which support the most popular social networks. We give an overview of its architectural components and technologies and we evaluate its performance and scalability using different query types over various cluster sizes. Using the web or mobile clients, users are allowed to register themselves with their own social network credentials, perform socially enhanced queries for POIs, browse the results and explore the automatic blog creation functionality that is extracted by analyzing already collected GPS traces.
Ioannis Mytilinis, Ioannis Giannakopoulos, Ioannis Konstantinou, Katerina Doka, Dimitrios Tsitsigkos, Manolis Terrovitis, Lampros Giampouras, Nectarios Koziris
SIGMOD Conference4
2014 MoDisSENSE: A distributed platform for social networking services over mobile devices
abstract
In this work we present MoDisSENSE, a distributed analytics platform for social networking services over mobile devices. MoDisSENSE collects and stores various types of data from heterogeneous sources, such as GPS traces from cell phones, user profile information and comments from social networks connected to the platform. These are combined through spatio-temporal and textual analysis, performed in a distributed fashion, in order to extract knowledge, make smart suggestions and leverage user experience. The datastore follows a hybrid approach to handle both raw and processed data, simultaneously covering the need for scalability and fast query processing. Thus, the platform is able to resolve complex, multi-parameter, socially charged queries over Points of Interest in the order of milliseconds even under heavy load.
Ioannis Mytilinis, Ioannis Giannakopoulos, Ioannis Konstantinou, Katerina Doka, Nectarios Koziris
IEEE BigData4
2012 Exploiting the Social and Semantic Web for Guided Web Archiving
Thomas Risse 0001, Stefan Dietze, Wim Peters, Katerina Doka, Yannis Stavrakas, Pierre Senellart
TPDL4
2011 Online querying of d-dimensional hierarchies
Katerina Doka, Dimitrios Tsoumakos, Nectarios Koziris
J. Parallel Distributed Comput.1
2011 Brown Dwarf: A fully-distributed, fault-tolerant data warehousing system
Katerina Doka, Dimitrios Tsoumakos, Nectarios Koziris
J. Parallel Distributed Comput.1
2010 Brown dwarf: a P2P data-warehousing system
abstract
In this demonstration we present the Brown Dwarf, a distributed system designed to efficiently store, query and update multidimensional data. Deployed on any number of commodity nodes, our system manages to distribute large volumes of data over network peers on-the-fly and process queries and updates on-line through cooperating nodes that hold parts of a materialized cube. Moreover, it adapts its resources according to demand and hardware failures and is cost-effective both over the required hardware and software components. All the aforementioned functionality will be tested using various datasets and query loads.
Katerina Doka, Dimitrios Tsoumakos, Nectarios Koziris
CIKM1
2010 Distributing the power of OLAP
abstract
In this paper we present the Brown Dwarf, a distributed system designed to efficiently store, query and update multidimensional data over an unstructured Peer-to-Peer overlay, without the use of any proprietary tool. Brown Dwarf manages to distribute a highly effective centralized structure among peers on-the-fly. Both point and aggregate queries are then naturally answered on-line through cooperating nodes that hold parts of a fully or partially materialized data cube. Updates are also performed on-line, eliminating the usually costly over-night process. Our initial evaluation on an actual testbed proves that Brown Dwarf manages to distribute the structure across the overlay nodes incurring only a small storage overhead compared to the centralized algorithm. Moreover, it accelerates cube creation up to 5 times and querying up to several tens of times by exploiting the capabilities of the available network nodes working in parallel.
Katerina Doka, Dimitrios Tsoumakos, Nectarios Koziris
HPDC1
2009 A grid middleware for data management exploiting peer-to-peer techniques
Athanasia Asiki, Katerina Doka, Ioannis Konstantinou, Antonis Zissimos, Dimitrios Tsoumakos, Nectarios Koziris, Panayiotis Tsanakas
Future Gener. Comput. Syst.2
2008 Support for Concept Hierarchies in DHTs
abstract
Concept hierarchies greatly help in the organization and reuse of information and are widely used in a variety of applications, such as data warehouses. In this paper, we describe a method for efficiently storing and querying data organized into concept hierarchies and dispersed over a DHT. In our method, peers individually decide on the level of indexing according to the incoming queries. Roll-up and drill-down operations are performed on a per-node basis in order to minimize the number of floods for answering queries on varying levels of granularity. Initial experimental results support this argument on a variety of workloads.
Athanasia Asiki, Katerina Doka, Dimitrios Tsoumakos, Nectarios Koziris
Peer-to-Peer Computing2