Damianos Chatziantoniou

dblp:c/DChatziantoniou · DBLP profile ↗
← Back
22ranked-venue papers
16as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 19 · 14 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 A Model-Based Approach for Simple Construction and Efficient Evaluation of Dataframes
Konstantina Zouni, Ioanna Moraiti, Sotirios Angelopoulos, Damianos Chatziantoniou, Verena Kantere
DEXA (2)4
2023 Data-driven and On-Demand Conceptual Modeling
Damianos Chatziantoniou, Verena Kantere
DaWaK1
2022 Data Virtual Machines: Simplifying Data Sharing, Exploration & Querying in Big Data Environments
abstract
Today’s analytics environments are characterized by a high degree of heterogeneity in terms of data systems, formats and types of analysis. Many occasions call for rapid, ad hoc, on demand construction of a data model that represents (parts of) the data infrastructure of an organization, including ML tasks. This data model is given to data scientists to play with (express reports, build ML models, explore, etc.) We present a novel graph-based conceptual model, the Data Virtual Machine (DVM) representing data (persistent, transient, derived) of an organization. A DVM can be built quickly and agilely, offering schema flexibility. It is amenable to visual interfaces for schema and query management. Dataframing, a frequent querying/preprocessing task in analytics applications, is usually carried out by experienced data engineers employing SQL (in the presence of a relational data warehouse) or Python/R: a procedural approach with all the known drawbacks. Dataframes over DVMs are expressed declaratively - and visually, via a simple and intuitive tool. This way, non-IT experts can be involved in dataframing. In addition, query evaluation takes place within an algebraic framework with all the known benefits. I.e. a DVM enables the delegation of data engineering tasks to simpler users. We have seen analogous cases in the past, e.g. with the introduction of SQL. Finally, a DVM offers a formalism that facilitates data sharing, data portability and a single view of any entity – because a DVM’s node is an attribute and an entity at the same time. In this respect, DVMs can excellently serve as a data virtualization technique, an emerging trend in the industry. We argue that DVMs can have a significant practical impact in today’s big data environments.
Damianos Chatziantoniou, Verena Kantere, Nikos Antoniou, Angeliki Gantzia
IEEE Big Data1
2021 DataMingler: A Novel Approach to Data Virtualization
abstract
A Data Virtual Machine (DVM) is a novel graph-based conceptual model, similar to the entity-relationship model, representing existing data (persistent, transient, derived) of an organization. A DVM can be built quickly, agilely, offering schematic flexibility to data engineers. Data scientists can visually define complex dataframe queries in an intuitive and simple manner, which are evaluated within an algebraic framework. A DVM can be easily materialized in any logical data model and can be "reoriented'' around any node, offering a "single view of any entity''. In this paper we demonstrate DataMingler, a tool implementing DVMs. We argue that DVMs can have a significant practical impact in analytics environments.
Damianos Chatziantoniou, Verena Kantere
SIGMOD Conference1
2018 Enabling Global Big Data Computations
Damianos Chatziantoniou, Panagiotis Louridas
DOLAP1
2011 Tagged MapReduce: Efficiently Computing Multi-analytics Using MapReduce
Andreas Williams, Pavlos Mitsoulis-Ntompos, Damianos Chatziantoniou
DaWaK3
2011 theta-Constrained multi-dimensional aggregation
Michael O. Akinde, Michael H. Böhlen, Damianos Chatziantoniou, Johann Gamper
Inf. Syst.3
2011 Supporting real-time supply chain decisions based on RFID data streams
Damianos Chatziantoniou, Katerina Pramatari, Yannis Sotiropoulos
J. Syst. Softw.1
2008 A session-oriented approach in modeling hierarchies of streams
abstract
Abstract Data stream systems became recently the focus of intense research activity. Sensor measurements, financial data, network packets and log files can all be seen as data streams. The ability to query data streams is of increasing importance and has been identified as a crucial element for modern organizations and agencies. This article identifies an interesting class of applications where stream sessions may be organized into a hierarchical fashion—i.e. sessions may consist of subsessions. For example, log streams from call centers belong to different call sessions and call sessions consist of services subsessions. We may want to monitor statistics and perform accounting at any level on this hierarchy, relative to any other (higher) level (e.g. monitoring the average service session per call vs the average service session for the entire system). We argue that data streams of this kind have rich procedural semantics—i.e. behaviour—and therefore a semantically rich model should be used: a session may be defined by opening and closing conditions, may have data and methods and may consist of subsessions. We propose a simple conceptual model based on the notion of ‘session’—similar to a class in an object‐oriented environment having lifetime semantics. Queries on top of this schema can be formulated via HSA (hierarchical stream aggregate) expressions. We give an algorithm describing how stream data ‘flow down’ session hierarchies and discuss potential evaluation and optimization techniques for HSAs. We finally present NESTREAM, a prototype implementation incorporating many of these concepts. Copyright © 2007 John Wiley & Sons, Ltd.
Damianos Chatziantoniou, Achilleas Anagnostopoulos
Softw. Pract. Exp.1
2007 Using grouping variables to express complex decision support queries
Damianos Chatziantoniou
Data Knowl. Eng.1
2007 Partitioned optimization of complex queries
Damianos Chatziantoniou, Kenneth A. Ross
Inf. Syst.1
2005 Decision support queries on a tape-resident data warehouse
Damianos Chatziantoniou, Theodore Johnson
Inf. Syst.1
2004 Hierarchical Stream Aggregates: Querying Nested Stream Sessions
Damianos Chatziantoniou, Achilleas Anagnostopoulos
SSDBM1
2001 The MD-join: An Operator for Complex OLAP
abstract
OLAP queries (i.e. group-by or cube-by queries with aggregation) have proven to be valuable for data analysis and exploration. Many decision support applications need very complex OLAP queries, requiring a fine degree of control over both the group definition and the aggregates that are computed. For example, suppose that the user has access to a data cube whose measure attribute is Sum(Sales). Then the user might wish to compute the sum of sales in New York and the sum of sales in California for those data cube entries in which Sum(Sales)>$1,000,000. This type of complex OLAP query is often difficult to express and difficult to optimize using standard relational operators (including standard aggregation operators). In this paper, we propose the MD-join operator for complex OLAP queries. The MD-join provides a clean separation between group definition and aggregate computation, allowing great flexibility in the expression of OLAP queries. In addition, the MD-join has a simple and easily optimizable implementation, while the equivalent relational algebra expression is often complex and difficult to optimize. We present several algebraic transformations that allow relational algebra queries that include MD-joins to be optimized.
Damianos Chatziantoniou, Michael O. Akinde, Theodore Johnson, Samuel Kim
ICDE1
1999 Extended SQL for manipulating clinical warehouse data
Stephen B. Johnson, Damianos Chatziantoniou
AMIA2
1999 Extending Complex Ad-Hoc OLAP
abstract
Large scale data analysis and mining activities require sophisticated information extraction queries. Many queries require complex aggregation, and many of these aggregates are non-distributive. Conventional solutions to this problem involve defining User Defined Aggregate Functions (UDAFs). However, the use of UDAFs entails several problems. Defining a new UDAF can be a significant burden for the user, and optimizing queries involving UDAFs is difficult because of the “black box” nature of the UDAF.
Theodore Johnson, Damianos Chatziantoniou
CIKM2
1999 Ad Hoc OLAP: Expression and Evaluation
abstract
Users frequently formulate complex data analysis queries in order to identify interesting trends, make unusual patterns stand out, or verify hypotheses. Being able to express these data mining queries concisely is of major importance not only from the user's, but also from the system's point of view. Recent research in OLAP has focused on datacubes and their applications; however expression and processing of ad hoc decision support queries has been given very little attention. We present an appropriate framework for these queries and introduce a syntactic construct to support it. This SQL extension allows most OLAP queries, such as pivoting, complex intra- and inter-group comparisons, trends and hierarchical comparisons, to be expressed in a compact, intuitive and simple manner. This succinct representation of a complex OLAP query translates immediately to a novel, simple and efficient evaluation algorithm. We show how to optimize, analyze and parallelize this algorithm and discuss issues such as multiple query analysis and scaling. We present several experimental results of real life queries that show orders of magnitude of performance improvement. We argue that this tight coupling between representation and algorithm is essential to efficient processing of ad hoc OLAP queries.
Damianos Chatziantoniou
ICDE1
1999 The PanQ Tool and EMF SQL for Complex Data Management
Damianos Chatziantoniou
KDD1
1999 Evaluation of Ad Hoc OLAP: In-Place Computation
abstract
Large-scale data analysis and mining activities, such as identifying interesting trends, making unusual patterns stand out and verifying hypotheses, require sophisticated information extraction queries. Being able to express these data mining queries concisely is of major importance not only from the user's, but also from the system's point of view. Recent research in OLAP has focused on data cubes and their applications; however, the expression and processing of ad-hoc decision support queries has been given very little attention. In this paper, we present an appropriate framework for these queries and introduce a syntactic construct to support it. This SQL extension allows most OLAP queries, such as complex intra- and inter-group comparisons, trends and hierarchical comparisons, to be expressed in a compact, intuitive and simple manner. However, this syntactic extension is not the focus of this paper. This succinct representation of a complex OLAP query translates immediately to a novel, simple and efficient evaluation algorithm. We show how to optimize, analyze and parallelize this algorithm and discuss issues such as multiple query analysis and scaling. This algorithm constitutes the main contribution of this paper. Finally, we introduce our implementation on top of a commercial system and present several experimental results of real-life queries that show orders of magnitude of performance improvement in certain cases. We argue that this tight coupling between representation and algorithm is essential to efficient processing of ad-hoc OLAP queries.
Damianos Chatziantoniou
SSDBM1
1998 Complex Aggregation at Multiple Granularities
Kenneth A. Ross, Divesh Srivastava, Damianos Chatziantoniou
EDBT3
1997 Groupwise Processing of Relational Queries
Damianos Chatziantoniou, Kenneth A. Ross
VLDB1
1996 Querying Multiple Features of Groups in Relational Databases
Damianos Chatziantoniou, Kenneth A. Ross
VLDB1