EDBT 2026 Demo / reviewers in the wild / expert
Yannis Sismanis
dblp:65/5468
· DBLP profile ↗
24ranked-venue papers
6as first author
1since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 24 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
16 papers |
Query processing and optimization · 26% Data stream processing · 13% Distributed and cloud data management · 12% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 95% Computational complexity · 5% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Distributed systems · 44% Parallel and multicore computing · 28% High-performance computing · 28% |
Topics — the 28 heaviest of 38, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Graph data management › graph analytics
dynamic graph analysis |
0.2 | 1 | 2015 | Dynamic interaction graphs with probabilistic edge decay · ICDE 2015 |
Distributed and cloud data management › mapreduce
mapreduce query processing |
0.2 | 1 | 2015 | Groupwise analytics via adaptive MapReduce · ICDE 2015 |
Web and social media mining
social influence analysis |
0.2 | 1 | 2014 | Scalable topic-specific influence analysis on microblogs · WSDM 2014 |
Query processing and optimization
OLAP |
0.2 | 2 | 2008 | DBPubs: multidimensional exploration of database publications · Proc. VLDB Endow. 2008 Towards keyword-driven analytical processing · SIGMOD Conference 2007 |
Recommender systems › collaborative filtering
matrix factorization |
0.1 | 1 | 2011 | Large-scale matrix factorization with distributed stochastic gradient descent · KDD 2011 |
Mathematical optimization › stochastic optimization › stochastic gradient methods
stochastic gradient descent |
0.1 | 1 | 2011 | Large-scale matrix factorization with distributed stochastic gradient descent · KDD 2011 |
Mathematical optimization
stochastic optimization |
0.1 | 1 | 2011 | Large-scale matrix factorization with distributed stochastic gradient descent · KDD 2011 |
Data integration and cleaning
entity resolution |
0.1 | 1 | 2009 | Resolution-Aware Query Answering for Business Intelligence · ICDE 2009 |
Query processing and optimization › OLAP › data cube
data cube computation |
0.1 | 2 | 2004 | The Complexity of Fully Materialized Coalesced Cubes · VLDB 2004 Dwarf: shrinking the PetaCube · SIGMOD Conference 2002 |
Data mining
clustering |
0.1 | 1 | 2008 | Discovering topical structures of databases · SIGMOD Conference 2008 |
Data mining › clustering › ensemble clustering
meta-clustering |
0.1 | 1 | 2008 | Discovering topical structures of databases · SIGMOD Conference 2008 |
Information retrieval
search engines |
0.1 | 1 | 2008 | DBPubs: multidimensional exploration of database publications · Proc. VLDB Endow. 2008 |
Data mining › clustering
table clustering |
0.1 | 1 | 2008 | Discovering topical structures of databases · SIGMOD Conference 2008 |
Query processing and optimization
cardinality estimation |
0.1 | 1 | 2007 | On synopses for distinct-value estimation under multiset operations · SIGMOD Conference 2007 |
Query processing and optimization › cardinality estimation
distinct element counting |
0.1 | 1 | 2007 | On synopses for distinct-value estimation under multiset operations · SIGMOD Conference 2007 |
Distributed systems
distributed graph processing |
0.1 | 1 | 2015 | Dynamic interaction graphs with probabilistic edge decay · ICDE 2015 |
Data stream processing
stream summarization |
0.1 | 1 | 2005 | Maintaining Implicated Statistics in Constrained Environments · ICDE 2005 |
Indexing and storage engines
compressed data structures |
0.0 | 1 | 2002 | Dwarf: shrinking the PetaCube · SIGMOD Conference 2002 |
Query processing and optimization › OLAP
data cube |
0.0 | 1 | 2002 | Dwarf: shrinking the PetaCube · SIGMOD Conference 2002 |
High-performance computing › data-intensive computing
large-scale data analytics |
0.0 | 1 | 2010 | Ricardo: integrating R and Hadoop · SIGMOD Conference 2010 |
High-performance computing
scientific computing systems |
0.0 | 1 | 2010 | Ricardo: integrating R and Hadoop · SIGMOD Conference 2010 |
Query processing and optimization
aggregate query processing |
0.0 | 1 | 2009 | Resolution-Aware Query Answering for Business Intelligence · ICDE 2009 |
Web and social media mining
citation network analysis |
0.0 | 1 | 2008 | DBPubs: multidimensional exploration of database publications · Proc. VLDB Endow. 2008 |
Data integration and cleaning
data warehouse |
0.0 | 1 | 1999 | The Active MultiSync Controller of the Cubetree Storage Organization · SIGMOD Conference 1999 |
Indexing and storage engines
multidimensional indexing |
0.0 | 1 | 1999 | The Active MultiSync Controller of the Cubetree Storage Organization · SIGMOD Conference 1999 |
Data integration and cleaning
dependency discovery |
0.0 | 1 | 2006 | GORDIAN: Efficient and Scalable Discovery of Composite Keys · VLDB 2006 |
Network measurement and analytics
traffic characterization |
0.0 | 1 | 2005 | Maintaining Implicated Statistics in Constrained Environments · ICDE 2005 |
Computational complexity
query complexity |
0.0 | 1 | 2004 | The Complexity of Fully Materialized Coalesced Cubes · VLDB 2004 |
Methods — techniques the papers use, named apart from their topics
sampling · 0.4incremental updating · 0.4incremental sample generation · 0.4data structure sharing · 0.4bulk execution · 0.4adaptive mapreduce · 0.4regenerative process theory · 0.4stochastic approximation theory · 0.2network structure analysis · 0.2content analysis · 0.2OLAP rollup-drilldown · 0.2mapreduce · 0.1algorithm decomposition · 0.1randomized algorithm · 0.1memory-constrained monitoring · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Making Data Engineering Declarative
Michael Armbrust, Ali Ghodsi 0002, Reynold Xin, Vuk Ercegovac, Sourav Chatterji, Eun-Gyu Kim, Paul Lappas, Yannis Papakonstantinou, Yingyi Bu, Yijia Cui, Rahul Govind, Aakash Japi, Kiavash Kianfar, Jon Mio, Mukul Murthy, Supun Nakandala, Yannis Sismanis, Justin Tang, Joseph Torres |
CIDR | 21 |
| 2015 | Groupwise analytics via adaptive MapReduceabstractShared-nothing systems such as Hadoop vastly simplify parallel programming when processing disk-resident data whose size exceeds aggregate cluster memory. Such systems incur a significant performance penalty, however, on the important class of “groupwise set-valued analytics” (GSVA) queries in which the data is dynamically partitioned into groups and then a set-valued synopsis is computed for some or all of the groups. Key examples of synopses include top-k sets, bottom-k sets, and uniform random samples. Applications of GSVA queries include micro-marketing, root-cause analysis for problem diagnosis, and fraud detection. A naive approach to executing GSVA queries first reshuffles all of the data so that all records in a group are at the same node and then computes the synopsis for the group. This approach can be extremely inefficient when, as is typical, only a very small fraction of the records in each group actually contribute to the final groupwise synopsis, so that most of the shuffling effort is wasted. We show how to significantly speed up GSVA queries by slightly modifying the shared-nothing environment to allow tasks to occasionally access a small, common data structure; we focus on the Hadoop setting and use the “Adaptive MapReduce” infrastructure of Vernica et al. to implement the data structure. Our approach retains most of the advantages of a system such as Hadoop while significantly improving GSVA query performance, and also allows for incremental updating of query results. Experiments show speedups of up to 5x. Importantly, our new technique can potentially be applied to other shared-nothing systems with disk-resident data. Liping Peng, Vuk Ercegovac, Kai Zeng 0002, Peter J. Haas, Andrey Balmin, Yannis Sismanis |
ICDE | 6 |
| 2015 | Dynamic interaction graphs with probabilistic edge decayabstractA large scale network of social interactions, such as mentions in Twitter, can often be modeled as a “dynamic interaction graph” in which new interactions (edges) are continually added over time. Existing systems for extracting timely insights from such graphs are based on either a cumulative “snapshot” model or a “sliding window” model. The former model does not sufficiently emphasize recent interactions. The latter model abruptly forgets past interactions, leading to discontinuities in which, e.g., the graph analysis completely ignores historically important influencers who have temporarily gone dormant. We introduce TIDE, a distributed system for analyzing dynamic graphs that employs a new “probabilistic edge decay” (PED) model. In this model, the graph analysis algorithm of interest is applied at each time step to one or more graphs obtained as samples from the current “snapshot” graph that comprises all interactions that have occurred so far. The probability that a given edge of the snapshot graph is included in a sample decays over time according to a user specified decay function. The PED model allows controlled trade-offs between recency and continuity, and allows existing analysis algorithms for static graphs to be applied to dynamic graphs essentially without change. For the important class of exponential decay functions, we provide efficient methods that leverage past samples to incrementally generate new samples as time advances. We also exploit the large degree of overlap between samples to reduce memory consumption from O(N) to O(logN) when maintaining N sample graphs. Finally, we provide bulk-execution methods for applying graph algorithms to multiple sample graphs simultaneously without requiring any changes to existing graph-processing APIs. Experiments on a real Twitter dataset demonstrate the effectiveness and efficiency of our TIDE prototype, which is built on top of the Spark distributed computing framework. Wenlei Xie, Yuanyuan Tian 0001, Yannis Sismanis, Andrey Balmin, Peter J. Haas |
ICDE | 3 |
| 2015 | Shared-memory and shared-nothing stochastic gradient descent algorithms for matrix completion
Faraz Makari, Christina Teflioudi, Rainer Gemulla, Peter J. Haas, Yannis Sismanis |
Knowl. Inf. Syst. | 5 |
| 2014 | Scalable topic-specific influence analysis on microblogsabstractSocial influence analysis on microblog networks, such as Twitter, has been playing a crucial role in online advertising and brand management. While most previous influence analysis schemes rely only on the links between users to find key influencers, they omit the important text content created by the users. As a result, there is no way to differentiate the social influence in different aspects of life (topics). Although a few prior works do support topic-specific influence analysis, they either separate the analysis of content from the analysis of network structure, or assume that content is the only cause of links, which is clearly an inappropriate assumption for microblog networks. Bin Bi, Yuanyuan Tian 0001, Yannis Sismanis, Andrey Balmin, Junghoo Cho |
WSDM | 3 |
| 2013 | Eagle-eyed elephant: split-oriented indexing in HadoopabstractAn increasingly important analytics scenario for Hadoop involves multiple (often ad hoc) grouping and aggregation queries with selection predicates over a slowly changing dataset. These queries are typically expressed via high-level query languages such as Jaql, Pig, and Hive, and are used either directly for business-intelligence applications or to prepare the data for statistical model building and machine learning. In such scenarios it has been increasingly recognized that, as in classical databases, techniques for avoiding access to irrelevant data can dramatically improve query performance. Prior work on Hadoop, however, has simply ported classical techniques to the MapReduce setting, focusing on record-level indexing and key-based partition elimination. Unfortunately, record-level indexing only slightly improves overall query performance, because it does not minimize the number of mapper "waves", which is determined by the number of processed splits. Moreover, key-based partitioning requires data reorganization, which is usually impractical in Hadoop settings. We therefore need to re-envision how data access mechanisms are defined and implemented. To this end, we introduce the Eagle-Eyed Elephant (E3) framework for boosting the efficiency of query processing in Hadoop by avoiding accesses of data splits that are irrelevant to the query at hand. Using novel techniques involving inverted indexes over splits, domain segmentation, materialized views, and adaptive caching, E3 avoids accessing irrelevant splits even in the face of evolving workloads and data. Our experiments show that E3 can achieve up to 20x cost savings with small to moderate storage overheads. Mohamed Y. Eltabakh, Fatma Özcan 0001, Yannis Sismanis, Peter J. Haas, Hamid Pirahesh, Jan Vondrák |
EDBT | 3 |
| 2013 | Sparkler: supporting large-scale matrix factorizationabstractLow-rank matrix factorization has recently been applied with great success on matrix completion problems for applications like recommendation systems, link predictions for social networks, and click prediction for web search. However, as this approach is applied to increasingly larger datasets, such as those encountered in web-scale recommender systems like Netflix and Pandora, the data management aspects quickly become challenging and form a road-block. In this paper, we introduce a system called Sparkler to solve such large instances of low rank matrix factorizations. Sparkler extends Spark, an existing platform for running parallel iterative algorithms on datasets that fit in the aggregate main memory of a cluster. Sparkler supports distributed stochastic gradient descent as an approach to solving the factorization problem -- an iterative technique that has been shown to perform very well in practice. We identify the shortfalls of Spark in solving large matrix factorization problems, especially when running on the cloud, and solve this by introducing a novel abstraction called "Carousel Maps" (CMs). CMs are well suited to storing large matrices in the aggregate memory of a cluster and can efficiently support the operations performed on them during distributed stochastic gradient descent. We describe the design, implementation, and the use of CMs in Sparkler programs. Through a variety of experiments, we demonstrate that Sparkler is faster than Spark by 4x to 21x, with bigger advantages for larger problems. Equally importantly, we show that this can be done without imposing any changes to the ease of programming. We argue that Sparkler provides a convenient and efficient extension to Spark for solving matrix factorization problems on very large datasets. Boduo Li, Sandeep Tata, Yannis Sismanis |
EDBT | 3 |
| 2012 | CRSI: a compact randomized similarity index for set-valued featuresabstractWe propose a similarity index for set-valued features and study algorithms for executing various set similarity queries on it. Such queries are fundamental for many application areas, including data integration and cleaning, data profiling as well as near duplicate document detection. In this paper, we focus on Jaccard similarity and present estimators that work for arbitrary similarity thresholds based on a single similarity index. We show how to build this similarity index a-priori, without knowledge about query similarity thresholds, based on recently proposed synopses for multiset operations. The index is deployed using existing disk-based inverted indexing implementations and our algorithms exploit available techniques, like skip-lists, to further optimize the query performance. The index has provably small space footprints, is orders of magnitude smaller and faster to create/incrementally maintain than exact solutions, and the algorithms provide approximate answers, with an error that is controlled by a user-specified parameter. We prove the error bounds of our algorithms analytically, and, finally, we demonstrate the performance of the algorithms and verify their accuracy experimentally. Petros Venetis, Yannis Sismanis, Berthold Reinwald |
EDBT | 2 |
| 2011 | Large-scale matrix factorization with distributed stochastic gradient descentabstractWe provide a novel algorithm to approximately factor large matrices with millions of rows, millions of columns, and billions of nonzero elements. Our approach rests on stochastic gradient descent (SGD), an iterative stochastic optimization algorithm. We first develop a novel "stratified" SGD variant (SSGD) that applies to general loss-minimization problems in which the loss function can be expressed as a weighted sum of "stratum losses." We establish sufficient conditions for convergence of SSGD using results from stochastic approximation theory and regenerative process theory. We then specialize SSGD to obtain a new matrix-factorization algorithm, called DSGD, that can be fully distributed and run on web-scale datasets using, e.g., MapReduce. DSGD can handle a wide variety of matrix factorizations. We describe the practical techniques used to optimize performance in our DSGD implementation. Experiments suggest that DSGD converges significantly faster and has better scalability properties than alternative algorithms. Rainer Gemulla, Erik Nijkamp, Peter J. Haas, Yannis Sismanis |
KDD | 4 |
| 2010 | Ricardo: integrating R and HadoopabstractMany modern enterprises are collecting data at the most detailed level possible, creating data repositories ranging from terabytes to petabytes in size. The ability to apply sophisticated statistical analysis methods to this data is becoming essential for marketplace competitiveness. This need to perform deep analysis over huge data repositories poses a significant challenge to existing statistical software and data management systems. On the one hand, statistical software provides rich functionality for data analysis and modeling, but can handle only limited amounts of data; e.g., popular packages like R and SPSS operate entirely in main memory. On the other hand, data management systems - such as MapReduce-based systems - can scale to petabytes of data, but provide insufficient analytical functionality. We report our experiences in building Ricardo, a scalable platform for deep analytics. Ricardo is part of the eXtreme Analytics Platform (XAP) project at the IBM Almaden Research Center, and rests on a decomposition of data-analysis algorithms into parts executed by the R statistical analysis system and parts handled by the Hadoop data management system. This decomposition attempts to minimize the transfer of data across system boundaries. Ricardo contrasts with previous approaches, which try to get along with only one type of system, and allows analysts to work on huge datasets from within a popular, well supported, and powerful analysis environment. Because our approach avoids the need to re-implement either statistical or data-management functionality, it can be used to solve complex problems right now. Sudipto Das, Yannis Sismanis, Kevin S. Beyer, Rainer Gemulla, Peter J. Haas, John McPherson |
SIGMOD Conference | 2 |
| 2009 | Resolution-Aware Query Answering for Business IntelligenceabstractEntity uncertainty is an unavoidable problem in modern enterprise databases, resulting from integration of data over multiple sources. In traditional warehousing, the administrator, during an ETL process, manually and laboriously resolves inconsistent data records to discover "true'' entities(customers, products, etc.) and identify their "correct'' attribute values. At any time point, however, the current entity resolution is merely a best guess, and OLAP query results based on this resolution are inherently imprecise. We propose a new approach that maintains the data in an unresolved state, and dynamically deals with entity uncertainty at query time. We enhance the traditional OLAP model to return not a single query answer, but rather upper and lower bounds on each OLAP aggregate. This approach avoids expensive entity-resolution processing, and serves to identify potential risks when making business decisions based on the results of OLAP queries. By focusing on bounds, rather than probability distributions, we can easily and efficiently process roll-up and group-by aggregation queries over all of the core aggregation functions. Moreover, our approach can be readily implemented in an existing RDBMS using SQL queries, and does not require the user to specify explicit probabilities for alternative entity resolutions. Experiments show that the overhead of our new OLAP functionality is small over a wide range of scenarios. Yannis Sismanis, Ariel Fuxman, Peter J. Haas, Berthold Reinwald |
ICDE | 1 |
| 2008 | Discovering topical structures of databasesabstractThe increasing complexity of enterprise databases and the prevalent lack of documentation incur significant cost in both understanding and integrating the databases. Existing solutions addressed mining for keys and foreign keys, but paid little attention to more high-level structures of databases. In this paper, we consider the problem of discovering topical structures of databases to support semantic browsing and large-scale data integration. We describe iDisc, a novel discovery system based on a multi-strategy learning framework. iDisc exploits varied evidence in database schema and instance values to construct multiple kinds of database representations. It employs a set of base clusterers to discover preliminary topical clusters of tables from database representations, and then aggregate them into final clusters via meta-clustering. To further improve the accuracy, we extend iDisc with novel multiple-level aggregation and clusterer boosting techniques. We introduce a new measure on table importance and propose an approach to discovering cluster representatives to facilitate semantic browsing. An important feature of our framework is that it is highly extensible, where additional database representations and base clusterers may be easily incorporated into the framework. We have extensively evaluated iDisc using large real-world databases and results show that it discovers topical structures with a high degree of accuracy. Wensheng Wu, Berthold Reinwald, Yannis Sismanis, Rajesh Manjrekar |
SIGMOD Conference | 3 |
| 2008 | DBPubs: multidimensional exploration of database publicationsabstractDBPubs is a system for effectively analyzing and exploring the content of database publications by combining keyword search with OLAP-style aggregations, navigation, and reporting. DBPubs starts with keyword search over the content of publications. The publications' metadata such as title, authors, venues, year, and so on, provide traditional OLAP static dimensions, which are combined with dynamic dimensions discovered from the content of the publications in the search result, such as frequent phrases, relevant phrases, and topics. We compute publication ranks based on the link structure between documents, i.e., citations, and aggregate them to find seminal papers, discover trends, and rank authors. We deploy an OLAP tool for multidimensional content exploration through traditional OLAP rollup-drilldown operations on the static and dynamic dimensions, solutions for multi-cube analysis, dynamic navigation of the content, and highlighting of interesting dices of the multidimensional content dataspace. Akanksha Baid, Andrey Balmin, Heasoo Hwang, Erik Nijkamp, Jun Rao, Berthold Reinwald, Alkis Simitsis, Yannis Sismanis, Frank van Ham |
Proc. VLDB Endow. | 8 |
| 2008 | Multidimensional content eXplorationabstractContent Management Systems (CMS) store enterprise data such as insurance claims, insurance policies, legal documents, patent applications, or archival data like in the case of digital libraries. Search over content allows for information retrieval, but does not provide users with great insight into the data. A more analytical view is needed through analysis, aggregations, groupings, trends, pivot tables or charts, and so on. Multidimensional Content eXploration (MCX) is about effectively analyzing and exploring large amounts of content by combining keyword search with OLAP-style aggregation, navigation, and reporting. We focus on unstructured data or generally speaking documents or content with limited metadata, as it is typically encountered in CMS. We formally present how CMS content and metadata should be organized in a well-defined multidimensional structure, so that sophisticated queries can be expressed and evaluated. The CMS metadata provide traditional OLAP static dimensions that are combined with dynamic dimensions discovered from the analyzed keyword search result, as well as measures for document scores based on the link structure between the documents. In addition, we provide means for multidimensional content exploration through traditional OLAP rollupdrilldown operations on the static and dynamic dimensions, solutions for multi-cube analysis and dynamic navigation of the content. We present our prototype, called DBPubs, which stores research publications as documents that can be searched and -most importantly-- analyzed, and explored. Finally, we present experimental results of the efficiency and effectiveness of our approach. Alkis Simitsis, Akanksha Baid, Yannis Sismanis, Berthold Reinwald |
Proc. VLDB Endow. | 3 |
| 2007 | On synopses for distinct-value estimation under multiset operationsabstractThe task of estimating the number of distinct values (DVs) in a large dataset arises in a wide variety of settings in computer science and elsewhere. We provide DV estimation techniques that are designed for use within a flexible and scalable "synopsis warehouse" architecture. In this setting, incoming data is split into partitions and a synopsis is created for each partition; each synopsis can then be used to quickly estimate the number of DVs in its corresponding partition. By combining and extending a number of results in the literature, we obtain both appropriate synopses and novel DV estimators to use in conjunction with these synopses. Our synopses can be created in parallel, and can then be easily combined to yield synopses and DV estimates for arbitrary unions, intersections or differences of partitions. Our synopses can also handle deletions of individual partition elements. We use the theory of order statistics to show that our DV estimators are unbiased, and to establish moment formulas and sharp error bounds. Based on a novel limit theorem, we can exploit results due to Cohen in order to select synopsis sizes when initially designing the warehouse. Experiments and theory indicate that our synopses and estimators lead to lower computational costs and more accurate DV estimates than previous approaches. Kevin S. Beyer, Peter J. Haas, Berthold Reinwald, Yannis Sismanis, Rainer Gemulla |
SIGMOD Conference | 4 |
| 2007 | Towards keyword-driven analytical processingabstractGaining business insights from data has recently been the focus of research and product development. On Line-Analytical Processing (OLAP) tools provide elaborate query languages that allow users to group and aggregate data in various ways, and explore interesting trends and patterns in the data. However, the dynamic nature of today's data along with the overwhelming detail at which data is provided, make it nearly impossible to organize the data in a way that a business analyst needs for thinking about the data. In this paper, we introduce "Keyword-Driven Analytical Processing" (KDAP), which combines intuitive keyword-based search with the power of aggregation in OLAP without having to spend considerable effort in organizing the data in terms that the business analyst understands. Our design point is around a user mentality that we frequently encounter: "users don't know how to specify what they want, but they know it when they see it". We present our complete solution framework, which implements various phases from disambiguating the keyword terms to organizing and ranking the results in dynamic facets, that allow the user to explore efficiently the aggregation space. We address specific issues that analysts encounter, like joins, groupings and aggregations, and we provide efficient and scalable solutions. We show, how KDAP can handle both categorical and numerical data equally well and, finally, we demonstrate the generality and applicability of KDAP to two different aspects of OLAP, namely, finding exceptions or surprises in the data and finding bellwether regions where local aggregates are highly correlated with global aggregates, using various experiments on real data. Yannis Sismanis, Berthold Reinwald |
SIGMOD Conference | 2 |
| 2006 | GORDIAN: Efficient and Scalable Discovery of Composite Keys
Yannis Sismanis, Paul Brown, Peter J. Haas, Berthold Reinwald |
VLDB | 1 |
| 2005 | Maintaining Implicated Statistics in Constrained EnvironmentsabstractAggregated information regarding implicated entities is critical for online applications like network management, traffic characterization or identifying patters of resource consumption. Recently there has been a flurry of research for online aggregation on streams (like quantiles, hot items, hierarchical heavy hitters) but surprisingly the problem of summarizing implicated information in stream data has received no attention. As an example, consider an IP-network and the implication source /spl rarr/ destination. Flash crowds - such as those that follow recent sport events (like the Olympics) or seek information regarding catastrophic events - or denial of service attacks direct a large volume of traffic from a huge number of sources to a very small number of destinations. In this paper we present novel randomized algorithms for monitoring such implications with constraints in both memory and processing power for environments like network routers. Our experiments demonstrate several factors of improvements over straightforward approaches. Yannis Sismanis, Nick Roussopoulos |
ICDE | 1 |
| 2004 | The Complexity of Fully Materialized Coalesced Cubes
Yannis Sismanis, Nick Roussopoulos |
VLDB | 1 |
| 2003 | Hierarchical dwarfs for the rollup cubeabstractThe data cube operator exemplifies two of the most important aspects of OLAP queries: aggregation and dimension hierarchies. In earlier work we presented Dwarf, a highly compressed and clustered structure for creating, storing and indexing data cubes. Dwarf is a complete architecture that supports queries and updates, while also including a tunable granularity parameter that controls the amount of materialization performed. However, it does not directly support dimension hierarchies. Rollup and drilldown queries on dimension hierarchies that naturally arise in OLAP need to be handled externally and are, thus, very costly. In this paper we present extensions to the Dwarf architecture for incorporating rollup data cubes, i.e. cubes with hierarchical dimensions. We show that the extended Hierarchical Dwarf retains all its advantages both in terms of creation time and space while being able to directly and efficiently support aggregate queries on every level of a dimension's hierarchy. Yannis Sismanis, Antonios Deligiannakis, Yannis Kotidis, Nick Roussopoulos |
DOLAP | 1 |
| 2003 | Efficient Dissemination of Aggregate Data over the Wireless Web
Mohamed A. Sharaf, Yannis Sismanis, Alexandros Labrinidis, Panos K. Chrysanthis, Nick Roussopoulos |
WebDB | 2 |
| 2002 | Dwarf: shrinking the PetaCubeabstractDwarf is a highly compressed structure for computing, storing, and querying data cubes. Dwarf identifies prefix and suffix structural redundancies and factors them out by coalescing their store. Prefix redundancy is high on dense areas of cubes but suffix redundancy is significantly higher for sparse areas. Putting the two together fuses the exponential sizes of high dimensional full cubes into a dramatically condensed data structure. The elimination of suffix redundancy has an equally dramatic reduction in the computation of the cube because recomputation of the redundant suffixes is avoided. This effect is multiplied in the presence of correlation amongst attributes in the cube. A Petabyte 25-dimensional cube was shrunk this way to a 2.3GB Dwarf Cube, in less than 20 minutes, a 1:400000 storage reduction ratio. Still, Dwarf provides 100% precision on cube queries and is a self-sufficient structure which requires no access to the fact table. What makes Dwarf practical is the automatic discovery,in a single pass over the fact table, of the prefix and suffix redundancies without user involvement or knowledge of the value distributions.This paper describes the Dwarf structure and the Dwarf cube construction algorithm. Further optimizations are then introduced for improving clustering and query performance. Experiments with the current implementation include comparisons on detailed measurements with real and synthetic datasets against previously published techniques. The comparisons show that Dwarfs by far out-perform these techniques on all counts: storage space, creation time, query response time, and updates of cubes. Yannis Sismanis, Antonios Deligiannakis, Nick Roussopoulos, Yannis Kotidis |
SIGMOD Conference | 1 |
| 2001 | Shared Index Scans for Data Warehouses
Yannis Kotidis, Yannis Sismanis, Nick Roussopoulos |
DaWaK | 2 |
| 1999 | The Active MultiSync Controller of the Cubetree Storage OrganizationabstractThe Cubetree Storage Organization (CSO)1 logically and physically clusters materialized-views data, multi-dimensional indices on them, and computed aggregate values all in one compact and tight storage structure that uses a fraction of the conventional table-based space. This is a breakthrough technology for storing and accessing multi-dimensional data in terms of storage reduction, query performance and incremental bulk update speed. CSO has been extended with an Active MultiSync controller for synchronizing multiple concurrent access and continuous asynchronous online updates for a non-stop data warehouse. Nick Roussopoulos, Yannis Kotidis, Yannis Sismanis |
SIGMOD Conference | 3 |