VLDB 2026 Research / reviewers in the wild / expert
Umeshwar Dayal
dblp:76/2198
· DBLP profile ↗
117ranked-venue papers
21as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 93 · 19 first-authorArtificial intelligence and machine learning · 9 · 2 first-authorHuman-computer interaction and ubiquitous computing · 8Graphics, computer vision, multimedia, augmented reality and games · 7Software engineering, systems software and programming languages · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 1Computer networks · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
61 papers |
Data integration and cleaning · 22% Database system architecture and tuning · 21% Data mining · 16% | |
| Artificial intelligence
3 papers |
Information extraction and text analysis · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Cloud and datacenter computing · 52% Performance modeling and evaluation · 25% Embedded and real-time systems · 9% | |
| Computer graphics and multimedia
2 papers |
Visualization and visual analytics · 100% |
Topics — the 30 heaviest of 99, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Database system architecture and tuning
workload management |
0.4 | 4 | 2009 | A Testbed for Managing Dynamic Mixed Workloads · Proc. VLDB Endow. 2009 rFEED: A Mixed Workload Scheduler for Enterprise Data Warehouses · ICDE 2009 Predicting Multiple Metrics for Queries: Better Decisions Enabled by Machine Learning · ICDE 2009 |
Data integration and cleaning
data warehouse |
0.2 | 7 | 2007 | Abstract Process Data Warehousing · ICDE 2007 A Data-Warehouse/OLAP Framework for Scalable Telecommunication Tandem Traffic Analysis · ICDE 2000 Industrial Panel on Data Warehousing Technologies: Experiences, Challenges, and Directions · VLDB 1999 |
Information retrieval › text summarization
opinion summarization |
0.2 | 1 | 2013 | Ranking explanatory sentences for opinion summarization · SIGIR 2013 |
Information retrieval › ranking › text ranking
sentence ranking |
0.2 | 1 | 2013 | Ranking explanatory sentences for opinion summarization · SIGIR 2013 |
Data integration and cleaning › extract-transform-load
ETL process optimization |
0.2 | 2 | 2013 | Optimizing ETL workflows for fault-tolerance · ICDE 2010 HFMS: Managing the lifecycle and complexity of hybrid analytic data flows · ICDE 2013 |
Query processing and optimization › analytical query processing
analytic data flow optimization |
0.1 | 1 | 2012 | Optimizing analytic data flows for multiple execution engines · SIGMOD Conference 2012 |
Natural language and speech › Information extraction and text analysis › sentiment analysis
aspect-based sentiment analysis |
0.1 | 1 | 2011 | Automatic construction of a context-aware sentiment lexicon: an optimization approach · WWW 2011 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.1 | 1 | 2011 | Automatic construction of a context-aware sentiment lexicon: an optimization approach · WWW 2011 |
Natural language and speech › Information extraction and text analysis › sentiment analysis
sentiment lexicon construction |
0.1 | 1 | 2011 | Automatic construction of a context-aware sentiment lexicon: an optimization approach · WWW 2011 |
Data mining › text mining
sentiment analysis |
0.1 | 1 | 2011 | LCI: a social channel analysis platform for live customer intelligence · SIGMOD Conference 2011 |
Query processing and optimization
query optimization |
0.1 | 2 | 2010 | Data desensitization of customer data for use in optimizer performance experiments · ICDE 2010 Overview of an Ada Compatible Distributed Database Manager · SIGMOD Conference 1983 |
Data mining
pattern mining |
0.1 | 3 | 2004 | Mining Sequential Patterns by Pattern-Growth: The PrefixSpan Approach · IEEE Trans. Knowl. Data Eng. 2004 PrefixSpan: Mining Sequential Patterns by Prefix-Projected Growth · ICDE 2001 FreeSpan: frequent pattern-projected sequential pattern mining · KDD 2000 |
Data mining › pattern mining
sequential pattern mining |
0.1 | 3 | 2004 | Mining Sequential Patterns by Pattern-Growth: The PrefixSpan Approach · IEEE Trans. Knowl. Data Eng. 2004 PrefixSpan: Mining Sequential Patterns by Prefix-Projected Growth · ICDE 2001 FreeSpan: frequent pattern-projected sequential pattern mining · KDD 2000 |
Query processing and optimization
query scheduling |
0.1 | 2 | 2009 | rFEED: A Mixed Workload Scheduler for Enterprise Data Warehouses · ICDE 2009 Optimal Semijoin Schedules For Query Processing in Local Distributed Database Systems · SIGMOD Conference 1981 |
Database system architecture and tuning › database design
physical database design |
0.1 | 1 | 2009 | QuickStart: An Upfront Client-Based Design Advisor for Parallel Data Warehouses · ICDE 2009 |
Information retrieval › evaluation
query performance prediction |
0.1 | 1 | 2009 | Predicting Multiple Metrics for Queries: Better Decisions Enabled by Machine Learning · ICDE 2009 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2009 | Predicting Multiple Metrics for Queries: Better Decisions Enabled by Machine Learning · ICDE 2009 |
Data mining › pattern mining
pattern-growth |
0.0 | 1 | 2004 | Mining Sequential Patterns by Pattern-Growth: The PrefixSpan Approach · IEEE Trans. Knowl. Data Eng. 2004 |
Query processing and optimization
OLAP |
0.0 | 2 | 2000 | A Data-Warehouse/OLAP Framework for Scalable Telecommunication Tandem Traffic Analysis · ICDE 2000 Data Warehousing and OLAP for Decision Support (Tutorial) · SIGMOD Conference 1997 |
Data stream processing › document stream processing
social media stream |
0.0 | 1 | 2011 | LCI: a social channel analysis platform for live customer intelligence · SIGMOD Conference 2011 |
Visualization and visual analytics
multivariate data visualization |
0.0 | 1 | 2002 | Hierarchical Pixel Bar Charts · IEEE Trans. Vis. Comput. Graph. 2002 |
Services computing and microservices › business process management
business process integration |
0.0 | 1 | 2002 | Integrating Workflow Management Systems with Business-to-Business Interaction Standard · ICDE 2002 |
Electronic design automation
multi-objective optimization |
0.0 | 1 | 2010 | Optimizing ETL workflows for fault-tolerance · ICDE 2010 |
Embedded and real-time systems › real-time scheduling
admission control |
0.0 | 1 | 2009 | A Testbed for Managing Dynamic Mixed Workloads · Proc. VLDB Endow. 2009 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.0 | 1 | 2009 | A Testbed for Managing Dynamic Mixed Workloads · Proc. VLDB Endow. 2009 |
Data mining
multidimensional data analysis |
0.0 | 1 | 2000 | A Data-Warehouse/OLAP Framework for Scalable Telecommunication Tandem Traffic Analysis · ICDE 2000 |
Database system architecture and tuning
active database |
0.0 | 2 | 1991 | Active Database Systems (Abstract) · VLDB 1991 The Architecture Of An Active Data Base Management System · SIGMOD Conference 1989 |
Information retrieval
document retrieval |
0.0 | 1 | 1995 | Join Queries with External Text Sources: Execution and Optimization Techniques · SIGMOD Conference 1995 |
Query processing and optimization › query optimization › join ordering
join optimization |
0.0 | 1 | 1995 | Reducing Multidatabase Query Response Time by Tree Balancing · SIGMOD Conference 1995 |
Query processing and optimization
join processing |
0.0 | 1 | 1995 | Join Queries with External Text Sources: Execution and Optimization Techniques · SIGMOD Conference 1995 |
Methods — techniques the papers use, named apart from their topics
machine learning · 0.3string desensitization · 0.2numeric desensitization · 0.2regression · 0.2policy controller · 0.2automatic tuning · 0.2regularization · 0.1optimization · 0.1simulation · 0.1heuristics · 0.1cost calculation · 0.1value discretization · 0.1cell mapping · 0.1pseudo projection · 0.0workflow technology · 0.0pixel-based encoding · 0.0hierarchical partitioning · 0.0planning · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Advanced visual analytics interfaces for adverse drug event detectionabstractAdverse reactions to drugs are a major public health care issue. Currently, the Food and Drug Administration (FDA) publishes quarterly reports that typically contain on the order of 200,000 adverse incidents. In such numerous incidents, low frequency events that are clinically highly significant often remain undetected. In this paper, we introduce a visual analytics system to solve this problem using (1) high scalable interfaces for analyzing correlations between a number of complex variables (e.g., drug and reaction); (2) enhanced statistical computations and interactive relevance filters to quickly identify significant events including those with a low frequency; and (3) a tight integration of expert knowledge for detecting and validating adverse drug events. We applied these techniques to the FDA Adverse Event Reporting System and were able to identify important adverse drug events, such as the known association of the drug Avandia with myocardial infarction and Seroquel with diabetes mellitus, as well as low frequency events such as the association of Boniva with femur fracture. In our evaluation, we found over 90% of the adverse drug events that were published in the Institute for Safe Medication Practices (ISMP) reports from 2009 to 2012. In addition, our domain expert was able to identify some previously unknown adverse drug events. Sebastian Mittelstädt, Ming C. Hao, Umeshwar Dayal, Meichun Hsu, Joseph Terdiman, Daniel A. Keim |
AVI | 3 |
| 2013 | Compact explanatory opinion summarizationabstractIn this paper, we propose a novel opinion summarization problem called compact explanatory opinion summarization (CEOS) which aims to extract within-sentence explanatory text segments from input opinionated texts to help users better understand the detailed reasons of sentiments. We propose and study general methods for identifying candidate boundaries and scoring the explanatoriness of text segments using Hidden Markov Models. We create new data sets and use a new evaluation measure to evaluate CEOS. Experimental results show that the proposed methods are effective for generating an explanatory opinion summary, outperforming a standard text summarization method. Hyun Duk Kim, Malú Castellanos, Meichun Hsu, ChengXiang Zhai, Umeshwar Dayal, Riddhiman Ghosh |
CIKM | 5 |
| 2013 | HFMS: Managing the lifecycle and complexity of hybrid analytic data flowsabstractTo remain competitive, enterprises are evolving their business intelligence systems to provide dynamic, near realtime views of business activities. To enable this, they deploy complex workflows of analytic data flows that access multiple storage repositories and execution engines and that span the enterprise and even outside the enterprise. We call these multi-engine flows hybrid flows. Designing and optimizing hybrid flows is a challenging task. Managing a workload of hybrid flows is even more challenging since their execution engines are likely under different administrative domains and there is no single point of control. To address these needs, we present a Hybrid Flow Management System (HFMS). It is an independent software layer over a number of independent execution engines and storage repositories. It simplifies the design of analytic data flows and includes optimization and executor modules to produce optimized executable flows that can run across multiple execution engines. HFMS dispatches flows for execution and monitors their progress. To meet service level objectives for a workload, it may dynamically change a flow's execution plan to avoid processing bottlenecks in the computing infrastructure. We present the architecture of HFMS and describe its components. To demonstrate its potential benefit, we describe performance results for running sample batch workloads with and without HFMS. The ability to monitor multiple execution engines and to dynamically adjust plans enables HFMS to provide better service guarantees and better system utilization. Alkis Simitsis, Kevin Wilkinson, Umeshwar Dayal, Meichun Hsu |
ICDE | 3 |
| 2013 | Ranking explanatory sentences for opinion summarizationabstractWe introduce a novel sentence ranking problem called explanatory sentence extraction (ESE) which aims to rank sentences in opinionated text based on their usefulness for helping users understand the detailed reasons of sentiments (i.e., "explanatoriness"). We propose and study several general methods for scoring the explanatoriness of a sentence. We create new data sets and propose a new measure for evaluation. Experiment results show that the proposed methods are effective, outperforming a state of the art sentence ranking method for standard text summarization. Hyun Duk Kim, Malú Castellanos, Meichun Hsu, ChengXiang Zhai, Umeshwar Dayal, Riddhiman Ghosh |
SIGIR | 5 |
| 2013 | Hybrid Analytic Flows - the Case for OptimizationabstractTo remain competitive, enterprises are evolving in order to quickly respond to changing market conditions and customer needs. In this new environment, a single centralized data warehouse is no longer sufficient. Next generation business intelligence Alkis Simitsis, Kevin Wilkinson, Umeshwar Dayal |
Fundam. Informaticae | 3 |
| 2012 | Intention insider: discovering people's intentions in the social channelabstractThe rapid proliferation of online forums has made it possible for people to share their intentions, wishes and experiences by posting comments with the aim of getting advice from other members of the forum. Extracting intentions from these comments provides valuable insight for companies who can exploit it to get a competitive edge. However, given the very large amount of this kind of online comments, manually extracting intentions is impractical, time consuming and expensive. Companies need tools that analyze the text to extract intentions and details about them. In this paper we propose to demo one such tool called Intention Insider which has been developed at HP Labs in close collaboration with business units and a few selected customers. The tool can ingest content from online forums or from uploaded files and quickly sift through very large amounts of comments to extract intention information. This information is loaded into a data warehouse to be correlated with other structured data and queried to produce interactive reports and dynamic visualizations that facilitate its exploration at detailed and aggregate levels. Malú Castellanos, Meichun Hsu, Umeshwar Dayal, Riddhiman Ghosh, Mohamed Dekhil, Carlos Ceja Limon, Marcial Puchi, Perla Ruiz |
EDBT | 3 |
| 2012 | Of Cubes, DAGs and Hierarchical Correlations: A Novel Conceptual Model for Analyzing Social Media Data
Umeshwar Dayal, Chetan Gupta 0001, Malú Castellanos, Song Wang 0001, Manolo García-Solaco |
ER | 1 |
| 2012 | Optimizing analytic data flows for multiple execution enginesabstractNext generation business intelligence involves data flows that span different execution engines, contain complex functionality like data/text analytics, machine learning operations, and need to be optimized against various objectives. Creating correct analytic data flows in such an environment is a challenging task and is both labor-intensive and time-consuming. Optimizing these flows is currently an ad-hoc process where the result is largely dependent on the abilities and experience of the flow designer. Our previous work addressed analytic flow optimization for multiple objectives over a single execution engine. This paper focuses on optimizing flows for a single objective, namely performance, over multiple execution engines. We consider flows that span a DBMS, a Map-Reduce engine, and an orchestration engine (e.g., an ETL tool or scripting language). This configuration is emerging as a common paradigm used to combine analysis of unstructured data with analysis of structured data (e.g., NoSQL plus SQL). We present flow transformations that model data shipping, function shipping, and operation decomposition and we describe how flow graphs are generated for multiple engines. Performance results for various configurations demonstrate the benefit of optimization. Alkis Simitsis, Kevin Wilkinson, Malú Castellanos, Umeshwar Dayal |
SIGMOD Conference | 4 |
| 2012 | Optimizing Flows for Real Time Operations Management
Alkis Simitsis, Chetan Gupta 0001, Kevin Wilkinson, Umeshwar Dayal |
SSDBM | 4 |
| 2012 | A platform for situational awareness in operational BI
Malú Castellanos, Chetan Gupta 0001, Song Wang 0001, Umeshwar Dayal, Miguel Durazo |
Decis. Support Syst. | 4 |
| 2012 | Feature-Based Visual Sentiment Analysis of Text Document StreamsabstractThis article describes automatic methods and interactive visualizations that are tightly coupled with the goal to enable users to detect interesting portions of text document streams. In this scenario the interestingness is derived from the sentiment, temporal density, and context coherence that comments about features for different targets (e.g., persons, institutions, product attributes, topics, etc.) have. Contributions are made at different stages of the visual analytics pipeline, including novel ways to visualize salient temporal accumulations for further exploration. Moreover, based on the visualization, an automatic algorithm aims to detect and preselect interesting time interval patterns for different features in order to guide analysts. The main target group for the suggested methods are business analysts who want to explore time-stamped customer feedback to detect critical issues. Finally, application case studies on two different datasets and scenarios are conducted and an extensive evaluation is provided for the presented intelligent visual interface for feature-based sentiment exploration over time. Christian Rohrdantz, Ming C. Hao, Umeshwar Dayal, Lars-Erik Haug, Daniel A. Keim |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2011 | Modeling the Performance of the Hadoop Online PrototypeabstractMapReduce is an important paradigm to support modern data-intensive applications. In this paper we address the challenge of modeling performance of one implementation of MapReduce called Hadoop Online Prototype (HOP), with a specific target on the intra-job pipeline parallelism. We use a hierarchical model that combines a precedence model and a queuing network model to capture the intra-job synchronization constraints. We first show how to build a precedence graph that represents the dependencies among multiple tasks of the same job. We then apply it jointly with an approximate Mean Value Analysis (aMVA) solution to predict mean job response time and resource utilization. We validate our solution against a queuing network simulator in various scenarios, finding that our performance model presents a close agreement, with maximum relative difference under 15%. Emanuel Vianna, Giovanni Comarela, Tatiana Pontes, Jussara M. Almeida, Virgílio A. F. Almeida, Kevin Wilkinson, Harumi A. Kuno, Umeshwar Dayal |
SBAC-PAD | 8 |
| 2011 | LCI: a social channel analysis platform for live customer intelligenceabstractThe rise of Web 2.0 with its increasingly popular social sites like Twitter, Facebook, blogs and review sites has motivated people to express their opinions publicly and more frequently than ever before. This has fueled the emerging field known as sentiment analysis whose goal is to translate the vagaries of human emotion into hard data. LCI is a social channel analysis platform that taps into what is being said to understand the sentiment with the particular ability of doing so in near real-time. LCI integrates novel algorithms for sentiment analysis and a configurable dashboard with different kinds of charts including dynamic ones that change as new data is ingested. LCI has been researched and prototyped at HP Labs in close interaction with the Business Intelligence Solutions (BIS) Division and a few customers. This paper presents an overview of the architecture and some of its key components and algorithms, focusing in particular on how LCI deals with Twitter and illustrating its capabilities with selected use cases. Malú Castellanos, Umeshwar Dayal, Meichun Hsu, Riddhiman Ghosh, Mohamed Dekhil, Yue Lu 0002, Mark Schreiman |
SIGMOD Conference | 2 |
| 2011 | Automatic construction of a context-aware sentiment lexicon: an optimization approachabstractThe explosion of Web opinion data has made essential the need for automatic tools to analyze and understand people's sentiments toward different topics. In most sentiment analysis applications, the sentiment lexicon plays a central role. However, it is well known that there is no universally optimal sentiment lexicon since the polarity of words is sensitive to the topic domain. Even worse, in the same domain the same word may indicate different polarities with respect to different aspects. For example, in a laptop review, "large" is negative for the battery aspect while being positive for the screen aspect. In this paper, we focus on the problem of learning a sentiment lexicon that is not only domain specific but also dependent on the aspect in context given an unlabeled opinionated text collection. We propose a novel optimization framework that provides a unified and principled way to combine different sources of information for learning such a context-dependent sentiment lexicon. Experiments on two data sets (hotel reviews and customer feedback surveys on printers) show that our approach can not only identify new sentiment words specific to the given domain but also determine the different polarities of a word depending on the aspect in context. In further quantitative evaluation, our method is proved to be effective in constructing a high quality lexicon by comparing with a human annotated gold standard. In addition, using the learned context-dependent sentiment lexicon improved the accuracy in an aspect-level sentiment classification task. Yue Lu 0002, Malú Castellanos, Umeshwar Dayal, ChengXiang Zhai |
WWW | 3 |
| 2011 | A Visual Analytics Approach for Peak-Preserving Prediction of Large Seasonal Time SeriesabstractAbstract Time series prediction methods are used on a daily basis by analysts for making important decisions. Most of these methods use some variant of moving averages to reduce the number of data points before prediction. However, to reach a good prediction in certain applications (e.g., power consumption time series in data centers) it is important to preserve peaks and their patterns. In this paper, we introduce automated peak‐preserving smoothing and prediction algorithms, enabling a reliable long term prediction for seasonal data, and combine them with an advanced visual interface: (1) using high resolution cell‐based time series to explore seasonal patterns, (2) adding new visual interaction techniques (multi‐scaling, slider, and brushing & linking) to incorporate human expert knowledge, and (3) providing both new visual accuracy color indicators for validating the predicted results and certainty bands communicating the uncertainty of the prediction. We have integrated these techniques into a well‐fitted solution to support the prediction process, and applied and evaluated the approach to predict both power consumption and server utilization in data centers with 70–80% accuracy. Ming C. Hao, Halldór Janetzko, Sebastian Mittelstädt, W. Hill, Umeshwar Dayal, Daniel A. Keim, Manish Marwah, Ratnesh K. Sharma |
Comput. Graph. Forum | 5 |
| 2011 | Guest editorial: Business process management
Umeshwar Dayal, Johann Eder, Hajo A. Reijers |
Data Knowl. Eng. | 1 |
| 2010 | Leveraging Business Process Models for ETL Design
Kevin Wilkinson, Alkis Simitsis, Malú Castellanos, Umeshwar Dayal |
ER | 4 |
| 2010 | Data desensitization of customer data for use in optimizer performance experimentsabstractImproving the performance and functionality of database system optimizers requires experimentation on real customer data. Often these data are of sensitive nature and the only way to keep them is by applying a non-reversible transformation to obfuscate them. However, in order that the database optimizer generates exactly the same query plans as for the sensitive data, the transformation has to preserve the order and some important properties of the data distribution. Unfortunately, existing data obfuscation techniques do not preserve all of these properties and therefore are not applicable in this context. In this paper we present a Desensitizer tool that we have developed for optimizer performance experiments of HP's Neoview high availability data warehousing product. The tool is based on novel numeric and string desensitization algorithms which are agnostic to the database system. We explain the core concepts behind the algorithms, how they preserve the required data properties and important implementation considerations that were made. We present the architecture of the Desensitizer tool and results of the extensive validation that we conducted. Malú Castellanos, Bin Zhang 0004, Ivo Jimenez, Perla Ruiz, Miguel Durazo, Umeshwar Dayal, Lily Jow |
ICDE | 6 |
| 2010 | Optimizing ETL workflows for fault-toleranceabstractExtract-Transform-Load (ETL) processes play an important role in data warehousing. Typically, design work on ETL has focused on performance as the sole metric to make sure that the ETL process finishes within an allocated time window. However, other quality metrics are also important and need to be considered during ETL design. In this paper, we address ETL design for performance plus fault-tolerance and freshness. There are many reasons why an ETL process can fail and a good design needs to guarantee that it can be recovered within the ETL time window. How to make ETL robust to failures is not trivial. There are different strategies that can be used and they each have different costs and benefits. In addition, other metrics can affect the choice of a strategy; e.g., higher freshness reduces the time window for recovery. The design space is too large for informal, ad-hoc approaches. In this paper, we describe our QoX optimizer that considers multiple design strategies and finds an ETL design that satisfies multiple objectives. In particular, we define the optimizer search space, cost functions, and search algorithms. Also, we illustrate its use through several experiments and we show that it produces designs that are very near optimal. Alkis Simitsis, Kevin Wilkinson, Umeshwar Dayal, Malú Castellanos |
ICDE | 3 |
| 2010 | SIE-OBI: a streaming information extraction platform for operational business intelligenceabstractEmerging business intelligence (BI) applications aim to provide situational awareness, i.e., information about real-world events that might affect the business operations of an enterprise. For instance, an enterprise might want to know whether customers are posting positive or negative comments about a new product it has just introduced; or whether some natural disaster affects its contracted suppliers. It is difficult to develop such applications today because they require extracting and correlating facts from multiple streaming and stored data sources, typically including unstructured data, which is not well supported by BI platforms today. In this paper, we describe SIE-OBI, a system that we are developing to enable the development and execution of such applications. We describe the novel features of this system, including a declarative interface for rapidly developing such applications, and a platform for optimizing and executing the applications. We illustrate its applicability through two use cases. Malú Castellanos, Song Wang 0001, Umeshwar Dayal, Chetan Gupta 0001 |
SIGMOD Conference | 3 |
| 2010 | Application of Visual Analytics for Thermal State Management in Large Data CentresabstractAbstract Today's large data centres are the computational hubs of the next generation of IT services. With the advent of dynamic smart cooling and rack level sensing, the need for visual data exploration is growing. If administrators know the rack level thermal state changes and catch problems in real time, energy consumption can be greatly reduced. In this paper, we apply a cell‐based spatio‐temporal overall view with high‐resolution time series to simultaneously analyze complex thermal state changes over time across hundreds of racks. We employ cell‐based visualization techniques for trouble shooting and abnormal state detection. These techniques are based on the detection of sensor temperature relations and events to help identify the root causes of problems. In order to optimize the data centre cooling system performance, we derive new non‐overlapped scatter plots to visualize the correlations between the temperatures and chiller utilization. All these techniques have been used successfully to monitor various time‐critical thermal states in real‐world large‐scale production data centres and to derive cooling policies. We are starting to embed these visualization techniques into a handheld device to add mobile monitoring capability. Ming C. Hao, Ratnesh K. Sharma, Daniel A. Keim, Umeshwar Dayal, Chandrakant D. Patel, Ravigopal Vennelakanti |
Comput. Graph. Forum | 4 |
| 2010 | Special issue on conceptual modeling - 28th International Conference on Conceptual Modeling (ER 2009)
Alberto H. F. Laender, Silvana Castano, Umeshwar Dayal |
Data Knowl. Eng. | 3 |
| 2009 | A Collaboration and Productiveness Analysis of the BPM Community
Hajo A. Reijers, Minseok Song 0001, Heidi L. Romero, Umeshwar Dayal, Johann Eder, Jana Koehler |
BPM | 4 |
| 2009 | Automating the loading of business process data warehousesabstractBusiness processes drive the operations of an enterprise. In the past, the focus was primarily on business process design, modeling, and automation. Recently, enterprises have realized that they can benefit tremendously from analyzing the behavior of their business processes with the objective of optimizing or improving them. In our research, we address the problem of warehousing business process execution data so that we can analyze their behavior using the analytic and reporting tools that are available in data warehouse environments. We build upon our previous work that described the design and implementation of a generic process data warehouse for use with any business processes. In this paper, we show how to automate the population of the generic process warehouse by tracking business events from an application environment. Typically, the source data consists of event streams that Malú Castellanos, Alkis Simitsis, Kevin Wilkinson, Umeshwar Dayal |
EDBT | 4 |
| 2009 | Data integration flows for business intelligenceabstractBusiness Intelligence (BI) refers to technologies, tools, and practices for collecting, integrating, analyzing, and presenting large volumes of information to enable better decision making. Today's BI architecture typically consists of a data warehouse (or one or more data marts), which consolidates data from several operational databases, and serves a variety of front-end querying, reporting, and analytic tools. The back-end of the architecture is a data integration pipeline for populating the data warehouse by extracting data from distributed and usually heterogeneous operational sources; cleansing, integrating and transforming the data; and loading it into the data warehouse. Since BI systems have been used primarily for off-line, strategic decision making, the traditional data integration pipeline is a oneway, batch process, usually implemented by extract-transform-load (ETL) tools. The design and implementation of the ETL pipeline is largely a labor-intensive activity, and typically consumes a large fraction of the effort in data warehousing projects. Increasingly, as enterprises become more automated, data-driven, and real-time, the BI architecture is evolving to support operational decision making. This imposes additional requirements and tradeoffs, resulting in even more complexity in the design of data integration flows. These include reducing the latency so that near real-time data can be delivered to the data warehouse, extracting information from a wider variety of data sources, extending the rigidly serial ETL pipeline to more general data flows, and considering alternative physical implementations. We describe the requirements for data integration flows in this next generation of operational BI system, the limitations of current technologies, the research challenges in meeting these requirements, and a framework for addressing these challenges. The goal is to facilitate the design and implementation of optimal flows to meet business requirements. Umeshwar Dayal, Malú Castellanos, Alkis Simitsis, Kevin Wilkinson |
EDBT | 1 |
| 2009 | Fair, effective, efficient and differentiated scheduling in an enterprise data warehouseabstractA typical online Business Intelligence (BI) workload consists of a combination of short, less intensive queries, along with long, resource intensive queries. As such, the longest queries in a typical BI workload may take several orders of magnitude more time to execute, compared with the shortest queries in the workload. This makes it challenging to design a good Mixed Workload Scheduler (MWS). In this paper we first define the design criteria that make a 'good' MWS. We then use these criteria to design rFEED, a MWS that is fair, effective, efficient, and differentiated. We simulate real workloads and compare our rFEED MWS with models of the current best of breed commercial systems. We show that the rFEED MWS works extremely well. Chetan Gupta 0001, Abhay Mehta, Song Wang 0001, Umeshwar Dayal |
EDBT | 4 |
| 2009 | Managing long-running queriesabstractBusiness Intelligence query workloads that run against very large data warehouses contain queries whose execution times range, sometimes unpredictably, from seconds to hours. The presence of even a handful of long-running queries can significantly slow down a workload consisting of thousands of queries, creating havoc for queries that require a quick response. Long-running queries are a known problem in all commercial database products. However, we have not seen a thorough classification of long-running queries nor a systematic study of the most effective corrective actions. Stefan Krompass, Harumi A. Kuno, Janet L. Wiener, Kevin Wilkinson, Umeshwar Dayal, Alfons Kemper |
EDBT | 5 |
| 2009 | QuickStart: An Upfront Client-Based Design Advisor for Parallel Data WarehousesabstractQuickStart is a tool to automate the physical design of data warehouses for HP's Neoview system. It has been researched and prototyped at HP Labs with close interaction from Neoview design experts. It embodies heuristics and best practices of the experts to search for candidate physical features and uses cost calculations to recommend the features that result in good designs. It has some unique characteristics that differentiate it from other physical design advisors. In particular, it is the only advisor that is client-based, does not require a DBMS server installation and can work off a laptop by simply connecting to the customers' flat files. Another unique characteristic of QuickStart is that it provides the rationale for its recommendations. Malú Castellanos, Ivo Jimenez, Neal Coddington, Hansjörg Zeller, Steven Euijong Whang, Umeshwar Dayal |
ICDE | 6 |
| 2009 | Predicting Multiple Metrics for Queries: Better Decisions Enabled by Machine LearningabstractOne of the most challenging aspects of managing a very large data warehouse is identifying how queries will behave before they start executing. Yet knowing their performance characteristics - their runtimes and resource usage - can solve two important problems. First, every database vendor struggles with managing unexpectedly long-running queries. When these long-running queries can be identified before they start, they can be rejected or scheduled when they will not cause extreme resource contention for the other queries in the system. Second, deciding whether a system can complete a given workload in a given time period (or a bigger system is necessary) depends on knowing the resource requirements of the queries in that workload. We have developed a system that uses machine learning to accurately predict the performance metrics of database queries whose execution times range from milliseconds to hours. For training and testing our system, we used both real customer queries and queries generated from an extended set of TPC-DS templates. The extensions mimic queries that caused customer problems. We used these queries to compare how accurately different techniques predict metrics such as elapsed time, records used, disk I/Os, and message bytes. The most promising technique was not only the most accurate, but also predicted these metrics simultaneously and using only information available prior to query execution. We validated the accuracy of this machine learning technique on a number of HP Neoview configurations. We were able to predict individual query elapsed time within 20% of its actual time for 85% of the test queries. Most importantly, we were able to correctly identify both the short and long-running (up to two hour) queries to inform workload management and capacity planning. Archana Ganapathi, Harumi A. Kuno, Umeshwar Dayal, Janet L. Wiener, Armando Fox, Michael I. Jordan, David A. Patterson 0001 |
ICDE | 3 |
| 2009 | rFEED: A Mixed Workload Scheduler for Enterprise Data WarehousesabstractA typical online business intelligence (BI) workload consists of a combination of short, less intensive queries, along with long, resource intensive queries. As such, the longest queries in a typical BI workload may take several orders of magnitude more time to execute, compared with the shortest queries in the workload. This makes it challenging to design a good mixed workload scheduler (MWS). In this paper we first define the design criteria that make a 'good' MWS. We then use these criteria to design rFEED, a MWS that is fair, effective, efficient, and differentiated. We simulate real workloads and compare our rFEED MWS with models of the current best of breed commercial systems. We show that the rFEED MWS works extremely well. Abhay Mehta, Chetan Gupta 0001, Song Wang 0001, Umeshwar Dayal |
ICDE | 4 |
| 2009 | QoX-driven ETL design: reducing the cost of ETL consulting engagementsabstractAs business intelligence becomes increasingly essential for organizations and as it evolves from strategic to operational, the complexity of Extract-Transform-Load (ETL) processes grows. In consequence, ETL engagements have become very time consuming, labor intensive, and costly. At the same time, additional requirements besides functionality and performance need to be considered in the design of ETL processes. In particular, the design quality needs to be determined by an intricate combination of different metrics like reliability, maintenance, scalability, and others. Unfortunately, there are no methodologies, modeling languages or tools to support ETL design in a systematic, formal way for achieving these quality requirements. The current practice handles them with ad-hoc approaches only based on designers' experience. This results in either poor designs that do not meet the quality objectives or costly engagements that require several iterations to meet them. A fundamental shift that uses automation in the ETL design task is the only way to reduce the cost of these engagements while obtaining optimal designs. Towards this goal, we present a novel approach to ETL design that incorporates a suite of quality metrics, termed QoX, at all stages of the design process. We discuss the challenges and tradeoffs among QoX metrics and illustrate their impact on alternative designs. Alkis Simitsis, Kevin Wilkinson, Malú Castellanos, Umeshwar Dayal |
SIGMOD Conference | 4 |
| 2009 | Classification with Unknown Classes
Chetan Gupta 0001, Song Wang 0001, Umeshwar Dayal, Abhay Mehta |
SSDBM | 3 |
| 2009 | A Testbed for Managing Dynamic Mixed WorkloadsabstractWorkload management for operational business intelligence (BI) databases is difficult. Queries vary widely in length and objectives. Resource contention is difficult to predict and to control as dynamically-arriving, long, analyst queries compete for resources with ongoing online-transaction processing (OLTP) queries and batch report queries. Currently, administrators struggle to choose workload management policies and set their thresholds manually. The goal of our project is a software framework to make the management of such mixed workloads easier. Our framework includes a policy controller that tunes workload management policies automatically to meet workload objectives. This demonstration of our system illustrates (1) the difficulty of managing a BI database workload and (2) the benefits of tuning policies automatically and individually for each service class of queries in a workload. In addition, our demonstrator is a useful research tool for understanding how policies and a policy controller adapt as the system state changes under a mixed workload. In our demo, the participant plays the administrator and tunes the policies for a variety of difficult-to-manage workloads as they execute. These policies include admission control, scheduling, and execution control policies. We visualize the policies, the user objectives, and the load on the system components (CPUs, memory, disks) during execution, which helps the participant see whether objectives are being met and make appropriate policy decisions. At the end of each workload, the participant is given the opportunity to compare how their policies met workload objectives versus policies determined by our automatic policy controller. Stefan Krompass, Harumi A. Kuno, Janet L. Wiener, Kevin Wilkinson, Umeshwar Dayal, Alfons Kemper |
Proc. VLDB Endow. | 5 |
| 2008 | BI batch manager: a system for managing batch workloads on enterprise data-warehousesabstractModern enterprise data warehouses have complex workloads that are notoriously difficult to manage. An important problem in workload management is to run these complex workloads 'optimally'. Traditionally this problem has been studied in the OLTP (Online Transaction Processing) context where MPL (Multi Programming Level) is used as a knob to achieve optimality. However, MPL is a tricky knob in a BI (Business Intelligence) scenario, since a low MPL can easily result in underload and a high MPL can easily result in overload and 'thrashing'.In this work we present BI Batch Manager, a workload management system to run batches of queries 'optimally' on an Enterprise Data Warehouse (EDW). It is comprised of three components: an admission control component, a scheduler and an execution control component. In order to automatically avoid underload and overload, we introduce a novel execution control mechanism, PGM (Priority Gradient Multiprogramming). In PGM, a priority gradient is created for the workload, with each query running at a distinctly different priority level. We demonstrate that this stabilizes the execution of a workload across a wide operating range. We use memory as the controlling factor for our admission control policy -- admitting batches of queries such that their memory requirement equals the available memory on the system. Our scheduling policy of largest memory query as the highest priority query further stabilizes the execution.We validate our BI Batch Manager using varying workloads on a commercial, enterprise class DBMS. We show that it effectively avoids underload and overload (thrashing) and can automatically run BI workloads with 'optimal' performance. Abhay Mehta, Chetan Gupta 0001, Umeshwar Dayal |
EDBT | 3 |
| 2008 | Density Displays for Data Stream MonitoringabstractAbstract In many business applications, large data workloads such as sales figures or process performance measures need to be monitored in real‐time. The data analysts want to catch problems in flight to reveal the root cause of anomalies. Immediate actions need to be taken before the problems become too expensive or consume too many resources. In the meantime, analysts need to have the “big picture” of what the information is about. In this paper, we derive and analyze two real‐time visualization techniques for managing density displays: (1) circular overlay d isplays which visualize large volumes of data without data shift movements after the display is full, thus freeing the analyst from adjusting the mental picture of the data after each data shift; and (2) variable resolution density displays which allow users to get the entire view without cluttering. We evaluate these techniques with respect to a number of evaluation measures, such as constancy of the display and usage of display space, and compare them to conventional d isplays with periodic shifts. Our real time data monitoring system also provides advanced interactions such as a local root cause analysis for further exploration. The applications using a number of real‐world data sets show the wide applicability and usefulness of our ideas. Ming C. Hao, Daniel A. Keim, Umeshwar Dayal, Daniela Oelke, Chantal Tremblay |
Comput. Graph. Forum | 3 |
| 2008 | Guest Editors' message
Gustavo Alonso, David B. Lomet, Umeshwar Dayal |
VLDB J. | 3 |
| 2007 | Abstract Process Data WarehousingabstractThis paper discusses process analysis issues, especially in the context of outsourced processes. In particular, we describe a process warehousing solution developed at HP along with the challenges we had to face and the lessons we learned in implementing and deploying it. Fabio Casati, Malú Castellanos, Norman Salazar, Umeshwar Dayal |
ICDE | 4 |
| 2007 | Multi-Resolution Techniques for Visual Exploration of Large Time-Series DataabstractTime series are a data type of utmost importance in many domains such as business management and service monitoring. We address the problem of visualizing large time-related data sets which are difficult to visualize effectively with standard techniques given the limitations of current display devices. We propose a framework for intelligent time- and data-dependent visual aggregation of data along multiple resolution levels. This idea leads to effective visualization support for long time-series data providing both focus and context. The basic idea of the technique is that either data-dependent or application-dependent, display space is allocated in proportion to the degree of interest of data subintervals, thereby (a) guiding the user in perceiving important information, and (b) freeing required display space to visualize all the data. The automatic part of the framework can accommodate any time series analysis algorithm yielding a numeric degree of interest scale. We apply our techniques on real-world data sets, compare it with the standard visualization approach, and conclude the usefulness and scalability of the approach. Ming C. Hao, Umeshwar Dayal, Daniel A. Keim, Tobias Schreck |
EuroVis | 2 |
| 2007 | A Generic solution for Warehousing Business Process Data
Fabio Casati, Malú Castellanos, Umeshwar Dayal, Norman Salazar |
VLDB | 3 |
| 2007 | Dynamic Workload Management for Very Large Data Warehouses: Juggling Feathers and Bowling Balls
Stefan Krompass, Umeshwar Dayal, Harumi A. Kuno, Alfons Kemper |
VLDB | 2 |
| 2007 | Improving process models by discovering decision points
Sharmila Subramaniam, Vana Kalogeraki, Dimitrios Gunopulos, Fabio Casati, Malú Castellanos, Umeshwar Dayal, Mehmet Sayal |
Inf. Syst. | 6 |
| 2007 | Value-Cell Bar Charts for Visualizing Large Transaction Data SetsabstractOne of the common problems businesses need to solve is how to use large volumes of sales histories, Web transactions, and other data to understand the behavior of their customers and increase their revenues. Bar charts are widely used for daily analysis, but only show highly aggregated data. Users often need to visualize detailed multidimensional information reflecting the health of their businesses. In this paper, we propose an innovative visualization solution based on the use of value cells within bar charts to represent business metrics. The value of a transaction can be discretized into one or multiple cells: high-value transactions are mapped to multiple value cells, whereas many small-value transactions are combined into one cell. With value-cell bar charts, users can 1) visualize transaction value distributions and correlations, 2) identify high-value transactions and outliers at a glance, and 3) instantly display values at the transaction record level. Value-Cell Bar Charts have been applied with success to different sales and IT service usage applications, demonstrating the benefits of the technique over traditional charting techniques. A comparison with two variants of the well-known Treemap technique and our earlier work on Pixel Bar Charts is also included. Daniel A. Keim, Ming C. Hao, Umeshwar Dayal, Martha Lyons |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2006 | A Metric Definition, Computation, and Reporting Model for Business Operation Analysis
Fabio Casati, Malú Castellanos, Umeshwar Dayal, Ming-Chien Shan |
EDBT | 3 |
| 2006 | Enabling Outsourced Service Providers to Think Globally While Acting Locally
Kevin Wilkinson, Harumi A. Kuno, Kannan Govindarajan, Kei Yuasa, Kevin Smathers, Jyotirmaya Nanda, Umeshwar Dayal |
EDBT | 7 |
| 2005 | iBOM: A Platform for Intelligent Business Operation ManagementabstractAs IT systems become more and more complex and as business operations become increasingly automated, there is a growing need from business managers to have better control on business operations and on how these are aligned with business goals. This paper describes iBOM, a platform for business operation management developed by HP that allows users to i) analyze operations from a business perspective and manage them based on business goals; ii) define business metrics, perform intelligent analysis on them to understand causes of undesired metric values, and predict future values; iii) optimize operations to improve business metrics. A key aspect is that all this functionality is readily available almost at the click of the mouse. The description of the work proceeds from some specific requirements to the solution developed to address them. We also show that the platform is indeed general, as demonstrated by subsequent deployment domains other than finance. Malú Castellanos, Fabio Casati, Ming-Chien Shan, Umeshwar Dayal |
ICDE | 4 |
| 2005 | VisBiz: A Business Process Visualization Case StudyabstractBusiness process management involves many parameters and relationships and is modeled as complex business process workflows. A common way to analyze the process data is by using flowcharts. Visual analysis of a largescale chart, however, is too complex. In this case study, we employ a novel visualization technique, called VisBiz. VisBiz reduces data complexity by automatically analyzing operational data and abstracting the most critical parameters that influence business process. The basic idea is to select the most relevant parameters and layout them on a triple-attributes circular graph based on their relationships and user domain knowledge. VizBiz transforms the attributes to nodes and the process flows to lines. VisBiz derives a new process flow matrix to link the process of multiple circular graphs as the analyst introduces more parameters for further analysis. The results of the real-world credit card fraud study show the significant advantages of this technique in finding fraud distribution patterns and root causes of frauds. Ming C. Hao, Daniel A. Keim, Umeshwar Dayal, Jörn Schneidewind |
EuroVis | 3 |
| 2004 | Probabilistic, context-sensitive, and goal-oriented service selectionabstractIn this paper we propose a novel approach and a platform for dynamic service selection in composite Web services. The problem we try to solve is that of selecting, for each composite service execution and for each step in the execution, the service that maximizes the probability of reaching a user-defined goal. We first underline the limitations of a priori approaches based on having each service provider declare non-functional parameters and on trying to select services based on some utility functions over these parameters. Then, we propose an approach that overcomes these limitations by tackling the problem a posteriori: we analyze past executions of the composite service and build, using data mining techniques, a set of context-sensitive service selection models to be applied at each stage in the composite service execution. We show the architecture of a prototype that implements this approach and we discuss its benefits over the a priori approach. Fabio Casati, Malú Castellanos, Umeshwar Dayal, Ming-Chien Shan |
ICSOC | 3 |
| 2004 | Managing the Intelligent EnterpriseabstractOver the last few years, we have seen the transformation of the traditional monolithic enterprise, in which all operations were performed in-house, to the extended enterprise, which consists of a network of collaborating entities. Global operations, outsourcing, and increasing specialization have all contributed to this trend. One challenge facing the extended enterprise is how to reconnect the information flows and business processes that were disconnected as the enterprise disaggregated. The emergence of web services, service-oriented architectures, and business process modeling and execution standards are helping to address this challenge. Our contention is that the next phase of evolution is the rise of the intelligent enterprise, which is characterized by being able to adapt quickly to changes in its operating environment. The intelligent enterprise monitors its own business processes and its interactions with customers, partners, suppliers, and collaborators; it understands how this information relates to its business objectives; and it acts to control and optimize its operations to meet its business objectives. Decisions are made quickly and accurately to modify business processes on the fly, dynamically allocate resources, or change business partners (e.g., suppliers, service providers) and partnerships (e.g., establish new service level agreements). This talk will describe challenges in managing the business operations of an intelligent enterprise. While a plethora of tools exist for managing the IT infrastructure (servers, storage, and network resources) of the enterprise, there is little systematic support today for the closed loop management and control of business operations. We will describe technology approaches to intelligent business operations management that we are pursuing at HP Labs., the progress we have made, and some open research questions. Umeshwar Dayal |
ICWS | 1 |
| 2004 | VisBiz: A Simplified Visualization of Business OperationabstractIn this poster, we present a new technique, VisBiz, for interactively visualizing business operations. The basic idea of this technique is to visually mining relationships between important operation parameters (attributes) and to map the parameters into visualizations. VisBiz simplifies the complexity by partitioning the operation into multiple attribute circular graphs. VisBiz allows the analysis of business data as follows: Ming C. Hao, Daniel A. Keim, Umeshwar Dayal |
IEEE Visualization | 3 |
| 2004 | A Comprehensive and Automated Approach to Intelligent Business Processes Execution Analysis
Malú Castellanos, Fabio Casati, Umeshwar Dayal, Ming-Chien Shan |
Distributed Parallel Databases | 3 |
| 2004 | Mining Sequential Patterns by Pattern-Growth: The PrefixSpan ApproachabstractSequential pattern mining is an important data mining problem with broad applications. However, it is also a difficult problem since the mining may have to generate or examine a combinatorially explosive number of intermediate subsequences. Most of the previously developed sequential pattern mining methods, such as GSP, explore a candidate generation-and-test approach [R. Agrawal et al. (1994)] to reduce the number of candidates to be examined. However, this approach may not be efficient in mining large sequence databases having numerous patterns and/or long patterns. In this paper, we propose a projection-based, sequential pattern-growth approach for efficient mining of sequential patterns. In this approach, a sequence database is recursively projected into a set of smaller projected databases, and sequential patterns are grown in each projected database by exploring only locally frequent fragments. Based on an initial study of the pattern growth-based sequential pattern mining, FreeSpan [J. Han et al. (2000)], we propose a more efficient method, called PSP, which offers ordered growth and reduced projected databases. To further improve the performance, a pseudoprojection technique is developed in PrefixSpan. A comprehensive performance study shows that PrefixSpan, in most cases, outperforms the a priori-based algorithm GSP, FreeSpan, and SPADE [M. Zaki, (2001)] (a sequential pattern mining algorithm that adopts vertical data format), and PrefixSpan integrated with pseudoprojection is the fastest among all the tested algorithms. Furthermore, this mining methodology can be extended to mining sequential patterns with user-specified constraints. The high promise of the pattern-growth approach may lead to its further extension toward efficient mining of other kinds of frequent patterns, such as frequent substructures. Jian Pei 0001, Jiawei Han 0001, Behzad Mortazavi-Asl, Jianyong Wang 0001, Helen Pinto, Umeshwar Dayal, Meichun Hsu |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2004 | Guest Editors' Introduction
Elisa Bertino, Tok Wang Ling, Umeshwar Dayal |
World Wide Web | 3 |
| 2002 | Integrating Workflow Management Systems with Business-to-Business Interaction StandardabstractBusiness-to-business (B2B) e-commerce is emerging as a new market with tremendous potential. Organizations are trying to link services across organizational boundaries in order to electronically trade goods and services. Standards such as RosettaNet, CBL, ED1, OB1, and cXML, describe how electronic B2B interactions should be carried on so that dynamic trade partnerships can be established and transactions can be executed across organizations. While the development of standards is a fundamental step towards enabling e-business, the problem of linking B2B interactions with internal business processes is still a challenge. In addition, as the industry standards evolve continuously based on changing needs, organizations have to adopt new standards quickly. In this paper we describe how workflow technology can be extended in order to support B2B interactions and to link them with the internal workflows. The proposed framework can be used to speed up both the development of new business processes that support B2B interaction standards and the enhancement of the existing business processes by the addition of B2B interaction capability. Mehmet Sayal, Fabio Casati, Umeshwar Dayal, Ming-Chien Shan |
ICDE | 3 |
| 2002 | Business Process Cockpit
Mehmet Sayal, Fabio Casati, Umeshwar Dayal, Ming-Chien Shan |
VLDB | 3 |
| 2002 | Hierarchical Pixel Bar ChartsabstractSimple presentation graphics are intuitive and easy-to-use, but only show highly aggregated data. Bar charts, for example, only show a rather small number of data values and x-y-plots often have a high degree of overlap. Presentation techniques are often chosen depending on the considered data type, bar charts, for example, are used for categorical data and x-y plots are used for numerical data. We propose a combination of traditional bar charts and x-y-plots, which allows the visualization of large amounts of data with categorical and numerical data. The categorical data dimensions are used for the partitioning into the bars and the numerical data dimensions are used for the ordering arrangement within the bars. The basic idea is to use the pixels within the bars to present the detailed information of the data records. Our so-called pixel bar charts retain the intuitiveness of traditional bar charts while applying the principle of x-y charts within the bars. In many applications, a natural hierarchy is defined on the categorical data dimensions such as time, region, or product type. In hierarchical pixel bar charts, the hierarchy is exploited to split the bars for selected portions of the hierarchy. Our application to a number of real-world e-business and Web services data sets shows the wide applicability and usefulness of our new idea. Daniel A. Keim, Ming C. Hao, Umeshwar Dayal |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2001 | Multi-Dimensional Sequential Pattern MiningabstractSequential pattern mining, which finds the set of frequent subsequences in sequence databases, is an important data-mining task and has broad applications. Usually, sequence patterns are associated with different circumstances, and such circumstances form a multiple dimensional space. For example, customer purchase sequences are associated with region, time, customer group, and others. It is interesting and useful to mine sequential patterns associated with multi-dimensional information.In this paper, we propose the theme of multi-dimensional sequential pattern mining, which integrates the multidimensional analysis and sequential data mining. We also thoroughly explore efficient methods for multi-dimensional sequential pattern mining. We examine feasible combinations of efficient sequential pattern mining and multi-dimensional analysis methods, as well as develop uniform methods for high-performance mining. Extensive experiments show the advantages as well as limitations of these methods. Some recommendations on selecting proper method with respect to data set properties are drawn. Helen Pinto, Jiawei Han 0001, Jian Pei 0001, Ke Wang 0001, Umeshwar Dayal |
CIKM | 6 |
| 2001 | Conceptual Modeling for Collaborative E-business Processes
Umeshwar Dayal, Meichun Hsu |
ER | 2 |
| 2001 | E-Business Applications for Supply Chain Automation: Challenges and SolutionsabstractSupply-chain management is a crucial activity in every company. Surprisingly, today, most of the supply-chain activities are carried out manually, and IT support is often limited to having a set of (disconnected) data repositories. In addition, business-to-business (B2B) communications are performed via phone, fax or e-mail. Increasing the operational efficiency of the supply chain results in huge savings and is the key to remaining competitive or even gaining a competitive advantage. Furthermore, a more efficient supply chain also enables revenue growth, which is often impossible to sustain with the current manual operations. In this paper, we discuss the requirements and challenges for e-business applications that support supply-chain management. Then, we propose an architecture that meets the requirements and enables solutions that deliver results quickly and that evolve with the business and IT environment. Both the requirements and the architecture are the results of several different types of supply-chain automation projects in which we have been involved. Fabio Casati, Umeshwar Dayal, Ming-Chien Shan |
ICDE | 2 |
| 2001 | PrefixSpan: Mining Sequential Patterns by Prefix-Projected GrowthabstractSequential pattern mining is an important data mining problem with broad applications. It is challenging since one may need to examine a combinatorially explosive number of possible subsequence patterns. Most of the previously developed sequential pattern mining methods follow the methodology of \t which may substantially reduce the number of combinations to be examined. However, \t still encounters problems when a sequence database is large and/or when sequential patterns to be mined are numerous and/or long. In this paper, we propose a novel sequential pattern mining method, called PrefixSpan (i.e., Prefix-projected Sequential pattern mining), which explores prefixprojection in sequential pattern mining. PrefixSpan mines the complete set of patterns but greatly reduces the efforts of candidate subsequence generation. Moreover, prefix-projection substantially reduces the size of projected databases and leads to efficient processing. Our performance study shows that PrefixSpan outperforms both the -based GSP algorithm and another recently proposed method, FreeSpan, in mining large sequence databases. 1 Jian Pei 0001, Jiawei Han 0001, Behzad Mortazavi-Asl, Helen Pinto, Umeshwar Dayal, Meichun Hsu |
ICDE | 6 |
| 2001 | Towards an Architecture for Real-Time Decision Support Systems: Challenges and SolutionsabstractIn large enterprises, huge volumes of data are generated and consumed, and substantial fractions of the data change rapidly. Business managers need up-to-date information to make timely and sound business decisions. Unfortunately, conventional decision support systems do not provide the low latencies needed for decision making in this rapidly changing environment. The paper introduces the notion of real time decision support systems. It distills the requirements of such systems from two real-life IT outsourcing examples drawn from our extensive experience in developing and deploying such systems. We argue that real time decision support systems are complex because they must combine elements of several different types of technologies: enterprise integration real time systems, workflow systems, knowledge management, and data warehousing and data mining. We then describe an approach to addressing these challenges. The approach is based on the message brokering paradigm for enterprise integration, and combines this paradigm with workflow management, knowledge management, and dynamic data warehousing and analysis. We conclude with lessons learnt from building systems based on this architectural approach, and discuss some hard research problems that arise. Kemal A. Delic, Laurent Douillet, Umeshwar Dayal |
IDEAS | 3 |
| 2001 | Warehousing Workflow Data: Challenges and Opportunities
Angela Bonifati, Fabio Casati, Umeshwar Dayal, Ming-Chien Shan |
VLDB | 3 |
| 2001 | Improving Business Process Quality through Exception Understanding, Prediction, and Prevention
Daniela Grigori, Fabio Casati, Umeshwar Dayal, Ming-Chien Shan |
VLDB | 3 |
| 2001 | Business Process Coordination: State of the Art, Trends, and Open Issues
Umeshwar Dayal, Meichun Hsu, Rivka Ladin |
VLDB | 1 |
| 2001 | Are Web Services the Next Revolution in e-Commerce? (Panel)
Shalom Tsur, Serge Abiteboul, Rakesh Agrawal 0001, Umeshwar Dayal, Gerhard Weikum |
VLDB | 4 |
| 2000 | Multi-Agent Cooperative Transactions for E-Commerce
Umeshwar Dayal |
CoopIS | 2 |
| 2000 | An OLAP-based Scalable Web Access Analysis Engine
Umeshwar Dayal, Meichun Hsu |
DaWaK | 2 |
| 2000 | A Data-Warehouse/OLAP Framework for Scalable Telecommunication Tandem Traffic AnalysisabstractIn a telecommunication network, hundreds of millions of call detail records (CDRs) are generated daily. Applications such as tandem traffic analysis require the collection and mining of CDRs on a continuous basis. The data volumes and data flow rates pose serious scalability and performance challenges. This has motivated us to develop a scalable data-warehouse/OLAP framework, and based on this framework, tackle the issue of scaling the whole operation chain, including data cleansing, loading, maintenance, access and analysis. We introduce the notion of dynamic data warehousing for managing information at different aggregation levels with different life spans. We use OLAP servers, together with the associated multidimensional databases, as a computation platform for data caching, reduction and aggregation, in addition to data analysis. The framework supports parallel computation for scaling up data mining, and supports incremental OLAP for providing continuous data mining. A tandem traffic analysis engine is implemented on the proposed framework. In addition to the parallel and incremental computation architecture, we provide a set of application-specific optimization mechanisms for scaling performance. These mechanisms fit well into the above framework. Our experience demonstrates the practical value of the above framework in supporting an important class of telecommunication business intelligence applications. Meichun Hsu, Umeshwar Dayal |
ICDE | 3 |
| 2000 | FreeSpan: frequent pattern-projected sequential pattern miningabstractArticle FreeSpan: frequent pattern-projected sequential pattern mining Share on Authors: Jiawei Han Intelligent Database Systems Research Lab. School of Computing Science, Simon Fraser University, Burnaby, B.C., Canada V5A 1S6 Intelligent Database Systems Research Lab. School of Computing Science, Simon Fraser University, Burnaby, B.C., Canada V5A 1S6View Profile , Jian Pei Intelligent Database Systems Research Lab. School of Computing Science, Simon Fraser University, Burnaby, B.C., Canada V5A 1S6 Intelligent Database Systems Research Lab. School of Computing Science, Simon Fraser University, Burnaby, B.C., Canada V5A 1S6View Profile , Behzad Mortazavi-Asl Intelligent Database Systems Research Lab. School of Computing Science, Simon Fraser University, Burnaby, B.C., Canada V5A 1S6 Intelligent Database Systems Research Lab. School of Computing Science, Simon Fraser University, Burnaby, B.C., Canada V5A 1S6View Profile , Qiming Chen Hewlett-Packard Labs. 1501 Page Mills Road, P.O. Box 10490, Palo Alto, California Hewlett-Packard Labs. 1501 Page Mills Road, P.O. Box 10490, Palo Alto, CaliforniaView Profile , Umeshwar Dayal Hewlett-Packard Labs. 1501 Page Mills Road, P.O. Box 10490, Palo Alto, California Hewlett-Packard Labs. 1501 Page Mills Road, P.O. Box 10490, Palo Alto, CaliforniaView Profile , Mei-Chun Hsu Hewlett-Packard Labs. 1501 Page Mills Road, P.O. Box 10490, Palo Alto, California Hewlett-Packard Labs. 1501 Page Mills Road, P.O. Box 10490, Palo Alto, CaliforniaView Profile Authors Info & Claims KDD '00: Proceedings of the sixth ACM SIGKDD international conference on Knowledge discovery and data miningAugust 2000 Pages 355–359https://doi.org/10.1145/347090.347167Online:01 August 2000Publication History 503citation3,103DownloadsMetricsTotal Citations503Total Downloads3,103Last 12 Months105Last 6 weeks11 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Jiawei Han 0001, Jian Pei 0001, Behzad Mortazavi-Asl, Umeshwar Dayal, Meichun Hsu |
KDD | 5 |
| 1999 | A Distributed OLAP Infrastructure for E-CommerceabstractWarehousing and mining sales transaction data to generate summary information, customer profiles, and business rules has become increasingly important in e-commerce. Such summary information and rules have to be extracted from very large collections of transaction data gathered at many distributed sites. This is challenging data mining, both in terms of the magnitude of data involved, and the need to incrementally adapt the mined patterns and rules as new data is collected. This paper describes a distributed and cooperative data warehousing, OLAP, and data mining infrastructure that addresses these challenges. Our contributions are as follows. First, we define various new classes of multidimensional and multi-level association rules (scoped multidimensional, with conjoint items, and functional) that can be extracted from customer profiles and are used for e-commerce applications. Then, we show how customer profiles and different classes of association rules can be computed in a distributed, cooperative manner using OLAP tools. Finally, we show how the summaries, profiles, and rules can be incrementally updated as new transaction data is collected. This infrastructure has been prototyped at HP Labs. Umeshwar Dayal, Meichun Hsu |
CoopIS | 2 |
| 1999 | OLAP-based Scalable Profiling of Customer Behavior
Umeshwar Dayal, Meichun Hsu |
DaWaK | 2 |
| 1999 | Dynamic Data Warehousing (abstract)
Umeshwar Dayal, Meichun Hsu |
DaWaK | 1 |
| 1999 | Industrial Panel on Data Warehousing Technologies: Experiences, Challenges, and Directions
Umeshwar Dayal |
VLDB | 1 |
| 1999 | Dynamic AgentsabstractWe claim that a dynamic agent infrastructure can provide a shift from static distributed computing to dynamic distributed computing, and we have developed an infrastructure to realize such a shift. We shall compare this infrastructure with other distributed computing infrastructures such as CORBA and DCOM, and demonstrate its value in highly dynamic system integration, service provisioning and distributed applications such as data mining on the Web. The infrastructure is Java-based, light-weight, and extensible. It differs from other agent platforms and client/server infrastructures in its support of dynamic behavior modification of agents. A dynamic agent is not designed to have a fixed set of predefined functions, but instead, to carry application-specific actions, which can be loaded and modified on the fly. This allows a dynamic agent to adjust its capability to accommodate changes in the environment and requirements, and play different roles across multiple applications. The above features are supported by the light-weight, built-in management facilities of dynamic agents, which can be commonly used by the "carried" application programs to communicate, manage resources and modify their problem-solving capabilities. Therefore, the proposed infrastructure allows application-specific multi-agent systems to be developed easily on top of it, provides "nuts and bolts" for run-time system integration, and supports dynamic service construction, modification and movement. A prototype has been developed at HP Labs and made available to several external research groups. Parvathi Chundi, Umeshwar Dayal, Meichun Hsu |
Int. J. Cooperative Inf. Syst. | 3 |
| 1998 | Management of Work in Progress in Relational Information SystemsabstractIn many complex applications, there is a need to manage work-in-progress. Typically, this requires that each user has a private and non-volatile workspace, in which multiple pieces of work are finished and stored before it is appropriate for these data changes to be made globally accessible to other users. These requirements reflect current work practices in paper-based systems, which ensure security, persistence, privacy and accountability for work-in-progress. This paper describes a technique for implementing work-in-progress that requires no extensions to existing relational database management systems. The semantics of private workspace and work-in-progress are implemented by augmenting the database schema and modifying query and update operations against this augmented schema. Rafi Ahmed, Umeshwar Dayal |
CoopIS | 2 |
| 1998 | Dynamic-Agents for Dynamic Service ProvisioningabstractWe claim that a dynamic-agent infrastructure can provide a shift from static distributed computing to dynamic distributed computing, and we have developed such an infrastructure to realize such a shift. We shall show its impact on software engineering through a comparison with other distributed object-oriented systems such as CORBA and DCOM, and demonstrate its value in highly dynamic system integration and service provisioning. The infrastructure is Java-based, light-weight, and extensible. It differs from other agent platforms and client/server infrastructures in its support of dynamic behavior modification of agents. A dynamic-agent is not designed to have a fixed set of predefined functions but instead, to carry application-specific actions, which can be loaded and modified on theory. This allows a dynamic-agent to adjust its capability for accommodating environment and requirement changes, and play different roles across multiple applications. The above features are supported by the light-weight, built-in management facilities of dynamic-agents, which can be commonly used by the "carried" application programs to communicate, manage resources and modify their problem solving capabilities. Therefore, the proposed infrastructure allows application-specific multi-agent systems to be developed easily on top of it, provides "nuts and bolts" for run-time system integration, and supports dynamic service construction, modification and movement. A prototype has been developed at HP Labs and made available to several external research groups. Parvathi Chundi, Umeshwar Dayal, Meichun Hsu |
CoopIS | 3 |
| 1997 | Failure Handling for Transaction HierarchiesabstractPreviously, failure recovery mechanisms have been developed separately for nested transactions and for transactional workflows specified as "flat" flow graphs. The paper develops unified techniques for complex business processes modeled as cooperative transaction hierarchies. Multiple cooperative transaction hierarchies often have operational dependencies, thus a failure occurring in one transaction hierarchy may need to be transferred to another. The existing transaction models do not support failure handling across transaction hierarchies. The authors introduce the notion of transaction execution history tree which allows one to develop a unified hierarchical failure recovery mechanism applicable to both nested and flat transaction structures. They also develop a cross-hierarchy undo mechanism for determining failure scopes and supporting backward and forward failure recovery over multiple transaction hierarchies. These mechanisms form a structured and unified approach for handling failures in flat transactional workflows, along a transaction hierarchy, and across transaction hierarchies. Umeshwar Dayal |
ICDE | 2 |
| 1997 | Data Warehousing and OLAP for Decision Support (Tutorial)abstractOn-Line Analytical Processing (OLAP) and Data Warehousing are decision support technologies. Their goal is to enable enterprises to gain competitive advantage by exploiting the ever-growing amount of data that is collected and stored in corporate databases and files for better and faster decision making. Over the past few years, these technologies have experienced explosive growth, both in the number of products and services offered, and in the extent of coverage in the trade press. Vendors, including all database companies, are paying increasing attention to all aspects of decision support. Surajit Chaudhuri, Umeshwar Dayal |
SIGMOD Conference | 2 |
| 1996 | Commit Scope Control in Nested Transactions
Umeshwar Dayal |
EDBT | 2 |
| 1996 | Database Research: Lead, Follow, or Get Out of the Way? - Panel Abstract
Surajit Chaudhuri, Ashok K. Chandra, Umeshwar Dayal, Jim Gray 0001, Michael Stonebraker, Gio Wiederhold, Moshe Y. Vardi |
ICDE | 3 |
| 1996 | A Transactional Nested Process Management SystemabstractProviding flexible transaction semantics and incorporating activities, data and agents are the key issues in workflow system development. Unfortunately, most of the commercial workflow systems lack the advanced features of transaction models, and an individual transaction model with specific emphasis lacks sufficient coverage for business process management. This report presents our solutions to the above problems in developing the Open Process Management System (OPMS) at HP Laboratories. OPMS is based on nested activity modeling with the following extensions and constraints: 'in-process open nesting' for extending closed/open nesting to accommodate applications that require improved process-wide concurrency without sacrificing top-level atomicity; 'confined open' as a constraint on both open and in-process open activities for avoiding the semantic inconsistencies in activity triggering and compensation; and 'two-phase remedy' as a generalized hierarchical approach for handling failures. Umeshwar Dayal |
ICDE | 2 |
| 1996 | From User Access Patterns to Dynamic Hypertext Linking
Tak Woon Yan, Matthew Jacobsen, Hector Garcia-Molina, Umeshwar Dayal |
Comput. Networks | 4 |
| 1995 | Join Queries with External Text Sources: Execution and Optimization TechniquesabstractText is a pervasive information type, and many applications require querying over text sources in addition to structured data. This paper studies the problem of query processing in a system that loosely integrates an extensible database system and a text retrieval system. We focus on a class of conjunctive queries that include joins between text and structured data, in addition to selections over these two types of data. We adapt techniques from distributed query processing and introduce a novel class of join methods based on probing that is especially useful for joins with text systems, and we present a cost model for the various alternative query processing methods. Experimental results confirm the utility of these methods. The space of query plans is extended due to the additional techniques, and we describe an optimization algorithm for searching this extended space. The techniques we describe in this paper are applicable to other types of external data managers loosely integrated with a database system. Surajit Chaudhuri, Umeshwar Dayal, Tak W. Yan |
SIGMOD Conference | 2 |
| 1995 | Reducing Multidatabase Query Response Time by Tree BalancingabstractExecution of multidatabase queries differs from that of traditional queries in that sort merge and hash joins are more often favored, as nested loop join requires repeated accesses to external data sources. As a consequence, left deep join trees obtained by traditional (e.g., System-R style) optimizers for multidatabase queries are often suboptimal, with respect to response time, due to the long delay for a sort merge (or hash) join node to produce its last result after the subordinate join node did. In this paper, we present an optimization strategy that first produces an optimal left deep join tree and then reduces the response time using simple tree transformations. This strategy has the advantages of guaranteed minimum total resource usage, improved response time, and low optimization overhead. We describe a class of basic transformations that is the cornerstone of our approach. Then we present algorithms that effectively apply basic transformations to balance a left deep join tree, and discuss how the technique can be incorporated into existing query optimizers. Weimin Du, Ming-Chien Shan, Umeshwar Dayal |
SIGMOD Conference | 3 |
| 1994 | An Overview of Repository Technology
Philip A. Bernstein, Umeshwar Dayal |
VLDB | 2 |
| 1993 | Third Generation TP Monitors: A Database ChallengeabstractIn a 1976 book, “Algorithms + Data Structures = Programs” [15], Niklaus Wirth defined programs to be algorithms and data structures. Of course, by now we know that man does not live from programs alone, and that there is a second fundamental computer science equation: “Programs + Databases = Information Systems.” Umeshwar Dayal, Hector Garcia-Molina, Meichun Hsu, Ben Kao, Ming-Chien Shan |
SIGMOD Conference | 1 |
| 1993 | Control of an Extensible Query Optimizer: A Planning-Based Approach
Gail Mitchell, Umeshwar Dayal, Stanley B. Zdonik |
VLDB | 2 |
| 1993 | DBMS Research at a Crossroads: The Vienna Update
Michael Stonebraker, Rakesh Agrawal 0001, Umeshwar Dayal, Erich J. Neuhold, Andreas Reuter 0001 |
VLDB | 3 |
| 1992 | A Uniform Model for Temporal Object-Oriented DatabasesabstractA temporal object-oriented model and query language that supports the modeling and manipulation of complex temporal or versioned objects is developed. The authors show that the approach not only provides a richer model than the relational for capturing the semantics of complex temporal objects, but also requires no special constructs in the query language. Consequently, the retrieval of temporal and non-temporal information is uniformly expressed. By allowing variables and quantifiers to range over time, can be formulated that require special operators in other languages. Temporal aggregation queries, which are not easily expressed in other models, are expressed using the same aggregation operators as for nontemporal data.> Gene T. J. Wuu, Umeshwar Dayal |
ICDE | 2 |
| 1992 | A Uniform Approach to Processing Temporal Queries
Umeshwar Dayal, Gene T. J. Wuu |
VLDB | 1 |
| 1991 | Active Database Systems (Abstract)
Umeshwar Dayal, Klaus R. Dittrich |
VLDB | 1 |
| 1991 | A Transactional Model for Long-Running Activities
Umeshwar Dayal, Meichun Hsu, Rivka Ladin |
VLDB | 1 |
| 1991 | Database Technologies for the 90's and Beyond (Panel)
Magdi N. Kamel, Umeshwar Dayal, Rakesh Agrawal 0001, Douglas Tolbert, Gilbert Vidal |
VLDB | 2 |
| 1990 | Organizing Long-Running Activities with Triggers and TransactionsabstractThis paper addresses the problem of organising and controlling activities that involve multiple steps of processing and that typically are of long duration. We explore the use of triggers and transactions to specify and organize such long-running activities. Triggers offer data- or event-driven specification of control flow, and thus provide a flexible and modular framework with which the control structures of the activities can be extended or modified. We describe a model based on event-condition-action rules and coupling modes. The execution of these rules is governed by an extended nested transaction model. Through a detailed example, we illustrate the utility of the various features of the model for chaining related steps without sacrificing concurrency, for enforcing integrity constraints, and for providing flexible failure and exception handling. Umeshwar Dayal, Meichun Hsu, Rivka Ladin |
SIGMOD Conference | 1 |
| 1989 | Time-Critical Database Scheduling: A Framework For Integrating Real-Time Scheduling and Concurrency ControlabstractA framework is presented for analysis of time-critical scheduling algorithms. The main assumptions are analyzed behind real-time scheduling and concurrency control algorithms, and a unified approach is proposed. Two main classes of schedulers are identified according to the availability of information about resource requirements and execution times: conflict-resolving schedulers resolve conflicts at run-time, and hence can only produce a sequence of operations satisfying task priorities and resource constraints; and conflict-avoiding schedulers determine resource requirements and expected execution times through offline transaction-class preanalysis and produce a complete time-critical schedule satisfying both timing and resource constraints. For the latter case, the resolution of overload is essential. Examples are given to illustrate the framework and the main classes of scheduling algorithms.> Alejandro P. Buchmann, Dennis R. McCarthy, Meichun Hsu, Umeshwar Dayal |
ICDE | 4 |
| 1989 | The Architecture Of An Active Data Base Management SystemabstractThe HiPAC project is investigating active, time-constrained database management. An active DBMS is one which automatically executes specified actions when specified conditions arise. HiPAC has proposed Event-Condition-Action (ECA) rules as a formalism for active database capabilities. We have also developed an execution model that specifies how these rules are processed in the context of database transactions. The additional functionality provided by ECA rules makes new demands on the design of an active DBMS. In this paper we propose an architecture for an active DBMS that supports ECA rules. This architecture provides new forms of interaction, in support of ECA rules, between application programs and the DBMS. This leads to a new paradigm for constructing database applications. Dennis R. McCarthy, Umeshwar Dayal |
SIGMOD Conference | 2 |
| 1987 | An Object-Oriented Approach to Data Management: Why Design Databases Need ItabstractAn object-oriented approach to management of engineering design data requires object persistence, object-specific rules for concurrency control and recovery, views, complex objects and derived data, and specialized treatment of operations, constraints, relationships and type descriptions. We discuss object-orientation as more than an implementation paradigm, and show how an object-oriented approach simplifies both use and implementation of engineering design systems. Sandra Heiler, Umeshwar Dayal, Jack A. Orenstein, Susan Radke-Sproull |
DAC | 2 |
| 1987 | Of Nests and Trees: A Unified Approach to Processing Queries That Contain Nested Subqueries, Aggregates, and Quantifiers
Umeshwar Dayal |
VLDB | 1 |
| 1987 | An Ada-compatible distributed database management systemabstractAdaplex is an integrated language for programming database applications. It results from the embedding of the database sublanguage Daplex in the general-purpose programming language Ada [1]. This paper describes the design of DDM, a general-purpose distributed database management system implemented in Ada that supports the use of Adaplex as interface language. There are two novel aspects in the design of this system. First, this is the first full-scale distributed database system to support a semantically rich, functional data model. DDM goes beyond systems like Distributed INGRES and R*(which are based on the relational technology) in providing advanced data modeling capabilities and ease of use. Second, this is the first full-function distributed DBMS designed to be compatible with the Ada programming environment. The coupling between Ada and Daplex has been achieved at the expression level which is much tighter than the statement level integration attained in previous systems. This tight coupling poses new implementation problems but also creates new opportunities for optimization. The current paper highlights the Adaplex language and discusses innovative aspects in DDM's design that are intended to meet the dual objectives of good performance and high data availability. Arvola Chan, Umeshwar Dayal, Stephen Fox |
Proc. IEEE | 2 |
| 1986 | Traversal Recursion: A Practical Approach to Supporting Recursive ApplicationsabstractMany capabilities that are needed for recursive applications in engineering and project management are not well supported by the usual formulations of recursion. We identify a class of recursions called “traversal recursions” (which model traversals of a directed graph) that have two important properties they can supply the necessary capabilities and efficient processing algorithms have been defined for them. First we present a taxonomy of traversal recursions based on properties of the recursion on graph structure and on unusual types of metadata. This taxonomy is exploited to identify solvable recursions and to select an execution algorithm. We show how graph traversal can sometimes outperform the more general iteration algorithm. Finally we show how a conventional query optimizer architecture can be extended to handle recursive queries and views. Arnon Rosenthal, Sandra Heiler, Umeshwar Dayal, Frank Manola |
SIGMOD Conference | 3 |
| 1984 | Using Semiouterjoins to Process Queries in Multidatabase SystemsabstractA multidatabase system provides a logically integrated view of existing, possibly inconsistent, databases. Logical integration is achieved primarily through the use of generalization, which can be modelled algebraically as a sequence of outerjoin and aggregation operations. Conventional distributed query processing techniques are inadequate for processing queries over views defined by outerjoins and aggregates. In a conventional distributed database system, selections and projections are inexpensive to process; hence joins have been the rocus of most previous research. In a multidatabase system, however, even selections and projections can be as expensive as joins. The semiouterjoin operation can potentially reduce query processing costs. In general, there may be many different strategies based on semiouterjoins for processing a given query. The query optimization problem is to choose the most profitable of these strategies. This paper studies the query optimization problem for selection and projection queries. It develops linear-time solutions to the problem, and then extends these solutions to provide heuristics for joins and conjunctive queries. Hai-Yann Hwang, Umeshwar Dayal, Mohamed G. Gouda |
PODS | 2 |
| 1984 | View Definition and Generalization for Database Integration in a Multidatabase SystemabstractAccess to a heterogeneous distributed collection of databases can be simplified by providing users with a logically integrated interface or global view. There are two aspects to database integration. Firstly, the local schemas may model objects and relationships differently and, secondly, the databases may contain mutually inconsistent data. This paper identifies several kinds of structural and data inconsistencies that might exist. It describes a versatile view definition facility for the functional data model and illustrates the use of this facility for resolving inconsistencies. In particular, the concept of generalization is extended to this model, and its importance to database integration is emphasized. The query modification algorithm for the relational model is extended to the semantically richer functional data model with generalization. Umeshwar Dayal, Hai-Yann Hwang |
IEEE Trans. Software Eng. | 1 |
| 1983 | Issues in Multiple Database Environmemnts (Panel)
Yuri Breitbart, Umeshwar Dayal, James P. Fry, Witold Litwin, Amihai Motro |
ER | 2 |
| 1983 | Processing Queries with Quantifiers: A Horticultural ApproachabstractMost research on query processing has focussed on quantifier-free conjunctive queries. Existing techniques for processing queries with quantifiers either compile the query into a nested loop program or use variants of Codd's reduction from the Relational Calculus to the Relational Algebra. In this paper we propose an alternative technique that uses an algebra of graft and prune operations on trees. This technique provides a significant savings in space and time. We show how to transform a quantified conjunctive query into a sequence of graft and prune operations. Umeshwar Dayal |
PODS | 1 |
| 1983 | A Recovery Algorithm for a Distributed Database SystemabstractWe describe a reliability algorithm being considered for DDM, a distributed database system under development at Computer Corporation of America. The algorithm is designed to tolerate clean site failures in which sites simply stop running. The algorithm allows the system to reconfigure itself to run correctly as sites fail and recover. The algorithm solves the subproblems of atomic commit and replicated data handling in an integrated manner. Nathan Goodman, Dale Skeen, Arvola Chan, Umeshwar Dayal, Stephen Fox, Daniel R. Ries |
PODS | 4 |
| 1983 | Overview of an Ada Compatible Distributed Database ManagerabstractAdaplex is an integrated language for programming database applications. It results from the embedding of the database sublanguage DAPLEX in the general purpose programming language Ada. This paper provides an overview of the DDM: a distributed database manager (DDM) that supports the use of Adaplex as an interface language. The important technical innovations we have incorporated in the design of this system include:1. An advanced data model that captures more application semantics than conventional data models.2. Support for flexible data distribution options that improve locality of reference and efficiency of query processing.3. Extensive query optimization that combines compile time access path optimization with run time site selection.4. Efficient transaction management that reduces transaction conflicts and improves the resiliency of replicated data.5. Robust, incremental recovery management that provides for automatic recovery from certain "catastrophic" failure conditions. Arvola Chan, Umeshwar Dayal, Stephen Fox, Nathan Goodman, Daniel R. Ries, Dale Skeen |
SIGMOD Conference | 2 |
| 1983 | Supporting a Semantic Data Model in a Distributed Database System
Arvola Chan, Umeshwar Dayal, Stephen Fox, Daniel R. Ries |
VLDB | 2 |
| 1983 | Processing Queries Over Generalization Hierarchies in a Multidatabase System
Umeshwar Dayal |
VLDB | 1 |
| 1982 | An Extended Relational Algebra with Control over Duplicate EliminationabstractIn the pure relational model, duplicate tuples are automatically eliminated. Some real world languages such as DAPLEX, however, give users control over duplicate elimination. This paper extends the relational model to include multiset relations, i.e., relations with duplicate tuples. It considers three formalisms for expressing queries in this model: extended relational algebra, tableaux, and DAPLEX. It shows that, as in the original algebra, the equivalence problem for conjunctive expressions in the extended algebra can be solved using tableaux, and is NP-complete. Finally, it demonstrates that the extended algebra and DAPLEX have essentially the same expressiveness relative to conjunctive expressions. Umeshwar Dayal, Nathan Goodman, Randy H. Katz |
PODS | 1 |
| 1982 | Theory of Serializability for a Parallel Model of TransactionsabstractIn this paper we present a parallel program schema model of a transaction system and generalize the concept of serializability from the sequential two-step model to a parallel multi-step model. We define two classes of serializable executions, and for each class we discuss two problems: recognition and scheduling. It is shown that the results for the recognition and online scheduling problems for the sequential model generalize to the parallel model. But it is argued that online scheduling is not suitable for a parallel execution environment. Therefore, batch schedulers are defined and a minimal set of precedence constraints is derived. Finally, it is shown that any optimal batch scheduler that uses syntactic information alone cannot be efficient. Ravi Krishnamurthy, Umeshwar Dayal |
PODS | 2 |
| 1982 | Query Optimization for CODASYL Database SystemsabstractOne of the tasks of MULTIBASE, a system for integrated access to heterogeneous distributed databases, is to present a high-level query interface to navigational systems such as CODASYL. The interface compiles queries into efficient programs that implement the queries. The principal problem in constructing such an interface is access path optimization, i.e., the selection of an optimal sequence of access paths that must be traversed to process a given query. This paper identifies a class of queries for which efficient programs can be synthesized. It characterizes the strategies for processing a given query, and shows how to synthesize a program for implementing each strategy. It develops a model for estimating the cost of executing a program, and uses this model to find the optimal strategy for processing a given query. Umeshwar Dayal, Nathan Goodman |
SIGMOD Conference | 1 |
| 1982 | Semantics of Network Data Manipulation Languages: An Object-Oriented Approach
Dipayan Gangopadhyay, Umeshwar Dayal, James C. Browne |
VLDB | 2 |
| 1982 | On the updatability of network views-extending relational view theory to the network model
Umeshwar Dayal, Philip A. Bernstein |
Inf. Syst. | 1 |
| 1982 | On the Correct Translation of Update Operations on Relational ViewsabstractMost relational database systems provide a facility for supporting user views. Permitting this level of abstraction has the danger, however, that update requests issued by a user within the context of his view may not translate correctly into equivalent updates on the underlying database. The purpose of this paper is to formalize the notion of update translation and derive conditions under which translation procedures will produce correct translations of view updates. Umeshwar Dayal, Philip A. Bernstein |
ACM Trans. Database Syst. | 1 |
| 1981 | Using the Entity-Relationship Model for Implementing Multi-Model Database Systems
Hai-Yann Hwang, Umeshwar Dayal |
ER | 2 |
| 1981 | Optimal Semijoin Schedules For Query Processing in Local Distributed Database SystemsabstractSemijoin strategies are a technique for query processing in distributed database systems. In the past, methodologies for constructing minimum communication-cost strategies for solving tree queries have been developed. These assume point-to-point communication and ignore local processing costs and the limited communication capacity of the system. In this paper, query processing in bus or loop systems is considered. The definition of strategy is extended to allow for broadcast mode of communication. We then address the problem of finding the minimum response-time schedule for executing a given strategy in an m-bus system taking into account local processing and system capacity. It is shown that the problem is computationally intractable for general tree queries, even in a 1-bus system, and for special classes of tree queries in an m-bus system. However, there is a polynomial-time algorithm for simple queries in a 1-bus system. Mohamed G. Gouda, Umeshwar Dayal |
SIGMOD Conference | 2 |
| 1979 | Synthesizing Independent Database SchemasabstractWe study the following database design problem. Given a universal relation scheme 〈U, F〉 where F is a set of functional dependencies, find an in some way normalized database schema D = {〈X1, F1〉,..., 〈Xn, Fn〉} where Xi ⊂ U and Fi is inherited from F, such that D is an independent representation of the universal scheme 〈U, F〉. This means that D has both the lossless join property and the faithful closure property, (***** Fi)+ = F+, where + denotes the closure of a set of functional dependencies. We show that this goal can easily be achieved by an extension of the well-known synthetic approach of Bernstein and others to database design. We merely have to check whether the usual synthesis procedure has produced a key component 〈Xi, Fi〉 such that Xi → U ε F+; in case this is true the output of the synthesis procedure is actually an independent (and not only faithful) representation, otherwise we only have to add one further component, namely just a key. These claims are proved by a careful inspection of the Aho/Beeri/Ullman algorithm to test for losslessness. Finally, we show how to use our method to synthesize minimal independent third normal form schemas. Joachim Biskup, Umeshwar Dayal, Philip A. Bernstein |
SIGMOD Conference | 2 |
| 1978 | On the Updatability of Relational Views
Umeshwar Dayal, Philip A. Bernstein |
VLDB | 1 |