Shamkant B. Navathe

dblp:n/ShamkantBNavathe · DBLP profile ↗
← Back
119ranked-venue papers
27as first author
6since 2021 · last 2023
0000-0002-6666-0031ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 90 · 26 first-author · 4 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 10Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Systems, architecture and hardware · 4Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Computer networks · 3Graphics, computer vision, multimedia, augmented reality and games · 2Theory of computation · 1
YearPublicationVenuePosition
2023 DataCockpit: A Toolkit for Data Lake Navigation and Monitoring Utilizing Quality and Usage Information
abstract
Modern organizations amass their datasets into centralized repositories called data lakes, affording analytics as needed. The resultant scale and complexity of these data lakes, however, can make data navigation and monitoring challenging for users. We present DataCockpit, a Python toolkit that leverages datasets, usage logs, and associated meta-data to provision data usage and quality characteristics. DataCockpit computes these characteristics for each attribute (e.g., number of times it was queried for subsequent use in downstream applications) and record (e.g., number of non-missing, valid values) and aggregates them at the level of datasets. We develop a visual monitoring tool, powered by DataCockpit, and demonstrate how it can assist data / system administrators as well as end-users to effectively navigate and monitor a data lake. DataCockpit and the monitoring tool are available as open source software for developers to build custom monitoring applications on top of data lakes.
Arpit Narechania, Surya Chakraborty, Shivam Agarwal, Atanu R. Sinha, Ryan Rossi, Fan Du, Jane Hoffswell, Shunan Guo, Eunyee Koh, Alex Endert, Shamkant B. Navathe
IEEE Big Data11
2023 DataPilot: Utilizing Quality and Usage Information for Subset Selection during Visual Data Preparation
abstract
Selecting relevant data subsets from large, unfamiliar datasets can be difficult. We address this challenge by modeling and visualizing two kinds of auxiliary information: (1) quality – the validity and appropriateness of data required to perform certain analytical tasks; and (2) usage – the historical utilization characteristics of data across multiple users. Through a design study with 14 data workers, we integrate this information into a visual data preparation and analysis tool, DataPilot. DataPilot presents visual cues about “the good, the bad, and the ugly” aspects of data and provides graphical user interface controls as interaction affordances, guiding users to perform subset selection. Through a study with 36 participants, we investigate how DataPilot helps users navigate a large, unfamiliar tabular dataset, prepare a relevant subset, and build a visualization dashboard. We find that users selected smaller, effective subsets with higher quality and usage, and with greater success and confidence.
Arpit Narechania, Fan Du, Atanu R. Sinha, Ryan Rossi, Jane Hoffswell, Shunan Guo, Eunyee Koh, Shamkant B. Navathe, Alex Endert
CHI8
2022 SPES: A Symbolic Approach to Proving Query Equivalence Under Bag Semantics
abstract
In database-as-a-service platforms, automated ver-ification of query equivalence helps eliminate redundant computation in the form of overlapping sub-queries. Researchers have proposed two pragmatic techniques to tackle this problem. The first approach consists of reducing the queries to algebraic expressions and proving their equivalence using an algebraic theory. The limitations of this technique are threefold. It cannot prove the equivalence of queries with significant differences in the attributes of their relational operators (e.g., predicates in the filter operator). It does not support certain widely-used SQL features (e.g., NULL values). Its verification procedure is computationally intensive. The second approach transforms this problem to a constraint satisfaction problem and leverages a general-purpose solver to determine query equivalence. This technique consists of deriving the symbolic representation of the queries and proving their equivalence by determining the query containment relationship between the symbolic expressions. While the latter approach addresses all the limitations of the former technique, it only proves the equivalence of queries under set semantics (i.e., output tables must not contain duplicate tuples). However, in practice, database applications use bag semantics (i.e., output tables may contain duplicate tuples) In this paper, we introduce a novel symbolic approach for proving query equivalence under bag semantics. We transform the problem of proving query equivalence under bag semantics to that of proving the existence of a bijective, identity map between tuples returned by the queries on all valid inputs. We classify SQL queries into four categories, and propose a set of novel category-specific verification algorithms. We implement this symbolic approach in SPES and demonstrate that it proves the equivalence of a larger set of query pairs (95/232) under bag semantics compared to the SOTA tools based on algebraic (30/232) and symbolic approaches (67/232) under set and bag semantics, respectively. Furthermore, SPES is 3X faster than the symbolic tool that proves equivalence under set semantics.
Qi Zhou 0010, Joy Arulraj, Shamkant B. Navathe, William Harris, Jinpeng Wu
ICDE3
2021 SIA: Optimizing Queries using Learned Predicates
abstract
Predicate-centric rules for rewriting queries is a key technique in optimizing queries. These include pushing down the predicate below the join and aggregation operators, or optimizing the order of evaluating predicates. However, many of these rules are only applicable when the predicate uses a certain set of columns. For example, to move the predicate below the join operator, the predicate must only use columns from one of the joined tables. By generating a predicate that satisfies these column constraints and preserves the semantics of the original query, the optimizer may leverage additional predicate-centric rules that were not applicable before. Researchers have proposed syntax-driven rewrite rules and machine learning algorithms for inferring such predicates. However, these techniques suffer from two limitations. First, they do not let the optimizer constrain the set of columns that may be used in the learned predicate. Second, machine learning algorithms do not guarantee that the learned predicate preserves semantics. In this paper, we present SIA, a system for learning predicates while being guided by counter-examples and a verification technique, that addresses these limitations. The key idea is to leverage satisfiability modulo theories to generate counter-examples and use them to iteratively learn a valid, optimal predicate. We formalize this problem by proving the key properties of synthesized predicates. We implement our approach in SIA and evaluate its efficacy and efficiency. We demonstrate that it synthesizes a larger set of valid predicates compared to prior approaches. On a collection of 200 queries derived from the TPC-H benchmark, SIA successfully rewrites 114 queries with learned predicates. 66 of these rewritten queries exhibit more than 2X speed up.
Qi Zhou 0010, Joy Arulraj, Shamkant B. Navathe, William Harris, Jinpeng Wu
SIGMOD Conference3
2021 OmniFair: A Declarative System for Model-Agnostic Group Fairness in Machine Learning
abstract
Machine learning (ML) is increasingly being used to make decisions in our society. ML models, however, can be unfair to certain demographic groups (e.g., African Americans or females) according to various fairness metrics. Existing techniques for producing fair ML models either are limited to the type of fairness constraints they can handle (e.g., preprocessing) or require nontrivial modifications to downstream ML training algorithms (e.g., in-processing).
Hantian Zhang, Xu Chu 0002, Abolfazl Asudeh, Shamkant B. Navathe
SIGMOD Conference4
2021 High-utility and diverse itemset mining
Siddharth Dawar, Shamkant B. Navathe, Vikram Goyal
Appl. Intell.4
2019 The unified chart of mobility services: Towards a systemic approach to analyze service quality in smart mobility ecosystem
Antonella Longo, Marco Zappatore, Shamkant B. Navathe
J. Parallel Distributed Comput.3
2019 Automated Verification of Query Equivalence Using Satisfiability Modulo Theories
abstract
Database-as-a-service offerings enable users to quickly create and deploy complex data processing pipelines. In practice, these pipelines often exhibit significant overlap of computation due to redundant execution of certain sub-queries. It is challenging for developers and database administrators to manually detect overlap across queries since they may be distributed across teams, organization roles, and geographic locations. Thus, we require automated cloud-scale tools for identifying equivalent queries to minimize computation overlap. State-of-the-art algebraic approaches to automated verification of query equivalence suffer from two limitations. First, they are unable to model the semantics of widely-used SQL features, such as complex query predicates and three-valued logic. Second, they have a computationally intensive verification procedure. These limitations restrict their efficacy and efficiency in cloud-scale database-as-a-service offerings. This paper makes the case for an alternate approach to determining query equivalence based on symbolic representation. The key idea is to effectively transform a wide range of SQL queries into first order logic formulae and then use satisfiability modulo theories to efficiently verify their equivalence. We have implemented this symbolic representation-based approach in EQUITAS. Our evaluation shows that EQUITAS proves the semantic equivalence of a larger set of query pairs compared to algebraic approaches and reduces the verification time by 27X. We also demonstrate that on a set of 17,461 real-world SQL queries, it automatically identifies redundant execution across 11% of the queries. Our symbolic-representation based technique is currently deployed on Alibaba's MaxCompute database-as-a-service platform.
Qi Zhou 0010, Joy Arulraj, Shamkant B. Navathe, William Harris
Proc. VLDB Endow.3
2018 VIGOR: Interactive Visual Exploration of Graph Query Results
abstract
Finding patterns in graphs has become a vital challenge in many domains from biological systems, network security, to finance (e.g., finding money laundering rings of bankers and business owners). While there is significant interest in graph databases and querying techniques, less research has focused on helping analysts make sense of underlying patterns within a group of subgraph results. Visualizing graph query results is challenging, requiring effective summarization of a large number of subgraphs, each having potentially shared node-values, rich node features, and flexible structure across queries. We present VIGOR, a novel interactive visual analytics system, for exploring and making sense of query results. VIGOR uses multiple coordinated views, leveraging different data representations and organizations to streamline analysts sensemaking process. VIGOR contributes: (1) an exemplar-based interaction technique, where an analyst starts with a specific result and relaxes constraints to find other similar results or starts with only the structure (i.e., without node value constraints), and adds constraints to narrow in on specific results; and (2) a novel feature-aware subgraph result summarization. Through a collaboration with Symantec, we demonstrate how VIGOR helps tackle real-world problems through the discovery of security blindspots in a cybersecurity dataset with over 11,000 incidents. We also evaluate VIGOR with a within-subjects study, demonstrating VIGOR's ease of use over a leading graph database management system, and its ability to help analysts understand their results at higher speed and make fewer errors.
Robert S. Pienta, Fred Hohman, Alex Endert, Acar Tamersoy, Kevin A. Roundy, Christopher Gates 0002, Shamkant B. Navathe, Polo Chau
IEEE Trans. Vis. Comput. Graph.7
2017 Visual Graph Query Construction and Refinement
abstract
Locating and extracting subgraphs from large network datasets is a challenge in many domains, one that often requires learning new querying languages. We will present the first demonstration of VISAGE, an interactive visual graph querying approach that empowers analysts to construct expressive queries, without writing complex code (see our video: https://youtu.be/l2L7Y5mCh1s). VISAGE guides the construction of graph queries using a data-driven approach, enabling analysts to specify queries with varying levels of specificity, by sampling matches to a query during the analyst's interaction. We will demonstrate and invite the audience to try VISAGE on a popular film-actor-director graph from Rotten Tomatoes.
Robert S. Pienta, Fred Hohman, Acar Tamersoy, Alex Endert, Shamkant B. Navathe, Hanghang Tong, Polo Chau
SIGMOD Conference5
2017 Crowd-Sourced Data Collection for Urban Monitoring via Mobile Sensors
abstract
A considerable amount of research has addressed Internet of Things and connected communities. It is possible to exploit the sensing capabilities of connected communities, by leveraging the continuously growing use of cloud computing solutions and mobile devices. The pervasiveness of mobile sensors also enables the Mobile Crowd Sensing (MCS) paradigm, which aims at using mobile-embedded sensors to extend monitoring of multiple (environmental) phenomena in expansive urban areas. In this article, we discuss our approach with a cloud-based platform to pave the way for applying crowd sensing in urban scenarios. We have implemented a complete solution for environmental monitoring of several pollutants, like noise, air, electromagnetic fields, and so on in an urban area based on this paradigm. Through extensive experimentation, specifically on noise pollution, we show how the proposed infrastructure exhibits the ability to collect data from connected communities, and enables a seamless support of services needed for improving citizens’ quality of life and eventually helps city decision makers in urban planning.
Antonella Longo, Marco Zappatore, Mario A. Bochicchio, Shamkant B. Navathe
ACM Trans. Internet Techn.4
2016 VISAGE: Interactive Visual Graph Querying
abstract
Extracting useful patterns from large network datasets has become a fundamental challenge in many domains. We present Visage, an interactive visual graph querying approach that empowers users to construct expressive queries, without writing complex code (e.g., finding money laundering rings of bankers and business owners). Our contributions are as follows: (1) we introduce graph autocomplete, an interactive approach that guides users to construct and refine queries, preventing over-specification; (2) Visage guides the construction of graph queries using a data-driven approach, enabling users to specify queries with varying levels of specificity, from concrete and detailed (e.g., query by example), to abstract (e.g., with "wildcard" nodes of any types), to purely structural matching; (3) a twelve-participant, within-subject user study demonstrates Visage's ease of use and the ability to construct graph queries significantly faster than using a conventional query language; (4) Visage works on real graphs with over 468K edges, achieving sub-second response times for common queries.
Robert S. Pienta, Acar Tamersoy, Alex Endert, Shamkant B. Navathe, Hanghang Tong, Polo Chau
AVI4
2016 Constraint based temporal event sequence mining for Glioblastoma survival prediction
Kunal Malhotra, Shamkant B. Navathe, Polo Chau, Costas Hadjipanayis, Jimeng Sun 0001
J. Biomed. Informatics2
2016 Interactive Browsing and Navigation in Relational Databases
abstract
Although researchers have devoted considerable attention to helping database users formulate queries, many users still find it challenging to specify queries that involve joining tables. To help users construct join queries for exploring relational databases, we propose ETable , a novel presentation data model that provides users with a presentation-level interactive view. This view compactly presents one-to-many and many-to-many relationships within a single enriched table by allowing a cell to contain a set of entity references . Users can directly interact with this enriched table to incrementally construct complex queries and navigate databases on a conceptual entity-relationship level. In a user study, participants performed a range of database querying tasks faster with ETable than with a commercial graphical query builder. Subjective feedback about ETable was also positive. All participants found that ETable was easier to learn and helpful for exploring databases.
Minsuk Kahng, Shamkant B. Navathe, John T. Stasko, Polo Chau
Proc. VLDB Endow.2
2015 Mining Frequent Spatial-Textual Sequence Patterns
Krishan K. Arya, Vikram Goyal, Shamkant B. Navathe, Sushil K. Prasad
DASFAA (2)3
2015 An Extended ER Algebra to Support Semantically Richer Queries in ERDBMS
Moritz Wilfer, Shamkant B. Navathe
ER2
2014 Towards a Form Based Dynamic Database Schema Creation and Modification System
Kunal Malhotra, Shibani Medhekar, Shamkant B. Navathe, M. D. David Laborde
CAiSE3
2013 Inside insider trading: patterns & discoveries from a large scale exploratory analysis
abstract
How do company insiders trade? Do their trading behaviors differ based on their roles (e.g., CEO vs. CFO)? Do those behaviors change over time (e.g., impacted by the 2008 market crash)? Can we identify insiders who have similar trading behaviors? And what does that tell us?
Acar Tamersoy, Bo Xie 0002, Stephen L. Lenkey, Bryan R. Routledge, Polo Chau, Shamkant B. Navathe
ASONAM6
2013 Click traffic analysis of short URL spam on Twitter
abstract
With an average of 80% length reduction, the URL shorteners have become the norm for sharing URLs on Twitter, mainly due to the 140-character limit per message. Unfortunately, spammers have also adopted the URL shorteners to camouflage and improve the user click-through of their spam URLs. In this p
De Wang, Shamkant B. Navathe, Ling Liu 0001, Danesh Irani, Acar Tamersoy, Calton Pu
CollaborateCom2
2012 Systematic Modeling, Testing, and Monitoring of Information Integrity in Federated Ontology-driven Data Sources
Mijung Kim, Jake Cobb, Tahsin M. Kurç, Alessandro Orso, Mary Jean Harrold, Andrew R. Post, Shamkant B. Navathe, Joel H. Saltz
AMIA7
2012 OSQR: A framework for ontology-based semantic query routing in unstructured P2P networks
abstract
Efficient searching for information is an important goal in unstructured peer-to-peer (P2P) networks. While several P2P systems have been proposed for data sharing purposes, many support only semantics-free keyword searches or coarser grained file name searches. In this paper, we present an ontology based semantic query routing algorithm that performs efficient semantic search in unstructured P2P overlay networks. In our proposed system, the queries are routed in the network by forwarding to peers with highly relevant content in their local storages. To aid in this semantic query routing, we propose a scheme where each peer in the network adheres to a global ontology and semantically tags its local document collection with concepts in the ontology. Based on the semantic tags, peer level semantic summaries are generated, exchanged with neighboring peers and propagated along search paths which aid in efficient local query processing and overlay query routing. An extensive set of simulations performed to evaluate the effectiveness of the system on P2P networks show 380% and 717% improvement in average recall rate, and 410% and 725% improvement in average precision over Ontology Index based Query Routing [20] and Random Walk [18], respectively for dynamic networks at comparable message overheads. Thus, our approach represents a significant advance in practical terms.
D. M. Rasanjalee Himali, Shamkant B. Navathe, Sushil K. Prasad
HiPC2
2012 Efficient regression testing of ontology-driven systems
abstract
To manage and integrate information gathered from heterogeneous databases, an ontology is often used. Like all systems, ontology-driven systems evolve over time and must be regression tested to gain confidence in the behavior of the modified system. Because rerunning all existing tests can be extremely expensive, researchers have developed regression-test-selection (RTS) techniques that select a subset of the available tests that are affected by the changes, and use this subset to test the modified system. Existing RTS techniques have been shown to be effective, but they operate on the code and are unable to handle changes that involve ontologies. To address this limitation, we developed and present in this paper a novel RTS technique that targets ontology-driven systems. Our technique creates representations of the old and new ontologies, compares them to identify entities affected by the changes, and uses this information to select the subset of tests to rerun. We also describe in this paper OntoRetest, a tool that implements our technique and that we used to empirically evaluate our approach on two biomedical ontology-driven database systems. The results of our evaluation show that our technique is both efficient and effective in selecting tests to rerun and in reducing the overall time required to perform regression testing.
Mijung Kim, Jake Cobb, Mary Jean Harrold, Tahsin M. Kurç, Alessandro Orso, Joel H. Saltz, Andrew R. Post, Kunal Malhotra, Shamkant B. Navathe
ISSTA9
2011 DTMBIO 2011: international workshop on data and textmining in biomedical informatics
abstract
ACM Fifth International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 11) organizers are pleased to announce that the fifth DTMBIO will be held in conjunction with CIKM, one of the largest data and text mining conferences. While CIKM presents the state-of-the-art research in informatics with the primary focus on data and text mining, the main focus of DTMBIO is on biomedical informatics. DTMBIO delegates will bring forth interesting applications of up-to-date informatics in the context of biomedical research.
Sophia Ananiadou, Doheon Lee, Shamkant B. Navathe, Min Song 0001
CIKM3
2010 DTMBIO workshop summary
abstract
No abstract available.
Hagit Shatkay, Doheon Lee, Min Song 0001, Shamkant B. Navathe
CIKM4
2008 Discovering semantic biomedical relations utilizing the Web
abstract
To realize the vision of a Semantic Web for Life Sciences, discovering relations between resources is essential. It is very difficult to automatically extract relations from Web pages expressed in natural language formats. On the other hand, because of the explosive growth of information, it is difficult to manually extract the relations. In this paper we present techniques to automatically discover relations between biomedical resources from the Web. For this purpose we retrieve relevant information from Web Search engines and Pubmed database using various lexico-syntactic patterns as queries over SOAP web services. The patterns are initially handcrafted but can be progressively learnt. The extracted relations can be used to construct and augment ontologies and knowledge bases. Experiments are presented for general biomedical relation discovery and domain specific search to show the usefulness of our technique.
Saurav Sahay, Sougata Mukherjea, Eugene Agichtein, Ernest V. Garcia, Shamkant B. Navathe, Ashwin Ram 0001
ACM Trans. Knowl. Discov. Data5
2007 Text Mining and Ontology Applications in Bioinformatics and GIS
Shamkant B. Navathe
ICMLA1
2007 Improving Secure Communication Policy Agreements by Building Coalitions
abstract
In collaborative applications, participants agree on certain level of secure communication based on communication policy specifications. Given secure communication policy specifications of various group members at design time, the minimum set of resources for a pair, called resolved policy level agreement (RPLA) is translated into appropriate security service implementations, for the pair-wise communication to take place. We propose a novel idea that the members may extend pair-wise communication quality through other trusted nodes whose communication resources offer more security. We propose a heuristic algorithm which finds the best quality of protection (QoP), a measure of the resistance to an attack, path through coalition of trusted nodes. The results from our experiments indicate a significant improvement in QoP in the range of 13% to 48% over pair-wise communications.
Srilaxmi Malladi, Sushil K. Prasad, Shamkant B. Navathe
IPDPS3
2006 Optimizing Peer Virtualization and Load Balancing
Wanxia Xie, Shamkant B. Navathe, Sushil K. Prasad, David Fisher
DASFAA2
2006 A Two-Layered Software Architecture for Distributed Workflow Coordination over Web Services
abstract
The current state of the art of workflows over Web services employs a centralized composite process to coordinate the constituent Web services. Therefore, the coordinator process is complex, less scalable, and bulky. This paper introduces an architecture and a technique for distributing the centralized coordination logic of traditional workflows by (i) extending the stateless Web services into self-coordinating entities using coordinator proxy objects, and (ii) creating a workflow over these entities by interconnecting them into a distributed network of objects using Web bond primitives. Previously, we have developed Web bond primitives to enforce interdependencies among autonomous entities. We have designed and prototyped our BondFlow system, which provides a platform to configure such distributed workflows, producing coordination components with footprint small enough (around 150 KB) to be executed on a handheld
Janaka Balasooriya, Jaimini Joshi, Sushil K. Prasad, Shamkant B. Navathe
ICWS4
2006 Grammatical rules for specifying information for automated product data modeling
Ghang Lee, Charles M. Eastman, Rafael Sacks, Shamkant B. Navathe
Adv. Eng. Informatics4
2006 Semantic Integrity Constraint Checking for Multiple XML Databases
abstract
Global semantic integrity constraints ensure integrity and consistency of data spanning multiple databases. In this paper, we take initial steps towards representing global semantic integrity constraints for XML databases. We also provide a general framework for checking global semantic integrity constraints for XML databases. Furthermore, we set forth an efficient algorithm for checking global semantic integrity constraints across multiple XML databases. Our algorithm is efficient for three reasons: (1) the algorithm does not require the update statement to be executed before the constraint check is carried out; hence, we avoid any potential problems associated with rollbacks, (2) sub constraint checks are executed in parallel, and (3) most of the processing of algorithm could happen at compile time; hence, we save time spent at run-time. As a proof of concept, we present a prototype of the system implementing the ideas discussed in this paper.
Praveen Madiraju, Rajshekhar Sunderraman, Shamkant B. Navathe
J. Database Manag.3
2005 Filter Indexing: A Scalable Solution to Large Subscription Based Systems
Wanxia Xie, Shamkant B. Navathe, Sushil K. Prasad
DASFAA2
2005 Text Mining Biomedical Literature for Discovering Gene-to-Gene Relationships: A Comparative Study of Algorithms
abstract
Partitioning closely related genes into clusters has become an important element of practically all statistical analyses of microarray data. A number of computer algorithms have been developed for this task. Although these algorithms have demonstrated their usefulness for gene clustering, some basic problems remain. This paper describes our work on extracting functional keywords from MEDLINE for a set of genes that are isolated for further study from microarray experiments based on their differential expression patterns. The sharing of functional keywords among genes is used as a basis for clustering in a new approach called BEA-PARTITION in this paper. Functional keywords associated with genes were extracted from MEDLINE abstracts. We modified the Bond Energy Algorithm (BEA), which is widely accepted in psychology and database design but is virtually unknown in bioinformatics, to cluster genes by functional keyword associations. The results showed that BEA-PARTITION and hierarchical clustering algorithm outperformed k-means clustering and self-organizing map by correctly assigning 25 of 26 genes in a test set of four known gene groups. To evaluate the effectiveness of BEA-PARTITION for clustering genes identified by microarray profiles, 44 yeast genes that are differentially expressed during the cell cycle and have been widely studied in the literature were used as a second test set. Using established measures of cluster quality, the results produced by BEA-PARTITION had higher purity, lower entropy, and higher mutual information than those produced by k-means and self-organizing map. Whereas BEA-PARTITION and the hierarchical clustering produced similar quality of clusters, BEA-PARTITION provides clear cluster boundaries compared to the hierarchical clustering. BEA-PARTITION is simple to implement and provides a powerful approach to clustering genes or to any clustering problem where starting matrices are available from experimental observations.
Ying Liu 0007, Shamkant B. Navathe, Jorge Civera, Venu Dasigi, Ashwin Ram 0001, Brian J. Ciliax, Ray Dingledine
IEEE ACM Trans. Comput. Biol. Bioinform.2
2004 Genomic and Proteomic Databases and Applications: A Challenge for Database Technology
Shamkant B. Navathe, Upen Patil
DASFAA1
2004 SyD: A Middleware Testbed for Collaborative Applications over Small Heterogeneous Devices and Data Stores
Sushil K. Prasad, Vijay K. Madisetti, Shamkant B. Navathe, Rajshekhar Sunderraman, Erdogan Dogdu, Anu G. Bourgeois, Bing Liu 0003, Janaka Balasooriya, Arthi Hariharan, Wanxia Xie, Praveen Madiraju, Srilaxmi Malladi, Raghupathy Sivakumar, Alex Zelikovsky, Yan-Qing Zhang 0001, Yi Pan 0001, Saeid Belkasim
Middleware3
2003 Managing vulnerabilities of information systems to security incidents
abstract
Information security-conscious managers of organizations have the responsibility to advise their senior management of the level of risks faced by the information systems. This requires managers to conduct vulnerability assessment as the first step of a risk analysis approach. However, a lack of real world data classification of security threats and develops a three-axis view of the threat space. It develops a scheme for probabilistic evaluation of impact of the security threats and proposes a risk management system consisting of a five-step approach. The goal is to assess the expected damages due to attacks, and managing the risk of attacks.
Fariborz Farahmand, Shamkant B. Navathe, Philip H. Enslow Jr., Gunter P. Sharp
ICEC2
2003 Efficient data access to multi-channel broadcast programs
abstract
This paper studies fast access to data that are broadcast on multiple channels. Broadcast is a useful data dissemination technique because of its scalability, but is lacking when it comes to response time. Increasing the number of available broadcast channels is a logical way of increasing throughput. Little work, however, has considered the access structures necessary for making effective use of the additional channels. We propose various indexing schemes for a multi-channel broadcast program. We demonstrate the effectiveness of our techniques in decreasing response time and tuning time via extensive experiments over a wide range of parameters.
Wai Gen Yee, Shamkant B. Navathe
CIKM2
2003 Fast Data Access on Multiple Broadcast Channels
abstract
We give an overview on techniques for quickly finding items that are broadcast on multiple data channels. Previous works assume that throughput increases with the number of data channels because the total available bandwidth increases. This assumption, however, rests on the client's ability read data as soon as it is broadcast on any channel. Without this ability, the benefit of additional channels is compromised as the search space for data also increases. We describe the consequences of increasing the number of broadcast channels on search. We then offer candidate search techniques and show their effects on three metrics: search time, tuning time, and hop count.
Wai Gen Yee, Shamkant B. Navathe
ICDE2
2003 Mobile User Recovery in the Context of Internet Transactions
abstract
With the expansion of Web sites to include business functions, a user interfaces with e-businesses through an interactive and multistep process, which is often time-consuming. For mobile users accessing the Web over digital cellular networks, the failure of the wireless link, a frequent occurrence, can result in the loss of work accomplished prior to the disruption. This work must then be repeated upon subsequent reconnection - often at significant cost in time and computation. This "disconnection-reconnection-repeat work" cycle may cause mobile clients to incur substantial monetary as well as resource (such as battery power) costs. In this paper, we propose a protocol for "recovering" a user to an appropriate recent interaction state after such a failure. The objective is to minimize the amount of work that needs to be redone upon restart after failure. Whereas classical database recovery focuses on recovering the system, i.e., all transactions, our work considers the problem of recovering a particular user interaction with the system. This recovery problem encompasses several interesting subproblems: (1) modeling user interaction in a way that is useful for recovery, (2) characterizing a user's "recovery state", (3) determining the state to which a user should be recovered, and (4) defining a recovery mechanism. We describe the user interaction with one or more Web sites using intuitive and familiar concepts from database transactions. We call this interaction an Internet transaction (iTX), distinguish this notion from extant transaction models, and develop a model for it, as well as for a user's state on a Web site. Based on the twin foundations of our iTX and state models, we finally describe an effective protocol for recovering users to valid states in Internet interactions.
Debra E. VanderMeer, Anindya Datta, Kaushik Dutta, Krithi Ramamritham, Shamkant B. Navathe
IEEE Trans. Mob. Comput.5
2002 Bridging the Gap between Response Time and Energy-Efficiency in Broadcast Schedule Design
Wai Gen Yee, Shamkant B. Navathe, Edward Omiecinski, Chris Jermaine
EDBT2
2002 XML Schema Mappings for Heterogeneous Database Access
Samuel Robert Collins, Shamkant B. Navathe, Leo Mark
Inf. Softw. Technol.2
2002 An abbreviated concept-based query language and its exploratory evaluation
Vesper Owei, Shamkant B. Navathe, Hyeun-Suk Rhee
J. Syst. Softw.2
2002 Efficient Data Allocation over Multiple Channels at Broadcast Servers
abstract
Broadcast is a scalable way of disseminating data because broadcasting an item satisfies all outstanding client requests for it. However, because the transmission medium is shared, individual requests may have high response times. In this paper, we show how to minimize the average response time given multiple broadcast channels by optimally partitioning data among them. We also offer an approximation algorithm that is less complex than the optimal and show that its performance is near-optimal for a wide range of parameters. Finally, we briefly discuss the extensibility of our work with two simple, yet seldom researched extensions, namely, handling varying sized items and generating single channel schedules.
Wai Gen Yee, Shamkant B. Navathe, Edward Omiecinski, Chris Jermaine
IEEE Trans. Computers2
2001 Scaling Replica Maintenance In Intermittently Synchronized Mobile Databases
abstract
To avoid the high cost of continuous connectivity, a class of mobile applications employs replicas of shared data that are periodically updated. Updates to these replicas are typically performed on a client-by-client basis--that is, the server individually computes and transmits updates to each client--limiting scalability. By basing updates on replica groups (instead of clients), however, update generation complexity is no longer bound by client population size. Clients then download updates of pertinent groups. Proper group design reduces redundancies in server processing, disk usage and bandwidth usage, and dimininishes the tie between the complexity of updating replicas and the size of the client population. In this paper, we expand on previous work done on group design, include a detailed I/O cost model for update generation, and propose a heuristic-based greedy algorithm for group computation. Experimental results with an adapted commercial replication system demonstrate a significant increase in overall scalability over the client-centric approach.
Wai Gen Yee, Edward Omiecinski, Michael J. Donahoo, Shamkant B. Navathe
CIKM4
2001 A formal basis for an abbreviated concept-based query language
Vesper Owei, Shamkant B. Navathe
Data Knowl. Eng.2
2001 Enriching the conceptual basis for query formulation through relationship semantics in databases
Vesper Owei, Shamkant B. Navathe
Inf. Syst.2
2001 A Framework for Method-Specific Knowledge Compilation from Databases
J. William Murdock, Ashok K. Goel 0001, Michael J. Donahoo, Shamkant B. Navathe
J. Intell. Inf. Syst.4
2001 An architecture to support scalable online personalization on the Web
Anindya Datta, Kaushik Dutta, Debra E. VanderMeer, Krithi Ramamritham, Shamkant B. Navathe
VLDB J.5
2000 A Framework for Designing Update Objects to Improve Server Scalability in Intermittently Synchronized Databases
abstract
Article Free Access Share on A framework for designing update objects to improve server scalability in intermittently synchronized databases Authors: Wai Gen Yee College of Computing, Georgia Institute of Technology, 801 Atlantic Drive, Atlanta, GA College of Computing, Georgia Institute of Technology, 801 Atlantic Drive, Atlanta, GAView Profile , Michael J. Donahoo Baylor University, P.O. Box 97356, Waco, TX Baylor University, P.O. Box 97356, Waco, TXView Profile , Shamkant B. Navathe College of Computing, Georgia Institute of Technology, 801 Atlantic Drive, Atlanta, GA College of Computing, Georgia Institute of Technology, 801 Atlantic Drive, Atlanta, GAView Profile Authors Info & Claims CIKM '00: Proceedings of the ninth international conference on Information and knowledge managementNovember 2000 Pages 54–61https://doi.org/10.1145/354756.354787Published:06 November 2000Publication History 4citation388DownloadsMetricsTotal Citations4Total Downloads388Last 12 Months16Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Wai Gen Yee, Michael J. Donahoo, Shamkant B. Navathe
CIKM3
2000 Enabling scalable online personalization on the Web
abstract
Online personalization is of great interest to e-companies. Virtually all personalization technologies are based on the idea of storing as much historical customer session data as possible, and then querying the data store as customers navigate through a web site. The holy grail of on-line personalization is an environment where fine-grained, detailed historical session data can be queried based on current online navigation patterns to formulate real-time responses. Unfortunately, as more consumers become e-shoppers, the user load and the amount of historical data continue to increase, causing scalability-related problems for almost all current personalization technologies. This paper chronicles the development of a real-time interaction management engine through the integration of historical data and on-line visitation patterns of e-commerce site visitors. This paper describes the scientific underpinnings of the system, as well as the architecture and a performance evaluation....
Debra E. VanderMeer, Kaushik Dutta, Anindya Datta, Krithi Ramamritham, Shamkant B. Navathe
EC5
2000 A 20/20 Vision of the VLDB-2020?
S. Misbah Deen, Anant Jhingran, Shamkant B. Navathe, Erich J. Neuhold, Gio Wiederhold
VLDB3
2000 C-FAR, change favorable representation
Tal Cohen, Shamkant B. Navathe, Robert E. Fulton
Comput. Aided Des.2
1999 The Challenge of Genome Information Management: A Practical Experience
abstract
While the accumulation of information on human biology and health is a primary motivation for the human genome initiative, little attention has been paid to how such information will be organized or functionally related to the genomic sequence. This stems from two issues. First, genetic data is complex and multi-layered. Medical genetic information ranges from structural (nucleotide, nucleotide position, mutation position, gene organization, chromosomal assignment) to functional (type of gene, gene expression, biochemical studies), to population-based (frequencies of mutations in different racial/ethnic groups), to clinical-based (family studies, genotype/phenotype correlations) etc. The magnitude of such data will continue to push the limits of information management and integration systems. Second, no portion of the human nuclear genome yet sequenced has been studied extensively enough to generate the magnitude of data that would necessitate the development of an integrated genetic information system.This talk will present an overview of the challenges that confront biologists and the database specialists in terms of creating a useful repository of information that addresses both the above issues. We will describe the MITOMAP system (http://www.gen.emory.edu/mitomap.) which we have developed in collaboration with themolecular genetics department at Emory University by focusing on the mitochondrial genome 3/4 a very small yet very significant subset of the genome.We will discuss the modeling, design and implementation issues of creating this repository. and the issues we have dealt with and the enhancements we propose to make. The system is being developed as a tool for researchers in biomedical science and clinical medicine. It will serve the biological community at large and will be an aid for further work in practical applications such as pharmaceutics or gene therapy.
Shamkant B. Navathe
IDEAS1
1998 Grouping Techniques for Update Propagation in Intermittently Connected Databases
abstract
We consider an environment where one or more servers carry databases that are of interest to a community of clients. The clients are only intermittently connected to the server for brief periods of time. Clients carry a part of the database for their own processing and accumulate local updates while disconnected. We call this the Intermittently Connected Database (ICDB) environment. ICDBs have a wide variety of applications including sales force automation, insurance claim processing, and mobile workforces. Our focus is on the problem of update propagation at the server in ICDBs and the associated processing at the clients. The typical client-centric approach involves the communication and processing of updates and transactions on a per-client basis, ignoring the overlap of data between clients. The complexity of this approach is in the order of the number of connecting clients, thereby limiting the scalability of the server. We propose a data-centric approach which clusters data into groups and assigns to each client one or more of these groups. The proposed scheme results in server processing complexity on the order of the number of groups, which we control. We propose various techniques for grouping and discuss the processing required at the clients to enable the grouping approach. While the client-centric approach is expected to significantly degrade with the increasing number of clients, we expect that a properly designed grouping scheme will sustain a number of clients that is significantly larger. A prototype has been developed and performance studies are in progress.
Sameer Mahajan, Michael J. Donahoo, Shamkant B. Navathe, Mostafa H. Ammar, Sanjoy Malik
ICDE3
1998 Mining for Strong Negative Associations in a Large Database of Customer Transactions
abstract
Mining for association rules is considered an important data mining problem. Many different variations of this problem have been described in the literature. We introduce the problem of mining for negative associations. A naive approach to finding negative associations leads to a very large number of rules with low interest measures. We address this problem by combining previously discovered positive associations with domain knowledge to constrain the search space such that fewer but more interesting negative rules are mined. We describe an algorithm that efficiently finds all such negative associations and present the experimental results.
Ashok Savasere, Edward Omiecinski, Shamkant B. Navathe
ICDE3
1997 The Challenges of Modeling Biological Information for Genome Databases
Shamkant B. Navathe, Andreas M. Kogelnik
Conceptual Modeling1
1997 From Data to Knowledge: Method-Specific Transformations
Michael J. Donahoo, J. William Murdock, Ashok K. Goel 0001, Shamkant B. Navathe, Edward Omiecinski
ISMIS4
1996 Two Techniques for On-Line Index Modification in Shared Nothing Parallel Databases
abstract
Whenever data is moved across nodes in the parallel database system, the indexes need to be modified too. Index modification overhead can be quite severe because there can be a large number of indexes on a relation. In this paper, we study two alternatives to index modification, namely OAT (One-At-a-Time page movement) and BULK (bulk page movement). OAT and BULK are two extremes on the spectrum of the granularity of data movement. OAT and BULK differ in two respects: first, OAT uses very little additional disk space (at most one extra page), whereas BULK uses a large amount of disk space. Second, BULK uses sequential prefetch I/O to optimize on the number of I/Os during index modification, while OAT does not. Using an experimental testbed, we show that BULK is an order of magnitude faster than OAT. In terms of the impact on transaction performance during reorganization, BULK and OAT perform differently: when the number of indexes to be modified is either one or two, OAT has a lesser impact on the transaction performance degradation. However, when the number of indexes is greater than two, both techniques have the same impact on transaction performance.
Kiran J. Achyutuni, Edward Omiecinski, Shamkant B. Navathe
SIGMOD Conference3
1996 Optimal Redesign Policies to Support Dynamic Processing of Applications on a Distributed Relational Database System
Kamalakar Karlapalem, Shamkant B. Navathe, Mostafa H. Ammar
Inf. Syst.2
1995 An Efficient Algorithm for Mining Association Rules in Large Databases
Ashok Savasere, Edward Omiecinski, Shamkant B. Navathe
VLDB3
1994 Providing orthogonal persistence to C++ using forced inheritance
abstract
Orthogonal persistence, a property that any object can be made to persist independent of its type, is an important requirement of a persistent language. We present a new technique called forced inheritance for providing the orthogonal persistence to C++. In this technique, properties that make objects persist are attached as a header to an object or a value of any type that is desired to persist. Attaching the header gives the effect of inheriting these properties from a virtual persistent root class regardless of its type. This technique provides orthogonal persistence since attaching the header to an object can be done for any object. It also provides portability since it does not extend the language. The new, approach not only has advantages over existing ones, but also can be efficiently implemented.>
Chong-Mok Park, Kyu-Young Whang, Il-Yeol Song, Shamkant B. Navathe
COMPSAC4
1994 On Behavioral Schema Evolution in Object-Oriented Databases
Magdi M. A. Morsi, Shamkant B. Navathe, John J. Shilling
EDBT2
1994 An Objective Function for Vertically Partitioning Relations in Distributed Databases and its Analysis
Sharma Chakravarthy, Jaykumar Muthuraj, Ravi Varadarajan, Shamkant B. Navathe
Distributed Parallel Databases4
1994 A Logic-Based Approach to Query Processing in Federated Databases
Sharma Chakravarthy, Whan-Kyu Whang, Shamkant B. Navathe
Inf. Sci.3
1994 A Conceptual Clustering Algorithm for Database Schema Design
abstract
Conceptual clustering techniques based on current theories of categorization provide a way to design database schemas that more accurately represent classes. An approach is presented in which classes are treated as complex clusters of concepts rather than as simple predicates. An important service provided by the database is determining whether a particular instance is a member of a class. A conceptual clustering algorithm based on theories of categorization aids in building classes by grouping related instances and developing class descriptions. The resulting database schema addresses a number of properties of categories, including default values and prototypes, analogical reasoning, exception handling, and family resemblance. Class cohesion results from trying to resolve conflicts between building generalized class descriptions and accommodating members of the class that deviate from these descriptions. This is achieved by combining techniques from machine learning, specifically explanation-based learning and case-based reasoning. A subsumption function is used to compare two class descriptions. A realization function is used to determine whether an instance meets an existing class description. A new function, INTERSECT, is introduced to compare the similarity of two instances. INTERSECT is used in defining an exception condition. Exception handling results in schema modification. This approach is applied to the database problems of schema integration, schema generation, query processing, and view creation.>
Howard W. Beck, Tarek M. Anwar, Shamkant B. Navathe
IEEE Trans. Knowl. Data Eng.3
1993 Application and System Prototyping Via an Extensible Object-Oriented Environment
Magdi M. A. Morsi, Shamkant B. Navathe
ER2
1993 On Mapping ER Models into OO Schemas
Badri Narasimhan, Shamkant B. Navathe, Sundaresan Jayaraman
ER2
1993 Using Active Database Techniques For Real Time Engineering Applications
abstract
An active database model for representing engineering design, simulation and monitoring applications is described. The physical aspects of these applications are modeled by structural objects. The functions of the application are modeled by functional objects. The interaction between structures and functions is modeled by interaction objects. Events relate structures and functions such that any state change will cause the functional model to compute the new consistent state of the application. Ways to model event correlation, event recall, and ways to define schemas for continuous systems are presented. A parallel computing architecture and a set of guidelines on task distributing that make the model applicable to real-time application are discussed.>
Aloysius Cornelio, Shamkant B. Navathe
ICDE2
1993 Database Technology and Standards: Are we Getting Anywhere? (Panel Abstract)
Peter Dadam, Shamkant B. Navathe
ICDE2
1993 Voltaire: A Database Programming Language with a Single Execution Model for Evaluating Queries, Constraints amd Functions
abstract
Various features of the Voltaire database programming language are described. Unlike most other languages, Voltaire has a single execution model for evaluating queries, enforcing constraints and computing functions. Such a design also facilitates a bootstrapped implementation. Voltaire supports a synthesis of declarative querying primitives with imperative programming primitives so that constraints are also enforced. The notion of temporary instance creation, which allows an equivalent semantics to be given to classes and functions is reviewed. A data definition facility in which persistence is a property of the instances rather than classes, and the transparency provided between persistent and transient objects by defining a single set of operators for both kinds of objects are discussed.>
Sunit K. Gala, Shamkant B. Navathe, Manuel E. Bermudez
ICDE2
1993 Database Supported Cooperative Problem Solving
abstract
Cooperative problem solving can be viewed as a complex activity requiring harmonious and dynamic interaction between active agents and passive agents. This problem is currently being addressed by the research community at various levels of abstraction. Broadly, this paper analyzes the problem of cooperative problem solving from a database perspective and argues that recent advances in database technology facilitate development of a viable solution to the above problem. Specifically, in this paper, we first analyze the problem of cooperative problem solving to identify its underlying key characteristics. Based upon our analysis, we partition the problem space into classes along a spectrum and indicate the level of cooperation required. We propose near-term as well as long-term solutions for cooperative problem solving that progressively enhances the functionality of database systems by synthesizing appropriate abstractions and techniques. Furthermore, we identify problems, such as agent capability modeling that require further research to address the most general form of cooperative problem solving.
Sharma Chakravarthy, Kamalakar Karlapalem, Shamkant B. Navathe, Asterio Kiyoshi Tanaka
Int. J. Cooperative Inf. Syst.3
1993 On Automatic Reasoning for Schema Integration
abstract
Success in database schema integration depends on the ability to capture real world semantics of the schema objects, and to reason about the semantics. Earlier schema integration approaches mainly rely on heuristics and human reasoning. In this paper, we discuss an approach to automate a significant part of the schema integration process. Our approach consists of three phases. An attribute hierarchy is generated in the first phase. This involves identifying relationships (equality, disjointness and inclusion) among attributes. We discuss a strategy based on user-specified semantic clustering. In the second phase, a classification algorithm based on the semantics of class subsumption is applied to the class definitions and the attribute hierarchy to automatically generate a class taxonomy. This class taxonomy represents a partially integrated schema. In the third phase, the user may employ a set of well-defined comparison operators in conjunction with a set of restructuring operators, to further modify the schema. These operators as well as the automatic reasoning during the second phase are based on subsumption. The formal semantics and automatic reasoning utilized in the second phase is based on a terminological logic as adapted in the CANDIDE data model. Classes are completely defined in terms of attributes and constraints. Our observation is that the inability to completely define attributes and thus completely capture their real world semantics imposes a fundamental limitation on the possibility of automatically reasoning about attribute definitions. This necessitates human reasoning during the first phase of the integration approach.
Amit P. Sheth, Sunit K. Gala, Shamkant B. Navathe
Int. J. Cooperative Inf. Syst.3
1992 Adaptive and Automated Index Selection in RDBMS
Martin R. Frank, Edward Omiecinski, Shamkant B. Navathe
EDBT3
1992 The Next Ten Years of Modeling, Methodologies, and Tools
Shamkant B. Navathe
ER1
1992 Knowledge Mining by Imprecise Querying: A Classification-Based Approach
abstract
Knowledge mining is the process of discovering knowledge that is hitherto unknown. An approach to knowledge mining by imprecise querying that utilizes conceptual clustering techniques is presented. The query processor has both a deductive and an inductive component. The deductive component finds precise matches in the traditional sense, and the inductive component identifies ways in which imprecise matches may be considered similar. Ranking on similarity is done by using the database taxonomy, by which similar instances become members of the same class. Relative similarity is determined by depth in the taxonomy. The conceptual clustering algorithm, its use in query processing, and an example are presented.>
Tarek M. Anwar, Howard W. Beck, Shamkant B. Navathe
ICDE3
1992 An Extensible Object-Oriented Database Testbed
abstract
The authors describe the object-oriented design and implementation of an extensible schema manager for object-oriented databases. The open class hierarchy approach has been adopted to achieve the extensibility of the implementation. In this approach. the system meta information is implemented as objects of system classes. A graphical interface for an object-oriented database scheme environment, GOOSE, has been developed. GOOSE supports several advanced features which include schema evolution, schema versioning, and DAG (direct acyclic graph) rearrangement view of a class hierarchy. Schema evolution is the ability to make a variety of changes to a database scheme without reorganization. Schema versioning is the ability to define multiple scheme versions and to keep track of schema changes. A novel type of view for object-oriented databases, the DAG rearrangement view of a class hierarchy, is also supported.>
Magdi M. A. Morsi, Shamkant B. Navathe
ICDE2
1992 Integration of expert systems and database management systems--An extended disjunctive normal form approach
Kyu-Young Whang, Shamkant B. Navathe
Inf. Sci.2
1992 Scheduling file transfers in fully connected networks
abstract
Abstract We consider the problem of transferring a set of files from their given locations in a fully connected network to their respective target locations in minimum time. We show that this problem is NP‐hard even with the restriction that no file uses more than two edges in its route. We present an efficient algorithm to solve this problem in the case when there is only one source and one or more destinations. For the general case, we propose a two‐phase approach to find two‐edge schedules that are optimal or close‐to‐optimal. In Phase I, two‐edge routes are assigned to files; in Phase II, a schedule is determined for the use of the links in these routes. For Phase I, we present an exact solution that is based on integer programming formulation and also give theoretical bounds for approximate solution. We also propose a route assignment algorithm that attempts to assign routes of minimum congestion. For Phase II, we present an efficient algorithm that constructs a schedule from the solution obtained in the first phase.
Pedro I. Rivera-Vega, Ravi Varadarajan, Shamkant B. Navathe
Networks3
1991 Relational Database Organization Based on Views and Fragments
Günther Pernul, Kamalakar Karlapalem, Shamkant B. Navathe
DEXA3
1991 A Transaction Architecture for a General Purpose Semantic Data Model
Shamkant B. Navathe, A. Balaraman
ER1
1991 Active, Real-Time, and Heterogenous Database Systems
Shamkant B. Navathe, Sharma Chakravarthy
ER1
1991 ER-R: An Enhanced ER Model with Situation-Action Rules to Capture Application Semantics
Asterio Kiyoshi Tanaka, Shamkant B. Navathe, Sharma Chakravarthy, Kamalakar Karlapalem
ER2
1991 Version Management of Composite Objects in CAD Databases
abstract
article Free Access Share on Version management of composite objects in CAD databases Authors: Rafi Ahmed Hewlett-Packard Laboratories, Palo Alto, CA Hewlett-Packard Laboratories, Palo Alto, CAView Profile , Shamkant B. Navathe College of Computing, Georgia Institute of Technology, Atlanta, GA College of Computing, Georgia Institute of Technology, Atlanta, GAView Profile Authors Info & Claims ACM SIGMOD RecordVolume 20Issue 2June 1991pp 218–227https://doi.org/10.1145/119995.115825Published:01 April 1991Publication History 45citation532DownloadsMetricsTotal Citations45Total Downloads532Last 12 Months21Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Rafi Ahmed, Shamkant B. Navathe
SIGMOD Conference2
1991 Cooperative Database Design (Panel)
Stefano Spaccapietra, Shamkant B. Navathe, Erich J. Neuhold, Amit P. Sheth
VLDB2
1990 Modeling Physical Systems by Complex Structural Objects and Complex Functional Objects
Shamkant B. Navathe, Aloysius Cornelio
EDBT1
1990 Active and Heterogenous Database Systems
Shamkant B. Navathe, Sharma Chakravarthy
ER1
1990 Conceptual Design for Non-Database Experts with an Interactive Schema Tailoring Tool
Shamkant B. Navathe, Seong Geum, Dinesh K. Desai, Herman Lam
ER1
1990 Extending Object-Oriented Concepts to Support Engineering Applications
abstract
A presentation is made of a structure-function (S-F) paradigm for representing engineering designs. The S-F paradigm treats structures and functions as first-class objects by modeling the physical configuration of the design as structural objects and the behavior of the design as functional objects. The paradigm supports engineering designs by using the associative knowledge between structures and functions to help in the design process and to support simulation by providing for a run-time interaction between structures and functions. The S-F paradigm would be generally applicable to model the passive and active information in any complex system. It may be implemented as a layer on top of any object-oriented model.>
Aloysius Cornelio, Shamkant B. Navathe, Keith L. Doty
ICDE2
1990 Scheduling Data Redistribution in Distributed Databases
abstract
In a distributed database system there is a need for periodic changes in data distribution due to such factors as changes in query patterns and network topology. A proper redistribution of data is necessary to provide acceptable system performance, as measured by the average execution time of transactions. The problem of properly scheduling data transfers in order to complete this redistribution process in minimum possible time is investigated. This problem takes into account the constraints on communication resources of the system. The complexity of the problem is discussed, and a useful upper bound for optimal solutions is presented. Also given are procedures for finding optimal and approximate solutions to this problem.>
Pedro I. Rivera-Vega, Ravi Varadarajan, Shamkant B. Navathe
ICDE3
1989 Modeling parts and discrete assembly operations, using an object-oriented data model
abstract
An application of the object-oriented semantic association model (OSAM) concepts to the modeling of parts and their relationships, including assembly operations, is presented. It is shown that the semantic primitives (facilities) of OSAM make it possible to model the necessary structural relationships and constraints adequately. The proposed scheme can be used for simulation and task planning in the workcell environment.>
Constantinos Papaconstantinou, Keith L. Doty, Shamkant B. Navathe
COMPSAC3
1989 Classification as a Query Processing Technique in the CANDIDE Semantic Data Model
abstract
The use of classification and subsumption to process database queries is discussed. The data model, called CANDIDE, is essentially an extended version of the FL-1, KANDOR and BACK, frame-based knowledge representation languages. A novel feature of the approach is that the data-description language and data-manipulation language are identical, thus providing uniform treatment of data objects, query objects and view objects. The classification algorithm find the correct placement for a query object in a given object taxonomy. Tractability issues are explored, and the expressiveness of queries is compared with relational algebra. This data model has been implemented in POPLOG as the basis for a knowledge-base management system that includes an integrated natural-language query system.>
Howard W. Beck, Sunit K. Gala, Shamkant B. Navathe
ICDE3
1989 Vertical Partitioning for Database Design: A Graphical Algorithm
abstract
Vertical partitioning is the process of subdividing the attributes of a relation or a record type, creating fragments. Previous approaches have used an iterative binary partitioning method which is based on clustering algorithms and mathematical cost functions. In this paper, however, we propose a new vertical partitioning algorithm using a graphical technique. This algorithm starts from the attribute affinity matrix by considering it as a complete graph. Then, forming a linearly connected spanning tree, it generates all meaningful fragments simultaneously by considering a cycle as a fragment. We show its computational superiority. It provides a cleaner alternative without arbitrary objective functions and provides an improvement over our previous work on vertical partitioning.
Shamkant B. Navathe, Minyoung Ra
SIGMOD Conference1
1989 A Theory of Attribute Equivalence in Databases with Application to Schema Integration
abstract
The authors present a common foundation for integrating pairs of entity sets, pairs of relationship sets, and an entity set with a relationship set. This common foundation is based on the basic principle of integrating attributes. Any pair of objects whose identifying attributes can be integrated can themselves be integrated. Several definitions of attribute equivalence are presented. These definitions can be used to specify the exact nature of the relationship between a pair of attributes. Based on these definitions, several strategies for attribute integration are presented and evaluated.>
James A. Larson, Shamkant B. Navathe, Ramez Elmasri
IEEE Trans. Software Eng.2
1988 OOER: Toward Making the E-R Approach Object-Oriented
Shamkant B. Navathe, Mahan K. Pillalamarri
ER1
1988 Data Bases and Knowledge Bases: Which Approach is Good for What? - Panel
Yannis Vassiliou, Valeria De Antonellis, Klaus R. Dittrich, Michael V. Mannino, Shamkant B. Navathe
ER5
1988 A Tool for Integrating Conceptual Schemas and User Views
abstract
An interactive tool has been developed to assist database designers and administrators (DDA) in integrating schemas. It collects the information required for integration from a DDA, performs essential bookkeeping, and integrates schemas according to the semantics provided. The authors present the capabilities of this tool by discussing the integration methodology and the user interface of the tool.>
Amit P. Sheth, James A. Larson, Aloysius Cornelio, Shamkant B. Navathe
ICDE4
1988 A rule-based approach for merging generalization hierarchies
Michael V. Mannino, Shamkant B. Navathe, Wolfgang Effelsberg
Inf. Syst.2
1987 Abstracting Relational and Hierarchical Data with a Semantic Data Model
Shamkant B. Navathe, A. M. Awong
ER1
1987 Dealing with Temporal Schema Anomalies in History Databases
N. G. Martin, Shamkant B. Navathe, Rafi Ahmed
VLDB2
1987 An Extended Disjunctive Normal Form Approach for Optimizing Recursive Logic Queries in Loosely Coupled Environments
Kyu-Young Whang, Shamkant B. Navathe
VLDB2
1987 A Cost-Benefit Decision Model: Analysis, Comparison, and Selection of Data Management Systems
abstract
This paper describes a general cost-benefit decision model that is applicable to the evaluation, comparison, and selection of alternative products with a multiplicity of features, such as complex computer systems. The application of this model is explained and illustrated using the selection of data management systems as an example. The model has the following features: (1) it is mathematically based on an extended continuous logic and a theory of complex criteria; (2) the decision-making procedure is very general yet systematic, well-structured, and quantitative; (3) the technique is based on a comprehensive cost analysis and an elaborate analysis of benefits expressed in terms of the decision maker's preferences. The decision methodology, when applied to the problem of selecting a data management system, takes into consideration the life cycle of a DMS and the objectives and goals for the new systems under evaluation. It allows the cost and preference analyses to be carried out separately using two different models. The model for preference analysis makes use of comprehensive performance (or preference) parameters and allows what we call a “logic scoring of preferences” using continuous values between zero and one, to express the degree with which candidate systems satisfy stated requirements. It aggregates preference parameters based on their relative weights and logical relationships to compute a global performance (preference) score for each system. The cost model incorporates an aggregation of costs which may be estimated over different time horizons and discounted at appropriate discount rates. A procedure to establish an overall ranking of alternative systems based on their global preference scores and global costs is also discussed.
Stanley Y. W. Su, Jozo J. Dujmovic, Don S. Batory, Shamkant B. Navathe, Richard Elnicki
ACM Trans. Database Syst.4
1986 SA-ER: A Methodology that Links Structured Analysis and Entity-Relationship Modeling for Database Design
John L. Carswell Jr., Shamkant B. Navathe
ER2
1986 Role of data dictionaries in information resource management
Shamkant B. Navathe, Larry Kerschberg
Inf. Manag.1
1984 Object Integration in Logical Database Design
abstract
View integration is one of the important phases in logical database design. During this phase, the individual views designed by separate user groups within the organization are integrated into a conceptual schema for the entire organization. In this paper, we present rules for some of the aspects of view integration, namely integration of entity classes in different views. We then compare our rules with previous approaches, and discuss some of the problems which have to be solved for other aspects of view integration.
Ramez Elmasri, Shamkant B. Navathe
ICDE2
1984 Databases for Expert Systems
Shamkant B. Navathe, Reind P. van de Riet
VLDB1
1984 Relationship Merging in Schema Integration
Shamkant B. Navathe, T. Sashidhar, Ramez Elmasri
VLDB1
1984 Vertical Partitioning Algorithms for Database Design
abstract
This paper addresses the vertical partitioning of a set of logical records or a relation into fragments. The rationale behind vertical partitioning is to produce fragments, groups of attribute columns, that “closely match” the requirements of transactions. Vertical partitioning is applied in three contexts: a database stored on devices of a single type, a database stored in different memory levels, and a distributed database. In a two-level memory hierarchy, most transactions should be processed using the fragments in primary memory. In distributed databases, fragment allocation should maximize the amount of local transaction processing. Fragments may be nonoverlapping or overlapping. A two-phase approach for the determination of fragments is proposed; in the first phase, the design is driven by empirical objective functions which do not require specific cost information. The second phase performs cost optimization by incorporating the knowledge of a specific application environment. The algorithms presented in this paper have been implemented, and examples of their actual use are shown.
Shamkant B. Navathe, Stefano Ceri, Gio Wiederhold, Jinglie Dou
ACM Trans. Database Syst.1
1983 A Methodology for Database Schema Mapping from Extended Entity-Relationship Models into the Hierarchical Model
Shamkant B. Navathe, A. Cheng
ER1
1983 Distribution Design of Logical Database Schemas
abstract
The optimal distribution of a database schema over a number of sites in a distributed network is considered. The database is modeled in terms of objects (relations or record sets) and links (predefined joins or CODASYL sets). The design is driven by user-supplied information about data distribution. The inputs required by the optimization model are: 1) cardinality and size information about objects and links, 2) a set of candidate horizontal partitions of relations into fragments and the allocations of the fragments, and 3) the specification of all important transactions, their frequencies, and their sites of origin.
Stefano Ceri, Shamkant B. Navathe, Gio Wiederhold
IEEE Trans. Software Eng.2
1982 Databases for Computer Aided Design and Manufacturing
Shamkant B. Navathe
VLDB1
1982 A Methodology for View Inegration in Logical Database Design
Shamkant B. Navathe, Suresh G. Gadgil
VLDB1
1980 An Intuitive Approach to Normalize Network Structured Data
Shamkant B. Navathe
VLDB1
1980 Schema Analysis for Database Restructuring
abstract
The problem of generalized restructuring of databases has been addressed with two limitations: first, it is assumed that the restructuring user is able to describe the source and target databases in terms of the implicit data model of a particular methodology; second, the restructuring user is faced with the task of judging the scope and applicability of the defined types of restructuring to his database implementation and then of actually specifying his restructuring needs by translating them into the restructuring operations on a foreign data model. A certain amount of analysis of the logical and physical structure of databases must be performed, and the basic ingredients for such an analysis are developed here. The distinction between hierarchical and nonhierarchical data relationships is discussed, and a classification for database schemata is proposed. Examples are given to illustrate how these schemata arise in the conventional hierarchical and network systems. Application of the schema analysis methodology to restructuring specification is also discussed. An example is presented to illustrate the different implications of restructuring three seemingly identical database structures.
Shamkant B. Navathe
ACM Trans. Database Syst.1
1979 1978 New Orleans Data Base Design Workshop Report
Vincent Y. Lum, Sakti P. Ghosh, Mario Schkolnick, Robert W. Taylor, D. Jefferson, Stanley Y. W. Su, James P. Fry, Toby J. Teorey, B. Yao, D. S. Rund, Beverly K. Kahn, Shamkant B. Navathe, L. Aguilar, William J. Barr, P. E. Jones
VLDB12
1978 View Representation in Logical Database Design
abstract
The process of logical database design consists of four phases: view modeling, view integration, schema optimization and schema mapping. View modeling is defined as the modeling of the usage and information structure perspectives of the real world from the point of view of different users and/or applications. The view integration phase combines these views into a single community view which is subjected to further optimization and mapping. As a result, instances of users' model may be altered and application programs transformed.This paper proposes a scheme for view representation which will facilitate the process of view integration. This is done by enhancing the data abstraction framework proposed by Smith and Smith. It takes into account the instance-level interrelationships among data and the identification of instances via these interrelationships. The usage perspective is incorporated as rules and assertions about schema- and instance-level insertion and deletion.The problem of view integration is briefly addressed. Valid transformations of views are indicated as a part of the integration process.
Shamkant B. Navathe, Mario Schkolnick
SIGMOD Conference1
1977 Schema Analysis for Database Restructuring
Shamkant B. Navathe
VLDB1
1976 Restructuring for Large Data Bases: Three Levels of Abstraction
abstract
The development of a powerful restructuring function involves two important components—the unambiguous specification of the restructuring operations and the realization of these operations in a software system. This paper is directed to the first component in the belief that a precise specification will provide a firm foundation for the development of restructuring algorithms and, subsequently, their implementation. The paper completely defines the semantics of the restructuring of tree structured databases. The delineation of the restructuring function is accomplished by formulating three different levels of abstraction, with each level of abstraction representing successively more detailed semantics of the function. At the first level of abstraction, the schema modification, three types are identified—naming, combining, and relating; these three types are further divided into eight schema operations. The second level of abstraction, the instance operations, constitutes the transformations on the data instances; they are divided into group operations such as replication, factoring, union, etc., and group relation operations such as collapsing, refinement, fusion, etc. The final level, the item value operations, includes the actual item operations, such as copy value, delete value, or create a null value.
Shamkant B. Navathe, James P. Fry
ACM Trans. Database Syst.1
1975 Investigations into the Application of the Relational Model to Data Translation
abstract
Experience with data translation over the last two years has resulted in a definite set of requirements for the normalized representation of data when used as an intermediate form during the process. The logical requirements pertain to the ability to represent a class of data structures including networks. The implementation requirements include the specification of two physical representations termed as the Restructurer Internal Form and the Interchange Form.Recent investigations with the Normal Forms of the relational model have shown that1. they satisfy the logical requirements2. a complete normalization to the third Normal Form requires excessive dependency data and results in reducing the restructuring operation to an identity transformation.3. normalization is not conducive to localizing the restructuring operations in the translator4. the first Normal Form used as a vehicle for translation poses problems with specification of keys and in handling networks.To alleviate the difficulties of normalization, a Modified Normal Form, similar to the First Normal Form, was designed and investigated. On the basis of the accessing, manipulation and processing requirements in the current model of data translation, it is concluded that a software system must be designed to work with the relational Normal Forms and their modified versions.
Shamkant B. Navathe, Alan G. Merten
SIGMOD Conference1
1975 Restructuring for Large Data Bases: Three Levels of Abstraction
abstract
The development of a powerful restructuring function involves two important components - the unambiguous specification of the restructuring operations and the realization of these operations in a software system. We direct our efforts to the first component in the belief that a precise specification will provide a firm foundation for restructuring algorithms and implementations. This paper defines completely the semantics of the restructuring of tree structured data bases.
Shamkant B. Navathe, James P. Fry
VLDB1