VLDB 2026 Research / reviewers in the wild / expert
Sudarshan S. Chawathe
dblp:c/SSChawathe
· DBLP profile ↗
27ranked-venue papers
21as first author
0since 2021 · last 2020
0000-0001-6930-0325ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 14 · 10 first-authorSecurity and privacy · 5 · 4 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
12 papers |
Data models and query languages · 29% Data stream processing · 27% Query processing and optimization · 25% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data stream processing
XML stream processing |
0.1 | 2 | 2005 | XSQ: A streaming XPath engine · ACM Trans. Database Syst. 2005 XPath Queries on Streaming Data · SIGMOD Conference 2003 |
Data models and query languages
semistructured data |
0.1 | 3 | 2001 | VQBD: Exploring Semistructured Data · SIGMOD Conference 2001 Representing and Querying Changes in Semistructured Data · ICDE 1998 Representative Objects: Concise Representations of Semistructured, Hierarchial Data · ICDE 1997 |
Query processing and optimization › XML query processing
XPath query evaluation |
0.1 | 1 | 2005 | XSQ: A streaming XPath engine · ACM Trans. Database Syst. 2005 |
Query processing and optimization › XML query processing
streaming XPath evaluation |
0.0 | 1 | 2003 | Streaming XPath Queries in XSQ · ICDE 2003 |
Query processing and optimization
XML query processing |
0.0 | 1 | 2003 | Streaming XPath Queries in XSQ · ICDE 2003 |
Data mining › time series analysis
change point detection |
0.0 | 2 | 1997 | Meaningful Change Detection in Structured Data · SIGMOD Conference 1997 Change Detection in Hierarchically Structured Information · SIGMOD Conference 1996 |
Data models and query languages
hierarchical data |
0.0 | 2 | 1997 | Meaningful Change Detection in Structured Data · SIGMOD Conference 1997 Change Detection in Hierarchically Structured Information · SIGMOD Conference 1996 |
Data models and query languages › query language
semistructured query language |
0.0 | 1 | 2001 | VQBD: Exploring Semistructured Data · SIGMOD Conference 2001 |
Data models and query languages › data modeling
hierarchical data model |
0.0 | 1 | 1999 | Comparing Hierarchical Data in External Memory · VLDB 1999 |
Data integration and cleaning
schema inference |
0.0 | 1 | 1997 | Representative Objects: Concise Representations of Semistructured, Hierarchial Data · ICDE 1997 |
Data integration and cleaning
heterogeneous information systems |
0.0 | 1 | 1996 | A Toolkit for Constraint Management in Heterogeneous Information Systems · ICDE 1996 |
Database system architecture and tuning › database design › physical database design
index selection |
0.0 | 1 | 1994 | On Index Selection Schemes for Nested Object Hierarchies · VLDB 1994 |
Information retrieval › query formulation
query generation |
0.0 | 1 | 1997 | Representative Objects: Concise Representations of Semistructured, Hierarchial Data · ICDE 1997 |
Transaction processing and concurrency control
consistency |
0.0 | 1 | 1996 | A Toolkit for Constraint Management in Heterogeneous Information Systems · ICDE 1996 |
Methods — techniques the papers use, named apart from their topics
hierarchical finite state automata · 0.1schema discovery · 0.0minimum-cost edge cover · 0.0heuristic algorithm · 0.0tree edit distance algorithms · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Organizing and compressing collections of files using differencesabstractA collection of related files often exhibits strong similarities among its constituents. These similarities, and the dual differences, may be used for both compressing the collection and for organizing it in a manner that reveals human-readable structure and relationships. This paper motivates and studies methods for such organizing and compression of file collections using inter-file differences. It presents an algorithm based on computing a minimum-weight spanning tree of a graph that has vertices corresponding to files and edges with weights corresponding to the size of the difference between the documents of its incident vertices. It describes the design and implementation of a prototype system called diboc (for difference-based organization and compression) that uses these methods to enable both compression and graphical organization and interactive exploration of a file collection. It illustrates the benefits of this system by presenting examples of its operation on a widely deployed and publicly available corpus of file collections (collections of PPD files used to configure the CUPS printing system as packaged by the Debian GNU/Linux distribution). In addition to these qualitative measures, some quantitative experimental results of applying the methods to the same corpus are also presented. Sudarshan S. Chawathe |
IDEAS | 1 |
| 2020 | Efficient File Collections for Embedded DevicesabstractThis paper studies methods for efficiently transferring and storing collections of related files in embedded devices and other environments with limitations on storage, network, and energy use. Files in collections based on purpose (e.g., system configurations) or other aspects often exhibit substantial inter-file similarities. These similarities may be used to achieve significant reductions in the network resources required for transferring or updating the collection, as well as for the storage resources required on the embedded devices on which it is stored. Sudarshan S. Chawathe |
ISCC | 1 |
| 2019 | Indoor-Location Classification Using RF SignaturesabstractIndoor locations are classified using spatial signatures of RF signals that are obtained using a measurement grid with a spacing of approximately a wavelength. The classification method is evaluated using a publicly available dataset of detailed signal measurements in a real environment. The experimental results suggest not only that high accuracy is achievable using much simpler signatures than those in prior work but also that this accuracy is maintained as the grid is significantly coarsened. Sudarshan S. Chawathe |
NCA | 1 |
| 2019 | Cost-Based Query-Rewriting for DynamoDB: Work in ProgressabstractThe following topics are dealt with: telecommunication traffic; learning (artificial intelligence); computer network security; optimisation; resource allocation; software defined networking; cloud computing; Internet of Things; virtualisation; quality of service. Sudarshan S. Chawathe |
NCA | 1 |
| 2018 | Analysis of Burst Header Packets in Optical Burst Switching NetworksabstractOptical Burst Switching (OBS) networks provide a practical alternative to optical packet switching and optical circuit switching by separating control information from the primary data, sending the former on a separate control channel. However, this separation also renders OBS networks susceptible to a denial- or degradation-of-service attack (intentional or otherwise) when the data provisioned by a header packet on the control channel does not materialize. This paper addresses the problem of detecting and characterizing such problems and describes a method based on monitoring network traffic on the control and data channels. The method is evaluated on a publicly available dataset. Sudarshan S. Chawathe |
NCA | 1 |
| 2018 | Monitoring IoT Networks for Botnet ActivityabstractThe Internet of Things (IoT) has rapidly transitioned from a novelty to a common, and often critical, part of residential, business, and industrial environments. Security vulnerabilities and exploits in the IoT realm have been well documented. In many cases, improving the security of an IoT device by hardening its software is not a realistic option, especially in the cost-sensitive consumer market or in legacy-bound industrial settings. As part of a multifaceted defense against botnet activity on the IoT, this paper explores a method based on monitoring the network activity of IoT devices. A notable benefit of this approach is that it does not require any special access to the devices and adapts well to the addition of new devices. The method is evaluated on a publicly available dataset drawn from a real IoT network. Sudarshan S. Chawathe |
NCA | 1 |
| 2009 | Effective whitelisting for filesystem forensicsabstractForensic analysis of the large filesystems commonly found on current computers requires an effective method for categorizing and prioritizing files in order to avoid overwhelming the investigator. A key technique for this purpose is whitelisting files, i.e., skipping the detailed analysis of files that match files in a well known reference collection of files. Effective use of this technique requires an efficient method to match files, detecting not only exact matches, but also near matches or approximate matches. This paper outlines the requirements for such matching, formalizes them as the bounded best match and approximate bounded near-match problems, and describes methods to solve these problems. In particular, the approximate bounded near-match problem is mapped to the problem of finding near neighbors in a high-dimensional metric space and solved using locality-sensitive hashing. Sudarshan S. Chawathe |
ISI | 1 |
| 2007 | Organizing Hot-Spot Police Patrol RoutesabstractWe address the problem of planning patrol routes to maximize coverage of important locations (hot spots) at minimum cost (length of patrol route). We model a road network using an edge-weighted graph in which edges represent streets, vertices represent intersections, and weights represent importance of the corresponding streets. We describe efficient methods that use this input to determine the most important patrol routes. In addition to the importance of streets (edge weights), important routes are affected by the topology of the road network. Our methods permit automation of a labor-intensive stage of the patrol-planning process and aid dynamic adjustment of patrol routes in response to changes in the input graph (as a result of a developing situation, for instance). Sudarshan S. Chawathe |
ISI | 1 |
| 2006 | Tracking Changes in Healthcare DocumentsabstractWe present methods for monitoring a large, diverse, and autonomously modified collection of healthcare documents on the Web. Our methods do not require document-providers to offer any special services. They are based on explicating changes between document versions on a per-user basis by using differencing algorithms. These changes are presented to users in the context of the documents using special XML elements. In order to effectively browse changes in large document collections, we use a variable-resolution XML browser. A noteworthy feature of this browser is that it produces usable displays at any level of detail specified by a user Sudarshan S. Chawathe |
CBMS | 1 |
| 2006 | Strategic Web-Service AgreementsabstractThis paper addresses issues of strategy in Web-service composition, and inter-site collaboration in general. The results are useful for both human and artificial agents of a site who are responsible for determining Web-service agreements that are most profitable for that site. We discuss three specific questions in this area. (1) How should the profit resulting from a composition of Web-services be divided among the participants? (2) How can we counter the risk of service-providers misrepresenting their services in an attempt to gain a larger share of the profit? (3) Are stable configurations guaranteed or feasible when each service-provider's decisions on how to collaborate (permit compositions) are guided solely by the goal of maximizing its profit? Sudarshan S. Chawathe |
ICWS | 1 |
| 2006 | Distributing the Cost of Securing a Transportation Infrastructure
Sudarshan S. Chawathe |
ISI | 1 |
| 2005 | Differencing Data StreamsabstractWe present external-memory algorithms for differencing large hierarchical datasets. Our methods are especially suited to streaming data with bounded differences. For input sizes m and n and maximum output (difference) size e, the I/O, RAM, and CPU costs of our algorithm rdiff are, respectively, m + n, 4e + 8, and O(MN). That is, given 4e + 8 blocks of RAM, our algorithm performs no I/O operations other than those required to read both inputs. We also present a variant of the algorithm that uses only four blocks of RAM, with I/O cost 8me + 18m + n + 6e + 5 and CPU cost O(MN). Sudarshan S. Chawathe |
IDEAS | 1 |
| 2005 | XSQ: A streaming XPath engineabstractWe have implemented and released the XSQ system for evaluating XPath queries on streaming XML data. XSQ supports XPath features such as multiple predicates, closures, and aggregation, which pose interesting challenges for streaming evaluation. Our implementation is based on using a hierarchical arrangement of augmented finite state automata. A design goal of XSQ is buffering data for the least amount of time possible. We present a detailed experimental study that characterizes the performance of XSQ and related systems, and that illustrates the performance implications of XPath features such as closures. Sudarshan S. Chawathe |
ACM Trans. Database Syst. | 2 |
| 2004 | Privacy-Preserving Inter-database Operations
Gang Liang, Sudarshan S. Chawathe |
ISI | 2 |
| 2004 | Managing RFID Data
Sudarshan S. Chawathe, Venkat Krishnamurthy, Sridhar Ramachandran, Sanjay E. Sarma |
VLDB | 1 |
| 2003 | Streaming XPath Queries in XSQ
Sudarshan S. Chawathe |
ICDE | 2 |
| 2003 | Tracking Hidden Groups Using Communications
Sudarshan S. Chawathe |
ISI | 1 |
| 2003 | XPath Queries on Streaming DataabstractWe present the design and implementation of the XSQ system for querying streaming XML data using XPath 1.0. Using a clean design based on a hierarchical arrangement of pushdown transducers augmented with buffers, XSQ supports features such as multiple predicates, closures, and aggregation. XSQ not only provides high throughput, but is also memory efficient: It buffers only data that must be buffered by any streaming XPath processor. We also present an empirical study of the performance characteristics of XPath features, as embodied by XSQ and several other systems. Sudarshan S. Chawathe |
SIGMOD Conference | 2 |
| 2002 | SEuS: Structure Extraction Using Summaries
Shayan Ghazizadeh, Sudarshan S. Chawathe |
Discovery Science | 2 |
| 2001 | VQBD: Exploring Semistructured DataabstractNo abstract available. Sudarshan S. Chawathe, Thomas Baby, Jihwang Yeo |
SIGMOD Conference | 1 |
| 1999 | Comparing Hierarchical Data in External Memory
Sudarshan S. Chawathe |
VLDB | 1 |
| 1998 | Representing and Querying Changes in Semistructured DataabstractSemistructured data may be irregular and incomplete and does not necessarily conform to a fixed schema. As with structured data, it is often desirable to maintain a history of changes to data, and to query over both the data and the changes. Representing and querying changes in semistructured data is more difficult than in structured data due to the irregularity and lack of schema. We present a model for representing changes in semistructured data and a language for querying over these changes. An important feature of our approach is that we represent and query changes directly as annotations on the affected data, instead of indirectly as the difference between database states. We describe the implementation of our model and query language. We also describe the design and implementation of a query subscription service that permits users to subscribe to changes in semistructured information sources. Sudarshan S. Chawathe, Serge Abiteboul, Jennifer Widom |
ICDE | 1 |
| 1997 | Representative Objects: Concise Representations of Semistructured, Hierarchial DataabstractIntroduces the concept of representative objects, which uncover the inherent schema(s) in semi-structured, hierarchical data sources and provide a concise description of the structure of the data. Semi-structured data, unlike data stored in typical relational or object-oriented databases, does not have a fixed schema that is known in advance and stored separately from the data. With the rapid growth of the World Wide Web, semi-structured hierarchical data sources are becoming widely available to the casual user. The lack of external schema information currently makes browsing and querying these data sources inefficient at best, and impossible at worst. We show how representative objects make schema discovery efficient and facilitate the generation of meaningful queries over the data. Svetlozar Nestorov, Jeffrey D. Ullman, Janet L. Wiener, Sudarshan S. Chawathe |
ICDE | 4 |
| 1997 | Meaningful Change Detection in Structured DataabstractDetecting changes by comparing data snapshots is an important requirement for difference queries, active databases, and version and configuration management. In this paper we focus on detecting meaningful changes in hierarchically structured data, such as nested-object data. This problem is much more challenging than the corresponding one for relational or flat-file data. In order to describe changes better, we base our work not just on the traditional “atomic” insert, delete, update operations, but also on operations that move an entire sub-tree of nodes, and that copy an entire sub-tree. These operations allows us to describe changes in a semantically more meaningful way. Since this change detection problem is NP-hard, in this paper we present a heuristic change detection algorithm that yields close to “minimal” descriptions of the changes, and that has fewer restrictions than previous algorithms. Our algorithm is based on transforming the change detection problem to a problem of computing a minimum-cost edge cover of a bipartite graph. We study the quality of the solution produced by our algorithm, as well as the running time, both analytically and experimentally. Sudarshan S. Chawathe, Hector Garcia-Molina |
SIGMOD Conference | 1 |
| 1996 | A Toolkit for Constraint Management in Heterogeneous Information SystemsabstractWe present a framework and a toolkit to monitor and enforce distributed integrity constraints in loosely coupled heterogeneous information systems. Our framework enables and formalizes weakened notions of consistency, which are essential in such environments. Our framework is used to describe: interfaces provided by a database for the data items involved in inter site constraints; strategies for monitoring and enforcing such constraints; guarantees regarding the level of consistency the system can provide. Our toolkit uses this framework to provide a set of configurable modules that are used to monitor and enforce constraints spanning loosely coupled heterogeneous information systems. Sudarshan S. Chawathe, Hector Garcia-Molina, Jennifer Widom |
ICDE | 1 |
| 1996 | Change Detection in Hierarchically Structured InformationabstractDetecting and representing changes to data is important for active databases, data warehousing, view maintenance, and version and configuration management. Most previous work in change management has dealt with flat-file and relational data; we focus on hierarchically structured data. Since in many cases changes must be computed from old and new versions of the data, we define the hierarchical change detection problem as the problem of finding a "minimum-cost edit script" that transforms one data tree to another, and we present efficient algorithms for computing such an edit script. Our algorithms make use of some key domain characteristics to achieve substantially better performance than previous, generalpurpose algorithms. We study the performance of our algorithms both analytically and empirically, and we describe the application of our techniques to hierarchically structured documents. 1 Introduction We study the problem of detecting and representing changes to hierarchically stru... Sudarshan S. Chawathe, Anand Rajaraman, Hector Garcia-Molina, Jennifer Widom |
SIGMOD Conference | 1 |
| 1994 | On Index Selection Schemes for Nested Object Hierarchies
Sudarshan S. Chawathe, Ming-Syan Chen, Philip S. Yu |
VLDB | 1 |