EDBT 2026 Demo / reviewers in the wild / expert
Ramakrishnan Srikant
dblp:s/RamSrikant
· DBLP profile ↗
43ranked-venue papers
8as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 41 · 7 first-authorArtificial intelligence and machine learning · 9 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorComputer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
34 papers |
Information retrieval · 58% Data mining · 16% Recommender systems · 11% | |
| Network and information security
11 papers |
Privacy and data protection · 75% Cryptographic protocols and secure computation · 14% Cryptographic primitives and cryptanalysis · 11% | |
| Theoretical computer science
3 papers |
Algorithmic game theory and mechanism design · 95% Algorithms and data structures · 4% Computational geometry · 1% |
Topics — the 30 heaviest of 60, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › user behavior › search behavior
click model |
0.5 | 2 | 2019 | Micro-Browsing Models for Search Snippets · ICDE 2019 User browsing models: relevance versus examination · KDD 2010 |
Information retrieval › user behavior
user browsing models |
0.5 | 2 | 2019 | Micro-Browsing Models for Search Snippets · ICDE 2019 User browsing models: relevance versus examination · KDD 2010 |
Recommender systems
user modeling |
0.2 | 1 | 2015 | User Modeling for a Personal Assistant · WSDM 2015 |
Algorithmic game theory and mechanism design › online advertising
budget pacing |
0.2 | 1 | 2013 | Optimizing budget constrained spend in search advertising · WSDM 2013 |
Algorithmic game theory and mechanism design › auction theory › advertising auctions
sponsored search |
0.2 | 1 | 2013 | Optimizing budget constrained spend in search advertising · WSDM 2013 |
Data mining › pattern mining
association rule mining |
0.2 | 8 | 2003 | Limiting privacy breaches in privacy preserving data mining · PODS 2003 Privacy preserving mining of association rules · KDD 2002 Discovering Predictive Association Rules · KDD 1998 |
Data mining
pattern mining |
0.1 | 9 | 2002 | Privacy preserving mining of association rules · KDD 2002 Discovering Predictive Association Rules · KDD 1998 Mining Association Rules with Item Constraints · KDD 1997 |
Recommender systems
click-through rate prediction |
0.1 | 1 | 2019 | Micro-Browsing Models for Search Snippets · ICDE 2019 |
Information retrieval › search interfaces
search result presentation |
0.1 | 1 | 2019 | Micro-Browsing Models for Search Snippets · ICDE 2019 |
Information retrieval
retrieval models |
0.1 | 1 | 2010 | User browsing models: relevance versus examination · KDD 2010 |
Privacy and data protection › privacy-preserving data analysis
privacy-preserving data mining |
0.1 | 3 | 2003 | Limiting privacy breaches in privacy preserving data mining · PODS 2003 Privacy preserving mining of association rules · KDD 2002 Privacy-Preserving Data Mining · SIGMOD Conference 2000 |
Privacy and data protection
privacy-preserving data sharing |
0.1 | 2 | 2004 | Enabling Sovereign Information Sharing Using Web Services · SIGMOD Conference 2004 Information Sharing Across Private Databases · SIGMOD Conference 2003 |
Cryptographic protocols and secure computation
secure multiparty computation |
0.1 | 2 | 2004 | Enabling Sovereign Information Sharing Using Web Services · SIGMOD Conference 2004 Information Sharing Across Private Databases · SIGMOD Conference 2003 |
Privacy and data protection › privacy policy
p3p |
0.1 | 2 | 2003 | An SPath-based preference language for P3P · WWW 2003 Implementing P3P Using Database Technology · ICDE 2003 |
Information retrieval
indexing |
0.1 | 2 | 2003 | Searching with Numbers · IEEE Trans. Knowl. Data Eng. 2003 Searching with numbers · WWW 2002 |
Privacy and data protection
randomization |
0.1 | 2 | 2003 | Limiting privacy breaches in privacy preserving data mining · PODS 2003 Privacy preserving mining of association rules · KDD 2002 |
Information retrieval › similarity search
all-pairs similarity search |
0.1 | 1 | 2007 | Scaling up all pairs similarity search · WWW 2007 |
Information retrieval
similarity search |
0.1 | 1 | 2007 | Scaling up all pairs similarity search · WWW 2007 |
Privacy and data protection
privacy-preserving data analysis |
0.1 | 1 | 2005 | Privacy Preserving OLAP · SIGMOD Conference 2005 |
Privacy and data protection › privacy-preserving query processing
privacy-preserving OLAP |
0.1 | 1 | 2005 | Privacy Preserving OLAP · SIGMOD Conference 2005 |
Cryptographic primitives and cryptanalysis
encryption |
0.0 | 1 | 2004 | Order-Preserving Encryption for Numeric Data · SIGMOD Conference 2004 |
Cryptographic primitives and cryptanalysis › encryption › property-preserving encryption
order-preserving encryption |
0.0 | 1 | 2004 | Order-Preserving Encryption for Numeric Data · SIGMOD Conference 2004 |
Web and social media mining
social network analysis |
0.0 | 1 | 2003 | Mining newsgroups using networks arising from social behavior · WWW 2003 |
Privacy and data protection
privacy policy |
0.0 | 1 | 2003 | Implementing P3P Using Database Technology · ICDE 2003 |
Cryptographic protocols and secure computation
private set intersection |
0.0 | 1 | 2003 | Information Sharing Across Private Databases · SIGMOD Conference 2003 |
Query processing and optimization › similarity join
high-dimensional similarity join |
0.0 | 1 | 2002 | High-Dimensional Similarity Joins · IEEE Trans. Knowl. Data Eng. 2002 |
Privacy and data protection › privacy engineering
hippocratic databases |
0.0 | 1 | 2002 | Hippocratic Databases · VLDB 2002 |
Privacy and data protection
privacy-preserving data management |
0.0 | 1 | 2002 | Hippocratic Databases · VLDB 2002 |
Information retrieval
click-through rate |
0.0 | 1 | 2010 | User browsing models: relevance versus examination · KDD 2010 |
Information retrieval
evaluation |
0.0 | 1 | 2010 | User browsing models: relevance versus examination · KDD 2010 |
Methods — techniques the papers use, named apart from their topics
classification · 0.4context identification algorithm · 0.2randomization · 0.2online algorithms · 0.2linear programming · 0.2indexing · 0.1cosine similarity · 0.1probabilistic modeling · 0.1data perturbation · 0.1web services · 0.1database querying technology · 0.1amplification · 0.1secure multiparty protocols · 0.0pseudorandom generator · 0.0XPath · 0.0APPEL translation · 0.0randomization operators · 0.0epsilon-kdb tree · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Micro-Browsing Models for Search SnippetsabstractClick-through rate (CTR) is a key signal of relevance for search engine results, both organic and sponsored. CTR of a result has two core components: (a) the probability of examination of a result by a user, and (b) the perceived relevance of the result given that it has been examined by the user. There has been considerable work on user browsing models, to model and analyze both the examination and the relevance components of CTR. In this paper, we propose a novel formulation: a micro-browsing model for how users read result snippets. The snippet text of a result often plays a critical role in the perceived relevance of the result. We study how particular words within a line of snippet can influence user behavior. We validate this new micro-browsing user model by considering the problem of predicting which snippet will yield higher CTR, and show that classification accuracy is dramatically higher with our micro-browsing user model. The key insight in this paper is that varying relatively few words within a snippet, and even their location within a snippet, can have a significant influence on the clickthrough of a snippet. Muhammad Asiful Islam, Ramakrishnan Srikant, Sugato Basu |
ICDE | 2 |
| 2015 | User Modeling for a Personal AssistantabstractWe present a user modeling system that serves as the foundation of a personal assistant. The system ingests web search history for signed-in users, and identifies coherent contexts that correspond to tasks, interests, and habits. Unlike past work which focused on either in-session tasks or tasks over a few days, we look at several months of history in order to identify not just short-term tasks, but also long-term interests and habits. The features we use for identifying coherent contexts yield substantially higher precision and recall than past work. We also present an algorithm for identifying contexts that is 8 to 30 times faster than previous algorithms. The user modeling system has been deployed in production. It runs over hundreds of millions of users, and updates the models with a 10-minute latency. The contexts identified by the system serve as the foundation for generating recommendations in Google Now. Ramanathan V. Guha, Vineet Gupta 0001, Vivek Raghunathan, Ramakrishnan Srikant |
WSDM | 4 |
| 2013 | Optimizing budget constrained spend in search advertisingabstractSearch engine ad auctions typically have a significant fraction of advertisers who are budget constrained, i.e., if allowed to participate in every auction that they bid on, they would spend more than their budget. This yields an important problem: selecting the ad auctions which these advertisers participate, in order to optimize different system objectives such as the return on investment for advertisers, and the quality of ads shown to users. We present a system and algorithms for optimizing budget constrained spend. The system is designed be deployed in a large search engine, with hundreds of thousands of advertisers, millions of searches per hour, and with the query stream being only partially predictable. We have validated the system design by implementing it in the Google ads serving system and running experiments on live traffic. We have also compared our algorithm to previous work that casts this problem as a large linear programming problem limited to popular queries, and show that our algorithms yield substantially better results. Chinmay Karande, Aranyak Mehta, Ramakrishnan Srikant |
WSDM | 3 |
| 2010 | User browsing models: relevance versus examinationabstractThere has been considerable work on user browsing models for search engine results, both organic and sponsored. The click-through rate (CTR) of a result is the product of the probability of examination (will the user look at the result) times the perceived relevance of the result (probability of a click given examination). Past papers have assumed that when the CTR of a result varies based on the pattern of clicks in prior positions, this variation is solely due to changes in the probability of examination. Ramakrishnan Srikant, Sugato Basu, Daryl Pregibon |
KDD | 1 |
| 2007 | Scaling up all pairs similarity searchabstractGiven a large collection of sparse vector data in a high dimensional space, we investigate the problem of finding all pairs of vectors whose similarity score (as determined by a function such as cosine distance) is above a given threshold. We propose a simple algorithm based on novel indexing and optimization strategies that solves this problem without relying on approximation methods or extensive parameter tuning. We show the approach efficiently handles a variety of datasets across a wide setting of similarity thresholds, with large speedups over previous state-of-the-art approaches. Roberto J. Bayardo, Ramakrishnan Srikant |
WWW | 3 |
| 2005 | Privacy Preserving OLAPabstractWe present techniques for privacy-preserving computation of multidimensional aggregates on data partitioned across multiple clients. Data from different clients is perturbed (randomized) in order to preserve privacy before it is integrated at the server. We develop formal notions of privacy obtained from data perturbation and show that our perturbation provides guarantees against privacy breaches. We develop and analyze algorithms for reconstructing counts of subcubes over perturbed data. We also evaluate the tradeoff between privacy guarantees and reconstruction accuracy and show the practicality of our approach. Rakesh Agrawal 0001, Ramakrishnan Srikant, Dilys Thomas |
SIGMOD Conference | 2 |
| 2005 | XPref: a preference language for P3P
Rakesh Agrawal 0001, Jerry Kiernan, Ramakrishnan Srikant, Yirong Xu |
Comput. Networks | 3 |
| 2004 | An Implementation of P3P Using Database Technology
Rakesh Agrawal 0001, Jerry Kiernan, Ramakrishnan Srikant, Yirong Xu |
EDBT | 3 |
| 2004 | Enabling Sovereign Information Sharing Using Web ServicesabstractSovereign information sharing allows autonomous entities to compute queries across their databases in such a way that nothing apart from the result is revealed. We describe an implementation of this model using web services infrastructure. Each site participating in sovereign sharing offers a data service that allows database operations to be applied on the tables they own. Of particular interest is the provision for binary operations such as relational joins. Applications are developed by combining these data services. We present performance measurements that show the promise of a new breed of practical applications based on the paradigm of sovereign information integration. Rakesh Agrawal 0001, Dmitri Asonov, Ramakrishnan Srikant |
SIGMOD Conference | 3 |
| 2004 | Order-Preserving Encryption for Numeric DataabstractEncryption is a well established technology for protecting sensitive data. However, once encrypted, data can no longer be easily queried aside from exact matches. We present an order-preserving encryption scheme for numeric data that allows any comparison operation to be directly applied on encrypted data. Query results produced are sound (no false hits) and complete (no false drops). Our scheme handles updates gracefully and new values can be added without requiring changes in the encryption of other values. It allows standard databse indexes to be built over encrypted tables and can easily be integrated with existing database systems. The proposed scheme has been designed to be deployed in application environments in which the intruder can get access to the encrypted database, but does not have prior domain information such as the distribution of values and annot encrypt or decrypt arbitrary values of his choice. The encryption is robust against estimation of the true value in such environments. Rakesh Agrawal 0001, Jerry Kiernan, Ramakrishnan Srikant, Yirong Xu |
SIGMOD Conference | 3 |
| 2004 | Auditing Compliance with a Hippocratic Database
Rakesh Agrawal 0001, Roberto J. Bayardo, Christos Faloutsos, Jerry Kiernan, Ralf Rantzau, Ramakrishnan Srikant |
VLDB | 6 |
| 2004 | Whither Data Mining?
Rakesh Agrawal 0001, Ramakrishnan Srikant |
VLDB | 2 |
| 2004 | Privacy preserving mining of association rules
Alexandre V. Evfimievski, Ramakrishnan Srikant, Rakesh Agrawal 0001, Johannes Gehrke |
Inf. Syst. | 2 |
| 2003 | Implementing P3P Using Database TechnologyabstractPlatform for privacy preferences (P3P) is the most significant effort currently underway to enable Web users to gain control over their private information. P3P provides mechanisms for Web site owners to express their privacy policies in a standard format that a user can programmatically check against her privacy preferences to decide whether to release her data to the Web site. We discuss architectural alternatives for implementing P3P and present a server-centric implementation that reuses database querying technology, as opposed to the prevailing client-centric implementations based on specialized engines. Not only does the proposed implementation have qualitative advantages, our experiments indicate that it performs significantly better than the sole public-domain client-centric implementation and that the latency introduced by preference matching is small enough for real-world deployments of P3P. Rakesh Agrawal 0001, Jerry Kiernan, Ramakrishnan Srikant, Yirong Xu |
ICDE | 3 |
| 2003 | Limiting privacy breaches in privacy preserving data miningabstractThere has been increasing interest in the problem of building accurate data mining models over aggregate data, while protecting privacy at the level of individual records. One approach for this problem is to randomize the values in individual records, and only disclose the randomized values. The model is then built over the randomized data, after first compensating for the randomization (at the aggregate level). This approach is potentially vulnerable to privacy breaches: based on the distribution of the data, one may be able to learn with high confidence that some of the randomized records satisfy a specified property, even though privacy is preserved on average.In this paper, we present a new formulation of privacy breaches, together with a methodology, "amplification", for limiting them. Unlike earlier approaches, amplification makes it is possible to guarantee limits on privacy breaches without any knowledge of the distribution of the original data. We instantiate this methodology for the problem of mining association rules, and modify the algorithm from [9] to limit privacy breaches without knowledge of the data distribution. Next, we address the problem that the amount of randomization required to avoid privacy breaches (when mining association rules) results in very long transactions. By using pseudorandom generators and carefully choosing seeds such that the desired items from the original transaction are present in the randomized transaction, we can send just the seed instead of the transaction, resulting in a dramatic drop in communication and storage cost. Finally, we define new information measures that take privacy breaches into account when quantifying the amount of privacy preserved by randomization. Alexandre V. Evfimievski, Johannes Gehrke, Ramakrishnan Srikant |
PODS | 3 |
| 2003 | Information Sharing Across Private DatabasesabstractLiterature on information integration across databases tacitly assumes that the data in each database can be revealed to the other databases. However, there is an increasing need for sharing information across autonomous entities in such a way that no information apart from the answer to the query is revealed. We formalize the notion of minimal information sharing across private databases, and develop protocols for intersection, equijoin, intersection size, and equijoin size. We also show how new applications can be built using the proposed protocols. Rakesh Agrawal 0001, Alexandre V. Evfimievski, Ramakrishnan Srikant |
SIGMOD Conference | 3 |
| 2003 | An SPath-based preference language for P3PabstractThe Platform for Privacy Preferences (P3P) is the most significant effort currently underway to enable web users to gain control over their private information. The designers of P3P simultaneously designed a preference language called APPEL to allow users to express their privacy preferences, thus enabling automatic matching of privacy preferences against P3P policies. Unfortunately subtle interactions between P3P and APPEL result in serious problems when using APPEL: Users can only directly specify what is unacceptable in a policy, not what is acceptable; simple preferences are hard to express; and writing APPEL preferences is error prone. We show that these problems follow from a fundamental design choice made by APPEL, and cannot be solved without completely redesigning the language. Therefore we explore alternatives to APPEL that can overcome these problems. In particular, we show that XPath serves quite nicely as a preference language and solves all the above problems. We identify the minimal subset of XPath that is needed, thus allowing matching programs to potentially use a smaller memory footprint. We also give an APPEL to XPath translator that shows that XPath is as expressive as APPEL. Rakesh Agrawal 0001, Jerry Kiernan, Ramakrishnan Srikant, Yirong Xu |
WWW | 3 |
| 2003 | Mining newsgroups using networks arising from social behaviorabstractRecent advances in information retrieval over hyperlinked corpora have convincingly demonstrated that links carry less noisy information than text. We investigate the feasibility of applying link-based methods in new applications domains. The specific application we consider is to partition authors into opposite camps within a given topic in the context of newsgroups. A typical newsgroup posting consists of one or more quoted lines from another posting followed by the opinion of the author. This social behavior gives rise to a network in which the vertices are individuals and the links represent "responded-to" relationships. An interesting characteristic of many newsgroups is that people more frequently respond to a message when they disagree than when they agree. This behavior is in sharp contrast to the WWW link graph, where linkage is an indicator of agreement or common interest. By analyzing the graph structure of the responses, we are able to effectively classify people into opposite camps. In contrast, methods based on statistical analysis of text yield low accuracy on such datasets because the vocabulary used by the two sides tends to be largely identical, and many newsgroup postings consist of relatively few words of text. Rakesh Agrawal 0001, Sridhar Rajagopalan, Ramakrishnan Srikant, Yirong Xu |
WWW | 3 |
| 2003 | Searching with NumbersabstractA large fraction of the useful Web is comprised of specification documents that largely consist of (attribute name, numeric value) pairs embedded in text. Examples include product information, classified advertisements, resumes, etc. The approach taken in the past to search these documents by first establishing correspondences between values and their names has achieved limited success because of the difficulty of extracting this information from free text. We propose a new approach that does not require this correspondence to be accurately established. Provided the data has "low reflectivity", we can do effective search even if the values in the data have not been assigned attribute names and the user has omitted attribute names in the query. We give algorithms and indexing structures for implementing the search. We also show how hints (i.e., imprecise, partial correspondences) from automatic data extraction techniques can be incorporated into our approach for better accuracy on high reflectivity data sets. Finally, we validate our approach by showing that we get high precision in our answers on real data sets from a variety of domains. Rakesh Agrawal 0001, Ramakrishnan Srikant |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2002 | Privacy preserving mining of association rulesabstractWe present a framework for mining association rules from transactions consisting of categorical items where the data has been randomized to preserve privacy of individual transactions. While it is feasible to recover association rules and preserve privacy using a straightforward "uniform" randomization, the discovered rules can unfortunately be exploited to find privacy breaches. We analyze the nature of privacy breaches and propose a class of randomization operators that are much more effective than uniform randomization in limiting the breaches. We derive formulae for an unbiased support estimator and its variance, which allow us to recover itemset supports from randomized datasets, and show how to incorporate these formulae into mining algorithms. Finally, we present experimental results that validate the algorithm by applying it on real datasets. Alexandre V. Evfimievski, Ramakrishnan Srikant, Rakesh Agrawal 0001, Johannes Gehrke |
KDD | 2 |
| 2002 | Privacy Preserving Data Mining: Challenges and Opportunities
Ramakrishnan Srikant |
PAKDD | 1 |
| 2002 | Hippocratic Databases
Rakesh Agrawal 0001, Jerry Kiernan, Ramakrishnan Srikant, Yirong Xu |
VLDB | 3 |
| 2002 | Database Technologies for Electronic Commerce
Rakesh Agrawal 0001, Ramakrishnan Srikant, Yirong Xu |
VLDB | 2 |
| 2002 | Searching with numbersabstractA large fraction of the useful web comprises of specification documents that largely consist of hattribute name, numeric valuei pairs embedded in text. Examples include product information, classified advertisements, resumes, etc. The approach taken in the past to search these documents by first establishing correspondences between values and their names has achieved limited success because of the difficulty of extracting this information from free text. We propose a new approach that does not require this correspondence to be accurately established. Provided the data has "low reflectivity", we can do effective search even if the values in the data have not been assigned attribute names and the user has omitted attribute names in the query. We give algorithms and indexing structures for implementing the search. We also show how hints (i. e, imprecise, partial correspondences) from automatic data extraction techniques can be incorporated into our approach for better accuracy on high reflectivity datasets. Finally, we validate our approach by showing that we get high precision in our answers on real datasets from a variety of domains. Rakesh Agrawal 0001, Ramakrishnan Srikant |
WWW | 2 |
| 2002 | High-Dimensional Similarity JoinsabstractMany emerging data mining applications require a similarity join between points in a high-dimensional domain. We present a new algorithm that utilizes a new index structure, called the /spl epsi/ tree, for fast spatial similarity joins on high-dimensional points. This index structure reduces the number of neighboring leaf nodes that are considered for the join test, as well as the traversal cost of finding appropriate branches in the internal nodes. The storage cost for internal nodes is independent of the number of dimensions. Hence, the proposed index structure scales to high-dimensional data. We analyze the cost of the join for the /spl epsi/ tree and the R-tree family, and show that the /spl epsi/ tree will perform better for high-dimensional joins. Empirical evaluation, using synthetic and real-life data sets, shows that similarity join using the /spl epsi/ tree is twice to an order of magnitude faster than the R/sup +/ tree, with the performance gap increasing with the number of dimensions. We also discuss how some of the ideas of the /spl epsi/ tree can be applied to the R-tree family. These biased R-trees perform better than the corresponding traditional R-trees for high-dimensional similarity joins, but do not match the performance of the /spl epsi/ tree. Kyuseok Shim, Ramakrishnan Srikant, Rakesh Agrawal 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2001 | On integrating catalogsabstractWe address the problem of integrating documents from different sources into a master catalog. This problem is pervasive in web marketplaces and portals. Current technology for automating this process consists of building a classifier that uses the categorization of documents in the master catalog to construct a model for predicting the category of unknown documents. Our key insight is that many of the data sources have their own categorization, and classification accuracy can be improved by factoring in the implicit information in these source categorizations. We show how a Naive Bayes classification can be enhanced to incorporate the similarity information present in source catalogs. Our analysis and empirical evaluation show substantial improvement in the accuracy of catalog integration. Keywords: Classification, Categorization, Data Mining, Catalog Integration, Web Portals, Web Marketplaces 1. Rakesh Agrawal 0001, Ramakrishnan Srikant |
WWW | 2 |
| 2001 | Mining web logs to improve website organizationabstractMany websites have a hierarchical organization of content. This organization may be quite different from the organization expected by visitors to the website. In particular, it is often unclear where a specific document is located. In this paper, we propose an algorithm to automatically find pages in a website whose location is different from where visitors expect to find them. The key insight is that visitors will backtrack if they do not find the information where they expect it: the point from where they backtrack is the expected location for the page. We present an algorithm for discovering such expected locations that can handle page caching by the browser. Expected locations with a significant number of hits are then presented to the website administrator. We also present algorithms for selecting expected locations (for adding navigation links) to optimize the benefit to the website or the visitor. We ran our algorithm on the Wharton business school website and found that even on this small website, there were many pages with expected locations different from their actual location. 1. Ramakrishnan Srikant, Yinghui Yang 0001 |
WWW | 1 |
| 2000 | Athena: Mining-Based Interactive Management of Text Database
Rakesh Agrawal 0001, Roberto J. Bayardo, Ramakrishnan Srikant |
EDBT | 3 |
| 2000 | Privacy-Preserving Data Mining
Rakesh Agrawal 0001, Ramakrishnan Srikant |
SIGMOD Conference | 2 |
| 1998 | Discovering Predictive Association Rules
Nimrod Megiddo, Ramakrishnan Srikant |
KDD | 2 |
| 1997 | High-Dimensional Similarity JoinsabstractMany emerging data mining applications require a similarity join between points in a high-dimensional domain. We present a new algorithm that utilizes a new index structure, called the /spl epsiv/-kdB tree, for fast spatial similarity joins on high-dimensional points. This index structure reduces the number of neighboring leaf nodes that are considered for the join test, as well as the traversal cost of finding appropriate branches in the internal nodes. The storage cost for internal nodes is independent of the number of dimensions. Hence the proposed index structure scales to high-dimensional data. Empirical evaluation, using synthetic and real-life datasets, shows that similarity join using the /spl epsiv/-kdB tree is 2 to an order of magnitude faster than the R/sup +/ tree, with the performance gap increasing with the number of dimensions. Kyuseok Shim, Ramakrishnan Srikant, Rakesh Agrawal 0001 |
ICDE | 2 |
| 1997 | Partial Classification Using Association Rules
Kamal Ali, Stefanos Manganaris, Ramakrishnan Srikant |
KDD | 3 |
| 1997 | Discovering Trends in Text Databases
Brian Lent, Rakesh Agrawal 0001, Ramakrishnan Srikant |
KDD | 3 |
| 1997 | Mining Association Rules with Item Constraints
Ramakrishnan Srikant, Quoc Vu, Rakesh Agrawal 0001 |
KDD | 1 |
| 1997 | Range Queries in OLAP Data CubesabstractA range query applies an aggregation operation over all selected cells of an OLAP data cube where the selection is specified by providing ranges of values for numeric dimensions. We present fast algorithms for range queries for two types of aggregation operations: SUM and MAX. These two operations cover techniques required for most popular aggregation operations, such as those supported by SQL. C. T. Howard Ho, Rakesh Agrawal 0001, Nimrod Megiddo, Ramakrishnan Srikant |
SIGMOD Conference | 4 |
| 1997 | Mining generalized association rules
Ramakrishnan Srikant, Rakesh Agrawal 0001 |
Future Gener. Comput. Syst. | 1 |
| 1996 | Mining Sequential Patterns: Generalizations and Performance Improvements
Ramakrishnan Srikant, Rakesh Agrawal 0001 |
EDBT | 1 |
| 1996 | The Quest Data Mining System
Rakesh Agrawal 0001, Manish Mehta 0002, John C. Shafer, Ramakrishnan Srikant, Andreas Arning, Toni Bollinger |
KDD | 4 |
| 1996 | Mining Quantitative Association Rules in Large Relational TablesabstractWe introduce the problem of mining association rules in large relational tables containing both quantitative and categorical attributes. An example of such an association might be "10% of married people between age 50 and 60 have at least 2 cars". We deal with quantitative attributes by fine-partitioning the values of the attribute and then combining adjacent partitions as necessary. We introduce measures of partial completeness which quantify the information lost due to partitioning. A direct application of this technique can generate too many similar rules. We tackle this problem by using a "greater-than-expected-value" interest measure to identify the interesting rules in the output. We give an algorithm for mining such quantitative association rules. Finally, we describe the results of using this approach on a real-life dataset. Ramakrishnan Srikant, Rakesh Agrawal 0001 |
SIGMOD Conference | 1 |
| 1995 | Mining Sequential PatternsabstractWe are given a large database of customer transactions, where each transaction consists of customer-id, transaction time, and the items bought in the transaction. We introduce the problem of mining sequential patterns over such databases. We present three algorithms to solve this problem, and empirically evaluate their performance using synthetic data. Two of the proposed algorithms, AprioriSome and AprioriAll, have comparable performance, albeit AprioriSome performs a little better when the minimum number of customers that must support a sequential pattern is low. Scale-up experiments show that both AprioriSome and AprioriAll scale linearly with the number of customer transactions. They also have excellent scale-up properties with respect to the number of transactions per customer and the number of items in a transaction.> Rakesh Agrawal 0001, Ramakrishnan Srikant |
ICDE | 2 |
| 1995 | Mining Generalized Association Rules
Ramakrishnan Srikant, Rakesh Agrawal 0001 |
VLDB | 1 |
| 1994 | Quest: A Project on Database MiningabstractNo abstract available. Rakesh Agrawal 0001, Michael J. Carey 0001, Christos Faloutsos, Sakti P. Ghosh, Maurice A. W. Houtsma, Tomasz Imielinski, Balakrishna R. Iyer, A. Mahboob, H. Miranda, Ramakrishnan Srikant, Arun N. Swami |
SIGMOD Conference | 10 |
| 1994 | Fast Algorithms for Mining Association Rules in Large Databases
Rakesh Agrawal 0001, Ramakrishnan Srikant |
VLDB | 2 |