VLDB 2026 Research / reviewers in the wild / expert
Ganesh Ramesh
dblp:81/5709
· DBLP profile ↗
8ranked-venue papers
3as first author
0since 2021 · last 2011
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
2 papers |
Privacy and data protection · 100% | |
| Databases, data mining, and information retrieval
4 papers |
Data mining · 66% Database theory · 26% Query processing and optimization · 8% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Privacy and data protection › privacy-preserving computation
input privacy |
0.1 | 1 | 2007 | Preservation Of Patterns and Input-Output Privacy · ICDE 2007 |
Privacy and data protection › privacy-preserving data analysis
output privacy |
0.1 | 1 | 2007 | Preservation Of Patterns and Input-Output Privacy · ICDE 2007 |
Privacy and data protection
privacy-preserving data analysis |
0.1 | 1 | 2007 | Preservation Of Patterns and Input-Output Privacy · ICDE 2007 |
Privacy and data protection
anonymization |
0.1 | 1 | 2005 | To Do or Not To Do: The Dilemma of Disclosing Anonymized Data · SIGMOD Conference 2005 |
Database theory › conjunctive query
tree pattern query |
0.0 | 1 | 2004 | On Testing Satisfiability of Tree Pattern Queries · VLDB 2004 |
Data mining › pattern mining › itemset mining
frequent itemset mining |
0.0 | 1 | 2003 | Feasible itemset distributions in data mining: theory and application · PODS 2003 |
Data mining
pattern mining |
0.0 | 1 | 2003 | Feasible itemset distributions in data mining: theory and application · PODS 2003 |
Data mining › predictive modeling › classification
decision tree mining |
0.0 | 1 | 2007 | Preservation Of Patterns and Input-Output Privacy · ICDE 2007 |
Data mining › pattern mining › frequent pattern mining
frequent set mining |
0.0 | 1 | 2005 | To Do or Not To Do: The Dilemma of Disclosing Anonymized Data · SIGMOD Conference 2005 |
Methods — techniques the papers use, named apart from their topics
random perturbation · 0.1decision tree induction · 0.1heuristic estimation · 0.1belief functions · 0.1length distribution analysis · 0.1combinatorial bounds · 0.0combinatorial bound · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | k-Anonymization of Social Networks by Vertex Addition
Sean Chester, Bruce M. Kapron, Ganesh Ramesh, Gautam Srivastava 0001, Alex Thomo, S. Venkatesh 0001 |
ADBIS (2) | 3 |
| 2008 | On disclosure risk analysis of anonymized itemsets in the presence of prior knowledgeabstractDecision makers of companies often face the dilemma of whether to release data for knowledge discovery, vis-a-vis the risk of disclosing proprietary or sensitive information. Among the various methods employed for “sanitizing” the data prior to disclosure, we focus in this article on anonymization, given its widespread use in practice. We do due diligence to the question “just how safe is the anonymized data?” We consider both those scenarios when the hacker has no information and, more realistically, when the hacker may have partial information about items in the domain. We conduct our analyses in the context of frequent set mining and address the safety question at two different levels: (i) how likely of being cracked (i.e., re-identified by a hacker), are the identities of individual items and (ii) how likely are sets of items cracked? For capturing the prior knowledge of the hacker, we propose a belief function , which amounts to an educated guess of the frequency of each item. For various classes of belief functions which correspond to different degrees of prior knowledge, we derive formulas for computing the expected number of cracks of single items and for itemsets, the probability of cracking the itemsets. While obtaining, exact values for more general situations is computationally hard, we propose a series of heuristics called the O-estimates . They are easy to compute and are shown fairly accurate, justified by empirical results on real benchmark datasets. Based on the O-estimates, we propose a recipe for the decision makers to resolve their dilemma. Our recipe operates at two different levels, depending on whether the data owner wants to reason in terms of single items or sets of items (or both). Finally, we present techniques for ascertaining a hacker's knowledge of correlation in terms of co-occurrence of items likely. This information regarding the hacker's knowledge can be incorporated into our framework of disclosure risk analysis and we present experimental results demonstrating how this knowledge affects the heuristic estimates we have developed. Laks V. S. Lakshmanan, Raymond T. Ng, Ganesh Ramesh |
ACM Trans. Knowl. Discov. Data | 3 |
| 2007 | Preservation Of Patterns and Input-Output PrivacyabstractPrivacy preserving data mining so far has mainly focused on the data collector scenario where individuals supply their personal data to an untrusted collector in exchange for value. In this scenario, random perturbation has proved to be very successful. An equally compelling, but overlooked scenario, is that of a data custodian, which either owns the data or is explicitly entrusted with ensuring privacy of individual data. In this scenario, we show that it is possible to minimize disclosure while guaranteeing no outcome change. We conduct our investigation in the context of building a decision tree and propose transformations that preserve the exact decision tree. We show with a detailed set of experiments that they provide substantial protection to both input data privacy and mining output privacy. Shaofeng Bu, Laks V. S. Lakshmanan, Raymond T. Ng, Ganesh Ramesh |
ICDE | 4 |
| 2005 | Distribution-Based Synthetic Database Generation Techniques for Itemset MiningabstractThe resource requirements of frequent pattern mining algorithms depend mainly on the length distribution of the mined patterns in the database. Synthetic databases, which are used to benchmark performance of algorithms, tend to have distributions far different from those observed in real datasets. In this paper we focus on the problem of synthetic database generation and propose algorithms to effectively embed within the database, any given set of maximal pattern collections, and make the following contributions: 1. A database generation technique is presented which takes k maximal itemset collections as input, and constructs a database which produces these maximal collections as output, when mined at k levels of support. To analyze the efficiency of the procedure, upper bounds are provided on the number of transactions output in the generated database; 2. A compression method is used and extended to reduce the size of the output database. An optimization to the generation procedure is provided which could potentially reduce the number of transactions generated; 3. Preliminary experimental results are presented to demonstrate the feasibility of using the generation technique. Ganesh Ramesh, Mohammed J. Zaki, William Maniatty |
IDEAS | 1 |
| 2005 | To Do or Not To Do: The Dilemma of Disclosing Anonymized DataabstractDecision makers of companies often face the dilemma of whether to release data for knowledge discovery, vis a vis the risk of disclosing proprietary or sensitive information. While there are various "sanitization" methods, in this paper we focus on anonymization, given its widespread use in practice. We give due diligence to the question of "just how safe the anonymized data is", in terms of protecting the true identities of the data objects. We consider both the scenarios when the hacker has no information, and more realistically, when the hacker may have partial information about items in the domain. We conduct our analyses in the context of frequent set mining. We propose to capture the prior knowledge of the hacker by means of a belief function, where an educated guess of the frequency of each item is assumed. For various classes of belief functions, which correspond to different degrees of prior knowledge, we derive formulas for computing the expected number of "cracks". While obtaining the exact values for the more general situations is computationally hard, we propose a heuristic called the O-estimate. It is easy to compute, and is shown to be accurate empirically with real benchmark datasets. Finally, based on the O-estimates, we propose a recipe for the decision makers to resolve their dilemma. Laks V. S. Lakshmanan, Raymond T. Ng, Ganesh Ramesh |
SIGMOD Conference | 3 |
| 2004 | On Testing Satisfiability of Tree Pattern Queries
Laks V. S. Lakshmanan, Ganesh Ramesh, Wendy Hui Wang, Zheng (Jessica) Zhao |
VLDB | 2 |
| 2003 | Feasible itemset distributions in data mining: theory and applicationabstractComputing frequent itemsets and maximally frequent item-sets in a database are classic problems in data mining. The resource requirements of all extant algorithms for both problems depend on the distribution of frequent patterns, a topic that has not been formally investigated. In this paper, we study properties of length distributions of frequent and maximal frequent itemset collections and provide novel solutions for computing tight lower bounds for feasible distributions. We show how these bounding distributions can help in generating realistic synthetic datasets, which can be used for algorithm benchmarking. Ganesh Ramesh, William Maniatty, Mohammed J. Zaki |
PODS | 1 |
| 2002 | A Text-based for Detection and Filtering of Commercial Segments in Broadcast News
Ganesh Ramesh, Amit Bagga |
LREC | 1 |