EDBT 2026 Demo / reviewers in the wild / expert
Daniel Barbará
dblp:b/DBarbara
· DBLP profile ↗
81ranked-venue papers
42as first author
1since 2021 · last 2024
0000-0002-2830-1038ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 44 · 27 first-authorArtificial intelligence and machine learning · 14 · 8 first-authorSystems, architecture and hardware · 14 · 11 first-authorSecurity and privacy · 11 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5Software engineering, systems software and programming languages · 4Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
23 papers |
Data mining · 59% Web and social media mining · 21% Information retrieval · 7% | |
| Computer architecture, parallel and distributed computing, and storage systems
14 papers |
Distributed systems · 50% Performance modeling and evaluation · 19% Cloud and datacenter computing · 14% | |
| Artificial intelligence
2 papers |
Information extraction and text analysis · 98% Probabilistic and Bayesian machine learning · 2% |
Topics — the 30 heaviest of 83, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › anomaly detection
fraud detection |
0.2 | 2 | 2011 | Spectrum based fraud detection in social networks · ICDE 2011 Spectrum based fraud detection in social networks · CCS 2010 |
Web and social media mining › web spam detection
random link attack detection |
0.2 | 2 | 2011 | Spectrum based fraud detection in social networks · ICDE 2011 Spectrum based fraud detection in social networks · CCS 2010 |
Data mining
anomaly detection |
0.2 | 2 | 2011 | Spectrum based fraud detection in social networks · ICDE 2011 Detecting outliers using transduction and statistical testing · KDD 2006 |
Web and social media mining
social network analysis |
0.1 | 1 | 2011 | Spectrum based fraud detection in social networks · ICDE 2011 |
Data mining
clustering |
0.1 | 3 | 2006 | Categorization and Keyword Identification of Unlabeled Documents · ICDM 2005 Using the fractal dimension to cluster datasets · KDD 2000 Detecting outliers using transduction and statistical testing · KDD 2006 |
Natural language and speech › Information extraction and text analysis
topic model |
0.1 | 1 | 2008 | On-line LDA: Adaptive Topic Models for Mining Text Streams with Applications to Topic Detection and Tracking · ICDM 2008 |
Data mining › anomaly detection
outlier detection |
0.1 | 1 | 2006 | Detecting outliers using transduction and statistical testing · KDD 2006 |
Data mining › clustering
document clustering |
0.1 | 1 | 2005 | Categorization and Keyword Identification of Unlabeled Documents · ICDM 2005 |
Data mining › dimensionality reduction
feature selection |
0.1 | 1 | 2005 | Categorization and Keyword Identification of Unlabeled Documents · ICDM 2005 |
Information retrieval › text analysis
keyword extraction |
0.1 | 1 | 2005 | Categorization and Keyword Identification of Unlabeled Documents · ICDM 2005 |
Data mining
text mining |
0.1 | 1 | 2005 | Categorization and Keyword Identification of Unlabeled Documents · ICDM 2005 |
Data mining › dimensionality reduction › feature selection
unsupervised feature selection |
0.1 | 1 | 2005 | Categorization and Keyword Identification of Unlabeled Documents · ICDM 2005 |
Data mining › pattern mining
association rule mining |
0.1 | 2 | 2003 | Screening and interpreting multi-item associations based on log-linear modeling · KDD 2003 Mining Relevant Text from Unlabelled Documents · ICDM 2003 |
Data mining › density estimation
log-linear analysis |
0.0 | 1 | 2003 | Screening and interpreting multi-item associations based on log-linear modeling · KDD 2003 |
Data mining
pattern mining |
0.0 | 1 | 2003 | Screening and interpreting multi-item associations based on log-linear modeling · KDD 2003 |
Data mining › text mining
text classification |
0.0 | 1 | 2003 | Mining Relevant Text from Unlabelled Documents · ICDM 2003 |
Web and mobile security › online social network security
social network attack |
0.0 | 1 | 2010 | Spectrum based fraud detection in social networks · CCS 2010 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.0 | 1 | 2001 | Preserving QoS of e-commerce sites through self-tuning: a performance model approach · EC 2001 |
Performance modeling and evaluation
queueing models |
0.0 | 1 | 2001 | Preserving QoS of e-commerce sites through self-tuning: a performance model approach · EC 2001 |
Distributed systems
fault tolerance |
0.0 | 6 | 1994 | Aggressive Transmissions of Short Messages Over Redundant Paths · IEEE Trans. Parallel Distributed Syst. 1994 Increasing Availability Under Mutual Exclusion Constraints with Dynamic Vote Reassignment · ACM Trans. Comput. Syst. 1989 The Reliability of Voting Mechanisms · IEEE Trans. Computers 1987 |
Data mining
exploratory data analysis |
0.0 | 1 | 1999 | Using Approximations to Scale Exploratory Data Analysis in Datacubes · KDD 1999 |
Distributed and cloud data management
mobile data management |
0.0 | 1 | 1999 | Mobile Computing and Databases - A Survey · IEEE Trans. Knowl. Data Eng. 1999 |
Information retrieval › indexing › text indexing
full-text index |
0.0 | 1 | 1996 | The Gold Text Indexing Engine · ICDE 1996 |
Information retrieval › document retrieval › bibliographic retrieval
online catalog |
0.0 | 1 | 1996 | Electronic Catalogs - Panel · ICDE 1996 |
Information retrieval › indexing
text indexing |
0.0 | 1 | 1996 | The Gold Text Indexing Engine · ICDE 1996 |
Spatial and temporal data management › spatial query processing
proximity search |
0.0 | 1 | 1995 | Efficient Processing of Proximity Queries for Large Databases · ICDE 1995 |
Query processing and optimization
similarity query processing |
0.0 | 1 | 1995 | Efficient Processing of Proximity Queries for Large Databases · ICDE 1995 |
Memory systems
cache management |
0.0 | 1 | 1995 | Sleepers and Workaholics: Caching Strategies in Mobile Environments · VLDB J. 1995 |
Database system architecture and tuning
cache invalidation |
0.0 | 1 | 1994 | Sleepers and Workaholics: Caching Strategies in Mobile Environments · SIGMOD Conference 1994 |
Indexing and storage engines
caching |
0.0 | 1 | 1994 | Sleepers and Workaholics: Caching Strategies in Mobile Environments · SIGMOD Conference 1994 |
Methods — techniques the papers use, named apart from their topics
spectral coordinate analysis · 0.2spectral analysis · 0.1latent dirichlet allocation · 0.1empirical bayes · 0.1transductive confidence machines · 0.1hypothesis testing · 0.1locally adaptive clustering · 0.1frequent itemset mining · 0.1simulation · 0.0loglinear model fitting · 0.0graph-theoretical decomposition · 0.0bucketing · 0.0association rule mining · 0.0hill-climbing · 0.0analytic queuing models · 0.0survey · 0.0analytical modeling · 0.0reachability analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | From Code to EM Signals: A Generative Approach to Side Channel Analysis-based Anomaly DetectionabstractToday, it is possible to perform external anomaly detection by analyzing the involuntary EM emanations of digital device components. However, one of the most important challenges of these methods is the manual collection of EM signals for fingerprinting. Indeed, this procedure must be conducted by a human expert and requires high precision. In this work, we introduce a framework that alleviates this requirement by relying on synthetic EM signals that have been generated from assembly code. The signals are produced with the use of a Generative Adversarial Network (GAN) model. Experimentally, we identify that the synthetic EM signals are extremely similar to the real and thus, can be used for training anomaly detection models effectively. Through experimental assessments, we prove that the anomaly detection models are capable of recognizing even minute alterations to the code with high accuracy. Kurt A. Vedros, Constantinos Kolias, Daniel Barbará, Robert C. Ivans |
ARES | 3 |
| 2019 | Identifying Near-Native Protein Structures via Anomaly DetectionabstractDiscriminating biologically-active/native tertiary protein structures from non-native ones is an outstanding challenge in computational structural biology. Computationally, the task involves teasing out near-native structures out of several thousands generated in silico. In this paper we build on the concept of anomaly detection in machine learning and propose several methods for discriminating near-native structures. Evaluations on benchmark datasets demonstrate that the proposed methods advance the state of the art and warrant further research on adapting concepts and techniques from machine learning to improve recognition of near-native structures in template-free protein structure prediction. Sivani Tadepalli, Nasrin Akhter 0001, Daniel Barbará, Amarda Shehu |
BIBM | 3 |
| 2018 | On early detection of application-level resource exhaustion and starvation
Mohamed Elsabagh, Daniel Barbará, Daniel Fleck, Angelos Stavrou |
J. Syst. Softw. | 2 |
| 2017 | Detecting ROP with Statistical Learning of Program CharacteristicsabstractReturn-Oriented Programming (ROP) has emerged as one of the most widely used techniques to exploit software vulnerabilities. Unfortunately, existing ROP protections suffer from a number of shortcomings: they require access to source code and compiler support, focus on specific types of gadgets, depend on accurate disassembly and construction of Control Flow Graphs, or use hardware-dependent (microarchitectural) characteristics. In this paper, we propose EigenROP, a novel system to detect ROP payloads based on unsupervised statistical learning of program characteristics. We study, for the first time, the feasibility and effectiveness of using microarchitecture-independent program characteristics -- namely, memory locality, register traffic, and memory reuse distance -- for detecting ROP. We propose a novel directional statistics based algorithm to identify deviations from the expected program characteristics during execution. EigenROP works transparently to the protected program, without requiring debug information, source code or disassembly. We implemented a dynamic instrumentation prototype of EigenROP using Intel Pin and measured it against in-the-wild ROP exploits and on payloads generated by the ROP compiler ROPC. Overall, EigenROP achieved significantly higher accuracy than prior anomaly-based solutions. It detected the execution of the ROP gadget chains with 81% accuracy, 80% true positive rate, only 0.8% false positive rate, and incurred comparable overhead to similar Pin-based solutions. Mohamed Elsabagh, Daniel Barbará, Daniel Fleck, Angelos Stavrou |
CODASPY | 2 |
| 2015 | Using myoelectric signals to recognize grips and movements of the handabstractPeople want to live independently, but too often disabilities or advanced age robs them of the ability to do the necessary activities of daily living (ADLs). Finding relationships between electromyograms measured in the arm and movements of the hand and wrist needed to perform ADLs can help address performance deficits and be exploited in designing myoelectrical control systems for prosthetics and computer interfaces. This paper reports on several machine learning techniques employed to discover the electromyogram patterns present when using the hand to perform 14 typical fine motor functional activities used to accomplish ADLs. Classification and clustering techniques are employed. Improvements to accuracies are introduced, including the use of exponential smoothing and using a symbolic representation to approximate signal streams. Results show the patterns can be learned to an accuracy of approximately 77% for a 15 class problem and the symbolic representation shows the potential for future improvement in accuracies. Gene Shuman, Zoran Duric, Daniel Barbará, Jessica Lin 0001, Naomi Lynn Gerber |
BIBM | 3 |
| 2015 | Radmin: Early Detection of Application-Level Resource Exhaustion and Starvation Attacks
Mohamed Elsabagh, Daniel Barbará, Daniel Fleck, Angelos Stavrou |
RAID | 2 |
| 2015 | Continuous Authentication on Mobile Devices Using Power Consumption, Touch Gestures and Physical Movement of Users
Rahul Murmuria, Angelos Stavrou, Daniel Barbará, Daniel Fleck |
RAID | 3 |
| 2014 | transAD: An Anomaly Detection Network Intrusion Sensor for the Web
Sharath Hiremagalore, Daniel Barbará, Daniel Fleck, Walter Powell, Angelos Stavrou |
ISC | 2 |
| 2014 | Exploring representations of protein structure for automated remote homology detection and mapping of protein structure spaceabstractBACKGROUND: Due to rapid sequencing of genomes, there are now millions of deposited protein sequences with no known function. Fast sequence-based comparisons allow detecting close homologs for a protein of interest to transfer functional information from the homologs to the given protein. Sequence-based comparison cannot detect remote homologs, in which evolution has adjusted the sequence while largely preserving structure. Structure-based comparisons can detect remote homologs but most methods for doing so are too expensive to apply at a large scale over structural databases of proteins. Recently, fragment-based structural representations have been proposed that allow fast detection of remote homologs with reasonable accuracy. These representations have also been used to obtain linearly-reducible maps of protein structure space. It has been shown, as additionally supported from analysis in this paper that such maps preserve functional co-localization of the protein structure space. METHODS: Inspired by a recent application of the Latent Dirichlet Allocation (LDA) model for conducting structural comparisons of proteins, we propose higher-order LDA-obtained topic-based representations of protein structures to provide an alternative route for remote homology detection and organization of the protein structure space in few dimensions. Various techniques based on natural language processing are proposed and employed to aid the analysis of topics in the protein structure domain. RESULTS: We show that a topic-based representation is just as effective as a fragment-based one at automated detection of remote homologs and organization of protein structure space. We conduct a detailed analysis of the information content in the topic-based representation, showing that topics have semantic meaning. The fragment-based and topic-based representations are also shown to allow prediction of superfamily membership. CONCLUSIONS: This work opens exciting venues in designing novel representations to extract information about protein structures, as well as organizing and mining protein structure space with mature text mining tools. Kevin Molloy, M. Jennifer Van, Daniel Barbará, Amarda Shehu |
BMC Bioinform. | 3 |
| 2012 | LSH-Div: Species diversity estimation using locality sensitive hashingabstractMetagenome sequencing projects attempt to determine the collective DNA of organisms, co-existing as communities across different environments. Computational approaches analyze the large volumes of sequence data obtained from these ecological samples, to provide an understanding of the species diversity, content and abundance. In this work we present a scalable, species diversity estimation algorithm that achieves computational efficiency by use of a locality sensitive hashing algorithm (LSH). Using fixed-length, gapless subsequences, we improve the sensitivity of pairwise sequence comparisons. Using the LSH-based function, we first group similar sequences into bins commonly referred to as operational taxonomic units (OTUs) and then compute several species diversity/richness metrics. The performance of our algorithm is evaluated on synthetic data and eight targeted metagenome samples obtained from the seawater. We compare our results to three state-of-the-art diversity estimation algorithms. We demonstrate the strength of our approach in terms of computational runtime and effective OTU assignments. The source code for LSH-Div is available at the supplementary website under the GNU GPL license. Supplementary material is available at http://www.cs.gmu.edu/~mlbio/LSH-DIV. Zeehasham Rasheed, Huzefa Rangwala, Daniel Barbará |
BIBM | 3 |
| 2012 | Efficient Clustering of Metagenomic Sequences using Locality Sensitive HashingabstractThe new generation of genomic technologies have allowed researchers to determine the collective DNA of organisms (e.g., microbes) co-existing as communities across the ecosystem (e.g., within the human host). There is a need for the computational approaches to analyze and annotate the large volumes of available sequence data from such microbial communities (metagenomes). In this paper, we developed an efficient and accurate metagenome clustering approach that uses the locality sensitive hashing (LSH) technique to approximate the computational complexity associated with comparing sequences. We introduce the use of fixed-length, gapless subsequences for improving the sensitivity of the LSH-based similarity function. We evaluate the performance of our algorithm on two metagenome datasets associated with microbes existing across different human skin locations. Our empirical results show the strength of the developed approach in comparison to three state-of-the-art sequence clustering algorithms with regards to computational efficiency and clustering quality. We also demonstrate practical significance for the developed clustering algorithm, to compare bacterial diversity and structure across different skin locations. Zeehasham Rasheed, Huzefa Rangwala, Daniel Barbará |
SDM | 3 |
| 2011 | Spectrum based fraud detection in social networksabstractSocial networks are vulnerable to various attacks such as spam emails, viral marketing and the such. In this paper we develop a spectrum based detection framework to discover the perpetrators of these attacks. In particular, we focus on Random Link Attacks (RLAs) in which the malicious user creates multiple false identities and interactions among those identities to later proceed to attack the regular members of the network. We show that RLA attackers can be filtered by using their spectral coordinate characteristics, which are hard to hide even after the efforts by the attackers of resembling as much as possible the rest of the network. Experimental results show that our technique is very effective in detecting those attackers and outperforms techniques previously published. Xiaowei Ying, Xintao Wu, Daniel Barbará |
ICDE | 3 |
| 2010 | Spectrum based fraud detection in social networksabstractSocial networks are vulnerable to various attacks such as spam emails, viral marketing and the such. In this paper we develop a spectrum based detection framework to discover the perpetrators of these attacks. In particular, we focus on Random Link Attacks (RLAs) in which the malicious user creates multiple false identities and interactions among those identities to later proceed to attack the regular members of the network. We show that RLA attackers can be filtered by using their spectral coordinate characteristics, which are hard to hide even after the efforts by the attackers of resembling as much as possible the rest of the network. Experimental results show that our technique is very effective in detecting those attackers and outperforms techniques previously published. Xiaowei Ying, Xintao Wu, Daniel Barbará |
CCS | 3 |
| 2009 | Topic Significance Ranking of LDA Generative Models
Loulwa S. Al-Sumait, Daniel Barbará, James Gentle, Carlotta Domeniconi |
ECML/PKDD (1) | 2 |
| 2008 | On-line LDA: Adaptive Topic Models for Mining Text Streams with Applications to Topic Detection and TrackingabstractThis paper presents Online Topic Model (OLDA), a topic model that automatically captures the thematic patterns and identifies emerging topics of text streams and their changes over time. Our approach allows the topic modeling framework, specifically the Latent Dirichlet Allocation (LDA) model, to work in an online fashion such that it incrementally builds an up-to-date model (mixture of topics per document and mixture of words per topic) when a new document (or a set of documents) appears. A solution based on the Empirical Bayes method is proposed. The idea is to incrementally update the current model according to the information inferred from the new stream of data with no need to access previous data. The dynamics of the proposed approach also provide an efficient mean to track the topics over time and detect the emerging topics in real time. Our method is evaluated both qualitatively and quantitatively using benchmark datasets. In our experiments, the OLDA has discovered interesting patterns by just analyzing a fraction of data at a time. Our tests also prove the ability of OLDA to align the topics across the epochs with which the evolution of the topics over time is captured. The OLDA is also comparable to, and sometimes better than, the original LDA in predicting the likelihood of unseen documents. Loulwa S. Al-Sumait, Daniel Barbará, Carlotta Domeniconi |
ICDM | 2 |
| 2006 | Detecting outliers using transduction and statistical testingabstractOutlier detection can uncover malicious behavior in fields like intrusion detection and fraud analysis. Although there has been a significant amount of work in outlier detection, most of the algorithms proposed in the literature are based on a particular definition of outliers (e.g., density-based), and use ad-hoc thresholds to detect them. In this paper we present a novel technique to detect outliers with respect to an existing clustering model. However, the test can also be successfully utilized to recognize outliers when the clustering information is not available. Our method is based on Transductive Confidence Machines, which have been previously proposed as a mechanism to provide individual confidence measures on classification decisions. The test uses hypothesis testing to prove or disprove whether a point is fit to be in each of the clusters of the model. We experimentally demonstrate that the test is highly robust, and produces very few misdiagnosed points, even when no clustering information is available. Furthermore, our experiments demonstrate the robustness of our method under the circumstances of data contaminated by outliers. We finally show that our technique can be successfully applied to identify outliers in a noisy data set for which no information is available (e.g., ground truth, clustering structure, etc.). As such our proposed methodology is capable of bootstrapping from a noisy data set a clean one that can be used to identify future outliers. Daniel Barbará, Carlotta Domeniconi, James P. Rogers |
KDD | 1 |
| 2005 | Categorization and Keyword Identification of Unlabeled DocumentsabstractIn this paper, we first propose a global unsupervised feature selection approach for text, based on frequent itemset mining. As a result, each document is represented as a set of words that co-occur frequently in the given corpus of documents. We then introduce a locally adaptive clustering algorithm, designed to estimate (local) word relevance and, simultaneously, to group the documents. We present experimental results to demonstrate the feasibility of our approach. Furthermore, the analysis of the weights credited to terms provides evidence that the identified keywords can guide the process of label assignment to clusters. We take into consideration both spam email filtering and general classification datasets. Our analysis of the distribution of weights in the two cases provides insights on how the spam problem distinguishes from the general classification case. Ning Kang 0009, Carlotta Domeniconi, Daniel Barbará |
ICDM | 3 |
| 2004 | Self-Similar Mining of Time Association Rules
Daniel Barbará, Ping Chen 0001, Zohreh Nazeri |
PAKDD | 1 |
| 2004 | Classifying Documents Without LabelsabstractAutomatic classification of documents is an important area of research with many applications in the fields of document searching, forensics and others. Methods to perform classification of text rely on the existence of a sample of documents whose class labels are known. However, in many situations, obtaining this sample may not be an easy (or even possible) task. Consider for instance, a set of documents that is returned as a result of a query. If we want to separate the documents that are truly relevant to the query from those that are not, it is unlikely that we will have at hand labelled documents to train classification models to perform this task. In this paper we focus on the classification of an unlabelled set of documents into two classes: relevant and irrelevant, given a topic of interest. By dividing the set of documents into buckets (for instance, answers returned by different search engines), and using association rule mining to find common sets of words among the buckets, we can efficiently obtain a sample of documents that has a large percentage of relevant ones. (I.e., a high “purity”.) This sample can be used to train models to classify the entire set of documents. We prove, via experimentation, that our method is capable of filtering relevant documents even in adverse conditions where the percentage of irrelevant documents in the buckets is relatively high. Daniel Barbará, Carlotta Domeniconi, Ning Kang 0009 |
SDM | 1 |
| 2003 | Compressing High Dimensional Datasets by FractalsabstractSummary form only given. Fractal technique extensions to general datasets were proposed. The high-dimensional dataset can be viewed as a collection of cells that represent some measure on an integer n-D grid. The measure over different scales along the dimension's hierarchies may exhibit self-similarities. A two-phase searching strategy was applied to overcome the increased searching time caused by additional dimensions. The search scheme checks a small number of spatially close local domain chunks. The data structure used is defined by 2/sup n/-tree which is a natural extension of quadtree (for image) and cotree (for volume). Each node, corresponding to a range chunk or a domain chunk, contains the summary information used for local matching. The experimental results have shown that the performance of fractal compression is comparable with rivals such as nonlinear model. The experiments over synthetic datasets have shown that the scalability of fractal compression techniques displays self-similar characteristics. To overcome high time complexity caused by additional dimensions, approximate multi-dimensional nearest neighbors searching techniques were presented that run in expected logarithmic time. Xintao Wu, Daniel Barbará |
DCC | 2 |
| 2003 | Mining Relevant Text from Unlabelled DocumentsabstractAutomatic classification of documents is an important area of research with many applications in the fields of document searching, forensics and others. Methods to perform classification of text rely on the existence of a sample of documents whose class labels are known. However, in many situations, obtaining this sample may not be an easy (or even possible) task. We focus on the classification of unlabelled documents into two classes: relevant and irrelevant, given a topic of interest. By dividing the set of documents into buckets (for instance, answers returned by different search engines), and using association rule mining to find common sets of words among the buckets, we can efficiently obtain a sample of documents that has a large percentage of relevant ones. This sample can be used to train models to classify the entire set of documents. We prove, via experimentation, that our method is capable of filtering relevant documents even in adverse conditions where the percentage of irrelevant documents in the buckets is relatively high. Daniel Barbará, Carlotta Domeniconi, Ning Kang 0009 |
ICDM | 1 |
| 2003 | Screening and interpreting multi-item associations based on log-linear modelingabstractAssociation rules have received a lot of attention in the data mining community since their introduction. The classical approach to find rules whose items enjoy high support (appear in a lot of the transactions in the data set) is, however, filled with shortcomings. It has been shown that support can be misleading as an indicator of how interesting the rule is. Alternative measures, such as lift, have been proposed. More recently, a paper by DuMouchel et al. proposed the use of all-two-factor loglinear models to discover sets of items that cannot be explained by pairwise associations between the items involved. This approach, however, has its limitations, since it stops short of considering higher order interactions (other than pairwise) among the items. In this paper, we propose a method that examines the parameters of the fitted loglinear models to find all the significant association patterns among the items. Since fitting loglinear models for large data sets can be computationally prohibitive, we apply graph-theoretical results to divide the original set of items into components (sets of items) that are statistically independent from each other. We then apply loglinear modeling to each of the components and find the interesting associations among items in them. The technique is experimentally evaluated with a real data set (insurance data) and a series of synthetic data sets. The results show that the technique is effective in finding interesting associations among the items involved. Xintao Wu, Daniel Barbará |
KDD | 2 |
| 2003 | Using Self-Similarity to Cluster Large Data Sets
Daniel Barbará, Ping Chen 0001 |
Data Min. Knowl. Discov. | 1 |
| 2003 | A Checksum-based Corruption Detection TechniqueabstractWe consider the problem of malicious attacks that lead to corruption of files in a file system. A typical method to detect such corruption is to compute signatures of all the files and store these signatures in a secure place. A malicious modification of a file can be detected by verifying the sign ature. This method, however, leaves the system vulnerable to an attacker who has access to some of the files and the signatures (but not the signing transformation) and who replaces some of the files by their old versions and the corresponding signatures by the signatures of the old versions. In this paper, we present a technique called Check2 that also relies on signatures for detecting corruption of files. The novel feature of our approach is that we compute additional levels of signatures to guarantee that any change of a file and the corresponding signature will require an attacker to perform a very lengthy chain of precise changes to successfully complete the corruption in an undetected manner. If an attacker fails to complete all the required changes, Check2 can be used to pinpoint which files have been corrupted. Two alternative ways of implementing Check2 are offered, the first using a deterministic way of combining signatures and the second using a randomized scheme. Our results show that the overhead added to the system is minimal. Daniel Barbará, Rajni Goel, Sushil Jajodia |
J. Comput. Secur. | 1 |
| 2003 | An Approximate Median Polish Algorithm for Large Multidimensional Data Sets
Daniel Barbará, Xintao Wu |
Knowl. Inf. Syst. | 1 |
| 2002 | COOLCAT: an entropy-based algorithm for categorical clusteringabstractIn this paper we explore the connection between clustering categorical data and entropy: clusters of similar poi lower entropy than those of dissimilar ones. We use this connection to design an incremental heuristic algorithm, COOLCAT, which is capable of efficiently clustering large data sets of records with categorical attributes, and data streams. In contrast with other categorical clustering algorithms published in the past, COOLCAT's clustering results are very stable for different sample sizes and parameter settings. Also, the criteria for clustering is a very intuitive one, since it is deeply rooted on the well-known notion of entropy. Most importantly, COOLCAT is well equipped to deal with clustering of data streams(continuously arriving streams of data point) since it is an incremental algorithm capable of clustering new points without having to look at every point that has been clustered so far. We demonstrate the efficiency and scalability of COOLCAT by a series of experiments on real and synthetic data sets. Daniel Barbará, Julia Couto 0001 |
CIKM | 1 |
| 2002 | Modeling and Imputation of Large Incomplete Multidimensional Datasets
Xintao Wu, Daniel Barbará |
DaWaK | 2 |
| 2002 | Mining Malicious Corruption of Data with Hidden Markov Models
Daniel Barbará, Rajni Goel, Sushil Jajodia |
DBSec | 1 |
| 2002 | Characterizing E-business Workloads Using Fractal Methods
Daniel A. Menascé, Bruno D. Abrahao, Daniel Barbará, Virgílio A. F. Almeida, Flávia Ribeiro |
J. Web Eng. | 3 |
| 2001 | Detecting Novel Network Intrusions Using Bayes Estimatorsabstract1 Introduction From the first appearance of network attacks, the internet worm, to the most recent one in which the servers of several famous e-business companies were paralyzed for several hours, causing huge financial losses, network-based attacks have been increasing in frequency and severity. As a powerful weapon to protect networks, intrusion detection has been gaining a lot of attention. Daniel Barbará, Ningning Wu, Sushil Jajodia |
SDM | 1 |
| 2001 | Preserving QoS of e-commerce sites through self-tuning: a performance model approachabstractThe Quality of Service (QoS) of e-commerce sites plays a crucial role in attracting and retaining customers. The workload experienced by these sites tends to vary in a very dynamic way. The complexity of the sites combined with the large short-terms variations of the workload calls for automated methods for site configuration. This paper describes a method for dynamically monitoring and tuning e-commerce sites so that desired QoS levels are attained. Our approach uses hill climbing techniques combined with analytic queuing models to guide the search for the best combination of configuration parameters. We validate our approach in an experimental setting by comparing the QoS levels of a TPC-W e-commerce site with and without control. We showed that under increasing loads, the controlled system meets its QoS goals, while the uncontrolled site fails to do so. Daniel A. Menascé, Daniel Barbará, Ronald C. Dodge |
EC | 2 |
| 2001 | Finding Dense Clusters in Hyperspace: An Approach Based on Row Shuffling
Daniel Barbará, Xintao Wu |
WAIM | 1 |
| 2001 | Loglinear-Based Quasi Cubes
Daniel Barbará, Xintao Wu |
J. Intell. Inf. Syst. | 1 |
| 2000 | Supporting Online Queries in ROLAP
Daniel Barbará, Xintao Wu |
DaWaK | 1 |
| 2000 | Protecting File systems Against Corruption Using Checksums
Daniel Barbará, Rajni Goel, Sushil Jajodia |
DBSec | 1 |
| 2000 | Using Checksums to Detect Data Corruption
Daniel Barbará, Rajni Goel, Sushil Jajodia |
EDBT | 1 |
| 2000 | Using the fractal dimension to cluster datasetsabstractClustering is a widely used knowledge discovery technique. It helps uncovering structures in data that were not previously known. The clustering of large data sets has received a lot of attention in recent years, however, clustering is a still a challenging task since many published algorithms fail to do well in scaling with the size of the data set and the number of dimensions that describe the points, or in finding arbitrary shapes of clusters, or dealing effectively with the presence of noise. In this paper, we present a new clustering algorithm, based in the fractal properties of the data sets. The new algorithm, which we call Fractal Clustering (FC), places points incrementally in the cluster for which the change in the fractal dimension after adding the point is the least. This is a very natural way of clustering points, since points in the same cluster have a great degree of self-similarity among them (and much less self-similarity with respect to points in other clusters). FC requires one scan of the data, is suspendable at will, providing the best answer possible at that point, and is incremental. We show via experiments that FC effectively deals with large data sets, high-dimensionality and noise and is capable of recognizing clusters of arbitrary shape. Daniel Barbará, Ping Chen 0001 |
KDD | 1 |
| 1999 | Using Approximations to Scale Exploratory Data Analysis in DatacubesabstractExploratory Data Analysis is a widely used technique to determine which factors have the most influence on data values in a multi-way table, or which cells in the table can be considered anomalous with respect to the other cells. In particular, median polish is a simple, yet robust method to perform Exploratory Data Analysis. Median polish is resistant to holes in the table (cells that have no values), but it may require a lot of iterations through the data. This factor makes it difficult to apply median polish to large multidimensional tables, since the I/O requirements may be prohibitive. This paper describes a technique that uses median polish over an approximation of a datacube, easing the burden of I/O. The results obtained are tested for quality, using a variety of measures. The technique scales to large datacubes and proves to give a good approximation of the results that would have been obtained by median polish in the original data. 1 Introduction Exploratory Data... Daniel Barbará, Xintao Wu |
KDD | 1 |
| 1999 | The Characterization of Continuous QueriesabstractIn a world where the amount of electronic information available is constantly growing, techniques to select and filter information efficiently become increasingly important. Continuous queries are a tool that allows users to monitor one or more information sources, by giving the impression that the queries are being run continually over them. In this paper, we formalize the notion of continuous queries for a wide spectrum of environments. We consider both append-only data sources and systems that allow more general data manipulation. We examine the case where the database management software may be modified as well as where we must treat it as a black box. We study the classes of queries that can be supported in each case and present efficient implementation techniques for them. Daniel Barbará |
Int. J. Cooperative Inf. Syst. | 1 |
| 1999 | Supporting Electronic Ink Databases
Walid G. Aref, Daniel Barbará |
Inf. Syst. | 2 |
| 1999 | Mobile Computing and Databases - A SurveyabstractThe emergence of powerful portable computers, along with advances in wireless communication technologies, has made mobile computing a reality. Among the applications that are finding their way to the market of mobile computing-those that involve data management-hold a prominent position. In the past few years, there has been a tremendous surge of research in the area of data management in mobile computing. This research has produced interesting results in areas such as data dissemination over limited bandwidth channels, location-dependent querying of data, and advanced interfaces for mobile computers. This paper is an effort to survey these techniques and to classify this research in a few broad areas. Daniel Barbará |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1997 | The AudioWebabstractThis paper describes the idea of a wide area information system based entirely on audio information (Audioweb).We address two issues.First, we address the issue of adding hyperaudio links to audio documents non-intrusively (Le., without the need to modify the audio track).By doing so, audio documents can be linked together to form the AudioWeb.Secondly, we show how, based on the audio format discussed previously, a series of interfaces (browsers), that enable the user to navigate a web of audio documents, can be implemented.l'miissiw 10 Ill:llic digilrlflinrd copk ol'all or psrJ oJ*lhis m;~Jgal Jijr pclsollid or dnw~oni LIW is gmalccl \bilhool Ike providsd Jhrl 111~ copi* QW 1101 lllmk Or disJrihukd br pmlil or rwaimcrci;d nd\:l~lJ~g~.111~ c"I,yridil noliw.IIlk! lilk Ol'llle publiolion mid ils cleIe ;jpp:nr.:Ind flalicc is Mn 1~11 rop~riglil ishy pcnnisir~n 1~1'tlw ..\CII.111~.'lb Lumpy o[[lc:n+a.JO rCpllhlid1.It1 pOSl 011 WITUS or III redis~rihuJrr Ju liti. Daniel Barbará, Shamim A. Naqvi |
CIKM | 1 |
| 1997 | Certification Reports: Supporting Transactions in Wireless SystemsabstractThe emergence of small portable computers and the advances in wireless networking have made mobile computing today a reality. Information systems and databases are among the applications that make mobile computing attractive. While the topic of querying data in wireless and mobile systems has received a lot of attention, techniques to efficiently update data in these systems while providing transaction semantics are not fully developed. We present a novel protocol that uses the broadcast facility to help mobile units do some of the work of verifying if the transactions being run by them need to be aborted. Only when the mobile unit cannot detect any conflict is the server involved in completing the verification. Of course, if the transaction can commit, the server will install the valves in the central database and notify the mobile units (again, using the broadcast channel). The protocol uses a modified version of optimistic control. We study the performance of the protocol by means of a detailed simulation. Daniel Barbará |
ICDCS | 1 |
| 1996 | The Gold Text Indexing EngineabstractThe proliferation of electronic communication including computer mail, faxes, voice mail, and net news has led to a variety of disjoint applications and usage paradigms that forces users to deal with multiple different user interfaces and access related information arriving over the different communication media separately. To enable users to cope with the overload of information arriving over heterogeneous communication media, we have developed the Gold document handling system that allows users to access all of these forms of communication at once, or to intermix them. The Gold system provides users with an integrated way to send and recieve messages using different media, efficiently store the messages, retrieve the messages based on their contents, and to access a variety of other sources of useful information. At the center of the Gold document handling system is the Gold Text Indexing Engine (GTIE) that provides a full text index over the documents. The paper describes our implementation of GTIE and the concurrency control protocols to ensure consistency of the index in the presence of concurrent operations. Daniel Barbará, Sharad Mehrotra, Padmavathi Vallabhaneni |
ICDE | 1 |
| 1996 | Electronic Catalogs - Panel
Arthur M. Keller, Don Brown, Anna-Lena Neches, Sherif Danish, Daniel Barbará |
ICDE | 5 |
| 1996 | Guest Editors' Introduction
Daniel Barbará, Ravi Jain, Narayanan Krishnakumar |
Distributed Parallel Databases | 1 |
| 1995 | Efficient Processing of Proximity Queries for Large DatabasesabstractEmerging multimedia applications require database systems to provide support for new types of objects and to process queries that may have no parallel in traditional database applications. One such important class of queries are the proximity queries that aims to retrieve objects in the database that are related by a distance metric in a way that is specified by the query. The importance of proximity queries has earlier been realized in developing constructs for visual languages. In this paper, we present algorithms for answering a class of proximity queries-fixed-radius nearest-neighbor queries over point object. Processing proximity queries using existing query processing techniques results in high CPU and I/O costs. We develop new algorithms to answer proximity queries over objects that lie in the one-dimensional space (e.g., words in a document). The algorithms exploit query semantics to reduce the CPU and I/O costs, and hence improve performance. We also show how our algorithms can be generalized to handle d-dimensional objects.> Walid G. Aref, Daniel Barbará, Sharad Mehrotra |
ICDE | 2 |
| 1995 | The Handwritten Trie: Indexing Electronic InkabstractThe emergence of the pen as the main interface device for personal digital assistants and pen-computers has made handwritten text, and more generally ink, a first-class object. As for any other type of data, the need of retrieval is a prevailing one. Retrieval of handwritten text is more difficult than that of conventional data since it is necessary to identify a handwritten word given slightly different variations in its shape. The current way of addressing this is by using handwriting recognition, which is prone to errors and limits the expressiveness of ink. Alternatively, one can retrieve from the database handwritten words that are similar to a query handwritten word using techniques borrowed from pattern and speech recognition. In particular, Hidden Markov Models (HMM) can be used as representatives of the handwritten words in the database. However, using HMM techniques to match the input against every item in the database (sequential searching) is unacceptably slow and does not scale up for large ink databases. In this paper, an indexing technique based on HMMs is proposed. The new index is a variation of the trie data structure that uses HMMs and a new search algorithm to provide approximate matching. Each node in the tree contains handwritten letters, where each letter is represented by an HMM. Branching in the trie is based on the ranking of matches given by the HMMs. The new search algorithm is parametrized so that it provides means for controlling the matching quality of the search process via a time-based budget. The index dramatically improves the search time in a database of handwritten words. Due to the variety of platforms for which this work is aimed, ranging from personal digital assistants to desktop computers, we implemented both main-memory and disk-based systems. The implementations are reported in this paper, along with performance results that show the practicality of the technique under a variety of conditions. Walid G. Aref, Daniel Barbará, Padmavathi Vallabhaneni |
SIGMOD Conference | 2 |
| 1995 | Sleepers and Workaholics: Caching Strategies in Mobile Environments
Daniel Barbará, Tomasz Imielinski |
VLDB J. | 1 |
| 1994 | Sleepers and Workaholics: Caching Strategies in Mobile EnvironmentsabstractIn the mobile wireless computing environment of the future a large number of users equipped with low powered palm-top machines will query databases over the wireless communication channels. Palmtop based units will often be disconnected for prolonged periods of time due to the battery power saving measures; palmtops will also frequencly relocate between different cells and connect to different data servers at different times. Caching of frequently accessed data items will be an important technique that will reduce contention on the narrow bandwidth wireless channel. However, cache invalidation strategies will be severely affected by the disconnection and mobility of the clients. The server may no longer know which clients are currently residing under its cell and which of them are currently on. We propose a taxonomy of different cache invalidation strategies and study the impact of client's disconnection times on their performance. We determine that for the units which are often disconnected (sleepers) the best cache invalidation strategy is based on signatures previously used for efficient file comparison. On the other hand, for units which are connected most of the time (workaholics), the best cache invalidation strategy is based on the periodic broadcast of changed data items. Daniel Barbará, Tomasz Imielinski |
SIGMOD Conference | 1 |
| 1994 | Aggressive Transmissions of Short Messages Over Redundant PathsabstractFault-tolerant computer systems have redundant paths connecting their components. Given these paths, it is possible to use aggressive techniques to reduce the average value and variability of the response time for short, critical messages. One technique is to send a copy of a packet over an alternate path before it is known whether the first copy failed or was delayed. A second technique is to split a single stream of packets over multiple paths. The authors analyze both approaches and show that they can provide significant improvements over conventional, conservative mechanisms.> Ben Kao, Hector Garcia-Molina, Daniel Barbará |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 1994 | The Demarcation Protocol: A Technique for Maintaining Constraints in Distributed Database Systems
Daniel Barbará, Hector Garcia-Molina |
VLDB J. | 1 |
| 1993 | Information Brokers: Sharing Knowledge in a Heterogeneous Distributed System
Daniel Barbará, Chris Clifton |
DEXA | 1 |
| 1993 | The Gold MailerabstractThe Gold Mailer, a system that provides users with an integrated way to send and receive messages using different media, efficiently store and retrieve these messages, and access a variety of sources of other useful information, is described. The mailer solves the problems of information overload, organization of messages and multiple interfaces. By providing good storage and retrieval facilities, it can be used as a powerful information processing engine covering a range of useful office information. The Gold Mailer's query language, indexing engine, file organization, data structures, and support of mail message data and multimedia documents are discussed.> Daniel Barbará, Chris Clifton, Fred Douglis, Hector Garcia-Molina, Ben Kao, Sharad Mehrotra, Jens Tellefsen, Rosemary Walsh |
ICDE | 1 |
| 1992 | The Demarcation Protocol: A Technique for Maintaining Linear Arithmetic Constraints in Distributed Database Systems
Daniel Barbará, Hector Garcia-Molina |
EDBT | 1 |
| 1992 | Probabilistic Dignosis of Hot SpotsabstractThe authors present several techniques to identify, or diagnose, hot spots in a database. All of them are probabilistic in the sense that they will classify the items as hot or cold and exhibit a non-zero probability of false diagnoses. Each technique is analysed to identify the tradeoffs of time and space involved in maintaining a low probability of false diagnosis. Each of the techniques is presented. The analyses of the techniques is considered to determine how likely they are to diagnose without error. The techniques are compared. A numerical comparison based on the analyses is included.> Kenneth Salem, Daniel Barbará, Richard J. Lipton |
ICDE | 2 |
| 1992 | The Management of Probabilistic DataabstractIt is often desirable to represent in a database, entities whose properties cannot be deterministically classified. The authors develop a data model that includes probabilities associated with the values of the attributes. The notion of missing probabilities is introduced for partially specified probability distributions. This model offers a richer descriptive language allowing the database to more accurately reflect the uncertain real world. Probabilistic analogs to the basic relational operators are defined and their correctness is studied. A set of operators that have no counterpart in conventional relational systems is presented.> Daniel Barbará, Hector Garcia-Molina, Daryl Porter |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1991 | Aggressive transmissions over redundant pathsabstractFault-tolerant computer systems have redundant paths connecting their components. Given these paths, it is possible to use aggressive techniques to reduce the average value and variability of the response time for critical messages. One technique is to send a copy of a packet over an alternate path before it is known if the first copy failed or was delayed. A second technique is to split a single stream of packets over multiple paths. The authors analyze both approaches and show that these techniques can provide significant improvements over conventional, conservative mechanisms.> Hector Garcia-Molina, Ben Kao, Daniel Barbará |
ICDCS | 3 |
| 1991 | Data Sharing in a Large Heterogeneous EnvironmentabstractThe issues involved in sharing information among a large collection of independent databases is explored. Some of the distinguishing features that characterize such large-scale environments (such as size, autonomy, and heterogeneity) are discussed. A multistep information sharing process for those systems is outlined and an architecture supporting that exchange is presented. A detailed description of a working prototype based on this architecture and some measurements of its performance are provided.> Rafael Alonso, Daniel Barbará, Steve Chon |
ICDE | 2 |
| 1991 | A Class of Randomized Strategies for Low-Cost Comparison of File CopiesabstractA class of algorithms that use randomized signatures to compare remotely located file copies is presented. A simple technique that sends on the order of 4/sup f/log(n) bits, where f is the number of differing pages that are to be diagnosed and n is the number of pages in the file, is described. A method to improve the bound in the number of bits sent, making them grow with f as flog(f) and with n as log(n)log(log(n)), and a class of algorithms in which the number of signatures grows with f as fr/sup f/, where r can be made to approach 1, are also presented. A comparison of these techniques is discussed.> Daniel Barbará, Richard J. Lipton |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1990 | A Probalilistic Relational Data Model
Daniel Barbará, Hector Garcia-Molina, Daryl Porter |
EDBT | 1 |
| 1990 | Using Stashing to Increase Node Autonomy in Distributed File SystemsabstractThe authors present an enhancement to distributed file systems that allows the users of the system to keep local copies of important files, decreasing the dependency over file servers. Using the notions of stashing and quasi-copies, the system allows users to tune up the quality of the service they want to receive when the file server is not reachable. One of the key points of this work is the focus on the tradeoff between availability and degradation of service. The other main contribution is the design of a distributed file system which is ideally suited to very large distributed systems, in that it provides users with greater tolerance of network partitions and server failures. It is emphasized that the use of stashing does not preclude the use of other performance-enhancing or fault-tolerant techniques. The file system architecture has been implemented and FACE, a prototype of a file system service based on Sun's NFS, is described. Performance figures are reported. These figures show that the overhead of providing the service is negligible. Current plans also call for porting the FACE design to a number of other processors.> Rafael Alonso, Daniel Barbará, Luis L. Cova |
SRDS | 2 |
| 1990 | Data Caching Issues in an Information Retrieval SystemabstractCurrently, a variety of information retrieval systems are available to potential users.… While in many cases these systems are accessed from personal computers, typically no advantage is taken of the computing resources of those machines (such as local processing and storage). In this paper we explore the possibility of using the user's local storage capabilities to cache data at the user's site. This would improve the response time of user queries albeit at the cost of incurring the overhead required in maintaining multiple copies. In order to reduce this overhead it may be appropriate to allow copies to diverge in a controlled fashion.… Thus, we introduce the notion of quasi-copies , which embodies the ideas sketched above. We also define the types of deviations that seem useful, and discuss the available implementation strategies.—From the Authors' Abstract Rafael Alonso, Daniel Barbará, Hector Garcia-Molina |
ACM Trans. Database Syst. | 2 |
| 1989 | A Randomized Technique for Remote File ComparisonabstractA technique for file comparison is presented that is based in a set of signatures that are selected by a randomized algorithm. The sites performing the comparison agree on this randomized set of signatures before any comparison takes place. This technique proves to be very competitive with previously published algorithms. It has an advantage over previous techniques in that one can set up the algorithm to diagnose up to a given number of different pages. This is done by changing the total number of bits sent to guarantee that the expected number of falsely diagnosed pages remains under a given level. A metric for comparing the complexity of file comparison techniques is introduced, based on the number of bits that the algorithm needs to send in order to diagnose a given number of differing pages while keeping the probability of false diagnosis under a certain level of confidence.> Daniel Barbará, Richard J. Lipton |
ICDCS | 1 |
| 1989 | Negotiating Data Access in Federated Database SystemsabstractA protocol is presented for negotiating access to data in a federated database. This protocol deals with several important aspects of data sharing in an environment where databases belonging to different organizations coexist and cooperate. It is based on the concept of quasicopies. These are cached values that are allowed to deviate from the central value in a controlled fashion. The degree of consistency of a quasicopy is established by the entity where the quasicopy is to reside. Two techniques are presented that can be used for numerical data and for copies that rely on version numbers. It is believed that these two ideas encompass many of the interesting cases in practice. However, it is also possible to fine tune the estimation of this function once the exact semantics of the data involved are known.> Rafael Alonso, Daniel Barbará |
ICDE | 2 |
| 1989 | Increasing Availability Under Mutual Exclusion Constraints with Dynamic Vote ReassignmentabstractVoting is used commonly to enforce mutual exclusion in distributed systems. Each node is assigned a number of votes, and only the group with a majority of votes is allowed to perform a restricted operation. This paper describes techniques for dynamically reassigning votes upon node or link failure, in an attempt to make the system more resilient to future failures. We focus on autonomous methods for achieving this, that is, methods that allow the nodes to make independent choices about changing their votes and picking new vote values, rather than group consensus techniques that require tight coordination among the remaining nodes. Protocols are given which allow nodes to install new vote values while still maintaining mutual exclusion requirements. The lemmas and theorems to validate the protocols are presented. A simple example shows how to apply the method to a database object-locking scheme; the protocols, however, are versatile andgeneral purpose, and can be used foranyapplication requiring mutual exclusion. In addition, policies are presented that allow nodes to autonomously select their new vote values. Simulation results are presented comparing the autonomous methods to static vote assignments and to group consensus strategies. These results demonstrate that under high failure rates, dynamic vote reassignment shows great improvement over a static assignment of votes in terms of availability. In addition, many autonomous methods for determining a new vote assignment yield almost as much availability as a group consensus method and at the same time are faster and more flexible. Daniel Barbará, Hector Garcia-Molina, Annemarie Spauster |
ACM Trans. Comput. Syst. | 1 |
| 1988 | Quasi-Copies: Efficient Data Sharing for Information Retrieval Systems
Rafael Alonso, Daniel Barbará, Hector Garcia-Molina, Soraya Abad |
EDBT | 2 |
| 1988 | Exploiting Symmetries for Low-Cost Comparison of File CopiesabstractA novel technique for comparison of remotely located file copies is examined. With this technique up to two differing pages can be located, and any other number of multiple differing pages can be detected. It uses a communication overhead of O(log/sup 2/(N)), where N is the number of pages in the file. It is based on a set of symmetries of a hypercube with dimension log(N).> Daniel Barbará, Hector Garcia-Molina, Bernardo Feijoo |
ICDCS | 1 |
| 1988 | A Software Environment for the Specification and Analysis of Problems of Coordination and ConcurrencyabstractThe SPANNER software environment for the specification and analysis of concurrent process coordination and resource sharing coordination is described. In the SPANNER environment, one can formally produce a specification of a distributed computing problem, and then verify its validity through reachability analysis and simulation. SPANNER is based on a finite-state machine model called the selection/resolution model. The capabilities of SPANNER are illustrated by the analysis of two classical coordination problems: (1) the dining philosophers; and (2) Dijkstra's concurrent programming problem. In addition, some of the more recently implemented capabilities of the SPANNER system are discussed, such as process types and cluster variables.> Sudhir Aggarwal, Daniel Barbará, Kalman Z. Meth |
IEEE Trans. Software Eng. | 2 |
| 1987 | Maintaining Availability of Replicated Data in a Dynamic Failure Environment
Daniel Barbará, Hector Garcia-Molina, Boris Kogan |
SRDS | 1 |
| 1987 | The Reliability of Voting MechanismsabstractIn a faulty distributed system, voting is commonly used to achieve mutual exclusion among groups of isolated nodes. Each node is assigned a number of votes, and any group with a majority of votes can perform the critical operations. Vote assignments can have a significant impact on system reliability. In this paper we address the problem of selecting vote assignments in order to maximize the probability that the critical operations can be performed at a given time by some group of nodes. We suggest simple heuristics to assign votes, and show that they give good results in most cases. We also study three particular homogeneous topologies (fully connected, Ethernet, and ring networks), and derive analytical expressions for system reliability. These expressions provide useful insights into the reliability provided by voting mechanisms. Daniel Barbará, Hector Garcia-Molina |
IEEE Trans. Computers | 1 |
| 1987 | SPANNER: A Tool for the Specification, Analysis, and Evaluation of ProtocolsabstractSPANNER is a software package for the specification, analysis, and evaluation of protocols. It is based on a mathematical model of coordinating processes called the selection/resolution model. Sudhir Aggarwal, Daniel Barbará, Kalman Z. Meth |
IEEE Trans. Software Eng. | 2 |
| 1986 | Specifying and Analyzing Protocols with SPANNER
Sudhir Aggarwal, Daniel Barbará, Kalman Z. Meth |
ICC | 2 |
| 1986 | Policies for Dynamic Vote Reassignment
Daniel Barbará, Hector Garcia-Molina, Annemarie Spauster |
ICDCS | 1 |
| 1986 | Protocols for Dynamic Vote ReassignmentabstractArticle Protocols for dynamic vote reassignment Share on Authors: Daniel Barbara Computer Science Department, Princeton University, Princeton, New Jersey Computer Science Department, Princeton University, Princeton, New JerseyView Profile , Hector Garcia-Molina Computer Science Department, Princeton University, Princeton, New Jersey Computer Science Department, Princeton University, Princeton, New JerseyView Profile , Annemarie Spauster Computer Science Department, Princeton University, Princeton, New Jersey Computer Science Department, Princeton University, Princeton, New JerseyView Profile Authors Info & Claims PODC '86: Proceedings of the fifth annual ACM symposium on Principles of distributed computingNovember 1986 Pages 195–205https://doi.org/10.1145/10590.10607Online:01 November 1986Publication History 25citation329DownloadsMetricsTotal Citations25Total Downloads329Last 12 Months8Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Daniel Barbará, Hector Garcia-Molina, Annemarie Spauster |
PODC | 1 |
| 1986 | Mutual Exclusion in Partitioned Distributed Systems
Daniel Barbará, Hector Garcia-Molina |
Distributed Comput. | 1 |
| 1986 | The Vulnerabilty of Vote AssignmentsabstractIn a faulty distributed system, voting is commonly used to achieve mutual exclusion among groups of nodes. Each node is assigned a number of votes, and any group with a majority of votes can perform the critical operations. Vote assignments can have a significant impact on system reliability, and in this paper we study the vote assignment problem. To compare vote assignments we define two deterministic measures, node and edge vulnerability. We present various properties of these measures and discuss how they can be computed. For these measures we discuss the selection of the best assignment and propose heuristics to identify good candidate assignments. Daniel Barbará, Hector Garcia-Molina |
ACM Trans. Comput. Syst. | 1 |
| 1985 | How to Assign Votes in a Distributed SystemabstractIn a distributed system, one strategy for achieving mutual exclusion of groups of nodes without communication is to assign to each node a number of votes. Only a group with a majority of votes can execute the critical operations, and mutual exclusion is achieved because at any given time there is at most one such group. A second strategy, which appears to be similar to votes, is to define a priori a set of groups that intersect each other. Any group of nodes that finds itself in this set can perform the restricted operations. In this paper, both of these strategies are studied in detail and it is shown that they are not equivalent in general (although they are in some cases). In doing so, a number of other interesting properties are proved. These properties will be of use to a system designer who is selecting a vote assignment or a set of groups for a specific application. Hector Garcia-Molina, Daniel Barbará |
J. ACM | 2 |
| 1984 | Optimizing the Reliability Provided by Voting Mechanisms
Hector Garcia-Molina, Daniel Barbará |
ICDCS | 2 |
| 1982 | How Expensive is Data Replication? An Example
Daniel Barbará, Hector Garcia-Molina |
ICDCS | 1 |
| 1981 | The cost of data replicationabstractWith the advent of data communication networks, researchers have been looking at the possibility of placing copies of a database at two or more nodes of a network. Such data replication is interesting because it makes the database accessible even when some of the nodes in the system fail. Furthermore, transactions which only read data may get faster access to the data when multiple copies exist. Hector Garcia-Molina, Daniel Barbará |
SIGCOMM | 2 |