EDBT 2026 Demo / reviewers in the wild / expert
Lise Getoor
dblp:g/LiseGetoor
· DBLP profile ↗
167ranked-venue papers
18as first author
15since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 102 · 11 first-author · 13 since 2021Databases, data management, data science and information retrieval · 89 · 13 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-authorHuman-computer interaction and ubiquitous computing · 12 · 1 since 2021Theory of computation · 5 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Neural-Symbolic Architectural Axioms of Integration: A ManifestoabstractThe integration of neural and symbolic methods has long been viewed as a promising path toward more general, interpretable, and robust artificial intelligence. The past two decades have seen a rapid proliferation of neural-symbolic (NeSy) systems, spanning a wide range of architectures, reasoning strategies, and application domains. However, this growth has outpaced theoretical clarity: many existing approaches conflate the roles of learning, inference, and representation, leading to a fragmented field lacking principled foundations. In this work, we address this gap by proposing a set of architectural axioms of integration—formal, implementation-agnostic principles that define how neural and symbolic components can be coherently combined. These axioms abstract away from system-specific details and instead characterize the structural interface between neural perception and symbolic reasoning. Rather than introducing a new method, this work offers a foundation to organize, compare, and reason about the rapidly expanding space of NeSy approaches. Connor Pryor, Lise Getoor |
NeSy | 2 |
| 2024 | Convex and Bilevel Optimization for Neural-Symbolic Inference and LearningabstractWe leverage convex and bilevel optimization techniques to develop a general gradient-based parameter learning framework for neural-symbolic (NeSy) systems. We demonstrate our framework with NeuPSL, a state-of-the-art NeSy architecture. To achieve this, we propose a smooth primal and dual formulation of NeuPSL inference and show learning gradients are functions of the optimal dual variables. Additionally, we develop a dual block coordinate descent algorithm for the new formulation that naturally exploits warm-starts. This leads to over $100 \times$ learning runtime improvements over the current best NeuPSL inference method. Finally, we provide extensive empirical evaluations across $8$ datasets covering a range of tasks and demonstrate our learning framework achieves up to a $16$% point prediction performance improvement over alternative learning methods. Charles Dickens, Changyu Gao, Connor Pryor, Stephen J. Wright 0001, Lise Getoor |
ICML | 5 |
| 2023 | Using Domain Knowledge to Guide Dialog Structure Induction via Neural Probabilistic Soft LogicabstractConnor Pryor, Quan Yuan, Jeremiah Liu, Mehran Kazemi, Deepak Ramachandran, Tania Bedrax-Weiss, Lise Getoor. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Connor Pryor, Quan Yuan 0001, Jeremiah Z. Liu, Mehran Kazemi, Deepak Ramachandran, Tania Bedrax-Weiss, Lise Getoor |
ACL (1) | 7 |
| 2023 | ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object NavigationabstractThe ability to accurately locate and navigate to a specific object is a crucial capability for embodied agents that operate in the real world and interact with objects to complete tasks. Such object navigation tasks usually require large-scale training in visual environments with labeled objects, which generalizes poorly to novel objects in unknown environments. In this work, we present a novel zero-shot object navigation method, Exploration with Soft Commonsense constraints (ESC), that transfers commonsense knowledge in pre-trained models to open-world object navigation without any navigation experience nor any other training on the visual environments. First, ESC leverages a pre-trained vision and language model for open-world prompt-based grounding and a pre-trained commonsense language model for room and object reasoning. Then ESC converts commonsense knowledge into navigation actions by modeling it as soft logic predicates for efficient exploration. Extensive experiments on MP3D, HM3D, and RoboTHOR benchmarks show that our ESC method improves significantly over baselines, and achieves new state-of-the-art results for zero-shot object navigation (e.g., 288% relative Success Rate improvement than CoW on MP3D). Kaiwen Zhou 0002, Kaizhi Zheng, Connor Pryor, Yilin Shen, Hongxia Jin, Lise Getoor, Xin Wang 0061 |
ICML | 6 |
| 2023 | NeuPSL: Neural Probabilistic Soft LogicabstractIn this paper, we introduce Neural Probabilistic Soft Logic (NeuPSL), a novel neuro-symbolic (NeSy) framework that unites state-of-the-art symbolic reasoning with the low-level perception of deep neural networks. To model the boundary between neural and symbolic representations, we propose a family of energy-based models, NeSy Energy-Based Models, and show that they are general enough to include NeuPSL and many other NeSy approaches. Using this framework, we show how to seamlessly integrate neural and symbolic parameter learning and inference in NeuPSL. Through an extensive empirical evaluation, we demonstrate the benefits of using NeSy methods, achieving upwards of 30% improvement over independent neural network models. On a well-established NeSy task, MNIST-Addition, NeuPSL demonstrates its joint reasoning capabilities by outperforming existing NeSy approaches by up to 10% in low-data settings. Furthermore, NeuPSL achieves a 5% boost in performance over state-of-the-art NeSy methods in a canonical citation network task with up to a 40 times speed up. Connor Pryor, Charles Dickens, Eriq Augustine, Alon Albalak, William Yang Wang, Lise Getoor |
IJCAI | 6 |
| 2023 | Collective Grounding: Applying Database Techniques to Grounding Templated ModelsabstractThe process of instantiating, or "grounding", a first-order model is a fundamental component of reasoning in logic. It has been widely studied in the context of theorem proving, database theory, and artificial intelligence. Within the relational learning community, the concept of grounding has been expanded to apply to models that use more general templates in the place of first-order logical formulae. In order to perform inference, grounding of these templates is required for instantiating a distribution over possible worlds. However, because of the complex data dependencies stemming from instantiating generalized templates with interconnected data, grounding is often the key computational bottleneck to relational learning. While we motivate our work in the context of relational learning, similar issues arise in probabilistic databases, particularly those that do not make strong tuple independence assumptions. In this paper, we investigate how key techniques from relational database theory can be utilized to improve the computational efficiency of the grounding process. We introduce the notion of collective grounding which treats logical programs not as a collection of independent rules, but instead as a joint set of interdependent workloads that can be shared. We introduce the theoretical concept of collective grounding, the components necessary in a collective grounding system, implementations of these components, and show how to use database theory to speed up these components. We demonstrate collective groundings effectiveness on seven popular datasets, and show up to a 70% reduction in runtime using collective grounding. Our results are fully reproducible and all code, data, and experimental scripts are included. Eriq Augustine, Lise Getoor |
Proc. VLDB Endow. | 2 |
| 2022 | Multi-relational Affinity PropagationabstractThere is a growing need for clustering algorithms which can operate in complex settings where there are multiple entity types with potential dependencies captured in different kinds of links. In this work, we present a novel approach for multi-relational clustering based on both the similarity of the entities' features, along with the multi-relational structure of the network among the entities. Our approach extends the affinity propagation clustering algorithm to multi-relational domains and encodes a variety of relational constraints to capture the dependencies across different node types in the underlying network. In contrast to the original formulation of affinity propagation that relies on enforcing hard constraints on the output clusters, we model the relational dependencies as soft constraints, allowing control over how they influence the final clustering of the nodes. This formulation allows us to balance between the homogeneity of the entities within the resulting clusters and their connections to clusters of nodes of the same and differing types. This in turn facilitates the exploration of the middle ground between feature-based similarity clustering, community detection, and block modeling in multi-relational networks. We present results on clustering a sample from Digg.com, a richly structured online social news website. We show that our proposed algorithm outperforms other clustering approaches on a variety of evaluation measures. We also analyze the impact of different parameter settings on the clustering output, in terms of both the homogeneity and the connectedness of the resulting clusters. Hossam Sharara, Lise Getoor |
ASONAM | 2 |
| 2022 | FETA: A Benchmark for Few-Sample Task Transfer in Open-Domain DialogueabstractAlon Albalak, Yi-Lin Tuan, Pegah Jandaghi, Connor Pryor, Luke Yoffe, Deepak Ramachandran, Lise Getoor, Jay Pujara, William Yang Wang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Alon Albalak, Yi-Lin Tuan, Pegah Jandaghi, Connor Pryor, Luke Yoffe, Deepak Ramachandran, Lise Getoor, Jay Pujara, William Yang Wang |
EMNLP | 7 |
| 2022 | The Power of (Statistical) Relational ThinkingabstractTaking into account relational structure during data mining can lead to better results, both in terms of quality and computational efficiency. This structure may be captured in the schema, in links between entities (e.g., graphs) or in rules describing the domain (e.g., knowledge graphs). Further, for richly structured prediction problems, there is often a need for a mix of both logical reasoning and statistical inference. In this talk, I will give an introduction to the field of Statistical Relational Learning (SRL), and I'll identify useful tips and tricks for exploiting structure in both the input and output space. I'll describe our recent work on highly scalable approaches for statistical relational inference. I'll close by introducing a broader interpretation of relational thinking that reveals new research opportunities (and challenges!). Lise Getoor |
KDD | 1 |
| 2022 | Learning explainable templated graphical modelsabstractTemplated graphical models (TGMs) encode model structure using rules that capture recurring relationships between multiple random variables. While the rules in TGMs are interpretable, it is not clear how they can be used to generate explanations for the individual predictions of the model. Further, learning these rules from data comes with high computational costs: it typically requires an expensive combinatorial search over the space of rules and repeated optimization over rule weights. In this work, we propose a new structure learning algorithm, Explainable Structured Model Search (ESMS), that learns a templated graphical model and an explanation framework for its predictions. ESMS uses a novel search procedure to efficiently search the space of models and discover models that trade-off predictive accuracy and explainability. We introduce the notion of relational stability and prove that our proposed explanation framework is stable. Further, our proposed piecewise pseudolikelihood (PPLL) objective does not require re-optimizing the rule weights across models during each iteration of the search. In our empirical evaluation on three realworld datasets, we show that our proposed approach not only discovers models that are explainable, but also significantly outperforms existing state-out-the-art structure learning approaches. Varun Embar, Sriram Srinivasan 0004, Lise Getoor |
UAI | 3 |
| 2022 | A taxonomy of weight learning methods for statistical relational learningabstractAbstract Statistical relational learning (SRL) frameworks are effective at defining probabilistic models over complex relational data. They often use weighted first-order logical rules where the weights of the rules govern probabilistic interactions and are usually learned from data. Existing weight learning approaches typically attempt to learn a set of weights that maximizes some function of data likelihood; however, this does not always translate to optimal performance on a desired domain metric, such as accuracy or F1 score. In this paper, we introduce a taxonomy of search-based weight learning approaches for SRL frameworks that directly optimize weights on a chosen domain performance metric. To effectively apply these search-based approaches, we introduce a novel projection, referred to as scaled space (SS), that is an accurate representation of the true weight space. We show that SS removes redundancies in the weight space and captures the semantic distance between the possible weight configurations. In order to improve the efficiency of search, we also introduce an approximation of SS which simplifies the process of sampling weight configurations. We demonstrate these approaches on two state-of-the-art SRL frameworks: Markov logic networks and probabilistic soft logic. We perform empirical evaluation on five real-world datasets and evaluate them each on two different metrics. We also compare them against four other weight learning approaches. Our experimental results show that our proposed search-based approaches outperform likelihood-based approaches and yield up to a 10% improvement across a variety of performance metrics. Further, we perform an extensive evaluation to measure the robustness of our approach to different initializations and hyperparameters. The results indicate that our approach is both accurate and robust. Sriram Srinivasan 0004, Charles Dickens, Eriq Augustine, Golnoosh Farnadi, Lise Getoor |
Mach. Learn. | 5 |
| 2021 | Context-Aware Online Collective Inference for Templated Graphical ModelsabstractIn this work, we examine online collective inference, the problem of maintaining and performing inference over a sequence of evolving graphical models. We utilize templated graphical models (TGM), a general class of graphical models expressed via templates and instantiated with data. A key challenge is minimizing the cost of instantiating the updated model. To address this, we define a class of exact and approximate context-aware methods for updating an existing TGM. These methods avoid a full re-instantiation by using the context of the updates to only add relevant components to the graphical model. Further, we provide stability bounds for the general online inference problem and regret bounds for a proposed approximation. Finally, we implement our approach in probabilistic soft logic, and test it on several online collective inference tasks. Through these experiments we verify the bounds on regret and stability, and show that our approximate online approach consistently runs two to five times faster than the offline alternative while, surprisingly, maintaining the quality of the predictions. Charles Dickens, Connor Pryor, Eriq Augustine, Lise Getoor |
ICML | 5 |
| 2021 | Local Explanation of Dialogue Response GenerationabstractIn comparison to the interpretation of classification models, the explanation of sequence generation models is also an important problem, however it has seen little attention. In this work, we study model-agnostic explanations of a representative text generation task -- dialogue response generation. Dialog response generation is challenging with its open-ended sentences and multiple acceptable responses. To gain insights into the reasoning process of a generation model, we propose a new method, local explanation of response generation (LERG) that regards the explanations as the mutual interaction of segments in input and output sentences. LERG views the sequence prediction as uncertainty estimation of a human response and then creates explanations by perturbing the input and calculating the certainty change over the human response. We show that LERG adheres to desired properties of explanations for text generation including unbiased approximation, consistency and cause identification. Empirically, our results show that our method consistently improves other widely used methods on proposed automatic- and human- evaluation metrics for this new task by $4.4$-$12.8$\%. Our analysis demonstrates that LERG can extract both explicit and implicit relations between input and output segments. Yi-Lin Tuan, Connor Pryor, Wenhu Chen, Lise Getoor, William Yang Wang |
NeurIPS | 4 |
| 2021 | A comparison of statistical relational learning and graph neural networks for aggregate graph queriesabstractAbstract Statistical relational learning (SRL) and graph neural networks (GNNs) are two powerful approaches for learning and inference over graphs. Typically, they are evaluated in terms of simple metrics such as accuracy over individual node labels. Complexaggregate graph queries(AGQ) involving multiple nodes, edges, and labels are common in the graph mining community and are used to estimate important network properties such as social cohesion and influence. While graph mining algorithms support AGQs, they typically do not take into account uncertainty, or when they do, make simplifying assumptions and do not build full probabilistic models. In this paper, we examine the performance of SRL and GNNs on AGQs over graphs with partially observed node labels. We show that, not surprisingly, inferring the unobserved node labels as a first step and then evaluating the queries on the fully observed graph can lead to sub-optimal estimates, and that a better approach is to compute these queries as an expectation under the joint distribution. We propose a sampling framework to tractably compute the expected values of AGQs. Motivated by the analysis of subgroup cohesion in social networks, we propose a suite of AGQs that estimate the community structure in graphs. In our empirical evaluation, we show that by estimating these queries as an expectation, SRL-based approaches yield up to a 50-fold reduction in average error when compared to existing GNN-based approaches. Varun Embar, Sriram Srinivasan 0004, Lise Getoor |
Mach. Learn. | 3 |
| 2021 | A Structured and Linguistic Approach to Understanding Recovery and Relapse in AAabstractAlcoholism, also known as Alcohol Use Disorder (AUD), is a serious problem affecting millions of people worldwide. Recovery from AUD is known to be challenging and often leads to relapse at various points after enrolling in a rehabilitation program such as Alcoholics Anonymous (AA). In this work, we present a structured and linguistic approach using hinge-loss Markov random fields (HL-MRFs) to understand recovery and relapse from AUD using social media data. We evaluate our models on AA-attending users extracted from: (i) the Twitter social network and predict recovery at two different points—90 days and 1 year after the user joins AA, respectively, and (ii) the Reddit AA recovery forums and predict whether the participating user is currently sober. The two datasets present two facets of the same underlying problem of understanding recovery and relapse in AUD users. We flesh out different characteristics in both these datasets: (i) In the Twitter dataset, we focus on the social aspect of the users and the relationship with recovery and relapse, and (ii) in the Reddit dataset, we focus on modeling the linguistic topics and dependency structure to understand users’ recovery journey. We design a unified modeling framework using HL-MRFs that takes the different characteristics of both these platforms into account. Our experiments reveal that our structured and linguistic approach is helpful in predicting recovery in users in both these datasets. We perform extensive quantitative analysis of different groups of features and dependencies among them in both datasets. The interpretable and intuitive nature of our models and analysis is helpful in making meaningful predictions and can potentially be helpful in identifying and preventing relapse early. Shawn Bailey, Yue Zhang 0047, Arti Ramesh, Jennifer Golbeck, Lise Getoor |
ACM Trans. Web | 5 |
| 2020 | Tandem Inference: An Out-of-Core Streaming Algorithm for Very Large-Scale Relational InferenceabstractStatistical relational learning (SRL) frameworks allow users to create large, complex graphical models using a compact, rule-based representation. However, these models can quickly become prohibitively large and not fit into machine memory. In this work we address this issue by introducing a novel technique called tandem inference (ti). The primary idea of ti is to combine grounding and inference such that both processes happen in tandem. ti uses an out-of-core streaming approach to overcome memory limitations. Even when memory is not an issue, we show that our proposed approach is able to do inference faster while using less memory than existing approaches. To show the effectiveness of ti, we use a popular SRL framework called Probabilistic Soft Logic (PSL). We implement ti for PSL by proposing a gradient-based inference engine and a streaming approach to grounding. We show that we are able to run an SRL model with over 1B cliques in under nine hours and using only 10 GB of RAM; previous approaches required more than 800 GB for this model and are infeasible on common hardware. To the best of our knowledge, this is the largest SRL model ever run. Sriram Srinivasan 0004, Eriq Augustine, Lise Getoor |
AAAI | 3 |
| 2020 | BOWL: Bayesian Optimization for Weight Learning in Probabilistic Soft LogicabstractProbabilistic soft logic (PSL) is a statistical relational learning framework that represents complex relational models with weighted first-order logical rules. The weights of the rules in PSL indicate their importance in the model and influence the effectiveness of the model on a given task. Existing weight learning approaches often attempt to learn a set of weights that maximizes some function of data likelihood. However, this does not always translate to optimal performance on a desired domain metric, such as accuracy or F1 score. In this paper, we introduce a new weight learning approach called Bayesian optimization for weight learning (BOWL) based on Gaussian process regression that directly optimizes weights on a chosen domain performance metric. The key to the success of our approach is a novel projection that captures the semantic distance between the possible weight configurations. Our experimental results show that our proposed approach outperforms likelihood-based approaches and yields up to a 10% improvement across a variety of performance metrics. Further, we performed experiments to measure the scalability and robustness of our approach on various realworld datasets. Sriram Srinivasan 0004, Golnoosh Farnadi, Lise Getoor |
AAAI | 3 |
| 2020 | Joint Estimation of User And Publisher Credibility for Fake News DetectionabstractFast propagation, ease-of-access, and low cost have made social media an increasingly popular means for news consumption. However, this has also led to an increase in the preponderance of fake news. Widespread propagation of fake news can be detrimental to society, and this has created enormous interest in fake news detection on social media. Many approaches to fake news detection use the news content, social context, or both. In this work, we look at fake news detection as a problem of estimating the credibility of both the news publishers and users that propagate news articles. We introduce a new approach called the credibility score-based model that can jointly infer fake news and credibility scores for publishers and users. We use a state-of-the-art statistical relational learning framework called probabilistic soft logic to perform this joint inference effectively. We show that our approach is accurate at both fake news detection and inferring credibility scores. Further, our model can easily integrate any auxiliary information that can aid in fake news detection. Using the FakeNewsNet dataset, we show that our approach significantly outperforms previous approaches at fake news detection by up to 10% in recall and 4% in accuracy. Furthermore, the credibility scores learned for both publishers and users are representative of their true behavior. Rajdipa Chowdhury, Sriram Srinivasan 0004, Lise Getoor |
CIKM | 3 |
| 2020 | VMI-PSL: Visual Model Inspector for Probabilistic Soft LogicabstractHybrid recommender systems achieve state-of-the-art performance by integrating several different information sources along with multiple recommendation approaches. Probabilistic Soft Logic (PSL) has been shown to be an accessible and effective means of creating extensible hybrid recommenders [11]. PSL allows users to easily create intuitive models that incorporate background information and capture complex interactions. However these complex interactions can sometimes make PSL models difficult to inspect, debug, and understand. In this paper, we present a generic visual model inspector for PSL, and show how our inspector can be used on a hybrid recommender system to: debug errors in the model, analyze the performance of individual components of the model, and explain recommendations made by the model. Aaron Rodden, Tarun Salh, Eriq Augustine, Lise Getoor |
RecSys | 4 |
| 2020 | Causal Relational LearningabstractCausal inference is at the heart of empirical research in natural and social sciences and is critical for scientific discovery and informed decision making. The gold standard in causal inference is performing randomized controlled trials ; unfortunately these are not always feasible due to ethical, legal, or cost constraints. As an alternative, methodologies for causal inference from observational data have been developed in statistical studies and social sciences. However, existing methods critically rely on restrictive assumptions such as the study population consisting of homogeneous elements that can be represented in a single flat table, where each row is referred to as a unit. In contrast, in many real-world settings, the study domain naturally consists of heterogeneous elements with complex relational structure, where the data is naturally represented in multiple related tables. In this paper, we present a formal framework for causal inference from such relational data. We propose a declarative language called CARL for capturing causal background knowledge and assumptions, and specifying causal queries using simple Datalog-like rules. CARL provides a foundation for inferring causality and reasoning about the effect of complex interventions in relational domains. We present an extensive experimental evaluation on real relational data to illustrate the applicability of CARL in social sciences and healthcare. Babak Salimi, Harsh Parikh, Moe Kayali, Lise Getoor, Sudeepa Roy 0001, Dan Suciu |
SIGMOD Conference | 4 |
| 2020 | Generating and Understanding Personalized Explanations in Hybrid Recommender SystemsabstractRecommender systems are ubiquitous and shape the way users access information and make decisions. As these systems become more complex, there is a growing need for transparency and interpretability. In this article, we study the problem of generating and visualizing personalized explanations for recommender systems that incorporate signals from many different data sources. We use a flexible, extendable probabilistic programming approach and show how we can generate real-time personalized recommendations. We then turn these personalized recommendations into explanations. We perform an extensive user study to evaluate the benefits of explanations for hybrid recommender systems. We conduct a crowd-sourced user study where our system generates personalized recommendations and explanations for real users of the last.fm music platform. First, we evaluate the performance of the recommendations in terms of perceived accuracy and novelty. Next, we experiment with (1) different explanation styles (e.g., user-based, item-based), (2) manipulating the number of explanation styles presented, and (3) manipulating the presentation format (e.g., textual vs. visual). We also apply a mixed-model statistical analysis to consider user personality traits as a control variable and demonstrate the usefulness of our approach in creating personalized hybrid explanations with different style, number, and format. Finally, we perform a post analysis that shows different preferences for explanation styles between experienced and novice last.fm users. Pigi Kouki, James Schaffer, Jay Pujara, John O'Donovan, Lise Getoor |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2019 | Lifted Hinge-Loss Markov Random Fields
Sriram Srinivasan 0004, Behrouz Babaki, Golnoosh Farnadi, Lise Getoor |
AAAI | 4 |
| 2019 | Responsible Data ScienceabstractData science is an emerging discipline that offers both promise and peril. Responsible data science refers to efforts that address both the technical and societal issues in emerging data- driven technologies. How can data-driven systems reason effectively about complex dependencies and uncertainty? Furthermore, how do we understand the ethical and societal issues involved in data-driven decision-making? There is a pressing need to integrate algorithmic and statistical principles, social science theories, and basic humanist concepts so that we can think critically and constructively about the socio-technical systems we are building. In this talk, I will overview this emerging area. Lise Getoor |
IEEE BigData | 1 |
| 2019 | Identifying Facet Mismatches In Search Via MicrographsabstractE-commerce search engines are the primary means by which customers shop for products online. Each customer query contains multiple facets such as product type, color, brand, etc. A successful search engine retrieves products that are relevant to the query along each of these attributes. However, due to lexical (erroneous title, description, etc.) and behavioral irregularities (clicks or purchases of products that do not belong to the same facet as the query), some mismatched products are often included in search results. These irregularities can be detected using simple binary classifiers like gradient boosted decision trees or logistic regression. Typically, these binary classifiers use strong independence assumptions between the results and ignore structural relationships available in the data, such as the connections between products and queries. In this paper, we use the connections that exist between products and query to identify a special kind of structure we refer to as a micrograph. Further, we make use of Statistical Relational Learning (SRL) to incorporate these micrographs in the data and pose the problem as a structured prediction problem. We refer to this approach as structured mismatch classification (\SMC). In addition, we show that naive addition of structure does not improve the performance of the model and hence introduce a variation of \SMC, strong \SMC~(\SSMC), which improves over the baseline by passing information from high-confidence predictions to lower confidence predictions. In our empirical evaluation we show that our proposed approach outperforms the baseline classification methods by up to 12% in precision. Furthermore, we use quasi-Newton methods to make our method viable for real-time inference in a search engine and show that our approach is up to 150 times faster than existing ADMM-based solvers. Sriram Srinivasan 0004, Nikhil Rao 0001, Karthik Subbian, Lise Getoor |
CIKM | 4 |
| 2019 | The Power of Relational Learning (Invited Talk)abstractWe live in a richly interconnected world and, not surprisingly, we generate richly interconnected data. From smart cities to social media to financial networks to biological networks, data is relational. While database theory is built on strong relational foundations, the same is not true for machine learning. The majority of machine learning methods flatten data into a single table before performing any processing. Further, database theory is also built on a bedrock of declarative representations. The same is not true for machine learning, in particular deep learning, which results in black-box, uninterpretable and unexplainable models. In this talk, I will introduce the field of statistical relational learning, an alternative machine learning approach based on declarative relational representations paired with probabilistic models. I’ll describe our work on probabilistic soft logic, a probabilistic programming language that is ideally suited to richly connected, noisy data. Our recent results show that by building on state-of-the-art optimization methods in a distributed implementation, we can solve very large relational learning problems orders of magnitude faster than existing approaches. Lise Getoor |
ICDT | 1 |
| 2019 | Estimating Causal Effects of Tone in Online DebatesabstractStatistical methods applied to social media posts shed light on the dynamics of online dialogue. For example, users' wording choices predict their persuasiveness and users adopt the language patterns of other dialogue participants. In this paper, we estimate the causal effect of reply tones in debates on linguistic and sentiment changes in subsequent responses. The challenge for this estimation is that a reply's tone and subsequent responses are confounded by the users' ideologies on the debate topic and their emotions. To overcome this challenge, we learn representations of ideology using generative models of text. We study debates from 4Forums.com and compare annotated tones of replying such as emotional versus factual, or reasonable versus attacking. We show that our latent confounder representation reduces bias in ATE estimation. Our results suggest that factual and asserting tones affect dialogue and provide a methodology for estimating causal effects from text. Dhanya Sridhar, Lise Getoor |
IJCAI | 2 |
| 2019 | Personalized explanations for hybrid recommender systemsabstractRecommender systems have become pervasive on the web, shaping the way users see information and thus the decisions they make. As these systems get more complex, there is a growing need for transparency. In this paper, we study the problem of generating and visualizing personalized explanations for hybrid recommender systems, which incorporate many different data sources. We build upon a hybrid probabilistic graphical model and develop an approach to generate real-time recommendations along with personalized explanations. To study the benefits of explanations for hybrid recommender systems, we conduct a crowd-sourced user study where our system generates personalized recommendations and explanations for real users of the last.fm music platform. We experiment with 1) different explanation styles (e.g., user-based, item-based), 2) manipulating the number of explanation styles presented, and 3) manipulating the presentation format (e.g., textual vs. visual). We apply a mixed model statistical analysis to consider user personality traits as a control variable and demonstrate the usefulness of our approach in creating personalized hybrid explanations with different style, number, and format. Pigi Kouki, James Schaffer, Jay Pujara, John O'Donovan, Lise Getoor |
IUI | 5 |
| 2019 | Responsible Data ScienceabstractData science is an emerging discipline that offers both promise and peril. Responsible data science refers to efforts that address both the technical and societal issues in emerging data-driven technologies. How can machine learning and database systems reason effectively about complex dependencies and uncertainty? Furthermore, how do we understand the ethical and societal issues involved in data-driven decision-making? There is a pressing need to integrate algorithmic and statistical principles, social science theories, and basic humanist concepts so that we can think critically and constructively about the socio-technical systems we are building. In this talk, I will overview this emerging area, with an emphasis on relational learning. Lise Getoor |
SIGMOD Conference | 1 |
| 2019 | The Responsibility Challenge for DataabstractAs data science and artificial intelligence become ubiquitous, they have an increasing impact on society. While many of these impacts are beneficial, others may not be. So understanding and managing these impacts is required of every responsible data scientist. Nevertheless, most human decision-makers use algorithms for efficiency purposes and not to make a better (i.e., fairer) decisions. Even the task of risk assessment in the criminal justice system enables efficiency instead of (and often at the expense of) fairness. So we need to frame the problem with fairness, and other societal impacts, as primary objectives. In this context, most attention has been paid to the machine learning of a model for a task, such as recognition, prediction, or classification. However, issues arise in all parts of the data eco-system, from data acquisition to data presentation. For example, the majority of the population is not white and male, yet this demographic is over-represented in the training data. It is challenging for a data scientist to satisfactorily discharge this broad responsibility. H. V. Jagadish, Francesco Bonchi, Tina Eliassi-Rad, Lise Getoor, Krishna P. Gummadi, Julia Stoyanovich |
SIGMOD Conference | 4 |
| 2019 | Collective entity resolution in multi-relational familial networks
Pigi Kouki, Jay Pujara, Christopher Steven Marcum, Laura M. Koehly, Lise Getoor |
Knowl. Inf. Syst. | 5 |
| 2019 | A Collective, Probabilistic Approach to Schema Mapping Using Diverse Noisy EvidenceabstractWe propose a probabilistic approach to the problem of schema mapping. Our approach is declarative, scalable, and extensible. It builds upon recent results in both schema mapping and probabilistic reasoning and contributes novel techniques in both fields. We introduce the problem of schema mapping selection, that is, choosing the best mapping from a space of potential mappings, given both metadata constraints and a data example. As selection has to reason holistically about the inputs and the dependencies between the chosen mappings, we define a new schema mapping optimization problem which captures interactions between mappings as well as inconsistencies and incompleteness in the input. We then introduce Collective Mapping Discovery (CMD), our solution to this problem using state-of-the-art probabilistic reasoning techniques. Our evaluation on a wide range of integration scenarios, including several real-world domains, demonstrates that CMD effectively combines data and metadata information to infer highly accurate mappings even with significant levels of noise. Angelika Kimmig, Alex Memory, Renée J. Miller, Lise Getoor |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2018 | Fairness in Relational DomainsabstractAI and machine learning tools are being used with increasing frequency for decision making in domains that affect peoples' lives such as employment, education, policing and loan approval. These uses raise concerns about biases of algorithmic discrimination and have motivated the development of fairness-aware machine learning. However, existing fairness approaches are based solely on attributes of individuals. In many cases, discrimination is much more complex, and taking into account the social, organizational, and other connections between individuals is important. We introduce new notions of fairness that are able to capture the relational structure in a domain. We use first-order logic to provide a flexible and expressive language for specifying complex relational patterns of discrimination. Furthermore, we extend an existing statistical relational learning framework, probabilistic soft logic (PSL), to incorporate our definition of relational fairness. We refer to this fairness-aware framework FairPSL. FairPSL makes use of the logical definitions of fairnesss but also supports a probabilistic interpretation. In particular, we show how to perform maximum a posteriori(MAP) inference by exploiting probabilistic dependencies within the domain while avoiding violation of fairness guarantees. Preliminary empirical evaluation shows that we are able to make both accurate and fair decisions. Golnoosh Farnadi, Behrouz Babaki, Lise Getoor |
AIES | 3 |
| 2018 | A Socio-linguistic Model for Cyberbullying DetectionabstractCyberbullying is a serious threat to both the short and long-term well-being of social media users. Addressing this problem in online environments demands the ability to automatically detect cyberbullying and to identify the roles that participants assume in social interactions. As cyberbullying occurs within online communities, it is also vital to understand the group dynamics that support bullying behavior. To this end, we propose a socio-linguistic model which jointly detects cyberbullying content in messages, discovers latent text categories, identifies participant roles and exploits social interactions. While our method makes use of content that is labeled as bullying, it does not require category, role or relationship labels. Furthermore, as bullying labels are often subjective, noisy and inconsistent, an important contribution of our paper is effective methods for leveraging inconsistent labels. Rather than discard inconsistent labels, we evaluate different methods for learning from them, demonstrating that incorporating uncertainty allows for better generalization. Our proposed socio-linguistic model achieves an 18% improvement over state-of-the-art methods. Sabina Tomkins, Lise Getoor, Yunfei Chen 0003, Yi Zhang 0001 |
ASONAM | 2 |
| 2018 | The Impact of Environmental Stressors on Human TraffickingabstractSevere environmental events have extreme effects on all segments of society, including criminal activity. Extreme weather events, such as tropical storms, fires, and floods create instability in communities, and can be exploited by criminal organizations. Here we investigate the potential impact of catastrophic storms on the criminal activity of human trafficking. We propose three theories of how these catastrophic storms might impact trafficking and provide evidence for each. Researching human trafficking is made difficult by its illicit nature and the obscurity of high-quality data. Here, we analyze online advertisements for services which can be collected at scale and provide insights into traffickers' behavior. To successfully combine relevant heterogenous sources of information, as well as spatial and temporal structure, we propose a collective, probabilistic approach. We implement this approach with Probabilistic Soft Logic, a probabilistic programming framework which can flexibly model relational structure and for which inference of future locations is highly efficient. Furthermore, this framework can be used to model hidden structure, such as latent links between locations. Our proposed approach can model and predict how traffickers move. In addition, we propose a model which learns connections between locations. This model is then adapted to have knowledge of environmental events, and we demonstrate that incorporating knowledge of environmental events can improve prediction of future locations. While we have validated our models on the impact of severe weather on human trafficking, we believe our models can be generalized to a variety of other settings in which environmental events impact human behavior. Sabina Tomkins, Golnoosh Farnadi, Brian Amanatullah, Lise Getoor, Steven Minton |
ICDM | 4 |
| 2018 | Scalable Probabilistic Causal Structure DiscoveryabstractComplex causal networks underlie many real-world problems, from the regulatory interactions between genes to the environmental patterns used to understand climate change. Computational methods seek to infer these causal networks using observational data and domain knowledge. In this paper, we identify three key requirements for inferring the structure of causal networks for scientific discovery: (1) robustness to noise in observed measurements; (2) scalability to handle hundreds of variables; and (3) flexibility to encode domain knowledge and other structural constraints. We first formalize the problem of joint probabilistic causal structure discovery. We develop an approach using probabilistic soft logic (PSL) that exploits multiple statistical tests, supports efficient optimization over hundreds of variables, and can easily incorporate structural constraints, including imperfect domain knowledge. We compare our method against multiple well-studied approaches on biological and synthetic datasets, showing improvements of up to 20% in F1-score over the best performing baseline in realistic settings. Dhanya Sridhar, Jay Pujara, Lise Getoor |
IJCAI | 3 |
| 2018 | Scalable structured prediction for richly structured socio-behavioral dataabstractOnline recommender systems, content-provider sites, and social media platforms provide richly structured socio-behavioral data. However, using this noisy and incomplete data to make decisions and recommendations is challenging. It often requires complex forms of structured prediction that rely on both the logical structure in the domain and probabilistic dependencies among interlinked entities. In this talk, I will describe some common inference patterns that are useful for socio-behavioral networks and introduce probabilistic soft logic (PSL). PSL is a highly scalable open-source probabilistic programming language being developed within my group that is well-suited for structured prediction over socio-behavioral data. Finally, I will review some of our recent work using PSL for hybrid recommender systems, explanation, and fair decision making. Lise Getoor |
RecSys | 1 |
| 2018 | Sustainability at scale: towards bridging the intention-behavior gap with sustainable recommendationsabstractFinding sustainable products and evaluating their claims is a significant barrier facing sustainability-minded customers. Tools that reduce both these burdens are likely to boost the sale of sustainable products. However, it is difficult to determine the sustainability characteristics of these products --- there are a variety of certifications and definitions of sustainability, and quality labeling requires input from domain experts. In this paper, we propose a flexible probabilistic framework that uses domain knowledge to identify sustainable products and customers, and uses these labels to predict customer purchases. We evaluate our approach on grocery items from the Amazon catalog. Our proposed approach outperforms established recommender system models in predicting future purchases while jointly inferring sustainability scores for customers and products. Sabina Tomkins, Steven Isley, Ben London 0001, Lise Getoor |
RecSys | 4 |
| 2018 | Topic Evolution Models for Long-Running MOOCs
Arti Ramesh, Lise Getoor |
WISE (2) | 2 |
| 2018 | A Structured Approach to Understanding Recovery and Relapse in AAabstractAlcoholism, also known as Alcohol Use Disorder (AUD), is a serious problem affecting millions of people worldwide. Recovery from AUD is known to be challenging and often leads to relapse at various points after enrolling in a rehabilitation program such as Alcoholics Anonymous (AA). In this work, we take a structured approach to understand recovery and relapse from AUD using social media data. To do so, we combine linguistic and psychological attributes of users with relational features that capture useful structure in the user interaction network. We evaluate our models on AA-attending users extracted from the Twitter social network and predict recovery at two different points---90 days and 1 year after the user joins AA, respectively. Our experiments reveal that our structured approach is helpful in predicting recovery in these users. We perform extensive quantitative analysis of different groups of features and dependencies among them. Our analysis sheds light on the role of each feature group and how they combine to predict recovery and relapse. Finally, we present a qualitative analysis of the different reasons behind users relapsing to AUD. Our models and analysis are helpful in making meaningful predictions in scenarios where only a subset of features are available and can potentially be helpful in identifying and preventing relapse early. Yue Zhang 0047, Arti Ramesh, Jennifer Golbeck, Dhanya Sridhar, Lise Getoor |
WWW | 5 |
| 2017 | Sparsity and Noise: Where Knowledge Graph Embeddings Fall ShortabstractKnowledge graph (KG) embedding techniques use structured relationships between entities to learn lowdimensional representations of entities and relations.One prominent goal of these approaches is to improve the quality of knowledge graphs by removing errors and adding missing facts.Surprisingly, most embedding techniques have been evaluated on benchmark datasets consisting of dense and reliable subsets of human-curated KGs, which tend to be fairly complete and have few errors.In this paper, we consider the problem of applying embedding techniques to KGs extracted from text, which are often incomplete and contain errors.We compare the sparsity and unreliability of different KGs and perform empirical experiments demonstrating how embedding approaches degrade as sparsity and unreliability increase. Jay Pujara, Eriq Augustine, Lise Getoor |
EMNLP | 3 |
| 2017 | A Collective, Probabilistic Approach to Schema MappingabstractWe propose a probabilistic approach to the problem of schema mapping. Our approach is declarative, scalable, and extensible. It builds upon recent results in both schema mapping and probabilistic reasoning and contributes novel techniques in both fields. We introduce the problem of mapping selection, that is, choosing the best mapping from a space of potential mappings, given both metadata constraints and a data example. As selection has to reason holistically about the inputs and the dependencies between the chosen mappings, we define a new schema mapping optimization problem which captures interactions between mappings. We then introduce Collective Mapping Discovery (CMD), our solution to this problem using stateof- the-art probabilistic reasoning techniques, which allows for inconsistencies and incompleteness. Using hundreds of realistic integration scenarios, we demonstrate that the accuracy of CMD is more than 33% above that of metadata-only approaches already for small data examples, and that CMD routinely finds perfect mappings even if a quarter of the data is inconsistent. Angelika Kimmig, Alex Memory, Renée J. Miller, Lise Getoor |
ICDE | 4 |
| 2017 | Collective Entity Resolution in Familial NetworksabstractEntity resolution in settings with rich relational structure often introduces complex dependencies between co-references. Exploiting these dependencies is challenging - it requires seamlessly combining statistical, relational, and logical dependencies. One task of particular interest is entity resolution in familial networks. In this setting, multiple partial representations of a family tree are provided, from the perspective of different family members, and the challenge is to reconstruct a family tree from these multiple, noisy, partial views. This reconstruction is crucial for applications such as understanding genetic inheritance, tracking disease contagion, and performing census surveys. Here, we design a model that incorporates statistical signals, such as name similarity, relational information, such as sibling overlap, and logical constraints, such as transitivity and bijective matching, in a collective model. We show how to integrate these features using probabilistic soft logic, a scalable probabilistic programming framework. In experiments on real-world data, our model significantly outperforms state-of-the-art classifiers that use relational features but are incapable of collective reasoning. Pigi Kouki, Jay Pujara, Christopher Steven Marcum, Laura M. Koehly, Lise Getoor |
ICDM | 5 |
| 2017 | Disambiguating Energy Disaggregation: A Collective Probabilistic ApproachabstractReducing household energy usage is a priority for improving the resiliency and stability of the power grid and decreasing the negative impact of energy consumption on the environment and public health.Relevant and timely feedback about the power consumption of specific appliances can help household residents to reduce their energy demand. Given only a total energy reading, such as that collected from a residential meter, energy disaggregation strives to discover the consumption of individual appliances. Existing disaggregation algorithms are computationally inefficient and rely heavily on high-resolution ground truth data. We introduce a probabilistic framework which infers the energy consumption of individual appliances using a hinge-loss Markov random field (HL-MRF), which admits highly scalable inference. To further enhance efficiency, we introduce a temporal representation which leverages state duration. We also explore how contextual information impacts solution quality with low-resolution data. Our framework is flexible in its ability to incorporate additional constraints; by constraining appliance usage with context and duration we can better disambiguate appliances with similar energy consumption profiles. We demonstrate the effectiveness of our framework on two public real-world datasets, reducing the error relative to a previous state-of-the-art method by as much as 50%. Sabina Tomkins, Jay Pujara, Lise Getoor |
IJCAI | 3 |
| 2017 | Statistical Relational Learning: Unifying AI & DB Perspectives on Structured Probabilistic ModelsabstractMachine learning and database approaches to structured probabilistic models share many commonalities, yet exhibit certain important differences. Machine learning methods focus on learning probabilistic models from (certain) data and efficient learning and inference, whereas probabilistic database approaches focus on storing and efficiently querying uncertain data. Nonetheless, the structured probabilistic models that both use are often (almost) identical. In this tutorial, I will overview the field of statistical relational learning (SRL) [1] and survey common approaches. I'll make connections to work in probabilistic databases [2], and highlight commonalities and differences among them. I'll close by describing some of our recent work on probabilistic soft logic [3]. Lise Getoor |
PODS | 1 |
| 2017 | User Preferences for Hybrid ExplanationsabstractHybrid recommender systems combine several different sources of information to generate recommendations. These systems demonstrate improved accuracy compared to single-source recommendation strategies. However, hybrid recommendation strategies are inherently more complex than those that use a single source of information, and thus the process of explaining recommendations to users becomes more challenging. In this paper we describe a hybrid recommender system built on a probabilistic programming language, and discuss the benefits and challenges of explaining its recommendations to users. We perform a mixed model statistical analysis of user preferences for explanations in this system. Through an online user survey, we evaluate explanations for hybrid algorithms in a variety of text and visual, graph-based formats, that are either novel designs or derived from existing hybrid recommender systems. Pigi Kouki, James Schaffer, Jay Pujara, John O'Donovan, Lise Getoor |
RecSys | 5 |
| 2017 | Multi-relational influence models for online professional networksabstractProfessional networks are a specialized class of social networks that are particularly aimed at forming and strengthening professional connections and have become a vital component of professional success and growth. In this paper, we present a holistic model to jointly represent different heterogenous relationships between pairs of individuals, user actions and their respective propagations to characterize influence in online professional networks. Previous work on influence in social networks typically only consider a single action type in characterizing influence. Our model is capable of representing and combining different kinds of information users assimilate in the network and compute pairwise values of influence taking the different types of actions into account. We evaluate our models on data from the largest professional network, LinkedIn and show the effectiveness of the inferred influence scores in predicting user actions. We further demonstrate that modeling different user actions, node features, and edge relationships between users leads to around 20% increase in precision at top k in predicting user actions, when compared to the current state-of-the-art model. Arti Ramesh, Mario Rodriguez, Lise Getoor |
WI | 3 |
| 2017 | Probabilistic Visitor Stitching on Cross-Device Web LogsabstractPersonalization -- the customization of experiences, interfaces, and content to individual users -- has catalyzed user growth and engagement for many web services. A critical prerequisite to personalization is establishing user identity. However the variety of devices, including mobile phones, appliances, and smart watches, from which users access web services from both anonymous and logged-in sessions poses a significant obstacle to user identification. The resulting entity resolution task of establishing user identity across devices and sessions is commonly referred to as ``visitor stitching.'' We introduce a general, probabilistic approach to visitor stitching using features and attributes commonly contained in web logs. Using web logs from two real-world corporate websites, we motivate the need for probabilistic models by quantifying the difficulties posed by noise, ambiguity, and missing information in deployment. Next, we introduce our approach using probabilistic soft logic (PSL), a statistical relational learning framework capable of capturing similarities across many sessions and enforcing transitivity. We present a detailed description of model features and design choices relevant to the visitor stitching problem. Finally, we evaluate our PSL model on binary classification performance for two real-world visitor stitching datasets. Our model demonstrates significantly better performance than several state-of-the-art classifiers, and we show how this advantage results from collective reasoning across sessions. Sungchul Kim, Nikhil Kini, Jay Pujara, Eunyee Koh, Lise Getoor |
WWW | 5 |
| 2017 | Hinge-Loss Markov Random Fields and Probabilistic Soft LogicabstractA fundamental challenge in developing high-impact machine learning technologies is balancing the need to model rich, structured domains with the ability to scale to big data. Many important problem areas are both richly structured and large scale, from social and biological networks, to knowledge graphs and the Web, to images, video, and natural language. In this paper, we introduce two new formalisms for modeling structured data, and show that they can both capture rich structure and scale to big data. The first, hinge-loss Markov random fields (HL-MRFs), is a new kind of probabilistic graphical model that generalizes different approaches to convex inference. We unite three approaches from the randomized algorithms, probabilistic graphical models, and fuzzy logic communities, showing that all three lead to the same inference objective. We then define HL- MRFs by generalizing this unified objective. The second new formalism, probabilistic soft logic (PSL), is a probabilistic programming language that makes HL-MRFs easy to define using a syntax based on first-order logic. We introduce an algorithm for inferring most-probable variable assignments (MAP inference) that is much more scalable than general-purpose convex optimization methods, because it uses message passing to take advantage of sparse dependency structures. We then show how to learn the parameters of HL-MRFs. The learned HL-MRFs are as accurate as analogous discrete models, but much more scalable. Together, these algorithms enable HL-MRFs and PSL to model rich, structured data at scales not previously possible. Stephen H. Bach, Matthias Broecheler, Bert Huang, Lise Getoor |
J. Mach. Learn. Res. | 4 |
| 2017 | Soft quantification in statistical relational learning
Golnoosh Farnadi, Stephen H. Bach, Marie-Francine Moens, Lise Getoor, Martine De Cock |
Mach. Learn. | 4 |
| 2016 | Unsupervised models for predicting strategic relations between organizationsabstractMicroblogging sites like Twitter provide a platform for sharing ideas and expressing opinions. The widespread popularity of these platforms and the complex social structure that arises within these communities provides a unique opportunity to understand the interactions between users. The political domain, especially in a multi-party system, presents compelling challenges, as political parties have different levels of alignment based on their political strategies. We use Twitter to understand the nuanced relationships between differing political entities in Latin America. Our model incorporates diverse signals from the content of tweets and social context from retweets, mentions and hashtag usage. Since direct communications between entities are relatively rare, we explore models based on the posts of users who interact with multiple political organizations. We present a quantitative and qualitative analysis of the results of models using different features, and demonstrate that a model capable of using sentiment strength, social context, and issue alignment has superior performance to less sophisticated baselines. Shachi H. Kumar, Jay Pujara, Lise Getoor, David Mares, Dipak Gupta, Ellen Riloff |
ASONAM | 3 |
| 2016 | ASONAM 2016 panel: Social network analysis for social goodabstractNo abstract or record of the panel discussion was made available for publication as part of the conference proceedings. V. S. Subrahmanian, Lada A. Adamic, Lise Getoor, Evimaria Terzi, Brian Uzzi, Lisa Singh |
ASONAM | 3 |
| 2016 | ASONAM 2016 keynotes: Ideas and inventionsabstractSummary form only given. The complete presentations were not made available for publication as part of the conference proceedings.These Keynotes speeches the following: Ideas and Inventions; Scalable Collective Reasoning for Richly Structured Socio-Behavioral Data; Computational Methods for Team Formation and Network structure and information diffusion. Brian Uzzi, Lise Getoor, Evimaria Terzi, Lada A. Adamic |
ASONAM | 2 |
| 2016 | Predicting Post-Test Performance from Student Behavior: A High School MOOC Case Study
Sabina Tomkins, Arti Ramesh, Lise Getoor |
EDM | 3 |
| 2016 | SourceSight: Enabling Effective Source SelectionabstractRecently there has been a rapid increase in the number of data sources and data services, such as cloud-based data markets and data portals, that facilitate the collection, publishing and trading of data. Data sources typically exhibit large heterogeneity in the type and quality of data they provide. Unfortunately, when the number of data sources is large, it is difficult for users to reason about the actual usefulness of sources for their applications and the trade-offs between the benefits and costs of acquiring and integrating sources. In this demonstration we present \textsc{SourceSight}, a system that allows users to interactively explore a large number of heterogeneous data sources, and discover valuable sets of sources for diverse integration tasks. \textsc{SourceSight}~uses a novel multi-level source quality index that enables effective source selection at different granularity levels, and introduces a collection of new techniques to discover and evaluate relevant sources for integration. Theodoros Rekatsinas, Amol Deshpande, Xin Dong 0001, Lise Getoor, Divesh Srivastava |
SIGMOD Conference | 4 |
| 2016 | A probabilistic approach for collective similarity-based drug-drug interaction predictionabstractMOTIVATION: As concurrent use of multiple medications becomes ubiquitous among patients, it is crucial to characterize both adverse and synergistic interactions between drugs. Statistical methods for prediction of putative drug-drug interactions (DDIs) can guide in vitro testing and cut down significant cost and effort. With the abundance of experimental data characterizing drugs and their associated targets, such methods must effectively fuse multiple sources of information and perform inference over the network of drugs. RESULTS: We propose a probabilistic approach for jointly inferring unknown DDIs from a network of multiple drug-based similarities and known interactions. We use the highly scalable and easily extensible probabilistic programming framework Probabilistic Soft Logic We compare against two methods including a state-of-the-art DDI prediction system across three experiments and show best performing improvements of more than 50% in AUPR over both baselines. We find five novel interactions validated by external sources among the top-ranked predictions of our model. AVAILABILITY AND IMPLEMENTATION: Final versions of all datasets and implementations will be made publicly available. CONTACT: [email protected]. Dhanya Sridhar, Shobeir Fakhraei, Lise Getoor |
Bioinform. | 3 |
| 2016 | Stability and Generalization in Structured PredictionabstractStructured prediction models have been found to learn effectively from a few large examples--- sometimes even just one. Despite empirical evidence, canonical learning theory cannot guarantee generalization in this setting because the error bounds decrease as a function of the number of examples. We therefore propose new PAC-Bayesian generalization bounds for structured prediction that decrease as a function of both the number of examples and the size of each example. Our analysis hinges on the stability of joint inference and the smoothness of the data distribution. We apply our bounds to several common learning scenarios, including max-margin and soft-max training of Markov random fields. Under certain conditions, the resulting error bounds can be far more optimistic than previous results and can even guarantee generalization from a single large example. Ben London 0001, Bert Huang, Lise Getoor |
J. Mach. Learn. Res. | 3 |
| 2016 | Collective Graph IdentificationabstractData describing networks—such as communication networks, transaction networks, disease transmission networks, collaboration networks, etc.—are becoming increasingly available. While observational data can be useful, it often only hints at the actual underlying process that governs interactions and attributes. For example, an email communication network provides insight into its users and their relationships, but is not the same as the “real” underlying social network. In this article, we introduce the problem of graph identification , i.e., discovering the latent graph structure underlying an observed network. We cast the problem as a probabilistic inference task, in which we must infer the nodes, edges, and node labels of a hidden graph, based on evidence. This entails solving several canonical problems in network analysis: entity resolution (determining when two observations correspond to the same entity), link prediction (inferring the existence of links), and node labeling (inferring hidden attributes). While each of these subproblems has been well studied in isolation, here we consider them as a single, collective task. We present a simple, yet novel, approach to address all three subproblems simultaneously. Our approach, which we refer to as C 3 , consists of a collection of Coupled Collective Classifiers that are applied iteratively to propagate inferred information among the subproblems. We consider variants of C 3 using different learning and inference techniques and empirically demonstrate that C 3 is superior, both in terms of predictive accuracy and running time, to state-of-the-art probabilistic approaches on four real problems. Galileo Namata, Ben London 0001, Lise Getoor |
ACM Trans. Knowl. Discov. Data | 3 |
| 2015 | Planned Protest Modeling in News and Social MediaabstractCivil unrest (protests, strikes, and “occupy” events) is a common occurrence in both democracies and authoritarian regimes. The study of civil unrest is a key topic for political scientists as it helps capture an important mechanism by which citizenry express themselves. In countries where civil unrest is lawful, qualitative analysis has revealed that more than 75% of the protests are planned, organized, and/or announced in advance; therefore detecting future time mentions in relevant news and social media is a direct way to develop a protest forecasting system. We develop such a system in this paper, using a combination of key phrase learning to identify what to look for, probabilistic soft logic to reason about location occurrences in extracted results, and time normalization to resolve future tense mentions. We illustrate the application of our system to 10 countries in Latin America, viz. Argentina, Brazil, Chile, Colombia, Ecuador, El Salvador, Mexico, Paraguay, Uruguay, and Venezuela. Results demonstrate our successes in capturing significant societal unrest in these countries with an average lead time of 4.08 days. We also study the selective superiorities of news media versus social media (Twitter, Facebook) to identify relevant tradeoffs. Sathappan Muthiah, Bert Huang, Jaime Arredondo, David Mares, Lise Getoor, Graham Katz, Naren Ramakrishnan |
AAAI | 5 |
| 2015 | Weakly Supervised Models of Aspect-Sentiment for Online Course Discussion ForumsabstractArti Ramesh, Shachi H. Kumar, James Foulds, Lise Getoor. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Arti Ramesh, Shachi H. Kumar, James R. Foulds, Lise Getoor |
ACL (1) | 4 |
| 2015 | Joint Models of Disagreement and Stance in Online DebateabstractDhanya Sridhar, James Foulds, Bert Huang, Lise Getoor, Marilyn Walker. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Dhanya Sridhar, James R. Foulds, Bert Huang, Lise Getoor, Marilyn A. Walker |
ACL (1) | 4 |
| 2015 | Unifying Local Consistency and MAX SAT Relaxations for Scalable Inference with Rounding GuaranteesabstractWe prove the equivalence of first-order local consistency relaxations and the MAX SAT relaxation of Goemans and Williamson (1994) for a class of MRFs we refer to as logical MRFs. This allows us to combine the advantages of each into a single MAP inference technique: solving the local consistency relaxation with any of a number of highly scalable message-passing algorithms, and then obtaining a high-quality discrete solution via a guaranteed rounding procedure when the relaxation is not tight. Logical MRFs are a general class of models that can incorporate many common dependencies, such as logical implications and mixtures of supermodular and submodular potentials. They can be used for many structured prediction tasks, including natural language processing, computer vision, and computational social science. We show that our new inference technique can improve solution quality by as much as 20% without sacrificing speed on problems with over one million dependencies. Stephen H. Bach, Bert Huang, Lise Getoor |
AISTATS | 3 |
| 2015 | Finding Quality in Quantity: The Challenge of Discovering Valuable Sources for Integration
Theodoros Rekatsinas, Xin Dong 0001, Lise Getoor, Divesh Srivastava |
CIDR | 3 |
| 2015 | RELLY: Inferring Hypernym Relationships Between Relational PhrasesabstractRelational phrases (e.g., "got married to") and their hypernyms (e.g., "is a relative of") are central for many tasks including question answering, open information extraction, paraphrasing, and entailment detection.This has motivated the development of several linguistic resources (e.g.DIRT, PATTY, and WiseNet) which systematically collect and organize relational phrases.These resources have demonstrable practical benefits, but are each limited due to noise, sparsity, or size.We present a new general-purpose method, RELLY, for constructing a large hypernymy graph of relational phrases with high-quality subsumptions using collective probabilistic programming techniques.Our graph induction approach integrates small highprecision knowledge bases together with large automatically curated resources, and reasons collectively to combine these resources into a consistent graph.Using RELLY, we construct a high-coverage, high-precision hypernymy graph consisting of 20K relational phrases and 35K hypernymy links.Our evaluation indicates a hypernymy link precision of 78%, and demonstrates the value of this resource for a document-relevance ranking task. Adam Grycner, Gerhard Weikum, Jay Pujara, James R. Foulds, Lise Getoor |
EMNLP | 5 |
| 2015 | Paired-Dual Learning for Fast Training of Latent Variable Hinge-Loss MRFsabstractLatent variables allow probabilistic graphical models to capture nuance and structure in important domains such as network science, natural language processing, and computer vision. Naive approaches to learning such complex models can be prohibitively expensive—because they require repeated inferences to update beliefs about latent variables—so lifting this restriction for useful classes of models is an important problem. Hinge-loss Markov random fields (HL-MRFs) are graphical models that allow highly scalable inference and learning in structured domains, in part by representing structured problems with continuous variables. However, this representation leads to challenges when learning with latent variables. We introduce paired-dual learning, a framework that greatly speeds up training by using tractable entropy surrogates and avoiding repeated inferences. Paired-dual learning optimizes an objective with a pair of dual inference problems. This allows fast, joint optimization of parameters and dual variables. We evaluate on social-group detection, trust prediction in social networks, and image reconstruction, finding that paired-dual learning trains models as accurate as those trained by traditional methods in much less time, often before traditional methods make even a single parameter update. Stephen H. Bach, Bert Huang, Jordan L. Boyd-Graber, Lise Getoor |
ICML | 4 |
| 2015 | Latent Topic Networks: A Versatile Probabilistic Programming Framework for Topic ModelsabstractTopic models have become increasingly prominent text-analytic machine learning tools for research in the social sciences and the humanities. In particular, custom topic models can be developed to answer specific research questions. The design of these models requires a non-trivial amount of effort and expertise, motivating general-purpose topic modeling frameworks. In this paper we introduce latent topic networks, a flexible class of richly structured topic models designed to facilitate applied research. Custom models can straightforwardly be developed in our framework with an intuitive first-order logical probabilistic programming language. Latent topic networks admit scalable training via a parallelizable EM algorithm which leverages ADMM in the M-step. We demonstrate the broad applicability of the models with case studies on modeling influence in citation networks, and U.S. Presidential State of the Union addresses. James R. Foulds, Shachi H. Kumar, Lise Getoor |
ICML | 3 |
| 2015 | HawkesTopic: A Joint Model for Network Inference and Topic Modeling from Text-Based CascadesabstractUnderstanding the diffusion of information in social network and social media requires modeling the text diffusion process. In this work, we develop the HawkesTopic model (HTM) for analyzing text-based cascades, such as "retweeting a post" or "publishing a follow-up blog post". HTM combines Hawkes processes and topic modeling to simultaneously reason about the information diffusion pathways and the topics characterizing the observed textual information. We show how to jointly infer them with a mean-field variational inference algorithm and validate our approach on both synthetic and real-world data sets, including a news media dataset for modeling information diffusion, and an ArXiv publication dataset for modeling scientific influence. The results show that HTM is significantly more accurate than several baselines for both tasks. Xinran He, Theodoros Rekatsinas, James R. Foulds, Lise Getoor, Yan Liu 0002 |
ICML | 4 |
| 2015 | The Benefits of Learning with Strongly Convex Approximate InferenceabstractWe explore the benefits of strongly convex free energies in variational inference, providing both theoretical motivation and a new meta-algorithm. Using the duality between strong convexity and stability, we prove a high-probability bound on the error of learned marginals that is inversely proportional to the modulus of convexity of the free energy, thereby motivating free energies whose moduli are constant with respect to the size of the graph. We identify sufficient conditions for Ω(1)-strong convexity in two popular variational techniques: tree-reweighted and counting number entropies. Our insights for the latter suggest a novel counting number optimization framework, which guarantees strong convexity for any given modulus. Our experiments demonstrate that learning with a strongly convex free energy, using our optimization framework to guarantee a given modulus, results in substantially more accurate marginal probabilities, thereby validating our theoretical claims and the effectiveness of our framework. Ben London 0001, Bert Huang, Lise Getoor |
ICML | 3 |
| 2015 | Statistical Relational Learning with Soft Quantifiers
Golnoosh Farnadi, Stephen H. Bach, Marjon Blondeel, Marie-Francine Moens, Lise Getoor, Martine De Cock |
ILP | 5 |
| 2015 | Collective Spammer Detection in Evolving Multi-Relational Social NetworksabstractDetecting unsolicited content and the spammers who create it is a long-standing challenge that affects all of us on a daily basis. The recent growth of richly-structured social networks has provided new challenges and opportunities in the spam detection landscape. Motivated by the Tagged.com social network, we develop methods to identify spammers in evolving multi-relational social networks. We model a social network as a time-stamped multi-relational graph where vertices represent users, and edges represent different activities between them. To identify spammer accounts, our approach makes use of structural features, sequence modelling, and collective reasoning. We leverage relational sequence information using k-gram features and probabilistic modelling with a mixture of Markov models. Furthermore, in order to perform collective reasoning and improve the predictive power of a noisy abuse reporting system, we develop a statistical relational model using hinge-loss Markov random fields (HL-MRFs), a class of probabilistic graphical models which are highly scalable. We use Graphlab Create and Probabilistic Soft Logic (PSL) to prototype and experimentally evaluate our solutions on internet-scale data from Tagged.com. Our experiments demonstrate the effectiveness of our approach, and show that models which incorporate the multi-relational nature of the social network significantly gain predictive performance over those that do not. Shobeir Fakhraei, James R. Foulds, Madhusudana V. S. Shashanka, Lise Getoor |
KDD | 4 |
| 2015 | HyPER: A Flexible and Extensible Probabilistic Framework for Hybrid Recommender SystemsabstractAs the amount of recorded digital information increases, there is a growing need for flexible recommender systems which can incorporate richly structured data sources to improve recommendations. In this paper, we show how a recently introduced statistical relational learning framework can be used to develop a generic and extensible hybrid recommender system. Our hybrid approach, HyPER (HYbrid Probabilistic Extensible Recommender), incorporates and reasons over a wide range of information sources. Such sources include multiple user-user and item-item similarity measures, content, and social information. HyPER automatically learns to balance these different information signals when making predictions. We build our system using a powerful and intuitive probabilistic programming language called probabilistic soft logic, which enables efficient and accurate prediction by formulating our custom recommender systems with a scalable class of graphical models known as hinge-loss Markov random fields. We experimentally evaluate our approach on two popular recommendation datasets, showing that HyPER can effectively combine multiple information types for improved performance, and can significantly outperform existing state-of-the-art approaches. Pigi Kouki, Shobeir Fakhraei, James R. Foulds, Magdalini Eirinaki, Lise Getoor |
RecSys | 5 |
| 2015 | SourceSeer: Forecasting Rare Disease Outbreaks Using Multiple Data SourcesabstractRapidly increasing volumes of news feeds from diverse data sources, such as online newspapers, Twitter and online blogs are proving to be extremely valuable resources in helping anticipate, detect, and forecast outbreaks of rare diseases. This paper presents SourceSeer, a novel algorithmic framework that combines spatio-temporal topic models with sourcebased anomaly detection techniques to effectively forecast the emergence and progression of infectious rare diseases. SourceSeer is capable of discovering the location focus of each source allowing sources to be used as experts with varying degrees of authoritativeness. To fuse the individual source predictions into a final outbreak prediction we employ a multiplicative weights algorithm taking into account the accuracy of each source. We evaluate the performance of SourceSeer using incidence data for hantavirus syndromes in multiple countries of Latin America provided by HealthMap over a timespan of fifteen months. We demonstrate that SourceSeer makes predictions of increased accuracy compared to several baselines and is capable of forecasting disease outbreaks in a timely manner even when no outbreaks were previously reported. Theodoros Rekatsinas, Saurav Ghosh, Sumiko R. Mekaru, Elaine O. Nsoesie, John S. Brownstein, Lise Getoor, Naren Ramakrishnan |
SDM | 6 |
| 2015 | Budgeted Online Collective Inference
Jay Pujara, Ben London 0001, Lise Getoor |
UAI | 3 |
| 2015 | Lifted graphical models: a survey
Angelika Kimmig, Lilyana Mihalkova, Lise Getoor |
Mach. Learn. | 3 |
| 2014 | Learning Latent Engagement Patterns of Students in Online CoursesabstractMaintaining and cultivating student engagement is critical for learning. Understanding factors affecting student engagement will help in designing better courses and improving student retention. The large number of participants in massive open online courses (MOOCs) and data collected from their interaction with the MOOC open up avenues for studying student engagement at scale. In this work, we develop a framework for modeling and understanding student engagement in online courses based on student behavioral cues. Our first contribution is the abstraction of student engagement types using latent representations and using that in a probabilistic model to connect student behavior with course completion. We demonstrate that the latent formulation for engagement helps in predicting student survival across three MOOCs. Next, in order to initiate better instructor interventions, we need to be able to predict student survival early in the course. We demonstrate that we can predict student survival early in the course reliably using the latent model. Finally, we perform a closer quantitative analysis of user interaction with the MOOC and identify student activities that are good indicators for survival at different points in the course. Arti Ramesh, Dan Goldwasser, Bert Huang, Hal Daumé III, Lise Getoor |
AAAI | 5 |
| 2014 | PAC-Bayesian Collective StabilityabstractRecent results have shown that the generalization error of structured predictors decreases with both the number of examples and the size of each example, provided the data distribution has weak dependence and the predictor exhibits a smoothness property called collective stability. These results use an especially strong definition of collective stability that must hold uniformly over all inputs and all hypotheses in the class. We investigate whether weaker definitions of collective stability suffice. Using the PAC-Bayes framework, which is particularly amenable to our new definitions, we prove that generalization is indeed possible when uniform collective stability happens with high probability over draws of predictors (and inputs). We then derive a generalization bound for a class of structured predictors with variably convex inference, which suggests a novel learning objective that optimizes collective stability. Ben London 0001, Bert Huang, Ben Taskar, Lise Getoor |
AISTATS | 4 |
| 2014 | Subgraph pattern matching over uncertain graphs with identity linkage uncertaintyabstractThere is a growing need for methods that can represent and query uncertain graphs. These uncertain graphs are often the result of an information extraction and integration system that attempts to extract an entity graph or a knowledge graph from multiple unstructured sources [25], [7]. Such an integration typically leads to identity uncertainty, as different data sources may use different references to the same underlying real-world entities. Integration usually also introduces additional uncertainty on node attributes and edge existence. In this paper, we propose the notion of a probabilistic entity graph (PEG), a formal model that uniformly and systematically addresses these three types of uncertainty. A PEG is a probabilistic graph model that defines a distribution over possible graphs at the entity level. We introduce a general framework for constructing a PEG given uncertain data at the reference level and develop efficient algorithms to answer subgraph pattern matching queries in this setting. Our algorithms are based on two novel ideas: context-aware path indexing and reduction by join-candidates, which drastically reduce the query search space. A comprehensive experimental evaluation shows that our approach outperforms baseline implementations by orders of magnitude. Walaa Eldin Moustafa, Angelika Kimmig, Amol Deshpande, Lise Getoor |
ICDE | 4 |
| 2014 | 'Beating the news' with EMBERS: forecasting civil unrest using open source indicatorsabstractWe describe the design, implementation, and evaluation of EMBERS, an automated, 24x7 continuous system for forecasting civil unrest across 10 countries of Latin America using open source indicators such as tweets, news sources, blogs, economic indicators, and other data sources. Unlike retrospective studies, EMBERS has been making forecasts into the future since Nov 2012 which have been (and continue to be) evaluated by an independent T&E team (MITRE). Of note, EMBERS has successfully forecast the June 2013 protests in Brazil and Feb 2014 violent protests in Venezuela. We outline the system architecture of EMBERS, individual models that leverage specific data sources, and a fusion and suppression engine that supports trading off specific evaluation criteria. EMBERS also provides an audit trail interface that enables the investigation of why specific predictions were made along with the data utilized for forecasting. Through numerous evaluations, we demonstrate the superiority of EMBERS over baserate methods and its capability to forecast significant societal happenings. Naren Ramakrishnan, Patrick Butler, Sathappan Muthiah, Nathan Self, Rupinder Paul Khandpur, Parang Saraf, Wei Wang 0064, Jose Cadena, Anil Vullikanti, Gizem Korkmaz, Chris J. Kuhlman, Achla Marathe, Liang Zhao 0002, Ting Hua, Feng Chen 0001, Chang-Tien Lu, Bert Huang, Aravind Srinivasan, Khoa Trinh, Lise Getoor, Graham Katz, Andy Doyle, Chris Ackermann, Ilya Zavorin, Jim Ford, Kristen Maria Summers, Youssef Fayed, Jaime Arredondo, Dipak Gupta, David Mares |
KDD | 20 |
| 2014 | Uncovering hidden engagement patterns for predicting learner performance in MOOCsabstractMaintaining and cultivating student engagement is a prerequisite for MOOCs to have broad educational impact. Understanding student engagement as a course progresses helps characterize student learning patterns and can aid in minimizing dropout rates, initiating instructor intervention. In this paper, we construct a probabilistic model connecting student behavior and class performance, formulating student engagement types as latent variables. We show that our model identifies course success indicators that can be used by instructors to initiate interventions and assist students. Arti Ramesh, Dan Goldwasser, Bert Huang, Hal Daumé III, Lise Getoor |
L@S | 5 |
| 2014 | Network-Based Drug-Target Interaction Prediction with Probabilistic Soft LogicabstractDrug-target interaction studies are important because they can predict drugs' unexpected therapeutic or adverse side effects. In silico predictions of potential interactions are valuable and can focus effort on in vitro experiments. We propose a prediction framework that represents the problem using a bipartite graph of drug-target interactions augmented with drug-drug and target-target similarity measures and makes predictions using probabilistic soft logic (PSL). Using probabilistic rules in PSL, we predict interactions with models based on triad and tetrad structures. We apply (blocking) techniques that make link prediction in PSL more efficient for drug-target interaction prediction. We then perform extensive experimental studies to highlight different aspects of the model and the domain, first comparing the models with different structures and then measuring the effect of the proposed blocking on the prediction performance and efficiency. We demonstrate the importance of rule weight learning in the proposed PSL model and then show that PSL can effectively make use of a variety of similarity measures. We perform an experiment to validate the importance of collective inference and using multiple similarity measures for accurate predictions in contrast to non-collective and single similarity assumptions. Finally, we illustrate that our PSL model achieves state-of-the-art performance with simple, interpretable rules and evaluate our novel predictions using online data sets. Shobeir Fakhraei, Bert Huang, Louiqa Raschid, Lise Getoor |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2014 | Topic Modeling for Wikipedia Link DisambiguationabstractMany articles in the online encyclopedia Wikipedia have hyperlinks to ambiguous article titles; these ambiguous links should be replaced with links to unambiguous articles, a process known as disambiguation. We propose a novel statistical topic model based on link text, which we refer to as the Link Text Topic Model (LTTM), that we use to suggest new link targets for ambiguous links. To evaluate our model, we describe a method for extracting ground truth for this link disambiguation task from edits made to Wikipedia in a specific time period. We use this ground truth to demonstrate the superiority of LTTM over other existing link- and content-based approaches to disambiguating links in Wikipedia. Finally, we build a web service that uses LTTM to make suggestions to human editors wanting to fix ambiguous links in Wikipedia. Bradley Skaggs, Lise Getoor |
ACM Trans. Inf. Syst. | 2 |
| 2013 | A hypergraph-partitioned vertex programming approach for large-scale consensus optimizationabstractIn modern data science problems, techniques for extracting value from big data require performing large-scale optimization over heterogenous, irregularly structured data. Much of this data is best represented as multi-relational graphs, making vertex-programming abstractions such as those of Pregel and GraphLab ideal fits for modern large-scale data analysis. In this paper, we describe a vertex-programming implementation of a popular consensus optimization technique known as the alternating direction method of multipliers (ADMM) [1]. ADMM consensus optimization allows the elegant solution of complex objectives such as inference in rich probabilistic models. We also introduce a novel hypergraph partitioning technique that improves over the state-of-the-art vertex programming framework and significantly reduces the communication cost by reducing the number of replicated nodes by an order of magnitude. We implement our algorithm in GraphLab and measure scaling performance on a variety of realistic bipartite graphs and a large synthetic voter-opinion analysis application. We show a 50% improvement in running time over the current GraphLab partitioning scheme. Hui Miao 0001, Bert Huang, Lise Getoor |
IEEE BigData | 4 |
| 2013 | Collective Stability in Structured Prediction: Generalization from One ExampleabstractStructured predictors enable joint inference over multiple interdependent output variables. These models are often trained on a small number of examples with large internal structure. Existing distribution-free generalization bounds do not guarantee generalization in this setting, though this contradicts a large body of empirical evidence from computer vision, natural language processing, social networks and other fields. In this paper, we identify a set of natural conditions – weak dependence, hypothesis complexity and a new measure, collective stability – that are sufficient for generalization from even a single example, without imposing an explicit generative model of the data. We then demonstrate that the complexity and stability conditions are satisfied by a broad class of models, including marginal inference in templated graphical models. We thus obtain uniform convergence rates that can decrease significantly faster than previous bounds, particularly when each structured example is sufficiently large and the number of training examples is constant, even one. Ben London 0001, Bert Huang, Ben Taskar, Lise Getoor |
ICML (3) | 4 |
| 2013 | Entity resolution for big dataabstractEntity resolution (ER), the problem of extracting, matching and resolving entity mentions in structured and unstructured data, is a long-standing challenge in database management, information retrieval, machine learning, natural language processing and statistics. Accurate and fast entity resolution has huge practical implications in a wide variety of commercial, scientific and security domains. Despite the long history of work on entity resolution, there is still a surprising diversity of approaches, and lack of guiding theory. Meanwhile, in the age of big data, the need for high quality entity resolution is growing, as we are inundated with more and more data, all of which needs to be integrated, aligned and matched, before further utility can be extracted. In this tutorial, we bring together perspectives on entity resolution from a variety of fields, including databases, information retrieval, natural language processing and machine learning, to provide, in one setting, a survey of a large body of work. We discuss both the practical aspects and theoretical underpinnings of ER. We describe existing solutions, current challenges and open research problems. In addition to giving attendees a thorough understanding of existing ER models, algorithms and evaluation methods, the tutorial will cover important research topics such as scalable ER, active and lightly supervised ER, and query-driven ER. Lise Getoor, Ashwin Machanavajjhala |
KDD | 1 |
| 2013 | Network samplingabstractNetwork data appears in various domains, including social, communication, and information sciences. Analysis of such data is crucial for making inferences and predictions about these networks, and moreover, for understanding the different processes that drive their evolution. However, a major bottleneck to perform such an analysis is the massive size of real-life networks, which makes modeling and analyzing these networks simply infeasible. Further, many networks, specifically those that belong to social and communication domains, are not visible to the public due to privacy concerns, and other networks, such as the Web, are only accessible via crawling. Therefore, to overcome the above challenges, researchers use network sampling overwhelmingly as a key statistical approach to select a sub-population of interest that can be studied thoroughly. Lise Getoor, Ashwin Machanavajjhala |
KDD | 1 |
| 2013 | Scalable text and link analysis with mixed-topic link modelsabstractMany data sets contain rich information about objects, as well as pairwise relations between them. For instance, in networks of websites, scientific papers, and other documents, each node has content consisting of a collection of words, as well as hyperlinks or citations to other nodes. In order to perform inference on such data sets, and make predictions and recommendations, it is useful to have models that are able to capture the processes which generate the text at each node and the links between them. In this paper, we combine classic ideas in topic modeling with a variant of the mixed-membership block model recently developed in the statistical physics community. The resulting model has the advantage that its parameters, including the mixture of topics of each document and the resulting overlapping communities, can be inferred with a simple and scalable expectation-maximization algorithm. We test our model on three data sets, performing unsupervised topic classification and link prediction. For both tasks, our model outperforms several existing state-of-the-art methods, achieving higher accuracy with significantly less computation, analyzing a data set with 1.3 million words and 44 thousand links in a few minutes. Yaojia Zhu, Xiaoran Yan, Lise Getoor, Cristopher Moore |
KDD | 3 |
| 2013 | Knowledge Graph Identification
Jay Pujara, Hui Miao 0001, Lise Getoor, William W. Cohen |
ISWC (1) | 3 |
| 2013 | GRDB: a system for declarative and interactive analysis of noisy information networksabstractThere is a growing interest in methods for analyzing data describing networks of all types, including biological, physical, social, and scientific collaboration networks. Typically the data describing these networks is observational, and thus noisy and incomplete; it is often at the wrong level of fidelity and abstraction for meaningful data analysis. This demonstration presents GrDB, a system that enables data analysts to write declarative programs to specify and combine different network data cleaning tasks, visualize the output, and engage in the process of decision review and correction if necessary. The declarative interface of GrDB makes it very easy to quickly write analysis tasks and execute them over data, while the visual component facilitates debugging the program and performing fine grained corrections. Walaa Eldin Moustafa, Hui Miao 0001, Amol Deshpande, Lise Getoor |
SIGMOD Conference | 4 |
| 2013 | Hinge-loss Markov Random Fields: Convex Inference for Structured Prediction
Stephen H. Bach, Bert Huang, Ben London 0001, Lise Getoor |
UAI | 4 |
| 2013 | TACI: Taxonomy-Aware Catalog IntegrationabstractA fundamental data integration task faced by online commercial portals and commerce search engines is the integration of products coming from multiple providers to their product catalogs. In this scenario, the commercial portal has its own taxonomy (the “master taxonomy”), while each data provider organizes its products into a different taxonomy (the “provider taxonomy”). In this paper, we consider the problem of categorizing products from the data providers into the master taxonomy, while making use of the provider taxonomy information. Our approach is based on a taxonomy-aware processing step that adjusts the results of a text-based classifier to ensure that products that are close together in the provider taxonomy remain close in the master taxonomy. We formulate this intuition as a structured prediction optimization problem. To the best of our knowledge, this is the first approach that leverages the structure of taxonomies in order to enhance catalog integration. We propose algorithms that are scalable and thus applicable to the large data sets that are typical on the web. We evaluate our algorithms on real-world data and we show that taxonomy-aware classification provides a significant improvement over existing approaches. Panagiotis Papadimitriou 0002, Panayiotis Tsaparas, Ariel Fuxman, Lise Getoor |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2012 | Ego-centric Graph Pattern CensusabstractThere is increasing interest in analyzing networks of all types including social, biological, sensor, computer, and transportation networks. Broadly speaking, we may be interested in global network-wide analysis (e.g., centrality analysis, community detection) where the properties of the entire network are of interest, or local ego-centric analysis where the focus is on studying the properties of nodes (egos) by analyzing their neighborhood sub graphs. In this paper we propose and study ego-centric pattern census queries, a new type of graph analysis query, where a given structural pattern is searched for in every node's neighborhood and the counts are reported or used in further analysis. This kind of analysis is useful in many domains in social network analysis including opinion leader identification, node classification, link prediction, and role identification. We propose an SQL-based declarative language to support this class of queries, and develop a series of efficient query evaluation algorithms for it. We evaluate our algorithms on a variety of synthetically generated graphs. We also show an application of our language in a real-world scenario for predicting future collaborations from DBLP data. Walaa Eldin Moustafa, Amol Deshpande, Lise Getoor |
ICDE | 3 |
| 2012 | Scaling MPE Inference for Constrained Continuous Markov Random Fields with Consensus OptimizationabstractProbabilistic graphical models are powerful tools for analyzing constrained, continuous domains. However, finding most-probable explanations (MPEs) in these models can be computationally expensive. In this paper, we improve the scalability of MPE inference in a class of graphical models with piecewise-linear and piecewise-quadratic dependencies and linear constraints over continuous domains. We derive algorithms based on a consensus-optimization framework and demonstrate their superior performance over state of the art. We show empirically that in a large-scale voter-preference modeling problem our algorithms scale linearly in the number of dependencies and constraints. Stephen H. Bach, Matthias Broecheler, Lise Getoor, Dianne P. O'Leary |
NIPS | 3 |
| 2012 | Local structure and determinism in probabilistic databasesabstractWhile extensive work has been done on evaluating queries over tuple-independent probabilistic databases, query evaluation over correlated data has received much less attention even though the support for correlations is essential for many natural applications of probabilistic databases, e.g., information extraction, data integration, computer vision, etc. In this paper, we develop a novel approach for efficiently evaluating probabilistic queries over correlated databases where correlations are represented using a factor graph, a class of graphical models widely used for capturing correlations and performing statistical inference. Our approach exploits the specific values of the factor parameters and the determinism in the correlations, collectively called local structure, to reduce the complexity of query evaluation. Our framework is based on arithmetic circuits, factorized representations of probability distributions that can exploit such local structure. Traditionally, arithmetic circuits are generated following a compilation process and can not be updated directly. We introduce a generalization of arithmetic circuits, called annotated arithmetic circuits, and a novel algorithm for updating them, which enables us to answer probabilistic queries efficiently. We present a comprehensive experimental analysis and show speed-ups of at least one order of magnitude in many cases. Theodoros Rekatsinas, Amol Deshpande, Lise Getoor |
SIGMOD Conference | 3 |
| 2012 | Entity Resolution: Theory, Practice & Open ChallengesabstractThis tutorial brings together perspectives on ER from a variety of fields, including databases, machine learning, natural language processing and information retrieval, to provide, in one setting, a survey of a large body of work. We discuss both the practical aspects and theoretical underpinnings of ER. We describe existing solutions, current challenges, and open research problems. Lise Getoor, Ashwin Machanavajjhala |
Proc. VLDB Endow. | 1 |
| 2012 | Organizing User Search HistoriesabstractUsers are increasingly pursuing complex task-oriented goals on the web, such as making travel arrangements, managing finances, or planning purchases. To this end, they usually break down the tasks into a few codependent steps and issue multiple queries around these steps repeatedly over long periods of time. To better support users in their long-term information quests on the web, search engines keep track of their queries and clicks while searching online. In this paper, we study the problem of organizing a user's historical queries into groups in a dynamic and automated fashion. Automatically identifying query groups is helpful for a number of different search engine components and applications, such as query suggestions, result ranking, query alterations, sessionization, and collaborative search. In our approach, we go beyond approaches that rely on textual similarity or time thresholds, and we propose a more robust approach that leverages search query logs. We experimentally study the performance of different techniques, and showcase their potential, especially when combined together. Heasoo Hwang, Hady Wirawan Lauw, Lise Getoor, Alexandros Ntoulas |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2011 | Differential Adaptive Diffusion: Understanding Diversity and Learning whom to Trust in Viral Marketing
Hossam Sharara, William Rand, Lise Getoor |
ICWSM | 3 |
| 2011 | Active Surveying: A Probabilistic Approach for Identifying Key Opinion Leaders
Hossam Sharara, Lise Getoor, Myra Norton |
IJCAI | 2 |
| 2011 | Collective graph identificationabstractData describing networks (communication networks, transaction networks, disease transmission networks, collaboration networks, etc.) is becoming increasingly ubiquitous. While this observational data is useful, it often only hints at the actual underlying social or technological structures which give rise to the interactions. For example, an email communication network provides useful insight but is not the same as the "real" social network among individuals. In this paper, we introduce the problem of graph identification, i.e., the discovery of the true graph structure underlying an observed network. We cast the problem as a probabilistic inference task, in which we must infer the nodes, edges, and node labels of a hidden graph, based on evidence provided by the observed network. This in turn corresponds to the problems of performing entity resolution, link prediction, and node labeling to infer the hidden graph. While each of these problems have been studied separately, they have never been considered together as a coherent task. We present a simple yet novel approach to address all three problems simultaneously. Our approach, called C3, consists of Coupled Collective Classifiers that are iteratively applied to propagate information among solutions to the problems. We empirically demonstrate that C3 is superior, in terms of both predictive accuracy and runtime, to state-of-the-art probabilistic approaches on three real-world problems. Galileo Namata, Stanley Kok, Lise Getoor |
KDD | 3 |
| 2011 | Learning statistical models from relational dataabstractStatistical Relational Learning (SRL) is a subarea of machine learning which combines elements from statistical and probabilistic modeling with languages which support structured data representations. In this survey, we will: 1) provide an introduction to SRL, 2) describe some of the distinguishing characteristics of SRL systems, including relational feature construction and collective classification, 3) describe three SRL systems in detail, 4) discuss applications of SRL techniques to important data management problems such as entity resolution, selectivity estimation, and information integration, and 5) discuss connections between SRL methods and existing database research such as probabilistic databases. Lise Getoor, Lilyana Mihalkova |
SIGMOD Conference | 1 |
| 2011 | Exploiting statistical and relational information on the web and in social mediaabstractThe popularity of Web 2.0, characterized by a proliferation of social media sites, and Web 3.0, with more richly semantically annotated objects and relationships, brings to light a variety of important prediction, ranking, and extraction tasks. The input to these tasks is often best seen as a (noisy) multi-relational graph, such as the click graph, defined by user interactions with Web sites; and the social graph, defined by friendships and affiliations on social media sites.This tutorial will provide an overview of statistical relational learning and inference techniques, motivating and illustrating them using web and social media applications. We will start by briefly surveying some of the sources of statistical and relational information on the web and in social media and will then dedicate most of the tutorial time to an introduction to representations and techniques for learning and reasoning with multi-relational information, viewing them through the lens of web and social media domains. We will end with a discussion of current trends and related fields, such as privacy in social networks. Lise Getoor, Lilyana Mihalkova |
WSDM | 1 |
| 2011 | Materializing multi-relational databases from the web using taxonomic queriesabstractRecently, much attention has been given to extracting tables from Web data. In this problem, the column definitions and tuples (such as what "company" is headquartered in what "city,") are extracted from Web text, structured Web data such as lists, or results of querying the deep Web, creating the table of interest. In this paper, we examine the problem of extracting and discovering multiple tables in a given domain, generating a truly multi-relational database as output. Beyond discovering the relations that define single tables, our approach discovers and leverages "within column" set membership relations, and discovers relations across the extracted tables (e.g., joins). By leveraging within-column relations our method can extract table instances that are ambiguous or rare, and by discovering joins, our method generates truly multi-relational output. Further, our approach uses taxonomic queries to bootstrap the extraction, rather than the more traditional "seed instances." Creating seeds often requires more domain knowledge than taxonomic queries, and previous work has shown that extraction methods may be sensitive to which input seeds they are given. We test our approach on two real world domains: NBA basketball and cancer information. Our results demonstrate that our approach generates databases of relevant tables from disparate Web information, and discovers the relations between them. Further, we show that by leveraging the "within column" relation our approach can identify a significant number of relevant tuples that would be difficult to do so otherwise. Matthew Michelson, Sofus A. Macskassy, Steven Minton, Lise Getoor |
WSDM | 4 |
| 2011 | A probabilistic approach for learning folksonomies from structured dataabstractLearning structured representations has emerged as an important problem in many domains, including document and Web data mining, bioinformatics, and image analysis. One approach to learning complex structures is to integrate many smaller, incomplete and noisy structure fragments. In this work, we present an unsupervised probabilistic approach that extends affinity propagation [7] to combine the small ontological fragments into a collection of integrated, consistent, and larger folksonomies. This is a challenging task because the method must aggregate similar structures while avoiding structural inconsistencies and handling noise. We validate the approach on a real-world social media dataset, comprised of shallow personal hierarchies specified by many individual users, collected from the photosharing website Flickr. Our empirical results show that our proposed approach is able to construct deeper and denser structures, compared to an approach using only the standard affinity propagation algorithm. Additionally, the approach yields better overall integration quality than a state-of-the-art approach based on incremental relational clustering. Anon Plangprasopchok, Kristina Lerman, Lise Getoor |
WSDM | 3 |
| 2011 | Value of Information Lattice: Exploiting Probabilistic Independence for Effective Feature Subset AcquisitionabstractWe address the cost-sensitive feature acquisition problem, where misclassifying an instance is costly but the expected misclassification cost can be reduced by acquiring the values of the missing features. Because acquiring the features is costly as well, the objective is to acquire the right set of features so that the sum of the feature acquisition cost and misclassification cost is minimized. We describe the Value of Information Lattice (VOILA), an optimal and efficient feature subset acquisition framework. Unlike the common practice, which is to acquire features greedily, VOILA can reason with subsets of features. VOILA efficiently searches the space of possible feature subsets by discovering and exploiting conditional independence properties between the features and it reuses probabilistic inference computations to further speed up the process. Through empirical evaluation on five medical datasets, we show that the greedy strategy is often reluctant to acquire features, as it cannot forecast the benefit of acquiring multiple features in combination. Mustafa Bilgic 0001, Lise Getoor |
J. Artif. Intell. Res. | 2 |
| 2011 | Dynamic Processing Allocation in VideoabstractLarge stores of digital video pose severe computational challenges to existing video analysis algorithms. In applying these algorithms, users must often trade off processing speed for accuracy, as many sophisticated and effective algorithms require large computational resources that make it impractical to apply them throughout long videos. One can save considerable effort by applying these expensive algorithms sparingly, directing their application using the results of more limited processing. We show how to do this for retrospective video analysis by modeling a video using a chain graphical model and performing inference both to analyze the video and to direct processing. We apply our method to problems in background subtraction and face detection, and show in experiments that this leads to significant improvements over baseline algorithms. Daozheng Chen, Mustafa Bilgic 0001, Lise Getoor, David Jacobs 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | Active Inference for Collective ClassificationabstractLabeling nodes in a network is an important problem that has seen a growing interest. A number of methods that exploit both local and relational information have been developed for this task. Acquiring the labels for a few nodes at inference time can greatly improve the accuracy, however the question of figuring out which node labels to acquire is challenging. Previous approaches have been based on simple structural properties. Here, we present a novel technique, which we refer to as reflect and correct,that can learn and predict when the underlying classification system is likely to make mistakes and it suggests acquisitions to correct those mistakes. Mustafa Bilgic 0001, Lise Getoor |
AAAI | 2 |
| 2010 | Active Learning for Networked Data
Mustafa Bilgic 0001, Lilyana Mihalkova, Lise Getoor |
ICML | 3 |
| 2010 | Graph Identification
Lise Getoor |
IDA | 1 |
| 2010 | Growing a tree in the forest: constructing folksonomies by integrating structured metadataabstractMany social Web sites allow users to annotate the content with descriptive metadata, such as tags, and more recently to organize content hierarchically. These types of structured metadata provide valuable evidence for learning how a community organizes knowledge. For instance, we can aggregate many personal hierarchies into a common taxonomy, also known as a folksonomy, that will aid users in visualizing and browsing social content, and also to help them in organizing their own content. However, learning from social metadata presents several challenges, since it is sparse, shallow, ambiguous, noisy, and inconsistent. We describe an approach to folksonomy learning based on relational clustering, which exploits structured metadata contained in personal hierarchies. Our approach clusters similar hierarchies using their structure and tag statistics, then incrementally weaves them into a deeper, bushier tree. We study folksonomy learning using social metadata extracted from the photo-sharing site Flickr, and demonstrate that the proposed approach addresses the challenges. Moreover, comparing to previous work, the approach produces larger, more accurate folksonomies, and in addition, scales better. Anon Plangprasopchok, Kristina Lerman, Lise Getoor |
KDD | 3 |
| 2010 | Computing Marginal Distributions over Continuous Markov Networks for Statistical Relational LearningabstractContinuous Markov random fields are a general formalism to model joint probability distributions over events with continuous outcomes. We prove that marginal computation for constrained continuous MRFs is #P-hard in general and present a polynomial-time approximation scheme under mild assumptions on the structure of the random field. Moreover, we introduce a sampling algorithm to compute marginal distributions and develop novel techniques to increase its efficiency. Continuous MRFs are a general purpose probabilistic modeling tool and we demonstrate how they can be applied to statistical relational learning. On the problem of collective classification, we evaluate our algorithm and show that the standard deviation of marginals serves as a useful measure of confidence. Matthias Broecheler, Lise Getoor |
NIPS | 2 |
| 2010 | Learning Algorithms for Link Prediction Based on Chance Constraints
Janardhan Rao Doppa, Prasad Tadepalli, Lise Getoor |
ECML/PKDD (1) | 4 |
| 2010 | Probabilistic Similarity Logic
Matthias Broecheler, Lilyana Mihalkova, Lise Getoor |
UAI | 3 |
| 2010 | Constructing folksonomies by integrating structured metadataabstractAggregating many personal hierarchies into a common taxonomy, also known as a folksonomy, presents several challenges due to its sparseness, ambiguity, noise, and inconsistency. We describe an approach to folksonomy learning based on relational clustering that addresses these challenges by exploiting structured metadata contained in personal hierarchies. Our approach clusters similar hierarchies using their structure and tag statistics, then incrementally weaves them into a deeper, bushier tree. We study folksonomy learning using social metadata extracted from the photo-sharing site Flickr. We evaluate the learned folksonomy quantitatively by automatically comparing it to a reference taxonomy created by the Open Directory Project. Our empirical results suggest that the proposed approach improves upon the state-of-the-art folksonomy learning method. Anon Plangprasopchok, Kristina Lerman, Lise Getoor |
WWW | 3 |
| 2010 | Indirect two-sided relative ranking: a robust similarity measure for gene expression dataabstractBACKGROUND: There is a large amount of gene expression data that exists in the public domain. This data has been generated under a variety of experimental conditions. Unfortunately, these experimental variations have generally prevented researchers from accurately comparing and combining this wealth of data, which still hides many novel insights. RESULTS: In this paper we present a new method, which we refer to as indirect two-sided relative ranking, for comparing gene expression profiles that is robust to variations in experimental conditions. This method extends the current best approach, which is based on comparing the correlations of the up and down regulated genes, by introducing a comparison based on the correlations in rankings across the entire database. Because our method is robust to experimental variations, it allows a greater variety of gene expression data to be combined, which, as we show, leads to richer scientific discoveries. CONCLUSIONS: We demonstrate the benefit of our proposed indirect method on several datasets. We first evaluate the ability of the indirect method to retrieve compounds with similar therapeutic effects across known experimental barriers, namely vehicle and batch effects, on two independent datasets (one private and one public). We show that our indirect method is able to significantly improve upon the previous state-of-the-art method with a substantial improvement in recall at rank 10 of 97.03% and 49.44%, on each dataset, respectively. Next, we demonstrate that our indirect method results in improved accuracy for classification in several additional datasets. These datasets demonstrate the use of our indirect method for classifying cancer subtypes, predicting drug sensitivity/resistance, and classifying (related) cell types. Even in the absence of a known (i.e., labeled) experimental barrier, the improvement of the indirect method in each of these datasets is statistically significant. Louis Licamele, Lise Getoor |
BMC Bioinform. | 2 |
| 2010 | Read-Once Functions and Query Evaluation in Probabilistic DatabasesabstractProbabilistic databases hold promise of being a viable means for large-scale uncertainty management, increasingly needed in a number of real world applications domains. However, query evaluation in probabilistic databases remains a computational challenge. Prior work on efficient exact query evaluation in probabilistic databases has largely concentrated on query-centric formulations (e.g., safe plans, hierarchical queries ), in that, they only consider characteristics of the query and not the data in the database. It is easy to construct examples where a supposedly hard query run on an appropriate database gives rise to a tractable query evaluation problem. In this paper, we develop efficient query evaluation techniques that leverage characteristics of both the query and the data in the database. We focus on tuple-independent databases where the query evaluation problem is equivalent to computing marginal probabilities of Boolean formulas associated with the result tuples. This latter task is easy if the Boolean formulas can be factorized into a form that has every variable appearing at most once (called read-once ). However, a naive approach that directly uses previously developed Boolean formula factorization algorithms is inefficient, because those algorithms require the input formulas to be in the disjunctive normal form (DNF). We instead develop novel, more efficient factorization algorithms that directly construct the read-once expression for a result tuple Boolean formula (if one exists), for a large subclass of queries (specifically, conjunctive queries without self-joins). We empirically demonstrate that (1) our proposed techniques are orders of magnitude faster than generic inference algorithms for queries where the result Boolean formulas can be factorized into read-once expressions, and (2) for the special case of hierarchical queries, they rival the efficiency of prior techniques specifically designed to handle such queries. Prithviraj Sen, Amol Deshpande, Lise Getoor |
Proc. VLDB Endow. | 3 |
| 2009 | The Dynamics of Actor Loyalty to Groups in Affiliation NetworksabstractIn this paper, we introduce a method for analyzing the temporal dynamics of affiliation networks. We define affiliation groups which describe temporally related subsets of actors and describe an approach for exploring changing memberships in these affiliation groups over time. To model the dynamic behavior in these networks, we consider the concept of loyalty and introduce a measure that captures an actorpsilas loyalty to an affiliation group as the degree of dasiacommitmentpsila an actor shows to the group over time. We evaluate our measure using two real world affiliation networks: a senate bill co-sponsorship network and a dolphin network. The results show how the behavior of actors in different affiliation groups change dynamically over time, reinforcing the utility of our measure for understanding the loyalty of actors to time-varying affiliation groups. Hossam Sharara, Lisa Singh, Lise Getoor, Janet Mann |
ASONAM | 3 |
| 2009 | Supervised and Unsupervised Methods in Employing Discourse Relations for Improving Opinion Polarity Classification
Swapna Somasundaran, Galileo Namata, Janyce Wiebe, Lise Getoor |
EMNLP | 4 |
| 2009 | Distinguishing Knowledge vs Social Capital in Social Media with Roles and Context
Vladimir Barash, Marc A. Smith, Lise Getoor, Howard T. Welser |
ICWSM | 3 |
| 2009 | Co-evolution of social and affiliation networksabstractIn the last few years, there has been a growing interest in studying online social and affiliation networks, leading to a new category of inference problems that consider the actor characteristics and their social environments. These problems have a variety of applications, from creating more effective marketing campaigns to designing better personalized services. Predictive statistical models allow learning hidden information automatically in these networks but also bring many privacy concerns. Three of the main challenges that I address in my thesis are understanding 1) how the complex observed and unobserved relationships among actors can help in building better behavior models, and in designing more accurate predictive algorithms, 2) what are the processes that drive the network growth and link formation, and 3) what are the implications of predictive algorithms on the privacy of users who share content online. The majority of previous work in prediction, evolution and privacy in online social networks has concentrated on the single-mode networks which form around user-user links, such as friendship and email communication. How- ever, single-mode networks often co-exist with two-mode affiliation networks in which users are linked to other entities, such as social groups, online content and events. I study the interplay between these two types of networks and show that analyzing these higher-order interactions can reveal dependencies that are difficult to extract from the pair-wise interactions alone. In particular, I present my contributions to the challenging problems of collective classification, link prediction, network evolution, and preserving privacy in social and affiliation networks. I evaluate my models on real-world data sets from well-known online social networks, such as Flickr, Facebook, Dogster and LiveJournal. PREDICTION, EVOLUTION AND PRIVACY Elena Zheleva, Hossam Sharara, Lise Getoor |
KDD | 3 |
| 2009 | On Maximum Coverage in the Streaming Model & Application to Multi-topic Blog-WatchabstractWe generalize the graph streaming model to hypergraphs. In this streaming model, hyperedges are arriving online and any computation has to be done on-the-fly using a small amount of space. Each hyperedge can be viewed as a set of elements (nodes), so we refer to our proposed model as the “set-streaming” model of computation. We consider the problem of “maximum coverage”, in which k sets have to be selected that maximize the total weight of the covered elements. In the set-streaming model of computation, we show that our algorithm for maximum-coverage achieves an approximation factor of ¼. When multiple passes are allowed, we also provide a ⊖(log n) approximation algorithm for the set-cover. We next consider a multi-topic blog-watch application, an extension of blog-alert like applications for handling simultaneous multiple-topic requests. We show how the problems of maximum-coverage and set-cover in the set-streaming model can be utilized to give efficient online solutions to this problem. We verify the effectiveness of our methods both on synthetic and real weblog data. Barna Saha, Lise Getoor |
SDM | 2 |
| 2009 | FutureRank: Ranking Scientific Articles by Predicting their Future PageRankabstractThe dynamic nature of citation networks makes the task of ranking scientific articles hard.Citation networks are continually evolving because articles obtain new citations every day.For ranking scientific articles, we can define the popularity or prestige of a paper based on the number of past citations at the user query time; however, we argue that what is most useful is the expected future references.We define a new measure, FutureRank, which is the expected future PageRank score based on citations that will be obtained in the future.In addition to making use of the citation network, FutureRank uses the authorship network and the publication time of the article in order to predict future citations.Our experiments compare FutureRank with existing approaches, and show that FutureRank is accurate and useful for finding and ranking publications. Hassan Sayyadi, Lise Getoor |
SDM | 2 |
| 2009 | Bisimulation-based Approximate Lifted Inference
Prithviraj Sen, Amol Deshpande, Lise Getoor |
UAI | 3 |
| 2009 | To join or not to join: the illusion of privacy in social networks with mixed public and private user profilesabstractIn order to address privacy concerns, many social media websites allow users to hide their personal profiles from the public. In this work, we show how an adversary can exploit an online social network with a mixture of public and private user profiles to predict the private attributes of users. We map this problem to a relational classification problem and we propose practical models that use friendship and group membership information (which is often not hidden) to infer sensitive attributes. The key novel idea is that in addition to friendship links, groups can be carriers of significant information. We show that on several well-known social media sites, we can easily and accurately recover the information of private-profile users. To the best of our knowledge, this is the first work that uses link-based and group-based classification to study privacy implications in social networks with mixed public and private user profiles. Elena Zheleva, Lise Getoor |
WWW | 2 |
| 2009 | Index Interactions in Physical Design Tuning: Modeling, Analysis, and ApplicationsabstractOne of the key tasks of a database administrator is to optimize the set of materialized indices with respect to the current workload. To aid administrators in this challenging task, commercial DBMSs provide advisors that recommend a set of indices based on a sample workload. It is left for the administrator to decide which of the recommended indices to materialize and when. This decision requires some knowledge of how the indices benefit the workload, which may be difficult to understand if there are any dependencies or interactions among indices. Unfortunately, advisors do not provide this crucial information as part of the recommendation. Motivated by this shortcoming, we propose a framework and associated tools that can help an administrator understand the interactions within the recommended set of indices. We formalize the notion of index interactions and develop a novel algorithm to identify the interaction relationships that exist within a set of indices. We present experimental results with a prototype implementation over IBM DB2 that demonstrate the efficiency of our approach. We also describe two new database tuning tools that utilize information about index interactions. The first tool visualizes interactions based on a partitioning of the index-set into non-interacting subsets, and the second tool computes a schedule that materializes the indices over several maintenance windows with maximal overall benefit. In both cases, we provide strong analytical results showing that index interactions can enable enhanced functionality. Karl Schnaitter, Neoklis Polyzotis, Lise Getoor |
Proc. VLDB Endow. | 3 |
| 2009 | Reflect and correct: A misclassification prediction approach to active inferenceabstractInformation diffusion, viral marketing, graph-based semi-supervised learning, and collective classification all attempt to model and exploit the relationships among nodes in a network to improve the performance of node labeling algorithms. However, sometimes the advantage of exploiting the relationships can become a disadvantage. Simple models like label propagation and iterative classification can aggravate a misclassification by propagating mistakes in the network, while more complex models that define and optimize a global objective function, such as Markov random fields and graph mincuts, can misclassify a set of nodes jointly. This problem can be mitigated if the classification system is allowed to ask for the correct labels for a few of the nodes during inference. However, determining the optimal set of labels to acquire is intractable under relatively general assumptions, which forces us to resort to approximate and heuristic techniques. We describe three such techniques in this article. The first one is based on directly approximating the value of the objective function of label acquisition and greedily acquiring the label that provides the most improvement. The second technique is a simple technique based on the analogy we draw between viral marketing and label acquisition. Finally, we propose a method, which we refer to as reflect and correct , that can learn and predict when the classification system is likely to make mistakes and suggests acquisitions to correct those mistakes. We empirically show on a variety of synthetic and real-world datasets that the reflect and correct method significantly outperforms the other two techniques, as well as other approaches based on network structural measures such as node degree and network clustering. Mustafa Bilgic 0001, Lise Getoor |
ACM Trans. Knowl. Discov. Data | 2 |
| 2009 | PrDB: managing and exploiting rich correlations in probabilistic databases
Prithviraj Sen, Amol Deshpande, Lise Getoor |
VLDB J. | 3 |
| 2008 | Effective label acquisition for collective classificationabstractInformation diffusion, viral marketing, and collective classification all attempt to model and exploit the relationships in a network to make inferences about the labels of nodes. A variety of techniques have been introduced and methods that combine attribute information and neighboring label information have been shown to be effective for collective labeling of the nodes in a network. However, in part because of the correlation between node labels that the techniques exploit, it is easy to find cases in which, once a misclassification is made, incorrect information propagates throughout the network. This problem can be mitigated if the system is allowed to judiciously acquire the labels for a small number of nodes. Unfortunately, under relatively general assumptions, determining the optimal set of labels to acquire is intractable. Here we propose an acquisition method that learns the cases when a given collective classification algorithm makes mistakes, and suggests acquisitions to correct those mistakes. We empirically show on both real and synthetic datasets that this method significantly outperforms a greedy approximate inference approach, a viral marketing approach, and approaches based on network structural measures such as node degree and network clustering. In addition to significantly improving accuracy with just a small amount of labeled data, our method is tractable on large networks. Mustafa Bilgic 0001, Lise Getoor |
KDD | 2 |
| 2008 | Learning Structured Bayesian Networks: Combining Abstraction Hierarchies and Tree-Structured Conditional Probability TablesabstractContext‐specific independence representations, such as tree‐structured conditional probability distributions, capture local independence relationships among the random variables in a Bayesian network (BN). Local independence relationships among the random variables can also be captured by using attribute‐value hierarchies to find an appropriate abstraction level for the values used to describe the conditional probability distributions. Capturing this local structure is important because it reduces the number of parameters required to represent the distribution. This can lead to more robust parameter estimation and structure selection, more efficient inference algorithms, and more interpretable models. In this paper, we introduce Tree‐Abstraction‐Based Search (TABS), an approach for learning a data distribution by inducing the graph structure and parameters of a BN from training data. TABS combines tree structure and attribute‐value hierarchies to compactly represent conditional probability tables. To construct the attribute‐value hierarchies, we investigate two data‐driven techniques: a global clustering method, which uses all of the training data to build the attribute‐value hierarchies, and can be performed as a preprocessing step; and a local clustering method, which uses only the local network structure to learn attribute‐value hierarchies. We present empirical results for three real‐world domains, finding that (1) combining tree structure and attribute‐value hierarchies improves the accuracy of generalization, while providing a significant reduction in the number of parameters in the learned networks, and (2) data‐derived hierarchies perform as well or better than expert‐provided hierarchies. Marie desJardins, Priyang Rathod, Lise Getoor |
Comput. Intell. | 3 |
| 2008 | Cost-sensitive learning with conditional Markov networks
Prithviraj Sen, Lise Getoor |
Data Min. Knowl. Discov. | 2 |
| 2008 | Structured machine learning: the next ten years
Thomas G. Dietterich, Pedro M. Domingos, Lise Getoor, Stephen H. Muggleton, Prasad Tadepalli |
Mach. Learn. | 3 |
| 2008 | Exploiting shared correlations in probabilistic databasesabstractThere has been a recent surge in work in probabilistic databases, propelled in large part by the huge increase in noisy data sources --- from sensor data, experimental data, data from uncurated sources, and many others. There is a growing need for database management systems that can efficiently represent and query such data. In this work, we show how data characteristics can be leveraged to make the query evaluation process more efficient. In particular, we exploit what we refer to as shared correlations where the same uncertainties and correlations occur repeatedly in the data. Shared correlations occur mainly due to two reasons: (1) Uncertainty and correlations usually come from general statistics and rarely vary on a tuple-to-tuple basis; (2) The query evaluation procedure itself tends to re-introduce the same correlations. Prior work has shown that the query evaluation problem on probabilistic databases is equivalent to a probabilistic inference problem on an appropriately constructed probabilistic graphical model (PGM). We leverage this by introducing a new data structure, called the random variable elimination graph (rv-elim graph) that can be built from the PGM obtained from query evaluation. We develop techniques based on bisimulation that can be used to compress the rv-elim graph exploiting the presence of shared correlations in the PGM, the compressed rv-elim graph can then be used to run inference. We validate our methods by evaluating them empirically and show that even with a few shared correlations significant speed-ups are possible. Prithviraj Sen, Amol Deshpande, Lise Getoor |
Proc. VLDB Endow. | 3 |
| 2008 | Trusting spam reporters: A reporter-based reputation system for email filteringabstractSpam is a growing problem; it interferes with valid email and burdens both email users and service providers. In this work, we propose a reactive spam-filtering system based on reporter reputation for use in conjunction with existing spam-filtering techniques. The system has a trust-maintenance component for users, based on their spam-reporting behavior. The challenge that we consider is that of maintaining a reliable system, not vulnerable to malicious users, that will provide early spam-campaign detection to reduce the costs incurred by users and systems. We report on the utility of a reputation system for spam filtering that makes use of the feedback of trustworthy users. We evaluate our proposed framework, using actual complaint feedback from a large population of users, and validate its spam-filtering performance on a collection of real email traffic over several weeks. To test the broader implication of the system, we create a model of the behavior of malicious reporters, and we simulate the system under various assumptions using a synthetic dataset. Elena Zheleva, Alek Kolcz, Lise Getoor |
ACM Trans. Inf. Syst. | 3 |
| 2008 | Interactive Entity Resolution in Relational Data: A Visual Analytic Tool and Its EvaluationabstractDatabases often contain uncertain and imprecise references to real-world entities. Entity resolution, the process of reconciling multiple references to underlying real-world entities, is an important data cleaning process required before accurate visualization or analysis of the data is possible. In many cases, in addition to noisy data describing entities, there is data describing the relationships among the entities. This relational data is important during the entity resolution process; it is useful both for the algorithms which determine likely database references to be resolved and for visual analytic tools which support the entity resolution process. In this paper, we introduce a novel user interface, D-Dupe, for interactive entity resolution in relational data. D-Dupe effectively combines relational entity resolution algorithms with a novel network visualization that enables users to make use of an entity's relational context for making resolution decisions. Since resolution decisions often are interdependent, D-Dupe facilitates understanding this complex process through animations which highlight combined inferences and a history mechanism which allows users to inspect chains of resolution decisions. An empirical study with 12 users confirmed the benefits of the relational context visualization on the performance of entity resolution tasks in relational data in terms of time as well as users' confidence and satisfaction. Hyunmo Kang, Lise Getoor, Ben Shneiderman, Mustafa Bilgic 0001, Louis Licamele |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2007 | Online Collective Entity Resolution
Indrajit Bhattacharya, Lise Getoor |
AAAI | 2 |
| 2007 | VOILA: Efficient Feature-value Acquisition for Classification
Mustafa Bilgic 0001, Lise Getoor |
AAAI | 2 |
| 2007 | Relationship Identification for Social Network Discovery
Christopher P. Diehl, Galileo Namata, Lise Getoor |
AAAI | 3 |
| 2007 | A dual-view approach to interactive network visualizationabstractVisualizing network data, from tree structures to arbitrarily connected graphs, is a difficult problem in information visualization. A large part of the problem is that in network data, users not only have to visualize the attributes specific to each data item, but also the links specifying how those items are connected to each other. Past approaches to resolving these difficulties focus on zooming, clustering, filtering and applying various methods of laying out nodes and edges. Such approaches, however, focus only on optimizing a network visualization in a single view, limiting the amount of information that can be shown and explored in parallel. Moreover, past approaches do not allow users to cross reference different subsets or aspects of large, complex networks. In this paper, we propose an approach to these limitations using multiple coordinated views of a given network. To illustrate our approach, we implement a tool called DualNet and evaluate the tool with a case study using an email communication network. We show how using multiple coordinated views improves navigation and provides insight into large networks with multiple node and link properties and types. Galileo Namata, Brian Staats, Lise Getoor, Ben Shneiderman |
CIKM | 3 |
| 2007 | GeoDDupe: A Novel Interface for Interactive Entity Resolution in Geospatial DataabstractDue to the growing interest in geospatial data mining and analysis, data cleaning and integration in geospatial data is becoming an important issue. Geospatial entity resolution is the process of reconciling multiple location references to the same real world location within a single data source (deduplication) or across multiple data sources (integration). In this paper, we introduce an interactive tool called GeoDDupe which effectively combines automatic data mining algorithms for geospatial entity resolution with a novel network visualization supporting users' resolution analysis and decisions. We illustrate the GeoDDupe interface with an example geospatial dataset and show how users can efficiently and accurately resolve location entities. Finally, the case study with two real-world geospatial datasets demonstrates the potential of GeoDDupe. Hyunmo Kang, Vivek Sehgal, Lise Getoor |
IV | 3 |
| 2007 | Visual Mining of Multi-Modal Social Networks at Different Abstraction LevelsabstractSocial networks continue to become more and more feature rich. Using local and global structural properties and descriptive attributes are necessary for more sophisticated social network analysis and support for visual mining tasks. While a number of visualization tools for social network applications have been developed, most of them are limited to uni-modal graph representations. Some of the tools support a wide range of visualization options, including interactive views. Others have better support for calculating structural graph properties such as the density of the graph or deploying traditional statistical social network analysis. We present Invenio, a new tool for visual mining of socials. Invenio integrates a wide range of interactive visualization options from Prefuse, with graph mining algorithm support from JUNG. While the integration expands the breadth of functionality within the core engine of the tool, our goal is to interactively explore multi-modal, multi-relational social networks. Invenio also supports construction of views using both database operations and basic graph mining operations. Lisa Singh, Mitchell Beard, Lise Getoor, M. Brian Blake |
IV | 3 |
| 2007 | Leveraging data and structure in ontology integrationabstractThere is a great deal of research on ontology integration which makes use of rich logical constraints to reason about the structural and logical alignment of ontologies. There is also considerable work on matching data instances from heterogeneous schema or ontologies. However, little work exploits the fact that ontologies include both data and structure. We aim to close this gap by presenting a new algorithm (ILIADS) that tightly integrates both data matching and logical reasoning to achieve better matching of ontologies. We evaluate our algorithm on a set of 30 pairs of OWL Lite ontologies with the schema and data matchings found by human reviewers. We compare against two systems-the ontology matching tool FCA-merge [28] and the schema matching tool COMA++ [1]. ILIADS shows an average improvement of 25 % in quality over FCA-merge and a 11% improvement in recall over COMA++. Octavian Udrea, Lise Getoor, Renée J. Miller |
SIGMOD Conference | 2 |
| 2007 | Features generated for computational splice-site prediction correspond to functional elementsabstractBACKGROUND: Accurate selection of splice sites during the splicing of precursors to messenger RNA requires both relatively well-characterized signals at the splice sites and auxiliary signals in the adjacent exons and introns. We previously described a feature generation algorithm (FGA) that is capable of achieving high classification accuracy on human 3' splice sites. In this paper, we extend the splice-site prediction to 5' splice sites and explore the generated features for biologically meaningful splicing signals. RESULTS: We present examples from the observed features that correspond to known signals, both core signals (including the branch site and pyrimidine tract) and auxiliary signals (including GGG triplets and exon splicing enhancers). We present evidence that features identified by FGA include splicing signals not found by other methods. CONCLUSION: Our generated features capture known biological signals in the expected sequence interval flanking splice sites. The method can be easily applied to other species and to similar classification problems, such as tissue-specific regulatory elements, polyadenylation sites, promoters, etc. Rezarta Islamaj Dogan, Lise Getoor, W. John Wilbur, Stephen M. Mount |
BMC Bioinform. | 2 |
| 2007 | Query-time Entity ResolutionabstractEntity resolution is the problem of reconciling database references corresponding to the same real-world entities. Given the abundance of publicly available databases that have unresolved entities, we motivate the problem of query-time entity resolution quick and accurate resolution for answering queries over such `unclean' databases at query-time. Since collective entity resolution approaches --- where related references are resolved jointly --- have been shown to be more accurate than independent attribute-based resolution for off-line entity resolution, we focus on developing new algorithms for collective resolution for answering entity resolution queries at query-time. For this purpose, we first formally show that, for collective resolution, precision and recall for individual entities follow a geometric progression as neighbors at increasing distances are considered. Unfolding this progression leads naturally to a two stage `expand and resolve' query processing strategy. In this strategy, we first extract the related records for a query using two novel expansion operators, and then resolve the extracted records collectively. We then show how the same strategy can be adapted for query-time entity resolution by identifying and resolving only those database references that are the most helpful for processing the query. We validate our approach on two large real-world publication databases where we show the usefulness of collective resolution and at the same time demonstrate the need for adaptive strategies for query processing. We then show how the same queries can be answered in real-time using our adaptive approach while preserving the gains of collective resolution. In addition to experiments on real datasets, we use synthetically generated data to empirically demonstrate the validity of the performance trends predicted by our analysis of collective entity resolution over a wide range of structural characteristics in the data. Indrajit Bhattacharya, Lise Getoor |
J. Artif. Intell. Res. | 2 |
| 2007 | Collective entity resolution in relational dataabstractMany databases contain uncertain and imprecise references to real-world entities. The absence of identifiers for the underlying entities often results in a database which contains multiple references to the same entity. This can lead not only to data redundancy, but also inaccuracies in query processing and knowledge extraction. These problems can be alleviated through the use of entity resolution . Entity resolution involves discovering the underlying entities and mapping each database reference to these entities. Traditionally, entities are resolved using pairwise similarity over the attributes of references. However, there is often additional relational information in the data. Specifically, references to different entities may cooccur. In these cases, collective entity resolution, in which entities for cooccurring references are determined jointly rather than independently, can improve entity resolution accuracy. We propose a novel relational clustering algorithm that uses both attribute and relational information for determining the underlying domain entities, and we give an efficient implementation. We investigate the impact that different relational similarity measures have on entity resolution quality. We evaluate our collective entity resolution algorithm on multiple real-world databases. We show that it improves entity resolution performance over both attribute-based baselines and over algorithms that consider relational information but do not resolve entities collectively. In addition, we perform detailed experiments on synthetically generated data to identify data characteristics that favor collective relational resolution over purely attribute-based algorithms. Indrajit Bhattacharya, Lise Getoor |
ACM Trans. Knowl. Discov. Data | 2 |
| 2007 | Probabilistic interval XMLabstractInterest in XML databases has been expanding rapidly over the last few years. In this paper, we study the problem of incorporating probabilistic information into XML databases. We propose the Probabilistic Interval XML ( PIXML for short) data model in this paper. Using this data model, users can express probabilistic information within XML markups. In addition, we provide two alternative formal model-theoretic semantics for PIXML data. The first semantics is a “global” semantics which is relatively intuitive, but is not directly amenable to computation. The second semantics is a “local” semantics which supports efficient computation. We prove several correspondence results between the two semantics. To our knowledge, this is the first formal model theoretic semantics for probabilistic interval XML. We then provide an operational semantics that may be used to compute answers to queries and that is correct for a large class of probabilistic instances. Edward Hung, Lise Getoor, V. S. Subrahmanian |
ACM Trans. Comput. Log. | 2 |
| 2006 | Entity resolution in geospatial data integrationabstractDue to the growing availability of geospatial data from a wide variety of sources, there is a pressing need for robust, accurate and automatic merging and matching techniques. Geospatial Entity Resolution is the process of determining, from a collection of database sources referring to geospatial locations, a single consolidated collection of 'true' locations. At the heart of this process is the problem of determining when two locations references match---i.e., when they refer to the same underlying location. In this paper, we introduce a novel method for resolving location entities in geospatial data. A typical geospatial database contains heterogeneous features such as location name, spatial coordinates, location type and demographic information. We investigate the use of all of these features in algorithms for geospatial entity resolution. Entity resolution is further complicated by the fact that the different sources may use different vocabularies for describing the location types and a semantic mapping is required. We propose a novel approach which learns how to combine the different features to perform accurate resolutions. We present experimental results showing that methods combining spatial and non-spatial features (e.g., location-name, location-type, etc.) together outperform methods based on spatial or name information alone. Vivek Sehgal, Lise Getoor, Peter Viechnicki |
GIS | 2 |
| 2006 | Social Capital in Friendship-Event NetworksabstractIn this paper, we examine a particular form of social network which we call a friendship-event network. A friendship-event network captures both the friendship relationship among a set of actors, and also the organizer and participation relationships of actors in a series of events. Within these networks, we formulate the notion of social capital based on the actor-organizer friendship relationship and the notion of benefit, based on event participation. We investigate appropriate definitions for the social capital of both a single actor and a collection of actors. We ground these definitions in a real-world example of academic collaboration networks, where the actors are researchers, the friendships are collaborations, the events are conferences, the organizers are program committee members and the participants are conference authors. We show that our definitions of capital and benefit capture interesting qualitative properties of event series. In addition, we show that social capital is a better publication predictor than publication history. Louis Licamele, Lise Getoor |
ICDM | 2 |
| 2006 | Cost-sensitive learning with conditional Markov networksabstractThere has been a recent, growing interest in classification and link prediction in structured domains. Methods such as CRFs (Lafferty et al., 2001) and RMNs (Taskar et al., 2002) support flexible mechanisms for modeling correlations due to the link structure. In addition, in many structured domains, there is an interesting structure in the risk or cost function associated with different misclassifications. There is a rich tradition of cost-sensitive learning applied to unstructured (IID) data. Here we propose a general framework which can capture correlations in the link structure and handle structured cost functions. We present a novel cost-sensitive structured classifier based on Maximum Entropy principles that directly determines the cost-sensitive classification. We contrast this with an approach which employs a standard 0/1 loss structured classifier followed by minimization of the expected cost of misclassification. We demonstrate the utility of our proposed classifier with experiments on both synthetic and real-world data. Prithviraj Sen, Lise Getoor |
ICML | 2 |
| 2006 | Query-time entity resolutionabstractThe goal of entity resolution is to reconcile database references corresponding to the same real-world entities. Given the abundance of publicly available databases where entities are not resolved, we motivate the problem of quickly processing queries that require resolved entities from such 'unclean' databases. We propose a two-stage collective resolution strategy for processing queries. We then show how it can be performed on-the-fly by adaptively extracting and resolving those database references that are the most helpful for resolving the query. We validate our approach on two large real-world publication databases where we show the usefulness of collective resolution and at the same time demonstrate the need for adaptive strategies for query processing. We then show how the same queries can be answered in real time using our adaptive approach while preserving the gains of collective resolution. Indrajit Bhattacharya, Lise Getoor, Louis Licamele |
KDD | 2 |
| 2006 | Is there a grand challenge or X-prize for data mining?abstractInternational audience Gregory Piatetsky-Shapiro, Robert Grossman, Chaabane Djeraba, Ronen Feldman, Lise Getoor, Mohammed J. Zaki |
KDD | 5 |
| 2006 | A Feature Generation Algorithm for Sequences with Application to Splice-Site Prediction
Rezarta Islamaj Dogan, Lise Getoor, W. John Wilbur |
PKDD | 2 |
| 2006 | A Latent Dirichlet Model for Unsupervised Entity ResolutionabstractEntity resolution has received considerable attention in recent years. Given many references to underlying entities, the goal is to predict which references correspond to the same entity. We show how to extend the Latent Dirichlet Allocation model for this task and propose a probabilistic model for collective entity resolution for relational domains where references are connected to each other. Our approach differs from other recently proposed entity resolution approaches in that it is a) generative, b) does not make pair-wise decisions and c) captures relations between entities through a hidden group variable. We propose a novel sampling algorithm for collective entity resolution which is unsupervised and also takes entity relations into account. Additionally, we do not assume the domain of entities to be known and show how to infer the number of entities from the data. We demonstrate the utility and practicality of our relational entity resolution approach for author resolution in two real-world bibliographic datasets. In addition, we present preliminary results on characterizing conditions under which relational information is useful. Indrajit Bhattacharya, Lise Getoor |
SDM | 2 |
| 2006 | Name Reference Resolution in Organizational Email ArchivesabstractOnline communications provide a rich resource for understanding social networks. Information about the actors, and their dynamic roles and relationships, can be inferred from both the communication content and traffic structure. A key component in the analysis of online communications such as email is the resolution of name references within the body of the message. Name reference resolution relies on the context of the message; both the content of the message and the sender and recipients' relationships can help to resolve a reference. Here we investigate a variety of approaches which make use of the email traffic network to disambiguate email name references. The email traffic network serves as a proxy for inferring relationships. These relationships in turn help us infer likely candidates for the name references. Our initial findings suggest that simple temporal models can help us effectively resolve name references. For the class of models proposed, performance is maximized by exploiting long-term traffic statistics to rank candidates. Christopher P. Diehl, Lise Getoor, Galileo Namata |
SDM | 2 |
| 2006 | PRL: A probabilistic relational language
Lise Getoor, John Grant |
Mach. Learn. | 1 |
| 2005 | Bayesian Network Learning with Abstraction Hierarchies and Context-Specific Independence
Marie desJardins, Priyang Rathod, Lise Getoor |
ECML | 3 |
| 2005 | D-Dupe: An Interactive Tool for Entity Resolution in Social Networks
Mustafa Bilgic 0001, Louis Licamele, Lise Getoor, Ben Shneiderman |
GD | 3 |
| 2005 | Pruning Social Networks Using Structural Properties and Descriptive AttributesabstractScale is often an issue with understanding and making sense of large social networks. Here we investigate methods for pruning social networks by determining the most relevant relationships. We measure importance in terms of predictive accuracy on a set of target attributes of the social network. Our goal is to create a pruned network that models only the most informative affiliations and relationships. We present methods for pruning networks based on both structural properties and descriptive attributes demonstrate it on a network of NASDAQ and NYSE businesses and on a bibliographic network. Lisa Singh, Lise Getoor, Louis Licamele |
ICDM | 2 |
| 2005 | Tutorial on Statistical Relational Learning
Lise Getoor |
ILP | 1 |
| 2004 | Unsupervised Sense Disambiguation Using Bilingual Probabilistic ModelsabstractWe describe two probabilistic models for unsupervised word-sense disambiguation using parallel corpora. The first model, which we call the Sense model, builds on the work of Diab and Resnik (2002) that uses both parallel text and a sense inventory for the target language, and recasts their approach in a probabilistic framework. The second model, which we call the Concept model, is a hierarchical model that uses a concept latent variable to relate different language specific sense labels. We show that both models improve performance on the word sense disambiguation task over previous unsupervised approaches, with the Concept model showing the largest improvement. Furthermore, in learning the Concept model, as a by-product, we learn a sense inventory for the parallel language. Indrajit Bhattacharya, Lise Getoor, Yoshua Bengio |
ACL | 2 |
| 2004 | Using the Structure of Web Sites for Automatic Segmentation of TablesabstractMany Web sites, especially those that dynamically generate HTML pages to display the results of a user's query, present information in the form of list or tables. Current tools that allow applications to programmatically extract this information rely heavily on user input, often in the form of labeled extracted records. The sheer size and rate of growth of the Web make any solution that relies primarily on user input is infeasible in the long term. Fortunately, many Web sites contain much explicit and implicit structure, both in layout and content, that we can exploit for the purpose of information extraction. This paper describes an approach to automatic extraction and segmentation of records from Web tables. Automatic methods do not require any user input, but rely solely on the layout and content of the Web source. Our approach relies on the common structure of many Web sites, which present information as a list or a table, with a link in each entry leading to a detail page containing additional information about that item. We describe two algorithms that use redundancies in the content of table and detail pages to aid in information extraction. The first algorithm encodes additional information provided by detail pages as constraints and finds the segmentation by solving a constraint satisfaction problem. The second algorithm uses probabilistic inference to find the record segmentation. We show how each approach can exploit the web site structure in a general, domain-independent manner, and we demonstrate the effectiveness of each algorithm on a set of twelve Web sites. Kristina Lerman, Lise Getoor, Steven Minton, Craig A. Knoblock |
SIGMOD Conference | 2 |
| 2004 | Understanding tuberculosis epidemiology using structured statistical models
Lise Getoor, Jeanne T. Rhee, Daphne Koller, Peter Small |
Artif. Intell. Medicine | 1 |
| 2003 | PXML: A Probabilistic Semistructured Data Model and AlgebraabstractDespite the recent proliferation of work on semistructured data models, there has been little work to date on supporting uncertainty in these models. We propose a model for probabilistic semistructured data (PSD). The advantage of our approach is that it supports a flexible representation that allows the specification of a wide class of distributions over semistructured instances. We provide two semantics for the model and show that the semantics are probabilistically coherent. Next, we develop an extension of the relational algebra to handle probabilistic semistructured data and describe efficient algorithms for answering queries that use this algebra. Finally, we present experimental results showing the efficiency of our algorithms. Edward Hung, Lise Getoor, V. S. Subrahmanian |
ICDE | 2 |
| 2003 | Probabilistic Interval XML
Edward Hung, Lise Getoor, V. S. Subrahmanian |
ICDT | 2 |
| 2003 | Link-based Classification
Lise Getoor |
ICML | 2 |
| 2002 | Learning Probabilistic Models of Link Structure
Lise Getoor, Nir Friedman, Daphne Koller, Ben Taskar |
J. Mach. Learn. Res. | 1 |
| 2001 | Learning Probabilistic Models of Relational Structure
Lise Getoor, Nir Friedman, Daphne Koller, Ben Taskar |
ICML | 1 |
| 2001 | Selectivity Estimation using Probabilistic ModelsabstractEstimating the result size of complex queries that involve selection on multiple attributes and the join of several relations is a difficult but fundamental task in database query processing. It arises in cost-based query optimization, query profiling, and approximate query answering. In this paper, we show how probabilistic graphical models can be effectively used for this task as an accurate and compact approximation of the joint frequency distribution of multiple attributes across multiple relations. Probabilistic Relational Models (PRMs) are a recent development that extends graphical statistical models such as Bayesian Networks to relational domains. They represent the statistical dependencies between attributes within a table, and between attributes across foreign-key joins. We provide an efficient algorithm for constructing a PRM front a database, and show how a PRM can be used to compute selectivity estimates for a broad class of queries. One of the major contributions of this work is a unified framework for the estimation of queries involving both select and foreign-key join operations. Furthermore, our approach is not limited to answering a small set of predetermined queries; a single model can be used to effectively estimate the sizes of a wide collection of potential queries across multiple tables. We present results for our approach on several real-world databases. For both single-table multi-attribute queries and a general class of select-join queries, our approach produces more accurate estimates than standard approaches to selectivity estimation, using comparable space and time. Lise Getoor, Ben Taskar, Daphne Koller |
SIGMOD Conference | 1 |
| 1999 | Learning Probabilistic Relational Models
Nir Friedman, Lise Getoor, Daphne Koller, Avi Pfeffer |
IJCAI | 2 |
| 1998 | Utility Elicitation as a Classification Problem
Urszula Chajewska, Lise Getoor, Joseph Norman, Yuval Shahar |
UAI | 2 |
| 1995 | Scope and Abstraction: Two Criteria for Localized Planning
Amy L. Lansky, Lise Getoor |
IJCAI | 2 |