EDBT 2026 Demo / reviewers in the wild / expert
Prithviraj Sen
dblp:58/1388
· DBLP profile ↗
31ranked-venue papers
11as first author
7since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 20 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 16 · 5 first-author · 6 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
14 papers |
Knowledge representation and reasoning · 57% Information extraction and text analysis · 17% Reinforcement learning · 13% | |
| Databases, data mining, and information retrieval
12 papers |
Data integration and cleaning · 55% Machine learning and data management · 23% Query processing and optimization · 10% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 70% High-performance computing · 30% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% |
Topics — the 30 heaviest of 52, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge incorporation › knowledge-infused learning › neuro-symbolic learning
logical neural networks |
1.1 | 2 | 2022 | Logical Neural Networks for Knowledge Base Completion with Embeddings & Rules · EMNLP 2022 Neuro-Symbolic Inductive Logic Programming with Logical Neural Networks · AAAI 2022 |
Data integration and cleaning › entity resolution
active learning for entity matching |
0.9 | 2 | 2021 | Deep Indexed Active Learning for Matching Heterogeneous Entity Representations · Proc. VLDB Endow. 2021 A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching · SIGMOD Conference 2020 |
Data integration and cleaning
entity resolution |
0.9 | 2 | 2021 | Deep Indexed Active Learning for Matching Heterogeneous Entity Representations · Proc. VLDB Endow. 2021 SystemER: A Human-in-the-loop System for Explainable Entity Resolution · Proc. VLDB Endow. 2019 |
Machine learning and data management › machine learning systems
declarative machine learning |
0.8 | 3 | 2018 | On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML · Proc. VLDB Endow. 2018 SystemML: Declarative Machine Learning on Spark · Proc. VLDB Endow. 2016 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML · Proc. VLDB Endow. 2014 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
abstract meaning representation |
0.7 | 1 | 2023 | Learning Symbolic Rules over Abstract Meaning Representations for Textual Reinforcement Learning · ACL (1) 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.7 | 1 | 2023 | Are Human Explanations Always Helpful? Towards Objective Evaluation of Human Natural Language Explanations · ACL (1) 2023 |
Machine learning › Reinforcement learning
textual reinforcement learning |
0.7 | 1 | 2023 | Learning Symbolic Rules over Abstract Meaning Representations for Textual Reinforcement Learning · ACL (1) 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic programming
inductive logic programming |
0.6 | 1 | 2022 | Neuro-Symbolic Inductive Logic Programming with Logical Neural Networks · AAAI 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph reasoning
knowledge base completion |
0.6 | 1 | 2022 | Logical Neural Networks for Knowledge Base Completion with Embeddings & Rules · EMNLP 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
logic-based reasoning |
0.6 | 1 | 2022 | Neuro-Symbolic Inductive Logic Programming with Logical Neural Networks · AAAI 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › rule learning
differentiable rule learning |
0.5 | 1 | 2021 | Neuro-Symbolic Approaches for Text-Based Policy Learning · EMNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis
entity linking |
0.5 | 1 | 2021 | LNN-EL: A Neuro-Symbolic Approach to Short-text Entity Linking · ACL/IJCNLP (1) 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
neuro-symbolic reasoning |
0.5 | 1 | 2021 | Neuro-Symbolic Approaches for Text-Based Policy Learning · EMNLP (1) 2021 |
Data integration and cleaning › entity resolution
blocking |
0.5 | 1 | 2021 | Deep Indexed Active Learning for Matching Heterogeneous Entity Representations · Proc. VLDB Endow. 2021 |
Machine learning and data management › scalable machine learning
distributed learning |
0.4 | 2 | 2016 | SystemML: Declarative Machine Learning on Spark · Proc. VLDB Endow. 2016 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML · Proc. VLDB Endow. 2014 |
Natural language and speech › Information extraction and text analysis › text classification
sentence classification |
0.4 | 1 | 2020 | Learning Explainable Linguistic Expressions with Neural Inductive Logic Programming for Sentence Classification · EMNLP (1) 2020 |
Data integration and cleaning
entity matching |
0.4 | 1 | 2020 | A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching · SIGMOD Conference 2020 |
Data integration and cleaning › entity resolution
explainable entity matching |
0.4 | 1 | 2019 | SystemER: A Human-in-the-loop System for Explainable Entity Resolution · Proc. VLDB Endow. 2019 |
Query processing and optimization › query compilation
operator fusion |
0.3 | 1 | 2018 | On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML · Proc. VLDB Endow. 2018 |
Database theory
probabilistic databases |
0.3 | 4 | 2010 | Read-Once Functions and Query Evaluation in Probabilistic Databases · Proc. VLDB Endow. 2010 PrDB: managing and exploiting rich correlations in probabilistic databases · VLDB J. 2009 Representing and Querying Correlated Tuples in Probabilistic Databases · ICDE 2007 |
Knowledge graphs
knowledge graph construction |
0.3 | 1 | 2017 | Creation and Interaction with Large-scale Domain-Specific Knowledge Bases · Proc. VLDB Endow. 2017 |
Compilers and program optimization › deep learning compiler
operator fusion |
0.2 | 1 | 2016 | SystemML: Declarative Machine Learning on Spark · Proc. VLDB Endow. 2016 |
Query processing and optimization
probabilistic query processing |
0.2 | 2 | 2010 | Read-Once Functions and Query Evaluation in Probabilistic Databases · Proc. VLDB Endow. 2010 Exploiting shared correlations in probabilistic databases · Proc. VLDB Endow. 2008 |
Parallel and multicore computing
parallel programming models |
0.2 | 1 | 2014 | Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML · Proc. VLDB Endow. 2014 |
Parallel and multicore computing › parallel programming models
task and data parallelism |
0.2 | 1 | 2014 | Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML · Proc. VLDB Endow. 2014 |
High-performance computing
cluster computing |
0.2 | 2 | 2018 | On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML · Proc. VLDB Endow. 2018 SystemML: Declarative Machine Learning on Spark · Proc. VLDB Endow. 2016 |
Parallel and multicore computing › data-parallel programming
data-parallel frameworks |
0.2 | 2 | 2018 | On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML · Proc. VLDB Endow. 2018 SystemML: Declarative Machine Learning on Spark · Proc. VLDB Endow. 2016 |
Natural language and speech › Information extraction and text analysis › entity linking › entity disambiguation
collective entity disambiguation |
0.1 | 1 | 2012 | Collective context-aware topic models for entity disambiguation · WWW 2012 |
Natural language and speech › Information extraction and text analysis › entity linking
entity disambiguation |
0.1 | 1 | 2012 | Collective context-aware topic models for entity disambiguation · WWW 2012 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge base construction |
0.1 | 1 | 2012 | Collective context-aware topic models for entity disambiguation · WWW 2012 |
Methods — techniques the papers use, named apart from their topics
cost-based optimization · 2.1active learning · 1.3neural network · 1.0code generation · 1.0linear algebra · 0.8symbolic rule learning · 0.7simulatability · 0.7fine-tuning · 0.7logical rules · 0.6gradient-based optimization · 0.6embedding · 0.6pre-trained transformer language models · 0.5neuro-symbolic reasoning · 0.5logic neural network · 0.5index-by-committee · 0.5end-to-end differentiable rule learning · 0.5supervised learning · 0.4ParFOR · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Learning Symbolic Rules over Abstract Meaning Representations for Textual Reinforcement LearningabstractSubhajit Chaudhury, Sarathkrishna Swaminathan, Daiki Kimura, Prithviraj Sen, Keerthiram Murugesan, Rosario Uceda-Sosa, Michiaki Tatsubori, Achille Fokoue, Pavan Kapanipathi, Asim Munawar, Alexander Gray. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Subhajit Chaudhury, Sarathkrishna Swaminathan, Daiki Kimura, Prithviraj Sen, Keerthiram Murugesan, Rosario Uceda-Sosa, Michiaki Tatsubori, Achille Fokoue, Pavan Kapanipathi, Asim Munawar, Alexander G. Gray |
ACL (1) | 4 |
| 2023 | Are Human Explanations Always Helpful? Towards Objective Evaluation of Human Natural Language ExplanationsabstractHuman-annotated labels and explanations are critical for training explainable NLP models.However, unlike human-annotated labels whose quality is easier to calibrate (e.g., with a majority vote), human-crafted free-form explanations can be quite subjective.Before blindly using them as ground truth to train ML models, a vital question needs to be asked: How do we evaluate a human-annotated explanation's quality?In this paper, we build on the view that the quality of a human-annotated explanation can be measured based on its helpfulness (or impairment) to the ML models' performance for the desired NLP tasks for which the annotations were collected.In comparison to the commonly used Simulatability score, we define a new metric that can take into consideration of the helpfulness of an explanation for model performance at both fine-tuning and inference.With the help of a unified dataset format, we evaluated the proposed metric on five datasets (e.g., e-SNLI) against two model architectures (T5 and BART), and the results show that our proposed metric can objectively evaluate the quality of human-annotated explanations, while Simulatability falls short. Bingsheng Yao, Prithviraj Sen, Lucian Popa 0001, James A. Hendler, Dakuo Wang |
ACL (1) | 2 |
| 2022 | Neuro-Symbolic Inductive Logic Programming with Logical Neural NetworksabstractRecent work on neuro-symbolic inductive logic programming has led to promising approaches that can learn explanatory rules from noisy, real-world data. While some proposals approximate logical operators with differentiable operators from fuzzy or real-valued logic that are parameter-free thus diminishing their capacity to fit the data, other approaches are only loosely based on logic making it difficult to interpret the learned ``rules". In this paper, we propose learning rules with the recently proposed logical neural networks (LNN). Compared to others, LNNs offer a strong connection to classical Boolean logic thus allowing for precise interpretation of learned rules while harboring parameters that can be trained with gradient-based optimization to effectively fit the data. We extend LNNs to induce rules in first-order logic. Our experiments on standard benchmarking tasks confirm that LNN rules are highly interpretable and can achieve comparable or higher accuracy due to their flexible parameterization. Prithviraj Sen, Breno W. Carvalho, Ryan Riegel, Alexander G. Gray |
AAAI | 1 |
| 2022 | Logical Neural Networks for Knowledge Base Completion with Embeddings & RulesabstractPrithviraj Sen, Breno William Carvalho, Ibrahim Abdelaziz, Pavan Kapanipathi, Salim Roukos, Alexander Gray. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Prithviraj Sen, Breno W. Carvalho, Ibrahim Abdelaziz, Pavan Kapanipathi, Salim Roukos, Alexander G. Gray |
EMNLP | 1 |
| 2021 | LNN-EL: A Neuro-Symbolic Approach to Short-text Entity LinkingabstractHang Jiang, Sairam Gurajada, Qiuhao Lu, Sumit Neelam, Lucian Popa, Prithviraj Sen, Yunyao Li, Alexander Gray. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Sairam Gurajada, Qiuhao Lu, Sumit Neelam, Lucian Popa 0001, Prithviraj Sen, Yunyao Li 0001, Alexander G. Gray |
ACL/IJCNLP (1) | 6 |
| 2021 | Neuro-Symbolic Approaches for Text-Based Policy LearningabstractText-Based Games (TBGs) have emerged as important testbeds for reinforcement learning (RL) in the natural language domain.Previous methods using LSTM-based action policies are uninterpretable and often overfit the training games showing poor performance to unseen test games.We present SymboLic Action policy for Textual Environments (SLATE), that learns interpretable action policy rules from symbolic abstractions of textual observations for improved generalization.We outline a method for end-to-end differentiable symbolic rule learning and show that such symbolic policies outperform previous stateof-the-art methods in text-based RL for the coin collector environment from 5 -10x fewer training games.Additionally, our method provides human-understandable policy rules that can be readily verified for their logical consistency and can be easily debugged.1 Subhajit Chaudhury, Prithviraj Sen, Masaki Ono, Daiki Kimura, Michiaki Tatsubori, Asim Munawar |
EMNLP (1) | 2 |
| 2021 | Deep Indexed Active Learning for Matching Heterogeneous Entity RepresentationsabstractGiven two large lists of records, the task in entity resolution (ER) is to find the pairs from the Cartesian product of the lists that correspond to the same real world entity. Typically, passive learning methods on such tasks require large amounts of labeled data to yield useful models. Active Learning is a promising approach for ER in low resource settings. However, the search space, to find informative samples for the user to label, grows quadratically for instance-pair tasks making active learning hard to scale. Previous works, in this setting, rely on hand-crafted predicates, pre-trained language model embeddings, or rule learning to prune away unlikely pairs from the Cartesian product. This blocking step can miss out on important regions in the product space leading to low recall. We propose DIAL, a scalable active learning approach that jointly learns embeddings to maximize recall for blocking and accuracy for matching blocked pairs. DIAL uses an Index-By-Committee framework, where each committee member learns representations based on powerful pre-trained transformer language models. We highlight surprising differences between the matcher and the blocker in the creation of the training data and the objective used to train their parameters. Experiments on five benchmark datasets and a multilingual record matching dataset show the effectiveness of our approach in terms of precision, recall and running time. Arjit Jain, Sunita Sarawagi, Prithviraj Sen |
Proc. VLDB Endow. | 3 |
| 2020 | Exploiting Node Content for Multiview Graph Convolutional Network and Adversarial RegularizationabstractNetwork representation learning (NRL) is crucial in the area of graph learning.Recently, graph autoencoders and its variants have gained much attention and popularity among various types of node embedding approaches.Most existing graph autoencoder-based methods aim to minimize the reconstruction errors of the input network while not explicitly considering the semantic relatedness between nodes.In this paper, we propose a novel network embedding method which models the consistency across different views of networks.More specifically, we create a second view from the input network which captures the relation between nodes based on node content and enforce the latent representations from the two views to be consistent by incorporating a multiview adversarial regularization module.The experimental studies on benchmark datasets prove the effectiveness of this method, and demonstrate that our method compares favorably with the state-of-the-art algorithms on challenging tasks such as link prediction and node clustering.We also evaluate our method on a real-world application, i.e., 30-day unplanned ICU readmission prediction, and achieve promising results compared with several baseline methods. Qiuhao Lu, Nisansa de Silva, Dejing Dou, Thien Huu Nguyen, Prithviraj Sen, Berthold Reinwald, Yunyao Li 0001 |
COLING | 5 |
| 2020 | Learning Explainable Linguistic Expressions with Neural Inductive Logic Programming for Sentence ClassificationabstractPrithviraj Sen, Marina Danilevsky, Yunyao Li, Siddhartha Brahma, Matthias Boehm, Laura Chiticariu, Rajasekar Krishnamurthy. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Prithviraj Sen, Marina Danilevsky, Yunyao Li 0001, Siddhartha Brahma, Matthias Boehm 0001, Laura Chiticariu, Rajasekar Krishnamurthy |
EMNLP (1) | 1 |
| 2020 | A Comprehensive Benchmark Framework for Active Learning Methods in Entity MatchingabstractEntity Matching (EM) is a core data cleaning task, aiming to identify different mentions of the same real-world entity. Active learning is one way to address the challenge of scarce labeled data in practice, by dynamically collecting the necessary examples to be labeled by an Oracle and refining the learned model (classifier) upon them. In this paper, we build a unified active learning benchmark framework for EM that allows users to easily combine different learning algorithms with applicable example selection algorithms. The goal of the framework is to enable concrete guidelines for practitioners as to what active learning combinations will work well for EM. Towards this, we perform comprehensive experiments on publicly available EM datasets from product and publication domains to evaluate active learning methods, using a variety of metrics including EM quality, #labels and example selection latencies. Our most surprising result finds that active learning with fewer labels can learn a classifier of comparable quality as supervised learning. In fact, for several of the datasets, we show that there is an active learning combination that beats the state-of-the-art supervised learning result. Our framework also includes novel optimizations that improve the quality of the learned model by roughly 9% in terms of F1-score and reduce example selection latencies by up to 10× without affecting the quality of the model. Venkata Vamsikrishna Meduri, Lucian Popa 0001, Prithviraj Sen, Mohamed Sarwat |
SIGMOD Conference | 3 |
| 2019 | Learning-Based Methods with Human-in-the-Loop for Entity ResolutionabstractThis tutorial is intended for researchers and practitioners working in the data integration area and, in particular, entity resolution (ER), which is a sub-area focused on linking entities across heterogeneous datasets. We outline the ideal requirements of modern ER systems: (1) capture domain knowledge via (minimal) human interaction, (2) provide as much automation as possible via machine learning techniques, and (3) achieve high explainability. We describe recent research trends towards bringing such ideal ER systems closer to reality. We begin with an overview of human-in-the-loop methods that are based on techniques such as crowdsourcing and active learning. We then dive into recent trends that involve deep learning techniques such as representation learning to automate feature engineering, and combinations of transfer and active learning to reduce the amount of user labels required. We also discuss how explainable AI relates to ER, and outline some of the recent advances towards explainable ER. Sairam Gurajada, Lucian Popa 0001, Kun Qian 0002, Prithviraj Sen |
CIKM | 4 |
| 2019 | SystemER: A Human-in-the-loop System for Explainable Entity ResolutionabstractEntity Resolution (ER) is the task of identifying different representations of the same real-world object. To achieve scalability and the desired level of quality, the typical ER pipeline includes multiple steps that may involve low-level coding and extensive human labor. We present SystemER, a tool for learning explainable ER models that reduces the human labor all throughout the stages of the ER pipeline. SystemER achieves explainability by learning rules that not only perform a given ER task but are human-comprehensible; this provides transparency into the learning process, and further enables verification and customization of the learned model by the domain experts. By leveraging a human in the loop and active learning, SystemER also ensures that a small number of labeled examples is sufficient to learn high-quality ER models. SystemER is a full-fledged tool that includes an easy to use interface, support for both flat files and semi-structured data, and scale-out capabilities by distributing computation via Apache Spark. Kun Qian 0002, Lucian Popa 0001, Prithviraj Sen |
Proc. VLDB Endow. | 3 |
| 2018 | On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemMLabstractMany machine learning (ML) systems allow the specification of ML algorithms by means of linear algebra programs, and automatically generate efficient execution plans. The opportunities for fused operators---in terms of fused chains of basic operators---are ubiquitous, and include fewer materialized intermediates, fewer scans of inputs, and sparsity exploitation across operators. However, existing fusion heuristics struggle to find good plans for complex operator DAGs or hybrid plans of local and distributed operations. In this paper, we introduce an exact yet practical cost-based optimization framework for fusion plans and describe its end-to-end integration into Apache SystemML. We present techniques for candidate exploration and selection of fusion plans, as well as code generation of local and distributed operations over dense, sparse, and compressed data. Our experiments in SystemML show end-to-end performance improvements of up to 22x, with negligible compilation overhead. Matthias Boehm 0001, Berthold Reinwald, Dylan Hutchison, Prithviraj Sen, Alexandre V. Evfimievski, Niketan Pansare |
Proc. VLDB Endow. | 4 |
| 2017 | SPOOF: Sum-Product Optimization and Operator Fusion for Large-Scale Machine Learning
Tarek Elgamal, Shangyu Luo, Matthias Boehm 0001, Alexandre V. Evfimievski, Shirish Tatikonda, Berthold Reinwald, Prithviraj Sen |
CIDR | 7 |
| 2017 | Active Learning for Large-Scale Entity ResolutionabstractEntity resolution (ER) is the task of identifying different representations of the same real-world object across datasets. Designing and tuning ER algorithms is an error-prone, labor-intensive process, which can significantly benefit from data-driven, automated learning methods. Our focus is on "big data'' scenarios where the primary challenges include 1) identifying, out of a potentially massive set, a small subset of informative examples to be labeled by the user, 2) using the labeled examples to efficiently learn ER algorithms that achieve both high precision and high recall, and 3) executing the learned algorithm to determine duplicates at scale. Recent work on learning ER algorithms has employed active learning to partially address the above challenges by aiming to learn ER rules in the form of conjunctions of matching predicates, under precision guarantees. While successful in learning a single rule, prior work has been less successful in learning multiple rules that are sufficiently different from each other, thus missing opportunities for improving recall. In this paper, we introduce an active learning system that learns, at scale, multiple rules each having significant coverage of the space of duplicates, thus leading to high recall, in addition to high-precision. We show the superiority of our system on real-world ER scenarios of sizes up to tens of millions of records, over state-of-the-art active learning methods that learn either rules or committees of statistical classifiers for ER, and even over sophisticated methods based on first-order probabilistic models. Kun Qian 0002, Lucian Popa 0001, Prithviraj Sen |
CIKM | 3 |
| 2017 | A Rectangle Mining Method for Understanding the Semantics of Financial TablesabstractFinancial statements report crucial information in tables with complex semantic structure, which are desirable, yet challenging, to interpret automatically. For example, in such tables a row of data cells is often explained by the headers of other rows. In a departure from prior art, we propose a rectangle mining framework for understanding complex tables, which considers rectangular regions rather than individual cells or pairs of cells in a table. We instantiate this framework with ReMine, an algorithm for extracting row header semantics of table, and show that it significantly outperforms prior pair-wise classification approaches on two datasets: (i) a set of manually labeled financial tables from multiple companies, and (ii) the ICDAR 2013 Table Competition dataset. Xilun Chen 0002, Laura Chiticariu, Marina Danilevsky, Alexandre V. Evfimievski, Prithviraj Sen |
ICDAR | 5 |
| 2017 | Creation and Interaction with Large-scale Domain-Specific Knowledge BasesabstractThe ability to create and interact with large-scale domain-specific knowledge bases from unstructured/semi-structured data is the foundation for many industry-focused cognitive systems. We will demonstrate the Content Services system that provides cloud services for creating and querying high-quality domain-specific knowledge bases by analyzing and integrating multiple (un/semi)structured content sources. We will showcase an instantiation of the system for a financial domain. We will also demonstrate both cross-lingual natural language queries and programmatic API calls for interacting with this knowledge base. Shreyas Bharadwaj, Laura Chiticariu, Marina Danilevsky, Samarth Dhingra, Samved Divekar, Arnaldo Carreno-Fuentes, Nitin Gupta 0005, Sang-Don Han, Mauricio A. Hernández, C. T. Howard Ho, Parag Jain, Salil Joshi 0001, Hima P. Karanam, Saravanan Krishnan, Rajasekar Krishnamurthy, Yunyao Li 0001, Satishkumaar Manivannan, Ashish R. Mittal, Fatma Özcan 0001, Abdul Quamar, Poornima Chozhiyath Raman, Diptikalyan Saha, Karthik Sankaranarayanan, Jaydeep Sen, Prithviraj Sen, Shivakumar Vaithyanathan, Mitesh Vasa, Huaiyu Zhu 0001 |
Proc. VLDB Endow. | 26 |
| 2016 | SystemML: Declarative Machine Learning on SparkabstractThe rising need for custom machine learning (ML) algorithms and the growing data sizes that require the exploitation of distributed, data-parallel frameworks such as MapReduce or Spark, pose significant productivity challenges to data scientists. Apache SystemML addresses these challenges through declarative ML by (1) increasing the productivity of data scientists as they are able to express custom algorithms in a familiar domain-specific language covering linear algebra primitives and statistical functions, and (2) transparently running these ML algorithms on distributed, data-parallel frameworks by applying cost-based compilation techniques to generate efficient, low-level execution plans with in-memory single-node and large-scale distributed operations. This paper describes SystemML on Apache Spark, end to end, including insights into various optimizer and runtime techniques as well as performance characteristics. We also share lessons learned from porting SystemML to Spark and declarative ML in general. Finally, SystemML is open-source, which allows the database community to leverage it as a testbed for further research. Matthias Boehm 0001, Michael Dusenberry, Deron Eriksson, Alexandre V. Evfimievski, Faraz Makari Manshadi, Niketan Pansare, Berthold Reinwald, Frederick Reiss 0001, Prithviraj Sen, Arvind Surve, Shirish Tatikonda |
Proc. VLDB Endow. | 9 |
| 2014 | Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemMLabstractSystemML aims at declarative, large-scale machine learning (ML) on top of MapReduce, where high-level ML scripts with R-like syntax are compiled to programs of MR jobs. The declarative specification of ML algorithms enables---in contrast to existing large-scale machine learning libraries---automatic optimization. SystemML's primary focus is on data parallelism but many ML algorithms inherently exhibit opportunities for task parallelism as well. A major challenge is how to efficiently combine both types of parallelism for arbitrary ML scripts and workloads. In this paper, we present a systematic approach for combining task and data parallelism for large-scale machine learning on top of MapReduce. We employ a generic Parallel FOR construct (ParFOR) as known from high performance computing (HPC). Our core contributions are (1) complementary parallelization strategies for exploiting multi-core and cluster parallelism, as well as (2) a novel cost-based optimization framework for automatically creating optimal parallel execution plans. Experiments on a variety of use cases showed that this achieves both efficiency and scalability due to automatic adaptation to ad-hoc workloads and unknown data characteristics. Matthias Boehm 0001, Shirish Tatikonda, Berthold Reinwald, Prithviraj Sen, Yuanyuan Tian 0001, Douglas Burdick, Shivakumar Vaithyanathan |
Proc. VLDB Endow. | 4 |
| 2013 | Community detection in content-sharing social networksabstractNetwork structure and content in microblogging sites like Twitter influence each other ---user A on Twitter follows user B for the tweets that B posts on the network, and A may then re-tweet the content shared by B to his/her own followers. In this paper, we propose a probabilistic model to jointly model link communities and content topics by leveraging both the social graph and the content shared by users. We model a community as a distribution over users, use it as a source for topics of interest, and jointly infer both communities and topics using Gibbs sampling. While modeling communities using the social graph, or modeling topics using content have received a great deal of attention, a few recent approaches try to model topics in content-sharing platforms using both content and social graph. Our work differs from the existing generative models in that we explicitly model the social graph of users along with the user-generated content, mimicking how the two entities co-evolve in content-sharing platforms. Recent studies have found Twitter to be more of a content-sharing network and less a social network, and it seems hard to detect tightly knit communities from the follower-followee links. Still, the question of whether we can extract Twitter communities using both links and content is open. In this paper, we answer this question in the affirmative. Our model discovers coherent communities and topics, as evinced by qualitative results on sub-graphs of Twitter users. Furthermore, we evaluate our model on the task of predicting follower-followee links. We show that joint modeling of links and content significantly improves link prediction performance on a sub-graph of Twitter (consisting of about 0.7 million users and over 27 million tweets), compared to generative models based on only structure or only content and paths-based methods such as Katz. Nagarajan Natarajan, Prithviraj Sen, Vineet Chaoji |
ASONAM | 2 |
| 2013 | Compiling machine learning algorithms with SystemMLabstractAnalytics on big data range from passenger volume prediction in transportation to customer satisfaction in automotive diagnostic systems, and from correlation analysis in social media data to log analysis in manufacturing. Expressing and running these analytics for varying data characteristics and at scale is challenging. To address these challenges, SystemML implements a declarative, high-level language using an R-like syntax extended with machine-learning-specific constructs, that is compiled to a MapReduce runtime [2]. The language is rich enough to express a wide class of statistical, predictive modeling and machine learning algorithms (Fig. 1). We chose robust algorithms that scale to large, and potentially sparse data with many features. Matthias Boehm 0001, Douglas Burdick, Alexandre V. Evfimievski, Berthold Reinwald, Prithviraj Sen, Shirish Tatikonda, Yuanyuan Tian 0001 |
SoCC | 5 |
| 2012 | Collective context-aware topic models for entity disambiguationabstractA crucial step in adding structure to unstructured data is to identify references to entities and disambiguate them. Such disambiguated references can help enhance readability and draw similarities across different pieces of running text in an automated fashion. Previous research has tackled this problem by first forming a catalog of entities from a knowledge base, such as Wikipedia, and then using this catalog to disambiguate references in unseen text. However, most of the previously proposed models either do not use all text in the knowledge base, potentially missing out on discriminative features, or do not exploit word-entity proximity to learn high-quality catalogs. In this work, we propose topic models that keep track of the context of every word in the knowledge base; so that words appearing within the same context as an entity are more likely to be associated with that entity. Thus, our topic models utilize all text present in the knowledge base and help learn high-quality catalogs. Our models also learn groups of co-occurring entities thus enabling collective disambiguation. Unlike most previous topic models, our models are non-parametric and do not require the user to specify the exact number of groups present in the knowledge base. In experiments performed on an extract of Wikipedia containing almost 60,000 references, our models outperform SVM-based baselines by as much as 18% in terms of disambiguation accuracy translating to an increment of almost 11,000 correctly disambiguated references. Prithviraj Sen |
WWW | 1 |
| 2011 | Entity disambiguation with hierarchical topic modelsabstractDisambiguating entity references by annotating them with unique ids from a catalog is a critical step in the enrichment of unstructured content. In this paper, we show that topic models, such as Latent Dirichlet Allocation (LDA) and its hierarchical variants, form a natural class of models for learning accurate entity disambiguation models from crowd-sourced knowledge bases such as Wikipedia. Our main contribution is a semi-supervised hierarchical model called Wikipedia-based Pachinko Allocation Model} (WPAM) that exploits: (1) All words in the Wikipedia corpus to learn word-entity associations (unlike existing approaches that only use words in a small fixed window around annotated entity references in Wikipedia pages), (2) Wikipedia annotations to appropriately bias the assignment of entity labels to annotated (and co-occurring unannotated) words during model learning, and (3) Wikipedia's category hierarchy to capture co-occurrence patterns among entities. We also propose a scheme for pruning spurious nodes from Wikipedia's crowd-sourced category hierarchy. In our experiments with multiple real-life datasets, we show that WPAM outperforms state-of-the-art baselines by as much as 16% in terms of disambiguation accuracy. Saurabh Kataria 0003, Krishnan S. Kumar, Rajeev Rastogi, Prithviraj Sen, Srinivasan H. Sengamedu |
KDD | 4 |
| 2011 | Web information extraction using markov logic networksabstractIn this paper, we consider the problem of extracting structured data from web pages taking into account both the content of individual attributes as well as the structure of pages and sites. We use Markov Logic Networks (MLNs) to capture both content and structural features in a single unified framework, and this enables us to perform more accurate inference. MLNs allow us to model a wide range of rich structural features like proximity, precedence, alignment, and contiguity, using first-order clauses. We show that inference in our information extraction scenario reduces to solving an instance of the maximum weight subgraph problem. We develop specialized procedures for solving the maximum subgraph variants that are far more efficient than previously proposed inference methods for MLNs that solve variants of MAX-SAT. Experiments with real-life datasets demonstrate the effectiveness of our MLN-based approach compared to existing state-of-the-art extraction methods. Sandeepkumar Satpal, Sahely Bhadra, Sundararajan Sellamanickam, Rajeev Rastogi, Prithviraj Sen |
KDD | 5 |
| 2010 | Read-Once Functions and Query Evaluation in Probabilistic DatabasesabstractProbabilistic databases hold promise of being a viable means for large-scale uncertainty management, increasingly needed in a number of real world applications domains. However, query evaluation in probabilistic databases remains a computational challenge. Prior work on efficient exact query evaluation in probabilistic databases has largely concentrated on query-centric formulations (e.g., safe plans, hierarchical queries ), in that, they only consider characteristics of the query and not the data in the database. It is easy to construct examples where a supposedly hard query run on an appropriate database gives rise to a tractable query evaluation problem. In this paper, we develop efficient query evaluation techniques that leverage characteristics of both the query and the data in the database. We focus on tuple-independent databases where the query evaluation problem is equivalent to computing marginal probabilities of Boolean formulas associated with the result tuples. This latter task is easy if the Boolean formulas can be factorized into a form that has every variable appearing at most once (called read-once ). However, a naive approach that directly uses previously developed Boolean formula factorization algorithms is inefficient, because those algorithms require the input formulas to be in the disjunctive normal form (DNF). We instead develop novel, more efficient factorization algorithms that directly construct the read-once expression for a result tuple Boolean formula (if one exists), for a large subclass of queries (specifically, conjunctive queries without self-joins). We empirically demonstrate that (1) our proposed techniques are orders of magnitude faster than generic inference algorithms for queries where the result Boolean formulas can be factorized into read-once expressions, and (2) for the special case of hierarchical queries, they rival the efficiency of prior techniques specifically designed to handle such queries. Prithviraj Sen, Amol Deshpande, Lise Getoor |
Proc. VLDB Endow. | 1 |
| 2009 | Bisimulation-based Approximate Lifted Inference
Prithviraj Sen, Amol Deshpande, Lise Getoor |
UAI | 1 |
| 2009 | PrDB: managing and exploiting rich correlations in probabilistic databases
Prithviraj Sen, Amol Deshpande, Lise Getoor |
VLDB J. | 1 |
| 2008 | Cost-sensitive learning with conditional Markov networks
Prithviraj Sen, Lise Getoor |
Data Min. Knowl. Discov. | 1 |
| 2008 | Exploiting shared correlations in probabilistic databasesabstractThere has been a recent surge in work in probabilistic databases, propelled in large part by the huge increase in noisy data sources --- from sensor data, experimental data, data from uncurated sources, and many others. There is a growing need for database management systems that can efficiently represent and query such data. In this work, we show how data characteristics can be leveraged to make the query evaluation process more efficient. In particular, we exploit what we refer to as shared correlations where the same uncertainties and correlations occur repeatedly in the data. Shared correlations occur mainly due to two reasons: (1) Uncertainty and correlations usually come from general statistics and rarely vary on a tuple-to-tuple basis; (2) The query evaluation procedure itself tends to re-introduce the same correlations. Prior work has shown that the query evaluation problem on probabilistic databases is equivalent to a probabilistic inference problem on an appropriately constructed probabilistic graphical model (PGM). We leverage this by introducing a new data structure, called the random variable elimination graph (rv-elim graph) that can be built from the PGM obtained from query evaluation. We develop techniques based on bisimulation that can be used to compress the rv-elim graph exploiting the presence of shared correlations in the PGM, the compressed rv-elim graph can then be used to run inference. We validate our methods by evaluating them empirically and show that even with a few shared correlations significant speed-ups are possible. Prithviraj Sen, Amol Deshpande, Lise Getoor |
Proc. VLDB Endow. | 1 |
| 2007 | Representing and Querying Correlated Tuples in Probabilistic DatabasesabstractProbabilistic databases have received considerable attention recently due to the need for storing uncertain data produced by many real world applications. The widespread use of probabilistic databases is hampered by two limitations: (1) current probabilistic databases make simplistic assumptions about the data (e.g., complete independence among tuples) that make it difficult to use them in applications that naturally produce correlated data, and (2) most probabilistic databases can only answer a re-stricted subset of the queries that can be expressed using traditional query languages. We address both these limitations by proposing a framework that can represent not only probabilistic tuples, but also correlations that may be present among them. Our proposed framework naturally lends itself to the possible world semantics thus preserving the precise query semantics extant in current probabilistic databases. We develop an effi-cient strategy for query evaluation over such probabilistic databases by casting the query processing problem as an inference problem in an ap-propriately constructed probabilistic graphical model. We present several optimizations specific to probabilistic databases that enable efficient query evaluation. We validate our approach by presenting an experimental eval-uation that illustrates the effectiveness of our techniques at answering various queries using real and synthetic datasets. 1 Prithviraj Sen, Amol Deshpande |
ICDE | 1 |
| 2006 | Cost-sensitive learning with conditional Markov networksabstractThere has been a recent, growing interest in classification and link prediction in structured domains. Methods such as CRFs (Lafferty et al., 2001) and RMNs (Taskar et al., 2002) support flexible mechanisms for modeling correlations due to the link structure. In addition, in many structured domains, there is an interesting structure in the risk or cost function associated with different misclassifications. There is a rich tradition of cost-sensitive learning applied to unstructured (IID) data. Here we propose a general framework which can capture correlations in the link structure and handle structured cost functions. We present a novel cost-sensitive structured classifier based on Maximum Entropy principles that directly determines the cost-sensitive classification. We contrast this with an approach which employs a standard 0/1 loss structured classifier followed by minimization of the expected cost of misclassification. We demonstrate the utility of our proposed classifier with experiments on both synthetic and real-world data. Prithviraj Sen, Lise Getoor |
ICML | 1 |