EDBT 2026 Demo / reviewers in the wild / expert
Prabhanjan Kambadur
dblp:82/7037
· DBLP profile ↗
16ranked-venue papers
4as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 1 since 2021Systems, architecture and hardware · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
4 papers |
Mathematical optimization · 50% Algorithms and data structures · 34% Information theory · 8% | |
| Artificial intelligence
3 papers |
Language models and text generation · 46% Information extraction and text analysis · 30% Probabilistic and Bayesian machine learning · 24% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 42% Data mining · 35% Knowledge graphs · 21% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Parallel and multicore computing · 94% High-performance computing · 4% Distributed systems · 2% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 76% Computing education · 24% |
Topics — the 30 heaviest of 38, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › pre-trained language model
domain-specific language model |
0.8 | 1 | 2024 | Academics Can Contribute to Domain-Specialized Language Models · EMNLP 2024 |
Mathematical optimization › combinatorial optimization
greedy algorithm |
0.5 | 2 | 2017 | Faster Greedy MAP Inference for Determinantal Point Processes · ICML 2017 Orthogonal Matching Pursuit for Sparse Quantile Regression · ICDM 2014 |
Parallel and multicore computing
parallel programming models |
0.4 | 3 | 2014 | X10 and APGAS at Petascale · PPoPP 2014 NIMBLE: a toolkit for the implementation of parallel data mining and machine learning algorithms on mapreduce · KDD 2011 PFunc: modern task parallelism for modern high performance computing · SC 2009 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.4 | 1 | 2019 | A Semi-Markov Structured Support Vector Machine Model for High-Precision Named Entity Recognition · ACL (1) 2019 |
Information retrieval › ranking
learning to rank |
0.3 | 1 | 2018 | Weakly-supervised Contextualization of Knowledge Graph Facts · SIGIR 2018 |
Information retrieval
ranking |
0.3 | 1 | 2018 | Weakly-supervised Contextualization of Knowledge Graph Facts · SIGIR 2018 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
determinantal point process |
0.3 | 1 | 2017 | Faster Greedy MAP Inference for Determinantal Point Processes · ICML 2017 |
Mathematical optimization › submodular optimization
submodular maximization |
0.3 | 1 | 2017 | Faster Greedy MAP Inference for Determinantal Point Processes · ICML 2017 |
Bioinformatics and computational biology › statistical genetics › gene-gene interaction
gene-gene interaction detection |
0.2 | 1 | 2015 | An Efficient Nonlinear Regression Approach for Genome-Wide Detection of Marginal and Interacting Genetic Variations · RECOMB 2015 |
Bioinformatics and computational biology › statistical genetics
genetic association study |
0.2 | 1 | 2015 | An Efficient Nonlinear Regression Approach for Genome-Wide Detection of Marginal and Interacting Genetic Variations · RECOMB 2015 |
Bioinformatics and computational biology
statistical genetics |
0.2 | 1 | 2015 | An Efficient Nonlinear Regression Approach for Genome-Wide Detection of Marginal and Interacting Genetic Variations · RECOMB 2015 |
Data mining
clustering |
0.2 | 1 | 2015 | Spectral Clustering via the Power Method - Provably · ICML 2015 |
Data mining › clustering
spectral clustering |
0.2 | 1 | 2015 | Spectral Clustering via the Power Method - Provably · ICML 2015 |
Algorithms and data structures
numerical linear algebra |
0.2 | 1 | 2015 | Spectral Clustering via the Power Method - Provably · ICML 2015 |
Algorithms and data structures › numerical linear algebra › eigenvalue computation
power iteration |
0.2 | 1 | 2015 | Spectral Clustering via the Power Method - Provably · ICML 2015 |
Parallel and multicore computing › parallel computing
parallel programming languages |
0.2 | 1 | 2014 | X10 and APGAS at Petascale · PPoPP 2014 |
Parallel and multicore computing › parallel programming models › distributed memory programming models
partitioned global address space |
0.2 | 1 | 2014 | X10 and APGAS at Petascale · PPoPP 2014 |
Information theory › signal processing › compressed sensing
orthogonal matching pursuit |
0.2 | 1 | 2014 | Orthogonal Matching Pursuit for Sparse Quantile Regression · ICDM 2014 |
Mathematical optimization › statistical estimation › regression
quantile regression |
0.2 | 1 | 2014 | Orthogonal Matching Pursuit for Sparse Quantile Regression · ICDM 2014 |
Mathematical optimization › statistical estimation › regression
sparse regression |
0.2 | 1 | 2014 | Orthogonal Matching Pursuit for Sparse Quantile Regression · ICDM 2014 |
Computational geometry
convex geometry |
0.2 | 1 | 2013 | Fast Conical Hull Algorithms for Near-separable Non-negative Matrix Factorization · ICML (1) 2013 |
Algorithms and data structures › numerical linear algebra
matrix factorization |
0.2 | 1 | 2013 | Fast Conical Hull Algorithms for Near-separable Non-negative Matrix Factorization · ICML (1) 2013 |
Algorithms and data structures › numerical linear algebra › matrix factorization › low-rank matrix factorization
nonnegative matrix factorization |
0.2 | 1 | 2013 | Fast Conical Hull Algorithms for Near-separable Non-negative Matrix Factorization · ICML (1) 2013 |
Data mining › big data analytics › large-scale data mining
parallel data mining |
0.1 | 1 | 2011 | NIMBLE: a toolkit for the implementation of parallel data mining and machine learning algorithms on mapreduce · KDD 2011 |
Parallel and multicore computing
load balancing |
0.1 | 1 | 2011 | Lifeline-based global load balancing · PPoPP 2011 |
Parallel and multicore computing › data-parallel programming
mapreduce |
0.1 | 1 | 2011 | NIMBLE: a toolkit for the implementation of parallel data mining and machine learning algorithms on mapreduce · KDD 2011 |
Parallel and multicore computing › load balancing › dynamic load balancing
work stealing |
0.1 | 1 | 2011 | Lifeline-based global load balancing · PPoPP 2011 |
Natural language and speech › Information extraction and text analysis
sequence labeling |
0.1 | 1 | 2019 | A Semi-Markov Structured Support Vector Machine Model for High-Precision Named Entity Recognition · ACL (1) 2019 |
Machine learning › Probabilistic and Bayesian machine learning
structured prediction |
0.1 | 1 | 2019 | A Semi-Markov Structured Support Vector Machine Model for High-Precision Named Entity Recognition · ACL (1) 2019 |
Parallel and multicore computing
parallel programming runtimes |
0.1 | 1 | 2009 | PFunc: modern task parallelism for modern high performance computing · SC 2009 |
Methods — techniques the papers use, named apart from their topics
stochastic trace estimation · 0.6log-determinant approximation · 0.6chebyshev expansion · 0.6power method · 0.4k-means · 0.4semi-Markov structured SVM · 0.4loss-augmented inference · 0.4quantile huber loss · 0.4orthogonal matching pursuit · 0.4interior point method · 0.4convex sparsity regularization · 0.4neural fact contextualization · 0.3distant supervision · 0.3mapreduce · 0.2building blocks · 0.2MPI · 0.2non-linear regression · 0.2asynchronous partitioned global address space · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Academics Can Contribute to Domain-Specialized Language ModelsabstractMark Dredze, Genta Indra Winata, Prabhanjan Kambadur, Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, David S Rosenberg, Sebastian Gehrmann. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Mark Dredze, Genta Indra Winata, Prabhanjan Kambadur, Ozan Irsoy, Steven Lu 0003, Vadim Dabravolski, David S. Rosenberg, Sebastian Gehrmann |
EMNLP | 3 |
| 2019 | A Semi-Markov Structured Support Vector Machine Model for High-Precision Named Entity RecognitionabstractNamed entity recognition (NER) is the backbone of many NLP solutions.F 1 score, the harmonic mean of precision and recall, is often used to select/evaluate the best models.However, when precision needs to be prioritized over recall, a state-of-the-art model might not be the best choice.There is little in the literature that directly addresses training-time modifications to achieve higher precision information extraction.In this paper, we propose a neural semi-Markov structured support vector machine model that controls the precisionrecall trade-off by assigning weights to different types of errors in the loss-augmented inference during training.The semi-Markov property provides more accurate phrase-level predictions, thereby improving performance.We empirically demonstrate the advantage of our model when high precision is required by comparing against strong baselines based on CRF.In our experiments with the CoNLL 2003 dataset, our model achieves a better precisionrecall trade-off at various precision levels.ORG PER LOC MISC ALL CRF P. 89.5 96.3 91.8 81.1 91.06 R. 87.7 95.4 93.8 81.3 90.88 F1 88.6 95.8 92.8 81.2 90.97 SSVM P. 90.0 95.7 91.0 80.4 90.75 R. 87.7 95.5 93.7 80.5 90.79 F1 88.8 95.6 92.4 80.4 90.77Semi.SSVM P. 89.3 96.0 92.3 80.1 90.92 R. 87.2 95.2 93.2 81.9 90.60 F1 88.2 95.6 92.8 81.0 90.76 Ravneet Arora, Chen-Tse Tsai, Ketevan Tsereteli, Prabhanjan Kambadur, Yi Yang 0038 |
ACL (1) | 4 |
| 2018 | Weakly-supervised Contextualization of Knowledge Graph FactsabstractKnowledge graphs (KGs) model facts about the world; they consist of nodes (entities such as companies and people) that are connected by edges (relations such as founderOf ). Facts encoded in KGs are frequently used by search applications to augment result pages. When presenting a KG fact to the user, providing other facts that are pertinent to that main fact can enrich the user experience and support exploratory information needs. \em KG fact contextualization is the task of augmenting a given KG fact with additional and useful KG facts. The task is challenging because of the large size of KGs; discovering other relevant facts even in a small neighborhood of the given fact results in an enormous amount of candidates. We introduce a neural fact contextualization method (\em NFCM ) to address the KG fact contextualization task. NFCM first generates a set of candidate facts in the neighborhood of a given fact and then ranks the candidate facts using a supervised learning to rank model. The ranking model combines features that we automatically learn from data and that represent the query-candidate facts with a set of hand-crafted features we devised or adjusted for this task. In order to obtain the annotations required to train the learning to rank model at scale, we generate training data automatically using distant supervision on a large entity-tagged text corpus. We show that ranking functions learned on this data are effective at contextualizing KG facts. Evaluation using human assessors shows that it significantly outperforms several competitive baselines. Nikos Voskarides, Edgar Meij, Ridho Reinanda, Abhinav Khaitan, Miles Osborne, Giorgio Stefanoni, Prabhanjan Kambadur, Maarten de Rijke |
SIGIR | 7 |
| 2017 | Faster Greedy MAP Inference for Determinantal Point ProcessesabstractDeterminantal point processes (DPPs) are popular probabilistic models that arise in many machine learning tasks, where distributions of diverse sets are characterized by determinants of their features. In this paper, we develop fast algorithms to find the most likely configuration (MAP) of large-scale DPPs, which is NP-hard in general. Due to the submodular nature of the MAP objective, greedy algorithms have been used with empirical success. Greedy implementations require computation of log-determinants, matrix inverses or solving linear systems at each iteration. We present faster implementations of the greedy algorithms by utilizing the orthogonal benefits of two log-determinant approximation schemes: (a) first-order expansions to the matrix log-determinant function and (b) high-order expansions to the scalar log function with stochastic trace estimators. In our experiments, our algorithms are orders of magnitude faster than their competitors, while sacrificing marginal accuracy. Insu Han, Prabhanjan Kambadur, KyoungSoo Park, Jinwoo Shin |
ICML | 2 |
| 2017 | Adaptive Submodular Ranking
Prabhanjan Kambadur, Viswanath Nagarajan, Fatemeh Navidi |
IPCO | 1 |
| 2016 | Geolocation for Twitter: Timing MattersabstractAutomated geolocation of social media messages can benefit a variety of downstream applications.However, these geolocation systems are typically evaluated without attention to how changes in time impact geolocation.Since different people, in different locations write messages at different times, these factors can significantly vary the performance of a geolocation system over time.We demonstrate cyclical temporal effects on geolocation accuracy in Twitter, as well as rapid drops as test data moves beyond the time period of training data.We show that temporal drift can effectively be countered with even modest online model updates. Mark Dredze, Miles Osborne, Prabhanjan Kambadur |
HLT-NAACL | 3 |
| 2015 | Spectral Clustering via the Power Method - ProvablyabstractSpectral clustering is one of the most important algorithms in data mining and machine intelligence; however, its computational complexity limits its application to truly large scale data analysis. The computational bottleneck in spectral clustering is computing a few of the top eigenvectors of the (normalized) Laplacian matrix corresponding to the graph representing the data to be clustered. One way to speed up the computation of these eigenvectors is to use the “power method” from the numerical linear algebra literature. Although the power method has been empirically used to speed up spectral clustering, the theory behind this approach, to the best of our knowledge, remains unexplored. This paper provides the first such rigorous theoretical justification, arguing that a small number of power iterations suffices to obtain near-optimal partitionings using the approximate eigenvectors. Specifically, we prove that solving the k-means clustering problem on the approximate eigenvectors obtained via the power method gives an additive-error approximation to solving the k-means problem on the optimal eigenvectors. Christos Boutsidis, Prabhanjan Kambadur, Alex Gittens |
ICML | 2 |
| 2015 | An Efficient Nonlinear Regression Approach for Genome-Wide Detection of Marginal and Interacting Genetic Variations
Seunghak Lee, Aurélie C. Lozano, Prabhanjan Kambadur, Eric P. Xing |
RECOMB | 3 |
| 2014 | Orthogonal Matching Pursuit for Sparse Quantile RegressionabstractWe consider new formulations and methods for sparse quantile regression in the high-dimensional setting. Quantile regression plays an important role in many data mining applications, including outlier-robust exploratory analysis in gene selection. In addition, the sparsity consideration in quantile regression enables the exploration of the entire conditional distribution of the response variable given the predictors and therefore yields a more comprehensive view of the important predictors. We propose a generalized Orthogonal Matching Pursuit algorithm for variable selection, taking the misfit loss to be either the traditional quantile loss or a smooth version we call quantile Huber, and compare the resulting greedy approaches with convex sparsity-regularized formulations. We apply a recently proposed interior point methodology to efficiently solve all formulations, provide theoretical guarantees of consistent estimation, and demonstrate the performance of our approach using empirical studies of simulated and genomic datasets. Aleksandr Y. Aravkin, Aurélie C. Lozano, Ronny Luss, Prabhanjan Kambadur |
ICDM | 4 |
| 2014 | X10 and APGAS at PetascaleabstractX10 is a high-performance, high-productivity programming language aimed at large-scale distributed and shared-memory parallel applications. It is based on the Asynchronous Partitioned Global Address Space (APGAS) programming model, supporting the same fine-grained concurrency mechanisms within and across shared-memory nodes. Olivier Tardieu, Benjamin Herta, David Cunningham, David Grove, Prabhanjan Kambadur, Vijay A. Saraswat, Avraham Shinnar, Mikio Takeuchi, Mandana Vaziri |
PPoPP | 5 |
| 2013 | A Parallel, Block Greedy Method for Sparse Inverse Covariance Estimation for Ultra-high DimensionsabstractDiscovering the graph structure of a Gaussian Markov Random Field is an important problem in application areas such as computational biology and atmospheric sciences. This task, which translates to estimating the sparsity pattern of the inverse covariance matrix, has been extensively studied in the literature. However, the existing approaches are unable to handle ultra-high dimensional datasets and there is a crucial need to develop methods that are both highly scalable and memory-efficient. In this paper, we present GINCO, a blocked greedy method for sparse inverse covariance matrix estimation. We also present detailed description of a highly-scalable and memory-efficient implementation of GINCO, which is able to operate on both shared- and distributed-memory architectures. Our implementation is able recover the sparsity pattern of 25,000 vertex random and chain graphs with 87% and 84% accuracy in \le 5 minutes using \le 10GB of memory on a single 8-core machine. Furthermore, our method is statistically consistent in recovering the sparsity pattern of the inverse covariance matrix, which we demonstrate through extensive empirical studies. Prabhanjan Kambadur, Aurélie C. Lozano |
AISTATS | 1 |
| 2013 | Fast Conical Hull Algorithms for Near-separable Non-negative Matrix FactorizationabstractThe separability assumption (Arora et al., 2012; Donoho & Stodden, 2003) turns non-negative matrix factorization (NMF) into a tractable problem. Recently, a new class of provably-correct NMF algorithms have emerged under this assumption. In this paper, we reformulate the separable NMF problem as that of finding the extreme rays of the conical hull of a finite set of vectors. From this geometric perspective, we derive new separable NMF algorithms that are highly scalable and empirically noise robust, and have several favorable properties in relation to existing methods. A parallel implementation of our algorithm scales excellently on shared and distributed-memory machines. Abhishek Kumar 0001, Vikas Sindhwani, Prabhanjan Kambadur |
ICML (1) | 3 |
| 2011 | NIMBLE: a toolkit for the implementation of parallel data mining and machine learning algorithms on mapreduceabstractIn the last decade, advances in data collection and storage technologies have led to an increased interest in designing and implementing large-scale parallel algorithms for machine learning and data mining (ML-DM). Existing programming paradigms for expressing large-scale parallelism such as MapReduce (MR) and the Message Passing Interface (MPI) have been the de facto choices for implementing these ML-DM algorithms. The MR programming paradigm has been of particular interest as it gracefully handles large datasets and has built-in resilience against failures. However, the existing parallel programming paradigms are too low-level and ill-suited for implementing ML-DM algorithms. To address this deficiency, we present NIMBLE, a portable infrastructure that has been specifically designed to enable the rapid implementation of parallel ML-DM algorithms. The infrastructure allows one to compose parallel ML-DM algorithms using reusable (serial and parallel) building blocks that can be efficiently executed using MR and other parallel programming models; it currently runs on top of Hadoop, which is an open-source MR implementation. We show how NIMBLE can be used to realize scalable implementations of ML-DM algorithms and present a performance evaluation. Amol Ghoting, Prabhanjan Kambadur, Edwin P. D. Pednault, Ramakrishnan Kannan |
KDD | 2 |
| 2011 | Lifeline-based global load balancingabstractOn shared-memory systems, Cilk-style work-stealing has been used to effectively parallelize irregular task-graph based applications such as Unbalanced Tree Search (UTS). There are two main difficulties in extending this approach to distributed memory. In the shared memory approach, thieves (nodes without work) constantly attempt to asynchronously steal work from randomly chosen victims until they find work. In distributed memory, thieves cannot autonomously steal work from a victim without disrupting its execution. When work is sparse, this results in performance degradation. In essence, a direct extension of traditional work-stealing to distributed memory violates the work-first principle underlying work-stealing. Further, thieves spend useless CPU cycles attacking victims that have no work, resulting in system inefficiencies in multi-programmed contexts. Second, it is non-trivial to detect active distributed termination (detect that programs at all nodes are looking for work, hence there is no work). This problem is well-studied and requires careful design for good performance. Unfortunately, in most existing languages/frameworks, application developers are forced to implement their own distributed termination detection. Vijay A. Saraswat, Prabhanjan Kambadur, Sreedhar B. Kodali, David Grove, Sriram Krishnamoorthy |
PPoPP | 2 |
| 2009 | Demand-driven execution of static directed acyclic graphs using task parallelismabstractThe dataflow model allows natural expression of parallelism in an application. Applications expressed in the dataflow model can be executed either using the data-driven or the demand-driven schemes. Although both these schemes have their utility in different scenarios, the realization of the demand-driven scheme is not adequately supported in the existing solutions for task parallelism. In this paper, we examine some of the requirements placed by the demand-driven execution scheme on task parallelism. We present PFunc, a new library-based solution for task parallelism that fully supports the demand-driven execution scheme. We compare the runtimes and peak memory consumption of an unsymmetric sparse LU factorization emulation parallelized using both the data- and demand-driven execution schemes. This comparison shows that the demand-driven model provides benefits that necessitate its full support in task parallelism. Prabhanjan Kambadur, Torsten Hoefler, Andrew Lumsdaine |
HiPC | 1 |
| 2009 | PFunc: modern task parallelism for modern high performance computingabstractHPC today faces new challenges due to paradigm shifts in both hardware and software. The ubiquity of multi-cores, many-cores, and GPGPUs is forcing traditional serial as well as distributed-memory parallel applications to be parallelized for these architectures. Emerging applications in areas such as informatics are placing unique requirements on parallel programming tools that have not yet been addressed. Although, of all the available parallel programming models, task parallelism appears to be the most promising in meeting these new challenges, current solutions for task parallelism are inadequate. In this paper, we introduce PFunc, a new library for task parallelism that extends the feature set of current solutions for task parallelism with custom task scheduling, task priorities, task affinities, multiple completion notifications and task groups. These features enable PFunc to naturally and efficiently parallelize a wide variety of modern HPC applications and to support the SPMD model of parallel programming. We present three case studies: demand-driven DAG execution, frequent pattern mining and iterative sparse solvers to demonstrate the utility of PFunc's new features. Prabhanjan Kambadur, Amol Ghoting, Haim Avron, Andrew Lumsdaine |
SC | 1 |