VLDB 2026 Research / reviewers in the wild / expert
Subhadeep Karan
dblp:185/0888
· DBLP profile ↗
4ranked-venue papers
4as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-authorSystems, architecture and hardware · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Probabilistic and Bayesian machine learning · 50% Learning theory · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning
bayesian network structure learning |
0.8 | 1 | 2024 | End-to-End Bayesian Networks Exact Learning in Shared Memory · IEEE Trans. Parallel Distributed Syst. 2024 |
Machine learning › Learning theory › computational learning theory
exact learning |
0.8 | 1 | 2024 | End-to-End Bayesian Networks Exact Learning in Shared Memory · IEEE Trans. Parallel Distributed Syst. 2024 |
Parallel and multicore computing › parallel programming models
shared-memory parallelization |
0.8 | 1 | 2024 | End-to-End Bayesian Networks Exact Learning in Shared Memory · IEEE Trans. Parallel Distributed Syst. 2024 |
Parallel and multicore computing › parallel programming models
task parallelism |
0.8 | 1 | 2024 | End-to-End Bayesian Networks Exact Learning in Shared Memory · IEEE Trans. Parallel Distributed Syst. 2024 |
Methods — techniques the papers use, named apart from their topics
shared-memory task parallelism · 2.3optimization techniques · 2.3dynamic programming · 2.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | End-to-End Bayesian Networks Exact Learning in Shared MemoryabstractBayesian networks are important Machine Learning models with many practical applications in, e.g., biomedicine and bioinformatics. The problem of Bayesian networks learning is$\mathcal {NP}$-hard and computationally challenging. In this article, we propose practical parallel exact algorithms to learn Bayesian networks from data. Our approach uses shared-memory task parallelism to realize exploration of dynamic programming lattices emerging in Bayesian networks structure learning, and introduces several optimization techniques to constraint and partition the underlying search space. Through extensive experimental testing we show that the resulting method is highly scalable, and it can be used to efficiently learn large globally optimal networks. Subhadeep Karan, Zainul Abideen Sayed, Jaroslaw Zola |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | Fast Counting in Machine Learning Applications
Subhadeep Karan, Matthew Eichhorn, Blake Hurlburt, Grant Iraci, Jaroslaw Zola |
UAI | 1 |
| 2017 | Scalable Exact Parent Sets Identification in Bayesian Networks Learning with Apache SparkabstractIn Machine Learning, the parent set identification problem is to find a set of random variables that best explain selected variable given the data and some predefined scoring function. This problem is a critical component to structure learning of Bayesian networks and Markov blankets discovery, and thus has many practical applications, ranging from fraud detection to clinical decision support. In this paper, we introduce a new distributed memory approach to the exact parent sets assignment problem. To achieve scalability, we derive theoretical bounds to constraint the search space when MDL scoring function is used, and we reorganize the underlying dynamic programming such that the computational density is increased and fine-grain synchronization is eliminated. We then design efficient realization of our approach in the Apache Spark platform. Through experimental results, we demonstrate that the method maintains strong scalability on a 500-core standalone Spark cluster, and it can be used to efficiently process data sets with 70 variables, far beyond the reach of the currently available solutions. Subhadeep Karan, Jaroslaw Zola |
HiPC | 1 |
| 2016 | Exact structure learning of Bayesian networks by optimal path extensionabstractBayesian networks are probabilistic graphical models often used in big data analytics. The problem of Bayesian network exact structure learning is to find a network structure that is optimal under certain scoring criteria. The problem is known to be NP-hard and the existing methods are both computationally and memory intensive. In this paper, we introduce a new approach for exact structure learning that leverages relationship between a partial network structure and the remaining variables to constrain the number of ways in which the partial network can be optimally extended. Via experimental results, we show that the method provides up to three times improvement in runtime, and orders of magnitude reduction in memory consumption over the current best algorithms. Subhadeep Karan, Jaroslaw Zola |
IEEE BigData | 1 |