VLDB 2026 Research / reviewers in the wild / expert
Pratik Karmakar
dblp:340/4276
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2025
0009-0008-1111-8801ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Discovering Voting Power for Ensemble Methods
Pratik Karmakar, Angelo Saadeh, Pierre Senellart, Stéphane Bressan |
DEXA (1) | 1 |
| 2025 | Expected Shapley Value is Shapley Value for Expected Utility Game
Pratik Karmakar, Antoine Gauquier, Pierre Senellart |
ECSQARU | 1 |
| 2025 | Using A Probabilistic Database in an Image Retrieval ApplicationabstractInternational audience Fajrian Yunus, Pratik Karmakar, Pierre Senellart, Talel Abdessalem, Stéphane Bressan |
EDBT | 2 |
| 2024 | Expected Shapley-Like Scores of Boolean functions: Complexity and Applications to Probabilistic DatabasesabstractShapley values, originating in game theory and increasingly prominent in explainable AI, have been proposed to assess the contribution of facts in query answering over databases, along with other similar power indices such as Banzhaf values. In this work we adapt these Shapley-like scores to probabilistic settings, the objective being to compute their expected value. We show that the computations of expected Shapley values and of the expected values of Boolean functions are interreducible in polynomial time, thus obtaining the same tractability landscape. We investigate the specific tractable case where Boolean functions are represented as deterministic decomposable circuits, designing a polynomial-time algorithm for this setting. We present applications to probabilistic databases through database provenance, and an effective implementation of this algorithm within the ProvSQL system, which experimentally validates its feasibility over a standard benchmark. Pratik Karmakar, Mikaël Monet, Pierre Senellart, Stéphane Bressan |
Proc. ACM Manag. Data | 1 |
| 2023 | Marich: A Query-efficient Distributionally Equivalent Model Extraction AttackabstractWe study design of black-box model extraction attacks that can *send minimal number of queries from* a *publicly available dataset* to a target ML model through a predictive API with an aim *to create an informative and distributionally equivalent replica* of the target.
First, we define *distributionally equivalent* and *Max-Information model extraction* attacks, and reduce them into a variational optimisation problem. The attacker sequentially solves this optimisation problem to select the most informative queries that simultaneously maximise the entropy and reduce the mismatch between the target and the stolen models. This leads to *an active sampling-based query selection algorithm*, Marich, which is *model-oblivious*. Then, we evaluate Marich on different text and image data sets, and different models, including CNNs and BERT. Marich extracts models that achieve $\sim 60-95\%$ of true model's accuracy and uses $\sim 1,000 - 8,500$ queries from the publicly available datasets, which are different from the private training datasets. Models extracted by Marich yield prediction distributions, which are $\sim2-4\times$ closer to the target's distribution in comparison to the existing active sampling-based attacks. The extracted models also lead to 84-96$\%$ accuracy under membership inference attacks. Experimental results validate that Marich is *query-efficient*, and capable of performing task-accurate, high-fidelity, and informative model extraction. Pratik Karmakar, Debabrota Basu |
NeurIPS | 1 |