VLDB 2026 Research / reviewers in the wild / expert
Jwala Dhamala
dblp:187/5905
· DBLP profile ↗
16ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-5396-9187ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 6 first-authorArtificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Trustworthy machine learning · 34% Language models and text generation · 29% Efficient and distributed learning · 18% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computing education · 100% |
Topics — the 10 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
1.4 | 3 | 2023 | Incorporating Fairness in Large Scale NLU Systems · WSDM 2023 Measuring Fairness of Text Classifiers via Prediction Sensitivity · ACL (1) 2022 Measures and Best Practices for Responsible AI · KDD 2021 |
Natural language and speech › Language models and text generation
natural language understanding |
1.3 | 2 | 2023 | Incorporating Fairness in Large Scale NLU Systems · WSDM 2023 Multi-VALUE: A Framework for Cross-Dialectal English NLP · ACL (1) 2023 |
Natural language and speech › Information extraction and text analysis
ambiguity resolution |
0.7 | 1 | 2023 | Resolving Ambiguities in Text-to-Image Generative Models · ACL (1) 2023 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.7 | 1 | 2023 | Incorporating Fairness in Large Scale NLU Systems · WSDM 2023 |
Machine learning › Efficient and distributed learning
model compression |
0.7 | 1 | 2023 | Incorporating Fairness in Large Scale NLU Systems · WSDM 2023 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.7 | 1 | 2023 | Resolving Ambiguities in Text-to-Image Generative Models · ACL (1) 2023 |
Natural language and speech › Language models and text generation
dialectal variation |
0.2 | 1 | 2023 | Multi-VALUE: A Framework for Cross-Dialectal English NLP · ACL (1) 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.2 | 1 | 2023 | Resolving Ambiguities in Text-to-Image Generative Models · ACL (1) 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.2 | 1 | 2023 | Multi-VALUE: A Framework for Cross-Dialectal English NLP · ACL (1) 2023 |
Natural language and speech › Information extraction and text analysis
text classification |
0.2 | 1 | 2022 | Measuring Fairness of Text Classifiers via Prediction Sensitivity · ACL (1) 2022 |
Methods — techniques the papers use, named apart from their topics
reward design analysis · 1.7benchmark checklist · 1.7rule-based translation · 0.7knowledge distillation · 0.7fine-tuning · 0.7fairness metrics · 0.7data augmentation · 0.7metrics design · 0.5dataset development · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Establishing Best Practices in Building Rigorous Agentic BenchmarksabstractBenchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to evaluate agents on complex, real-world tasks. These benchmarks typically measure agent capabilities by evaluating task outcomes via specific reward designs. However, we show that many agentic benchmarks have issues in task setup or reward design. For example, SWE-bench-Verified uses insufficient test cases, while $\tau$-bench counts empty responses as successes. Such issues can lead to under- or overestimation of agents’ performance by up to 100% in relative terms. To make agentic evaluation rigorous, we introduce the Agentic Benchmark Checklist (ABC), a set of guidelines that we synthesized from our benchmark-building experience, a survey of best practices, and previously reported issues. When applied to CVE-Bench, a benchmark with a particularly complex evaluation design, ABC reduces performance overestimation by 33%. Yuxuan Zhu 0003, Tengjun Jin, Yada Pruksachatkun, Andy Zhang, Sasha Cui, Sayash Kapoor, Shayne Longpre, Kevin Meng, Rebecca Weiss, Fazl Barez, Rahul Gupta 0001, Jwala Dhamala, Jacob Merizian, Mario Giulianelli, Harry Coppock, Cozmin Ududec, Antony Kellermann, Jasjeet S. Sekhon, Jacob Steinhardt, Sarah Schwettmann, Arvind Narayanan, Matei Zaharia, Ion Stoica, Percy Liang, Daniel Kang 0001 |
NeurIPS | 13 |
| 2024 | Tree-of-Traversals: A Zero-Shot Reasoning Algorithm for Augmenting Black-box Language Models with Knowledge GraphsabstractElan Markowitz, Anil Ramakrishna, Jwala Dhamala, Ninareh Mehrabi, Charith Peris, Rahul Gupta, Kai-Wei Chang, Aram Galstyan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Elan Markowitz, Anil Ramakrishna, Jwala Dhamala, Ninareh Mehrabi, Charith Peris, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan |
ACL (1) | 3 |
| 2023 | Resolving Ambiguities in Text-to-Image Generative ModelsabstractNinareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Varun Kumar, Qian Hu, Kai-Wei Chang, Richard Zemel, Aram Galstyan, Rahul Gupta. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ninareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Kai-Wei Chang 0001, Richard S. Zemel, Aram Galstyan, Rahul Gupta 0001 |
ACL (1) | 4 |
| 2023 | Multi-VALUE: A Framework for Cross-Dialectal English NLPabstractDialect differences caused by regional, social, and economic factors cause performance discrepancies for many groups of language technology users.Inclusive and equitable language technology must critically be dialect invariant, meaning that performance remains constant over dialectal shifts.Current systems often fall short of this ideal since they are designed and tested on a single dialect: Standard American English (SAE).We introduce a suite of resources for evaluating and achieving English dialect invariance.The resource is called Multi-VALUE, a controllable rule-based translation system spanning 50 English dialects and 189 unique linguistic features.Multi-VALUE maps SAE to synthetic forms of each dialect.First, we use this system to stress tests question answering, machine translation, and semantic parsing.Stress tests reveal significant performance disparities for leading models on nonstandard dialects.Second, we use this system as a data augmentation technique to improve the dialect robustness of existing systems.Finally, we partner with native speakers of Chicano and Indian English to release new goldstandard variants of the popular CoQA task.To execute the transformation code, run model checkpoints, and download both synthetic and gold-standard dialectal benchmark datasets, see http://value-nlp.org/. Caleb Ziems, William Barr Held, Jingfeng Yang 0001, Jwala Dhamala, Rahul Gupta 0001, Diyi Yang |
ACL (1) | 4 |
| 2023 | Incorporating Fairness in Large Scale NLU SystemsabstractNLU models power several user facing experiences such as conversations agents and chat bots. Building NLU models typically consist of 3 stages: a) building or finetuning a pre-trained model b) distilling or fine-tuning the pre-trained model to build task specific models and, c) deploying the task-specific model to production. In this presentation, we will identify fairness considerations that can be incorporated in the aforementioned three stages in the life-cycle of NLU model building: (i) selection/building of a large scale language model, (ii) distillation/fine-tuning the large model into task specific model and, (iii) deployment of the task specific model. We will present select metrics that can be used to quantify fairness in NLU models and fairness enhancement techniques that can be deployed in each of these stages. Finally, we will share some recommendations to successfully implement fairness considerations when building an industrial scale NLU system. Rahul Gupta 0001, Lisa Bauer, Kai-Wei Chang 0001, Jwala Dhamala, Aram Galstyan, Palash Goyal, Avni Khatri, Rohit Parimi, Charith Peris, Apurv Verma, Richard S. Zemel, Premkumar Natarajan |
WSDM | 4 |
| 2022 | Measuring Fairness of Text Classifiers via Prediction SensitivityabstractSatyapriya Krishna, Rahul Gupta, Apurv Verma, Jwala Dhamala, Yada Pruksachatkun, Kai-Wei Chang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Satyapriya Krishna, Rahul Gupta 0001, Apurv Verma, Jwala Dhamala, Yada Pruksachatkun, Kai-Wei Chang 0001 |
ACL (1) | 4 |
| 2022 | An Analysis of The Effects of Decoding Algorithms on Fairness in Open-Ended Language GenerationabstractSeveral prior works have shown that language models (LMs) can generate text containing harmful social biases and stereotypes. While decoding algorithms play a central role in determining properties of LM generated text, their impact on the fairness of the generations has not been studied. We present a systematic analysis of the impact of decoding algorithms on LM fairness, and analyze the trade-off between fairness, diversity and quality. Our experiments with top-p, top-k and temperature decoding algorithms, in open-ended language generation, show that fairness across demographic groups changes significantly with change in decoding algorithm's hyper-parameters. Notably, decoding algorithms that output more diverse text also output more texts with negative sentiment and regard. We present several findings and provide recommendations on standardized reporting of decoding details in fairness evaluations and optimization of decoding algorithms for fairness alongside quality and diversity. Jwala Dhamala, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan |
SLT | 1 |
| 2021 | Measures and Best Practices for Responsible AIabstractThe use of machine learning (ML) based systems has become ubiquitous including their usage in critical applications like medicine and assistive technologies. Therefore, it is important to determine the trustworthiness of these ML models and tasks. A key component in this determination is the development of task specific datasets, metrics, and best practices which are able to measure the various aspects of responsible model development and deployment including robustness, interpretability and fairness. Further, datasets are also key when training for a given task, be it coreference resolution in language modeling or facial recognition in computer vision. Imbalances and inadequate representation in datasets can have repercussions of an undesirable nature. Some common examples include how coreference resolution systems in NLU are often not all gender inclusive, discrepancies in the measurement of how robust and trustworthy machine predictions are in domains where the selective labels problem is prevalent, and discriminatory determination of pain or care levels of people belonging to different demographics in health science applications. Development of task specific datasets which do better in this regard is also extremely vital. In this workshop, we invite contributions towards different (i) datasets which help enhance task performance and inclusivity, (ii) measures and metrics which help in determining the trustworthiness of a model/dataset, (iii) assessment or remediation tools for fairer, more transparent, robust, and reliable models, and (iv) case studies describing responsible development and deployment of AI systems across fields such as healthcare, financial services, insurance, etc. The datasets, measures, mitigation techniques, and best practices could focus on different areas including (but not restricted to) the following: Fairness and Bias Robustness Reliability and Safety Interpretability Explainability Ethical AI Causal Inference Counterfactual Example Analysis They could also be focussed on the applications in diverse fields such as industry, finance, healthcare and beyond. Text based datasets can be in languages other than English as well. Sunipa Dev, Mehrnoosh Sameki, Jwala Dhamala, Cho-Jui Hsieh |
KDD | 3 |
| 2020 | Learning Geometry-Dependent and Physics-Based Inverse Image Reconstruction
Xiajun Jiang, Sandesh Ghimire, Jwala Dhamala, Zhiyuan Li 0007, Prashnna Gyawali |
MICCAI (6) | 3 |
| 2020 | Embedding high-dimensional Bayesian optimization via generative modeling: Parameter personalization of cardiac electrophysiological models
Jwala Dhamala, Pradeep Bajracharya, Hermenegild Arevalo, John L. Sapp, B. Milan Horácek, Katherine C. Wu, Natalia A. Trayanova |
Medical Image Anal. | 1 |
| 2019 | Bayesian Optimization on Large Graphs via a Graph Convolutional Generative Model: Application in Cardiac Model Personalization
Jwala Dhamala, Sandesh Ghimire, John L. Sapp, B. Milan Horácek |
MICCAI (2) | 1 |
| 2018 | High-Dimensional Bayesian Optimization of Personalized Cardiac Model Parameters via an Embedded Generative Model
Jwala Dhamala, Sandesh Ghimire, John L. Sapp, B. Milan Horácek |
MICCAI (2) | 1 |
| 2018 | Generative Modeling and Inverse Imaging of Cardiac Transmembrane Potential
Sandesh Ghimire, Jwala Dhamala, Prashnna Gyawali, John L. Sapp, B. Milan Horácek |
MICCAI (2) | 2 |
| 2018 | Quantifying the uncertainty in model parameters using Gaussian process-based Markov chain Monte Carlo in cardiac electrophysiology
Jwala Dhamala, Hermenegild Arevalo, John L. Sapp, B. Milan Horácek, Katherine C. Wu, Natalia A. Trayanova |
Medical Image Anal. | 1 |
| 2017 | Spatially Adaptive Multi-Scale Optimization for Local Parameter Estimation in Cardiac ElectrophysiologyabstractTo obtain a patient-specific cardiac electro-physiological (EP) model, it is important to estimate the 3-D distributed tissue properties of the myocardium. Ideally, the tissue property should be estimated at the resolution of the cardiac mesh. However, such high-dimensional estimation faces major challenges in identifiability and computation. Most existing works reduce this dimension by partitioning the cardiac mesh into a pre-defined set of segments. The resulting low-resolution solutions have a limited ability to represent the underlying heterogeneous tissue properties of varying sizes, locations, and distributions. In this paper, we present a novel framework that, going beyond a uniform low-resolution approach, is able to obtain a higher resolution estimation of tissue properties represented by spatially non-uniform resolution. This is achieved by two central elements: 1) a multi-scale coarse-to-fine optimization that facilitates higher resolution optimization using the lower resolution solution and 2) a spatially adaptive decision criterion that retains lower resolution in homogeneous tissue regions and allows higher resolution in heterogeneous tissue regions. The presented framework is evaluated in estimating the local tissue excitability properties of a cardiac EP model on both synthetic and real data experiments. Its performance is compared with optimization using pre-defined segments. Results demonstrate the feasibility of the presented framework to estimate local parameters and to reveal heterogeneous tissue properties at a higher resolution without using a high number of unknowns. Jwala Dhamala, Hermenegild Arevalo, John L. Sapp, B. Milan Horácek, Katherine C. Wu, Natalia A. Trayanova |
IEEE Trans. Medical Imaging | 1 |
| 2016 | Spatially-Adaptive Multi-scale Optimization for Local Parameter Estimation: Application in Cardiac Electrophysiological ModelsabstractThe estimation of local parameter values for a 3D cardiac model is important for revealing abnormal tissues with altered material properties and for building patient-specific models. Existing works in local parameter estimation typically represent the heart with a small number of pre-defined segments to reduce the dimension of unknowns. Such low-resolution approaches have limited ability to estimate tissues with varying sizes, locations, and distributions. We present a novel optimization framework to achieve a higher-resolution parameter estimation without using a high number of unknowns. It has two central elements: (1) a multi-scale coarse-to-fine optimization that uses low-resolution solutions to facilitate the higher-resolution optimization; and (2) a spatially-adaptive scheme that dedicates higher resolution to regions of heterogeneous tissue properties whereas retaining low resolution in homogeneous regions. Synthetic and real-data experiments demonstrate the ability of the presented framework to improve the accuracy of local parameter estimation in comparison to optimization based on fixed-segment models. Jwala Dhamala, John L. Sapp, B. Milan Horácek |
MICCAI (3) | 1 |