Jwala Dhamala

dblp:187/5905 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-5396-9187ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 6 first-authorArtificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 34% Language models and text generation · 29% Efficient and distributed learning · 18%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
1.432023
Incorporating Fairness in Large Scale NLU Systems · WSDM 2023
Measuring Fairness of Text Classifiers via Prediction Sensitivity · ACL (1) 2022
Measures and Best Practices for Responsible AI · KDD 2021
Natural language and speech › Language models and text generation
natural language understanding
1.322023
Incorporating Fairness in Large Scale NLU Systems · WSDM 2023
Multi-VALUE: A Framework for Cross-Dialectal English NLP · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
ambiguity resolution
0.712023
Resolving Ambiguities in Text-to-Image Generative Models · ACL (1) 2023
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.712023
Incorporating Fairness in Large Scale NLU Systems · WSDM 2023
Machine learning › Efficient and distributed learning
model compression
0.712023
Incorporating Fairness in Large Scale NLU Systems · WSDM 2023
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.712023
Resolving Ambiguities in Text-to-Image Generative Models · ACL (1) 2023
Natural language and speech › Language models and text generation
dialectal variation
0.212023
Multi-VALUE: A Framework for Cross-Dialectal English NLP · ACL (1) 2023
Machine learning › Trustworthy machine learning
interpretability
0.212023
Resolving Ambiguities in Text-to-Image Generative Models · ACL (1) 2023
Machine learning › Trustworthy machine learning
robustness
0.212023
Multi-VALUE: A Framework for Cross-Dialectal English NLP · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
text classification
0.212022
Measuring Fairness of Text Classifiers via Prediction Sensitivity · ACL (1) 2022

Methods — techniques the papers use, named apart from their topics

reward design analysis · 1.7benchmark checklist · 1.7rule-based translation · 0.7knowledge distillation · 0.7fine-tuning · 0.7fairness metrics · 0.7data augmentation · 0.7metrics design · 0.5dataset development · 0.5
YearPublicationVenuePosition
2025 Establishing Best Practices in Building Rigorous Agentic Benchmarks
abstract
Benchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to evaluate agents on complex, real-world tasks. These benchmarks typically measure agent capabilities by evaluating task outcomes via specific reward designs. However, we show that many agentic benchmarks have issues in task setup or reward design. For example, SWE-bench-Verified uses insufficient test cases, while $\tau$-bench counts empty responses as successes. Such issues can lead to under- or overestimation of agents’ performance by up to 100% in relative terms. To make agentic evaluation rigorous, we introduce the Agentic Benchmark Checklist (ABC), a set of guidelines that we synthesized from our benchmark-building experience, a survey of best practices, and previously reported issues. When applied to CVE-Bench, a benchmark with a particularly complex evaluation design, ABC reduces performance overestimation by 33%.
Yuxuan Zhu 0003, Tengjun Jin, Yada Pruksachatkun, Andy Zhang, Sasha Cui, Sayash Kapoor, Shayne Longpre, Kevin Meng, Rebecca Weiss, Fazl Barez, Rahul Gupta 0001, Jwala Dhamala, Jacob Merizian, Mario Giulianelli, Harry Coppock, Cozmin Ududec, Antony Kellermann, Jasjeet S. Sekhon, Jacob Steinhardt, Sarah Schwettmann, Arvind Narayanan, Matei Zaharia, Ion Stoica, Percy Liang, Daniel Kang 0001
NeurIPS13
2024 Tree-of-Traversals: A Zero-Shot Reasoning Algorithm for Augmenting Black-box Language Models with Knowledge Graphs
abstract
Elan Markowitz, Anil Ramakrishna, Jwala Dhamala, Ninareh Mehrabi, Charith Peris, Rahul Gupta, Kai-Wei Chang, Aram Galstyan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Elan Markowitz, Anil Ramakrishna, Jwala Dhamala, Ninareh Mehrabi, Charith Peris, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan
ACL (1)3
2023 Resolving Ambiguities in Text-to-Image Generative Models
abstract
Ninareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Varun Kumar, Qian Hu, Kai-Wei Chang, Richard Zemel, Aram Galstyan, Rahul Gupta. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ninareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Kai-Wei Chang 0001, Richard S. Zemel, Aram Galstyan, Rahul Gupta 0001
ACL (1)4
2023 Multi-VALUE: A Framework for Cross-Dialectal English NLP
abstract
Dialect differences caused by regional, social, and economic factors cause performance discrepancies for many groups of language technology users.Inclusive and equitable language technology must critically be dialect invariant, meaning that performance remains constant over dialectal shifts.Current systems often fall short of this ideal since they are designed and tested on a single dialect: Standard American English (SAE).We introduce a suite of resources for evaluating and achieving English dialect invariance.The resource is called Multi-VALUE, a controllable rule-based translation system spanning 50 English dialects and 189 unique linguistic features.Multi-VALUE maps SAE to synthetic forms of each dialect.First, we use this system to stress tests question answering, machine translation, and semantic parsing.Stress tests reveal significant performance disparities for leading models on nonstandard dialects.Second, we use this system as a data augmentation technique to improve the dialect robustness of existing systems.Finally, we partner with native speakers of Chicano and Indian English to release new goldstandard variants of the popular CoQA task.To execute the transformation code, run model checkpoints, and download both synthetic and gold-standard dialectal benchmark datasets, see http://value-nlp.org/.
Caleb Ziems, William Barr Held, Jingfeng Yang 0001, Jwala Dhamala, Rahul Gupta 0001, Diyi Yang
ACL (1)4
2023 Incorporating Fairness in Large Scale NLU Systems
abstract
NLU models power several user facing experiences such as conversations agents and chat bots. Building NLU models typically consist of 3 stages: a) building or finetuning a pre-trained model b) distilling or fine-tuning the pre-trained model to build task specific models and, c) deploying the task-specific model to production. In this presentation, we will identify fairness considerations that can be incorporated in the aforementioned three stages in the life-cycle of NLU model building: (i) selection/building of a large scale language model, (ii) distillation/fine-tuning the large model into task specific model and, (iii) deployment of the task specific model. We will present select metrics that can be used to quantify fairness in NLU models and fairness enhancement techniques that can be deployed in each of these stages. Finally, we will share some recommendations to successfully implement fairness considerations when building an industrial scale NLU system.
Rahul Gupta 0001, Lisa Bauer, Kai-Wei Chang 0001, Jwala Dhamala, Aram Galstyan, Palash Goyal, Avni Khatri, Rohit Parimi, Charith Peris, Apurv Verma, Richard S. Zemel, Premkumar Natarajan
WSDM4
2022 Measuring Fairness of Text Classifiers via Prediction Sensitivity
abstract
Satyapriya Krishna, Rahul Gupta, Apurv Verma, Jwala Dhamala, Yada Pruksachatkun, Kai-Wei Chang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Satyapriya Krishna, Rahul Gupta 0001, Apurv Verma, Jwala Dhamala, Yada Pruksachatkun, Kai-Wei Chang 0001
ACL (1)4
2022 An Analysis of The Effects of Decoding Algorithms on Fairness in Open-Ended Language Generation
abstract
Several prior works have shown that language models (LMs) can generate text containing harmful social biases and stereotypes. While decoding algorithms play a central role in determining properties of LM generated text, their impact on the fairness of the generations has not been studied. We present a systematic analysis of the impact of decoding algorithms on LM fairness, and analyze the trade-off between fairness, diversity and quality. Our experiments with top-p, top-k and temperature decoding algorithms, in open-ended language generation, show that fairness across demographic groups changes significantly with change in decoding algorithm's hyper-parameters. Notably, decoding algorithms that output more diverse text also output more texts with negative sentiment and regard. We present several findings and provide recommendations on standardized reporting of decoding details in fairness evaluations and optimization of decoding algorithms for fairness alongside quality and diversity.
Jwala Dhamala, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan
SLT1
2021 Measures and Best Practices for Responsible AI
abstract
The use of machine learning (ML) based systems has become ubiquitous including their usage in critical applications like medicine and assistive technologies. Therefore, it is important to determine the trustworthiness of these ML models and tasks. A key component in this determination is the development of task specific datasets, metrics, and best practices which are able to measure the various aspects of responsible model development and deployment including robustness, interpretability and fairness. Further, datasets are also key when training for a given task, be it coreference resolution in language modeling or facial recognition in computer vision. Imbalances and inadequate representation in datasets can have repercussions of an undesirable nature. Some common examples include how coreference resolution systems in NLU are often not all gender inclusive, discrepancies in the measurement of how robust and trustworthy machine predictions are in domains where the selective labels problem is prevalent, and discriminatory determination of pain or care levels of people belonging to different demographics in health science applications. Development of task specific datasets which do better in this regard is also extremely vital. In this workshop, we invite contributions towards different (i) datasets which help enhance task performance and inclusivity, (ii) measures and metrics which help in determining the trustworthiness of a model/dataset, (iii) assessment or remediation tools for fairer, more transparent, robust, and reliable models, and (iv) case studies describing responsible development and deployment of AI systems across fields such as healthcare, financial services, insurance, etc. The datasets, measures, mitigation techniques, and best practices could focus on different areas including (but not restricted to) the following: Fairness and Bias Robustness Reliability and Safety Interpretability Explainability Ethical AI Causal Inference Counterfactual Example Analysis They could also be focussed on the applications in diverse fields such as industry, finance, healthcare and beyond. Text based datasets can be in languages other than English as well.
Sunipa Dev, Mehrnoosh Sameki, Jwala Dhamala, Cho-Jui Hsieh
KDD3
2020 Learning Geometry-Dependent and Physics-Based Inverse Image Reconstruction
Xiajun Jiang, Sandesh Ghimire, Jwala Dhamala, Zhiyuan Li 0007, Prashnna Gyawali
MICCAI (6)3
2020 Embedding high-dimensional Bayesian optimization via generative modeling: Parameter personalization of cardiac electrophysiological models
Jwala Dhamala, Pradeep Bajracharya, Hermenegild Arevalo, John L. Sapp, B. Milan Horácek, Katherine C. Wu, Natalia A. Trayanova
Medical Image Anal.1
2019 Bayesian Optimization on Large Graphs via a Graph Convolutional Generative Model: Application in Cardiac Model Personalization
Jwala Dhamala, Sandesh Ghimire, John L. Sapp, B. Milan Horácek
MICCAI (2)1
2018 High-Dimensional Bayesian Optimization of Personalized Cardiac Model Parameters via an Embedded Generative Model
Jwala Dhamala, Sandesh Ghimire, John L. Sapp, B. Milan Horácek
MICCAI (2)1
2018 Generative Modeling and Inverse Imaging of Cardiac Transmembrane Potential
Sandesh Ghimire, Jwala Dhamala, Prashnna Gyawali, John L. Sapp, B. Milan Horácek
MICCAI (2)2
2018 Quantifying the uncertainty in model parameters using Gaussian process-based Markov chain Monte Carlo in cardiac electrophysiology
Jwala Dhamala, Hermenegild Arevalo, John L. Sapp, B. Milan Horácek, Katherine C. Wu, Natalia A. Trayanova
Medical Image Anal.1
2017 Spatially Adaptive Multi-Scale Optimization for Local Parameter Estimation in Cardiac Electrophysiology
abstract
To obtain a patient-specific cardiac electro-physiological (EP) model, it is important to estimate the 3-D distributed tissue properties of the myocardium. Ideally, the tissue property should be estimated at the resolution of the cardiac mesh. However, such high-dimensional estimation faces major challenges in identifiability and computation. Most existing works reduce this dimension by partitioning the cardiac mesh into a pre-defined set of segments. The resulting low-resolution solutions have a limited ability to represent the underlying heterogeneous tissue properties of varying sizes, locations, and distributions. In this paper, we present a novel framework that, going beyond a uniform low-resolution approach, is able to obtain a higher resolution estimation of tissue properties represented by spatially non-uniform resolution. This is achieved by two central elements: 1) a multi-scale coarse-to-fine optimization that facilitates higher resolution optimization using the lower resolution solution and 2) a spatially adaptive decision criterion that retains lower resolution in homogeneous tissue regions and allows higher resolution in heterogeneous tissue regions. The presented framework is evaluated in estimating the local tissue excitability properties of a cardiac EP model on both synthetic and real data experiments. Its performance is compared with optimization using pre-defined segments. Results demonstrate the feasibility of the presented framework to estimate local parameters and to reveal heterogeneous tissue properties at a higher resolution without using a high number of unknowns.
Jwala Dhamala, Hermenegild Arevalo, John L. Sapp, B. Milan Horácek, Katherine C. Wu, Natalia A. Trayanova
IEEE Trans. Medical Imaging1
2016 Spatially-Adaptive Multi-scale Optimization for Local Parameter Estimation: Application in Cardiac Electrophysiological Models
abstract
The estimation of local parameter values for a 3D cardiac model is important for revealing abnormal tissues with altered material properties and for building patient-specific models. Existing works in local parameter estimation typically represent the heart with a small number of pre-defined segments to reduce the dimension of unknowns. Such low-resolution approaches have limited ability to estimate tissues with varying sizes, locations, and distributions. We present a novel optimization framework to achieve a higher-resolution parameter estimation without using a high number of unknowns. It has two central elements: (1) a multi-scale coarse-to-fine optimization that uses low-resolution solutions to facilitate the higher-resolution optimization; and (2) a spatially-adaptive scheme that dedicates higher resolution to regions of heterogeneous tissue properties whereas retaining low resolution in homogeneous regions. Synthetic and real-data experiments demonstrate the ability of the presented framework to improve the accuracy of local parameter estimation in comparison to optimization based on fixed-segment models.
Jwala Dhamala, John L. Sapp, B. Milan Horácek
MICCAI (3)1