VLDB 2026 Research / reviewers in the wild / expert
Saloni Potdar
dblp:194/3158
· DBLP profile ↗
15ranked-venue papers
0as first author
10since 2021 · last 2025
0009-0006-0607-6104ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Do Large Language Models have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMsabstractCurrent Large Language Models (LLMs) are predominantly designed with English as the primary language, and even the few that are multilingual tend to exhibit strong English-centric biases.Much like speakers who might produce awkward expressions when learning a second language, LLMs often generate unnatural outputs in non-English languages, reflecting English-centric patterns in both vocabulary and grammar.Despite the importance of this issue, the naturalness of multilingual LLM outputs has received limited attention.In this paper, we address this gap by introducing novel automatic corpus-level metrics to assess the lexical and syntactic naturalness of LLM outputs in a multilingual context.Using our new metrics, we evaluate state-of-the-art LLMs on a curated benchmark in French and Chinese 1 , revealing a tendency towards English-influenced patterns.To mitigate this issue, we also propose a simple and effective alignment method to improve the naturalness of an LLM in a target language and domain, achieving consistent improvements in naturalness without compromising the performance on general-purpose benchmarks.Our work highlights the importance of developing multilingual metrics, resources and methods for the new wave of multilingual LLMs. * Work done during internship at Apple. Yanzhu Guo, Simone Conia, Zelin Zhou, Saloni Potdar, Henry Xiao |
ACL (1) | 5 |
| 2025 | KG-TRICK: Unifying Textual and Relational Information Completion of Knowledge for Multilingual Knowledge GraphsabstractMultilingual knowledge graphs (KGs) provide high-quality relational and textual information for various NLP applications, but they are often incomplete, especially in non-English languages. Previous research has shown that combining information from KGs in different languages aids either Knowledge Graph Completion (KGC), the task of predicting missing relations between entities, or Knowledge Graph Enhancement (KGE), the task of predicting missing textual information for entities. Although previous efforts have considered KGC and KGE as independent tasks, we hypothesize that they are interdependent and mutually beneficial. To this end, we introduce KG-TRICK, a novel sequence-to-sequence framework that unifies the tasks of textual and relational information completion for multilingual KGs. KG-TRICK demonstrates that: i) it is possible to unify the tasks of KGC and KGE into a single framework, and ii) combining textual information from multiple languages is beneficial to improve the completeness of a KG. As part of our contributions, we also introduce WikiKGE10++, the largest manually-curated benchmark for textual information completion of KGs, which features over 25,000 entities across 10 diverse languages. Zelin Zhou, Simone Conia, Shenglei Huang, Umar Farooq Minhas, Saloni Potdar, Henry Xiao, Yunyao Li 0001 |
COLING | 7 |
| 2024 | Towards Cross-Cultural Machine Translation with Retrieval-Augmented Generation from Multilingual Knowledge GraphsabstractTranslating text that contains entity names is a challenging task, as cultural-related references can vary significantly across languages.These variations may also be caused by transcreation, an adaptation process that entails more than transliteration and word-for-word translation.In this paper, we address the problem of cross-cultural translation on two fronts: (i) we introduce XC-Translate, the first large-scale, manually-created benchmark for machine translation that focuses on text that contains potentially culturally-nuanced entity names, and (ii) we propose KG-MT, a novel end-to-end method to integrate information from a multilingual knowledge graph into a neural machine translation model by leveraging a dense retrieval mechanism.Our experiments and analyses show that current machine translation systems and large language models still struggle to translate texts containing entity names, whereas KG-MT outperforms state-of-the-art approaches by a large margin, obtaining a 129% and 62% relative improvement compared to NLLB-200 and GPT-4, respectively. Simone Conia, Umar Farooq Minhas, Saloni Potdar, Yunyao Li 0001 |
EMNLP | 5 |
| 2024 | AGRaME: Any-Granularity Ranking with Multi-Vector EmbeddingsabstractRanking is a fundamental problem in search, however, existing ranking algorithms usually restrict the granularity of ranking to full passages or require a specific dense index for each desired level of granularity.Such lack of flexibility in granularity negatively affects many applications that can benefit from more granular ranking, such as sentence-level ranking for open-domain QA, or proposition-level ranking for attribution.In this work, we introduce the idea of any-granularity ranking 1 which leverages multi-vector embeddings to rank at varying levels of granularity while maintaining encoding at a single (coarser) level of granularity.We propose a multi-granular contrastive loss for training multi-vector approaches and validate its utility with both sentences and propositions as ranking units.Finally, we demonstrate the application of proposition-level ranking to post-hoc citation addition in retrievalaugmented generation, surpassing the performance of prompt-driven citation generation. Revanth Gangi Reddy, Omar Attia, Yunyao Li 0001, Heng Ji 0001, Saloni Potdar |
EMNLP | 5 |
| 2024 | The Third Workshop on Applied Machine Learning ManagementabstractMachine learning applications are rapidly adopted by industry leaders in any field.The growth of investment in AI-driven solutions,including the emerging field of General AI (GenAI), has created new challenges in managing Data Science and ML resources, people and projects as a whole.The discipline of managing applied machine learning teams, requires a healthy mix between agile product development tool-set and a long term research oriented mindset.The abilities of investing in deep research while at the same time connecting the outcomes to significant business results create a large knowledge based on management methods and best practices in the field.The Third KDD Workshop on Applied Machine Learning Management brings together applied research managers from various fields to share methodologies and case-studies on management of ML teams, products, and projects, achieving business impact with advanced AI-methods. Dmitri Goldenberg, Shir Meir Lador, Elena Sokolova, Lin Lee Cheong, Mohak Sukhwani, Saloni Potdar |
KDD | 6 |
| 2024 | Entity Disambiguation via Fusion Entity DecodingabstractJunxiong Wang, Ali Mousavi, Omar Attia, Ronak Pradeep, Saloni Potdar, Alexander Rush, Umar Farooq Minhas, Yunyao Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Junxiong Wang, Ali Mousavi 0003, Omar Attia, Ronak Pradeep, Saloni Potdar, Alexander M. Rush, Umar Farooq Minhas, Yunyao Li 0001 |
NAACL-HLT | 5 |
| 2023 | The Second Workshop on Applied Machine Learning ManagementabstractMachine learning applications are rapidly adopted by industry leaders in any field. The growth of investment in AI-driven solutions created new challenges in managing Data Science and ML resources, people and projects as a whole. The discipline of managing applied machine learning teams, requires a healthy mix between agile product development tool-set and a long term research oriented mindset. The abilities of investing in deep research while at the same time connecting the outcomes to significant business results create a large knowledge based on management methods and best practices in the field. The Second KDD Workshop on Applied Machine Learning Management brings together applied research managers from various fields to share methodologies and case-studies on management of ML teams, products, and projects, achieving business impact with advanced AI-methods. Dmitri Goldenberg, Chana Ross, Shir Meir Lador, Lin Lee Cheong, Elena Sokolova, Amit Mandelbaum, Irina Vasilinetc, Amit Weil Modlinger, Saloni Potdar |
KDD | 11 |
| 2022 | Improved Text Classification via Contrastive Adversarial TrainingabstractWe propose a simple and general method to regularize the fine-tuning of Transformer-based encoders for text classification tasks. Specifically, during fine-tuning we generate adversarial examples by perturbing the word embedding matrix of the model and perform contrastive learning on clean and adversarial examples in order to teach the model to learn noise-invariant representations. By training on both clean and adversarial examples along with the additional contrastive objective, we observe consistent improvement over standard fine-tuning on clean examples. On several GLUE benchmark tasks, our fine-tuned Bert_Large model outperforms Bert_Large baseline by 1.7% on average, and our fine-tuned Roberta_Large improves over Roberta_Large baseline by 1.3%. We additionally validate our method in different domains using three intent classification datasets, where our fine-tuned Roberta_Large outperforms Roberta_Large baseline by 1-2% on average. For the challenging low-resource scenario, we train our system using half of the training data (per intent) in each of the three intent classification datasets, and achieve similar performance compared to the baseline trained with full training data. Lin Pan 0003, Chung-Wei Hang, Avirup Sil, Saloni Potdar |
AAAI | 4 |
| 2021 | Multilingual BERT Post-Pretraining AlignmentabstractLin Pan, Chung-Wei Hang, Haode Qi, Abhishek Shah, Saloni Potdar, Mo Yu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Lin Pan 0003, Chung-Wei Hang, Haode Qi, Abhishek Shah, Saloni Potdar, Mo Yu |
NAACL-HLT | 5 |
| 2021 | Narrative Question Answering with Cutting-Edge Open-Domain QA Techniques: A Comprehensive StudyabstractAbstract Recent advancements in open-domain question answering (ODQA), that is, finding answers from large open-domain corpus like Wikipedia, have led to human-level performance on many datasets. However, progress in QA over book stories (Book QA) lags despite its similar task formulation to ODQA. This work provides a comprehensive and quantitative analysis about the difficulty of Book QA: (1) We benchmark the research on the NarrativeQA dataset with extensive experiments with cutting-edge ODQA techniques. This quantifies the challenges Book QA poses, as well as advances the published state-of-the-art with a ∼7% absolute improvement on ROUGE-L. (2) We further analyze the detailed challenges in Book QA through human studies.1 Our findings indicate that the event-centric questions dominate this task, which exemplifies the inability of existing QA models to handle event-oriented scenarios. Xiangyang Mou, Chenghao Yang 0001, Mo Yu, Bingsheng Yao, Saloni Potdar, Hui Su |
Trans. Assoc. Comput. Linguistics | 6 |
| 2019 | Extracting Multiple-Relations in One-Pass with Pre-Trained TransformersabstractThe state-of-the-art solutions for extracting multiple entity-relations from an input paragraph always require a multiple-pass encoding on the input.This paper proposes a new solution that can complete the multiple entityrelations extraction task with only one-pass encoding on the input corpus, and achieve a new state-of-the-art accuracy performance, as demonstrated in the ACE 2005 benchmark.Our solution is built on top of the pre-trained self-attentive models (Transformer).Since our method uses a single-pass to compute all relations at once, it scales to larger datasets easily; which makes it more usable in real-world applications.1 * Equal contributions from the corresponding authors: {wanghaoy,mingtan,yum}@us.ibm.com.Part of Haoyu Wang 0002, Mo Yu, Shiyu Chang, Dakuo Wang, Saloni Potdar |
ACL (1) | 8 |
| 2019 | Context-Aware Conversation Thread Detection in Multi-Party ChatabstractMing Tan, Dakuo Wang, Yupeng Gao, Haoyu Wang, Saloni Potdar, Xiaoxiao Guo, Shiyu Chang, Mo Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Dakuo Wang, Yupeng Gao, Haoyu Wang 0002, Saloni Potdar, Shiyu Chang, Mo Yu |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Out-of-Domain Detection for Low-Resource Text Classification TasksabstractMing Tan, Yang Yu, Haoyu Wang, Dakuo Wang, Saloni Potdar, Shiyu Chang, Mo Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yang Yu 0029, Haoyu Wang 0002, Dakuo Wang, Saloni Potdar, Shiyu Chang, Mo Yu |
EMNLP/IJCNLP (1) | 5 |
| 2018 | Diverse Few-Shot Text Classification with Multiple MetricsabstractMo Yu, Xiaoxiao Guo, Jinfeng Yi, Shiyu Chang, Saloni Potdar, Yu Cheng, Gerald Tesauro, Haoyu Wang, Bowen Zhou. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Mo Yu, Jinfeng Yi, Shiyu Chang, Saloni Potdar, Yu Cheng 0001, Gerald Tesauro, Haoyu Wang 0002 |
NAACL-HLT | 5 |
| 2017 | Neural Models for Sequence ChunkingabstractMany natural language understanding (NLU) tasks, such as shallow parsing (i.e., text chunking) and semantic slot filling, require the assignment of representative labels to the meaningful chunks in a sentence. Most of the current deep neural network (DNN) based methods consider these tasks as a sequence labeling problem, in which a word, rather than a chunk, is treated as the basic unit for labeling. These chunks are then inferred by the standard IOB (Inside-Outside- Beginning) labels. In this paper, we propose an alternative approach by investigating the use of DNN for sequence chunking, and propose three neural models so that each chunk can be treated as a complete unit for labeling. Experimental results show that the proposed neural sequence chunking models can achieve start-of-the-art performance on both the text chunking and slot filling tasks. Feifei Zhai, Saloni Potdar, Bing Xiang, Bowen Zhou 0006 |
AAAI | 2 |