VLDB 2026 Research / reviewers in the wild / expert
V. G. Vinod Vydiswaran
dblp:67/6469
· DBLP profile ↗
25ranked-venue papers
8as first author
5since 2021 · last 2024
0000-0002-3122-1936ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 5 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Knowledge representation and reasoning · 47% Language models and text generation · 24% Trustworthy machine learning · 24% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 77% Information retrieval · 23% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning
explanation generation |
0.5 | 1 | 2021 | LIREx: Augmenting Language Inference with Relevant Explanations · AAAI 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › explanation generation
justification |
0.5 | 1 | 2021 | LIREx: Augmenting Language Inference with Relevant Explanations · AAAI 2021 |
Machine learning › Trustworthy machine learning › interpretability
natural language explanation |
0.5 | 1 | 2021 | LIREx: Augmenting Language Inference with Relevant Explanations · AAAI 2021 |
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference |
0.5 | 1 | 2021 | LIREx: Augmenting Language Inference with Relevant Explanations · AAAI 2021 |
Data mining › network analysis
trust propagation |
0.1 | 1 | 2011 | Content-driven trust propagation framework · KDD 2011 |
Natural language and speech › Information extraction and text analysis
textual entailment |
0.1 | 1 | 2010 | "Ask Not What Textual Entailment Can Do for You..." · ACL 2010 |
Information retrieval › ranking › text ranking
news ranking |
0.0 | 1 | 2011 | Content-driven trust propagation framework · KDD 2011 |
Methods — techniques the papers use, named apart from their topics
rationale-enabled generation · 0.5instance selection · 0.5retrieval-based approach · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ChatGPT as an Attack Tool: Stealthy Textual Backdoor Attack via Blackbox Generative Model TriggerabstractJiazhao Li, Yijin Yang, Zhuofeng Wu, V.G. Vinod Vydiswaran, Chaowei Xiao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jiazhao Li, Yijin Yang, Zhuofeng Wu 0001, V. G. Vinod Vydiswaran, Chaowei Xiao |
NAACL-HLT | 4 |
| 2023 | Dementia and electronic health record phenotypes: a scoping review of available phenotypes and opportunities for future researchabstractOBJECTIVE: We performed a scoping review of algorithms using electronic health record (EHR) data to identify patients with Alzheimer's disease and related dementias (ADRD), to advance their use in research and clinical care. MATERIALS AND METHODS: Starting with a previous scoping review of EHR phenotypes, we performed a cumulative update (April 2020 through March 1, 2023) using Pubmed, PheKB, and expert review with exclusive focus on ADRD identification. We included algorithms using EHR data alone or in combination with non-EHR data and characterized whether they identified patients at high risk of or with a current diagnosis of ADRD. RESULTS: For our cumulative focused update, we reviewed 271 titles meeting our search criteria, 49 abstracts, and 26 full text papers. We identified 8 articles from the original systematic review, 8 from our new search, and 4 recommended by an expert. We identified 20 papers describing 19 unique EHR phenotypes for ADRD: 7 algorithms identifying patients with diagnosed dementia and 12 algorithms identifying patients at high risk of dementia that prioritize sensitivity over specificity. Reference standards range from only using other EHR data to in-person cognitive screening. CONCLUSION: A variety of EHR-based phenotypes are available for use in identifying populations with or at high-risk of developing ADRD. This review provides comparative detail to aid in choosing the best algorithm for research, clinical care, and population health projects based on the use case and available data. Future research may further improve the design and use of algorithms by considering EHR data provenance. Anne M. Walling, Joshua M. Pevnick, Antonia V. Bennett, V. G. Vinod Vydiswaran, Christine S. Ritchie |
J. Am. Medical Informatics Assoc. | 4 |
| 2022 | Validating Complex Phenotypes: A Structured Approach for Dementia
David A. Dorr, Nicole Gray Weiskopf, Michelle Bobo, MJ Dunne, Peijan Han, Jessica Kim, V. G. Vinod Vydiswaran |
AMIA | 7 |
| 2022 | IDPG: An Instance-Dependent Prompt Generation MethodabstractZhuofeng Wu, Sinong Wang, Jiatao Gu, Rui Hou, Yuxiao Dong, V.G.Vinod Vydiswaran, Hao Ma. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Zhuofeng Wu 0001, Sinong Wang, Jiatao Gu, Yuxiao Dong, V. G. Vinod Vydiswaran, Hao Ma 0001 |
NAACL-HLT | 6 |
| 2021 | LIREx: Augmenting Language Inference with Relevant ExplanationsabstractNatural language explanations (NLEs) are a special form of data annotation in which annotators identify rationales (most significant text tokens) when assigning labels to data instances, and write out explanations for the labels in natural language based on the rationales. NLEs have been shown to capture human reasoning better, but not as beneficial for natural language inference (NLI). In this paper, we analyze two primary flaws in the way NLEs are currently used to train explanation generators for language inference tasks. We find that the explanation generators do not take into account the variability inherent in human explanation of labels, and that the current explanation generation models generate spurious explanations. To overcome these limitations, we propose a novel framework, LIREx, that incorporates both a rationale-enabled explanation generator and an instance selector to select only relevant, plausible NLEs to augment NLI models. When evaluated on the standardized SNLI data set, LIREx achieved an accuracy of 91.87%, an improvement of 0.32 over the baseline and matching the best-reported performance on the data set. It also achieves significantly better performance than previous studies when transferred to the out-of-domain MultiNLI data set. Qualitative analysis shows that LIREx generates flexible, faithful, and relevant NLEs that allow the model to be more robust to spurious explanations. The code is available at https://github.com/zhaoxy92/LIREx. V. G. Vinod Vydiswaran |
AAAI | 2 |
| 2020 | Uncovering the relationship between food-related discussion on Twitter and neighborhood characteristicsabstractOBJECTIVE: Initiatives to reduce neighborhood-based health disparities require access to meaningful, timely, and local information regarding health behavior and its determinants. We examined the validity of Twitter as a source of information for neighborhood-level analysis of dietary choices and attitudes. MATERIALS AND METHODS: We analyzed the "healthiness" quotient and sentiment in food-related tweets at the census tract level, and associated them with neighborhood characteristics and health outcomes. We analyzed keywords driving the differences in food healthiness between the most and least-affluent tracts, and qualitatively analyzed contents of a random sample of tweets. RESULTS: Significant, albeit weak, correlations existed between healthiness and sentiment in food-related tweets and tract-level measures of affluence, disadvantage, race, age, U.S. density, and mortality from conditions associated with obesity. Analyses of keywords driving the differences in food healthiness revealed foods high in saturated fat (eg, pizza, bacon, fries) were mentioned more frequently in less-affluent tracts. Food-related discussion referred to activities (eating, drinking, cooking), locations where food was consumed, and positive (affection, cravings, enjoyment) and negative attitudes (dislike, personal struggles, complaints). DISCUSSION: Tweet-based healthiness scores largely correlated with offline phenomena in the expected directions. Social media offer less resource-intensive data collection methods than traditional surveys do. Twitter may assist in informing local health programs that focus on drivers of food consumption and could inform interventions focused on attitudes and the food environment. CONCLUSIONS: Twitter provided weak but significant signals concerning food-related behavior and attitudes at the neighborhood level, suggesting its potential usefulness for informing local health disparity reduction efforts. V. G. Vinod Vydiswaran, Daniel M. Romero, Deahan Yu, Iris N. Gomez-Lopez, Jin Xiu Lu, Bradley E. Iott, Ana Baylin, Erica C. Jansen, Philippa Clarke, Veronica J. Berrocal, Robert Goodspeed, Tiffany C. Veinot |
J. Am. Medical Informatics Assoc. | 1 |
| 2019 | Hybrid bag of approaches to characterize selection criteria for cohort identificationabstractOBJECTIVE: The 2018 National NLP Clinical Challenge (2018 n2c2) focused on the task of cohort selection for clinical trials, where participating systems were tasked with analyzing longitudinal patient records to determine if the patients met or did not meet any of the 13 selection criteria. This article describes our participation in this shared task. MATERIALS AND METHODS: We followed a hybrid approach combining pattern-based, knowledge-intensive, and feature weighting techniques. After preprocessing the notes using publicly available natural language processing tools, we developed individual criterion-specific components that relied on collecting knowledge resources relevant for these criteria and pattern-based and weighting approaches to identify "met" and "not met" cases. RESULTS: As part of the 2018 n2c2 challenge, 3 runs were submitted. The overall micro-averaged F1 on the training set was 0.9444. On the test set, the micro-averaged F1 for the 3 submitted runs were 0.9075, 0.9065, and 0.9056. The best run was placed second in the overall challenge and all 3 runs were statistically similar to the top-ranked system. A reimplemented system achieved the best overall F1 of 0.9111 on the test set. DISCUSSION: We highlight the need for a focused resource-intensive effort to address the class imbalance in the cohort selection identification task. CONCLUSION: Our hybrid approach was able to identify all selection criteria with high F1 performance on both training and test sets. Based on our participation in the 2018 n2c2 task, we conclude that there is merit in continuing a focused criterion-specific analysis and developing appropriate knowledge resources to build a quality cohort selection system. V. G. Vinod Vydiswaran, Asher Strayhorn, Phil Robinson, Mahesh Agarwal, Erin Bagazinski, Madia Essiet, Bradley E. Iott, Hyeon Joo, PingJui Ko, Dahee Lee, Jin Xiu Lu, Jinghui Liu, Adharsh Murali, Koki Sasagawa, Nalingna Yuan |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | Feasibility of Identifying Oral Anticancer Agent Toxicity Self-Reporting and Management Advice from Clinical Notes
V. G. Vinod Vydiswaran, Eun-Young Lee, Hyeon Joo, Anna Zheng, Marcelline R. Harris |
AMIA | 2 |
| 2018 | Supervised Learning Approach to Link Prediction in FDA Adverse Event Reporting System (FAERS) Database Network
Andy Jinseok Lee, Sunyang Fu, V. G. Vinod Vydiswaran |
AMIA | 3 |
| 2018 | "Bacon Bacon Bacon": Food-Related Tweets and Sentiment in Metro Detroit
V. G. Vinod Vydiswaran, Daniel M. Romero, Deahan Yu, Iris N. Gomez-Lopez, Jin Xiu Lu, Bradley E. Iott, Ana Baylin, Philippa Clarke, Veronica J. Berrocal, Robert Goodspeed, Tiffany C. Veinot |
ICWSM | 1 |
| 2018 | User acceptance of location-tracking technologies in health research: Implications for study design and data quality
Jean Hardy, Tiffany C. Veinot, Veronica J. Berrocal, Philippa Clarke, Robert Goodspeed, Iris N. Gomez-Lopez, Daniel M. Romero, V. G. Vinod Vydiswaran |
J. Biomed. Informatics | 9 |
| 2017 | Matching Consumer Health Vocabulary with Professional Medical Terms Through Concept Embedding
Yue Wang 0035, Jian Tang 0005, V. G. Vinod Vydiswaran, Kai Zheng 0002, Hua Xu 0001, Qiaozhu Mei |
AMIA | 3 |
| 2017 | HyDeXT: A Hybrid De-identification and Extraction Tool for Health Text
V. G. Vinod Vydiswaran |
AMIA | 2 |
| 2017 | Identifying Usage Expression Sentences in Consumer Product ReviewsabstractIn this paper we introduce the problem of identifying usage expression sentences in a consumer product review. We create a human-annotated gold standard dataset of 565 reviews spanning five distinct product categories. Our dataset consists of more than 3,000 annotated sentences. We further introduce a classification system to label sentences according to whether or not they describe some “usage”. The system combines lexical, syntactic, and semantic features in a product-agnostic fashion to yield good classification performance. We show the effectiveness of our approach using importance ranking of features, error analysis, and cross-product classification experiments. Shibamouli Lahiri, V. G. Vinod Vydiswaran, Rada Mihalcea |
IJCNLP(1) | 2 |
| 2017 | Development and empirical user-centered evaluation of semantically-based query recommendation for an electronic health record search engine
David A. Hanauer, Danny T. Y. Wu, Qiaozhu Mei, Katherine B. Murkowski-Steffy, V. G. Vinod Vydiswaran, Kai Zheng 0002 |
J. Biomed. Informatics | 6 |
| 2016 | Assessing the readability of ClinicalTrials.govabstractOBJECTIVE: ClinicalTrials.gov serves critical functions of disseminating trial information to the public and helping the trials recruit participants. This study assessed the readability of trial descriptions at ClinicalTrials.gov using multiple quantitative measures. MATERIALS AND METHODS: The analysis included all 165,988 trials registered at ClinicalTrials.gov as of April 30, 2014. To obtain benchmarks, the authors also analyzed 2 other medical corpora: (1) all 955 Health Topics articles from MedlinePlus and (2) a random sample of 100,000 clinician notes retrieved from an electronic health records system intended for conveying internal communication among medical professionals. The authors characterized each of the corpora using 4 surface metrics, and then applied 5 different scoring algorithms to assess their readability. The authors hypothesized that clinician notes would be most difficult to read, followed by trial descriptions and MedlinePlus Health Topics articles. RESULTS: Trial descriptions have the longest average sentence length (26.1 words) across all corpora; 65% of their words used are not covered by a basic medical English dictionary. In comparison, average sentence length of MedlinePlus Health Topics articles is 61% shorter, vocabulary size is 95% smaller, and dictionary coverage is 46% higher. All 5 scoring algorithms consistently rated CliniclTrials.gov trial descriptions the most difficult corpus to read, even harder than clinician notes. On average, it requires 18 years of education to properly understand these trial descriptions according to the results generated by the readability assessment algorithms. DISCUSSION AND CONCLUSION: Trial descriptions at CliniclTrials.gov are extremely difficult to read. Significant work is warranted to improve their readability in order to achieve CliniclTrials.gov's goal of facilitating information dissemination and subject recruitment. Danny T. Y. Wu, David A. Hanauer, Qiaozhu Mei, Patricia M. Clark, Lawrence C. An, Joshua Proulx, Qing T. Zeng, V. G. Vinod Vydiswaran, Kevyn Collins-Thompson, Kai Zheng 0002 |
J. Am. Medical Informatics Assoc. | 8 |
| 2015 | Identifying Patterns Indicative of Copying/Pasting Behavior in Patient Generated Online Content
Tera L. Reynolds, V. G. Vinod Vydiswaran, Qiaozhu Mei, David A. Hanauer, Kai Zheng 0002 |
AMIA | 2 |
| 2015 | Overcoming bias to learn about controversial topicsabstractDeciding whether a claim is true or false often requires a deeper understanding of the evidence supporting and contradicting the claim. However, when presented with many evidence documents, users do not necessarily read and trust them uniformly. Psychologists and other researchers have shown that users tend to follow and agree with articles and sources that hold viewpoints similar to their own, a phenomenon known as confirmation bias. This suggests that when learning about a controversial topic, human biases and viewpoints about the topic may affect what is considered “trustworthy” or credible. It is an interesting challenge to build systems that can help users overcome this bias and help them decide the truthfulness of claims. In this article, we study various factors that enable humans to acquire additional information about controversial claims in an unbiased fashion. Specifically, we designed a user study to understand how presenting evidence with contrasting viewpoints and source expertise ratings affect how users learn from the evidence documents. We find that users do not seek contrasting viewpoints by themselves, but explicitly presenting contrasting evidence helps them get a well‐rounded understanding of the topic. Furthermore, explicit knowledge of the credibility of the sources and the context in which the source provides the evidence document not only affects what users read but also whether they perceive the document to be credible. V. G. Vinod Vydiswaran, ChengXiang Zhai, Dan Roth 0001, Peter Pirolli |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2015 | Ease of adoption of clinical natural language processing software: An evaluation of five systemsabstractOBJECTIVE: In recognition of potential barriers that may inhibit the widespread adoption of biomedical software, the 2014 i2b2 Challenge introduced a special track, Track 3 - Software Usability Assessment, in order to develop a better understanding of the adoption issues that might be associated with the state-of-the-art clinical NLP systems. This paper reports the ease of adoption assessment methods we developed for this track, and the results of evaluating five clinical NLP system submissions. MATERIALS AND METHODS: A team of human evaluators performed a series of scripted adoptability test tasks with each of the participating systems. The evaluation team consisted of four "expert evaluators" with training in computer science, and eight "end user evaluators" with mixed backgrounds in medicine, nursing, pharmacy, and health informatics. We assessed how easy it is to adopt the submitted systems along the following three dimensions: communication effectiveness (i.e., how effective a system is in communicating its designed objectives to intended audience), effort required to install, and effort required to use. We used a formal software usability testing tool, TURF, to record the evaluators' interactions with the systems and 'think-aloud' data revealing their thought processes when installing and using the systems and when resolving unexpected issues. RESULTS: Overall, the ease of adoption ratings that the five systems received are unsatisfactory. Installation of some of the systems proved to be rather difficult, and some systems failed to adequately communicate their designed objectives to intended adopters. Further, the average ratings provided by the end user evaluators on ease of use and ease of interpreting output are -0.35 and -0.53, respectively, indicating that this group of users generally deemed the systems extremely difficult to work with. While the ratings provided by the expert evaluators are higher, 0.6 and 0.45, respectively, these ratings are still low indicating that they also experienced considerable struggles. DISCUSSION: The results of the Track 3 evaluation show that the adoptability of the five participating clinical NLP systems has a great margin for improvement. Remedy strategies suggested by the evaluators included (1) more detailed and operation system specific use instructions; (2) provision of more pertinent onscreen feedback for easier diagnosis of problems; (3) including screen walk-throughs in use instructions so users know what to expect and what might have gone wrong; (4) avoiding jargon and acronyms in materials intended for end users; and (5) packaging prerequisites required within software distributions so that prospective adopters of the software do not have to obtain each of the third-party components on their own. Kai Zheng 0002, V. G. Vinod Vydiswaran, Yang Liu 0019, Yue Wang 0035, Amber Stubbs, Özlem Uzuner, Anupama E. Gururaj, Samuel Bayer, John S. Aberdeen, Anna Rumshisky, Serguei V. S. Pakhomov, Hua Xu 0001 |
J. Biomed. Informatics | 2 |
| 2014 | Mining Consumer Health Vocabulary from Community-Generated Text
V. G. Vinod Vydiswaran, Qiaozhu Mei, David A. Hanauer, Kai Zheng 0002 |
AMIA | 1 |
| 2014 | User-Created Groups in Health Forums: What Makes Them Special?
V. G. Vinod Vydiswaran, Yang Liu 0019, Kai Zheng 0002, David A. Hanauer, Qiaozhu Mei |
ICWSM | 1 |
| 2012 | BiasTrust: teaching biased users about controversial topicsabstractDeciding whether a claim is true or false often requires understanding the evidence supporting and contradicting the claim. However, when learning about a controversial claim, human biases and viewpoints may affect which evidence documents are considered "trustworthy" or credible. It is important to overcome this bias and know both viewpoints to get a balanced perspective. In this paper, we study various factors that affect learning about the truthfulness of controversial claims. We designed a user study to understand the impact of these factors. Specifically, we studied the impact of presenting evidence with contrasting viewpoints and source expertise rating on how users accessed the evidence documents. This would help us optimize how to teach users about controversial topics in the most effective way, and to design better claim verification systems. We find that users do not seek contrasting viewpoints by themselves, but explicitly presenting contrasting evidence helps them get a well-rounded understanding of the topic. Furthermore, explicit knowledge of the source credibility and the context not only affects what users read, but also how credible they perceive the document to be. V. G. Vinod Vydiswaran, ChengXiang Zhai, Dan Roth 0001, Peter Pirolli |
CIKM | 1 |
| 2012 | Reliability Prediction of Webpages in the Medical Domain
Parikshit Sondhi, V. G. Vinod Vydiswaran, ChengXiang Zhai |
ECIR | 2 |
| 2011 | Content-driven trust propagation frameworkabstractExisting fact-finding models assume availability of structured data or accurate information extraction. However, as online data gets more unstructured, these assumptions are no longer valid. To overcome this, we propose a novel, content-based, trust propagation framework that relies on signals from the textual content to ascertain veracity of free-text claims and compute trustworthiness of their sources. We incorporate the quality of relevant content into the framework and present an iterative algorithm for propagation of trust scores. We show that existing fact finders on structured data can be modeled as specific instances of this framework. Using a retrieval-based approach to find relevant articles, we instantiate the framework to compute trustworthiness of news sources and articles. We show that the proposed framework helps ascertain trustworthiness of sources better. We also show that ranking news articles based on trustworthiness learned from the content-driven framework is significantly better than baselines that ignore either the content quality or the trust framework. V. G. Vinod Vydiswaran, ChengXiang Zhai, Dan Roth 0001 |
KDD | 1 |
| 2010 | "Ask Not What Textual Entailment Can Do for You..."
Mark Sammons, V. G. Vinod Vydiswaran, Dan Roth 0001 |
ACL | 2 |