VLDB 2026 Research / reviewers in the wild / expert
Suma Bhat
dblp:66/9013
· DBLP profile ↗
47ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0003-0324-5890ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource LanguagesabstractSaeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina, Ashwath Rao B, Parameswari Krishnamurthy, Muhammad Cendekia Airlangga, Rifo Ahmad Genadi, Nguyen Phan Gia Bao, Amir Hossein Yari, Hawau Olamide Toyin, Nurdaulet Mukhituly, Mena Attia, Besher Hassan, Ahmad Fathan Hidayatullah, Tatsuki Kuribayashi, Haonan Li, Suma Bhat, Fajri Koto. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina, Ashwath Rao, Parameswari Krishnamurthy, Muhammad Cendekia Airlangga, Rifo Ahmad Genadi, Nguyen Phan Gia Bao, Amir Hossein Yari, Hawau Olamide Toyin, Nurdaulet Mukhituly, Mena Attia, Besher Hassan, Ahmad Fathan Hidayatullah, Tatsuki Kuribayashi, Haonan Li 0002, Suma Bhat, Fajri Koto |
ACL (1) | 18 |
| 2026 | Examining Students' Code Comprehension with LLMs in Block- and Text-Based ProgrammingabstractUnderstanding how students reason about code is essential for providing tailored scaffolding in computer science (CS) education. Prior work has used think-aloud protocols with the Structure of the Observed Learning Outcomes (SOLO) taxonomy to examine students' code comprehension and programming levels. However, analyzing such data is labor-intensive and requires expert judgment. Recent advances in large language models (LLMs) offer a promising avenue for scaling this analysis, though their reliability for fine-grained coding remains uncertain. To address this gap, our study investigates the extent to which GPT-5 and 4o can classify SOLO levels and identify code-comprehension strategies from think-aloud transcripts of 27 high-school students working on block-based and text-based tasks. Results show modest alignment with human ratings for SOLO, with one-shot prompting improving agreement over zero-shot, though distinctions between adjacent lower levels (e.g., Prestructural 1 vs. 2) remained difficult. Strategy detection demonstrated stronger performance, achieving accuracies of 75–77% (block) and 62–67% (text), particularly for surface-visible strategies such as 'walkthroughs', 'control-structure identification', and 'pattern recognition', but weaker for less frequent, abstract, meta-cognitive strategies such as 'strategizing' (planning an approach) or 'thoroughness' (systematically checking work). These findings highlight both the potential and the limitations of using GPT-5 and 4o to analyze think-aloud data. While this work represents an initial step, with plans to examine more models, our preliminary results indicate that a human-in-the-loop approach is essential to ensure reliability and interpretive depth. Future work will extend this evaluation to other LLMs to better understand their role in supporting instructional decision-making. Shan Zhang 0003, Toni V. Earle-Randell, Priyadharshini Ganapathy Prasad, Zifeng Liu, Yang Shi 0004, Suma Bhat, Maya Israel, Anthony Botelho |
SIGCSE (2) | 6 |
| 2026 | Investigating High School Students' Code Comprehension and Strategy Use Across Block-Based and Text-Based ProgrammingabstractUnderstanding how students comprehend code is essential for designing effective instructional support in computer science (CS). While prior studies have often relied on written responses, few have examined students' reasoning processes through think-aloud data. In this study, we analyzed the verbal reasoning of 27 high school students as they completed block-based and text-based code comprehension tasks targeting loops and conditional statements. Using an adapted SOLO taxonomy framework, we found that most students were classified at lower levels, with performance declining as they transitioned from block-based to text-based code. Students' strategy use, informed by prior work on code comprehension, showed that walkthroughs and identifying program structures were the most common approaches. Text-based tasks more often led students to use pattern-recognition strategies, such as interpreting operators or identifying numerical patterns, whereas block-based tasks occasionally prompted them to articulate broader problem-solving approaches. Overall, these findings demonstrate the value of applying the SOLO taxonomy to evaluate students' programming levels and highlight how programming modality impacts both the depth of understanding and the strategies students employ during code comprehension. Shan Zhang 0003, Priyadharshini Ganapathy Prasad, Toni V. Earle-Randell, Yang Shi 0004, Suma Bhat, Maya Israel |
SIGCSE (2) | 5 |
| 2025 | An LLM-Based Framework for Simulating, Classifying, and Correcting Students' Programming Knowledge with the SOLO TaxonomyabstractNovice programmers often face challenges in designing computational artifacts and fixing code errors, which can lead to task abandonment and over-reliance on external support. While research has explored effective meta-cognitive strategies to scaffold novice programmers' learning, it is essential to first understand and assess students' conceptual, procedural, and strategic/conditional programming knowledge at scale. To address this issue, we propose a three-model framework that leverages Large Language Models (LLMs) to simulate, classify, and correct student responses to programming questions based on the SOLO Taxonomy. The SOLO Taxonomy provides a structured approach for categorizing student understanding into four levels: Pre-structural, Uni-structural, Multi-structural, and Relational. Our results showed that GPT-4o achieved high accuracy in generating and classifying responses for the Relational category, with moderate accuracy in the Uni-structural and Pre-structural categories, but struggled with the Multi-structural category. The model successfully corrected responses to the Relational level. Although further refinement is needed, these findings suggest that LLMs hold significant potential for supporting computer science education by assessing programming knowledge and guiding students toward deeper cognitive engagement. Shan Zhang 0003, Pragati Shuddhodhan Meshram, Priyadharshini Ganapathy Prasad, Maya Israel, Suma Bhat |
SIGCSE (2) | 5 |
| 2024 | The Relation Among Gender, Language, and Posting Type in Online Chemistry Course Discussion ForumsabstractThis study explored gendered language used in an online chemistry course’s discussion forums, to understand how using gendered language might help or hinder learning outcomes, while considering the goal of various posting structures required in the course. Findings revealed that although gendered-language use did not differ between men and women, gendered forms of language were widely used throughout the forums. The use of gendered language appeared strategic, however, and reliably varied by the goal of the discussion post (i.e., posting a solution to a homework problem, asking a question, or answering a question). Ultimately, gender, language and posting type were found to be related to final grade. Genevieve M. Henricks, Michelle Perry, Suma Bhat |
LAK | 3 |
| 2024 | No Context Needed: Contextual Quandary In Idiomatic Reasoning With Pre-Trained Language ModelsabstractReasoning in the presence of idiomatic expressions (IEs) remains a challenging frontier in natural language understanding (NLU).Unlike standard text, the non-compositional nature of an IE makes it difficult for model comprehension, as their figurative or non-literal meaning usually cannot be inferred from the constituent words alone.It stands to reason that in these challenging circumstances, pre-trained language models (PTLMs) should make use of the surrounding context to infer additional information about the IE.In this paper, we investigate the utilization of said context for idiomatic reasoning tasks, which is under-explored relative to arithmetic or commonsense reasoning (Liu et al., 2022;Yu et al., 2023).Preliminary findings point to a surprising observation: general purpose PTLMs are actually negatively affected by the context, as performance almost always increases with its removal.In these scenarios, models may see gains of up to 3.89%.As a result, we argue that only IE-aware models remain suitable for idiomatic reasoning tasks, given the unexpected and unexplainable manner in which general purpose PTLMs reason over IEs.Additionally, we conduct studies to examine how models utilize the context in various situations, as well as an in-depth analysis on dataset formation and quality. 1 Finally, we provide some explanations and insights into the reasoning process itself based on our results. Kellen Cheng, Suma Bhat |
NAACL-HLT | 2 |
| 2023 | CLCL: Non-compositional Expression Detection with Contrastive Learning and Curriculum LearningabstractNon-compositional expressions present a substantial challenge for natural language processing (NLP) systems, necessitating more intricate processing compared to general language tasks, even with large pre-trained language models.Their non-compositional nature and limited availability of data resources further compound the difficulties in accurately learning their representations.This paper addresses both of these challenges.By leveraging contrastive learning techniques to build improved representations it tackles the non-compositionality challenge.Additionally, we propose a dynamic curriculum learning framework specifically designed to take advantage of the scarce available data for modeling non-compositionality.Our framework employs an easy-to-hard learning strategy, progressively optimizing the model's performance by effectively utilizing available training data.Moreover, we integrate contrastive learning into the curriculum learning approach to maximize its benefits.Experimental results demonstrate the gradual improvement in the model's performance on idiom usage recognition and metaphor detection tasks.Our evaluation encompasses six datasets, consistently affirming the effectiveness of the proposed framework.Our models available at Jianing Zhou, Ziheng Zeng, Suma Bhat |
ACL (1) | 3 |
| 2023 | IEKG: A Commonsense Knowledge Graph for Idiomatic ExpressionsabstractIdiomatic expression (IE) processing and comprehension have challenged pre-trained language models (PTLMs) because their meanings are non-compositional.Unlike prior works that enable IE comprehension through finetuning PTLMs with sentences containing IEs, in this work, we construct IEKG, a commonsense knowledge graph for figurative interpretations of IEs.This extends the established ATOMIC 20 20 (Hwang et al., 2021) graph, converting PTLMs into knowledge models (KMs) that encode and infer commonsense knowledge related to IE use.Experiments show that various PTLMs can be converted into KMs with IEKG.We verify the quality of IEKG and the ability of the trained KMs with automatic and human evaluation.Through applications in natural language understanding, we show that a PTLM injected with knowledge from IEKG exhibits improved IE comprehension ability and can generalize to IEs unseen during training. Ziheng Zeng, Kellen Tan Cheng, Srihari Venkat Nanniyur, Jianing Zhou, Suma Bhat |
EMNLP | 5 |
| 2023 | CRISP: Curriculum based Sequential neural decoders for Polar code familyabstractPolar codes are widely used state-of-the-art codes for reliable communication that have recently been included in the $5^{\text{th}}$ generation wireless standards ($5$G). However, there still remains room for design of polar decoders that are both efficient and reliable in the short blocklength regime. Motivated by recent successes of data-driven channel decoders, we introduce a novel $\textbf{ C}$ur${\textbf{RI}}$culum based $\textbf{S}$equential neural decoder for $\textbf{P}$olar codes (CRISP). We design a principled curriculum, guided by information-theoretic insights, to train CRISP and show that it outperforms the successive-cancellation (SC) decoder and attains near-optimal reliability performance on the $\text{Polar}(32,16)$ and $\text{Polar}(64,22)$ codes. The choice of the proposed curriculum is critical in achieving the accuracy gains of CRISP, as we show by comparing against other curricula. More notably, CRISP can be readily extended to Polarization-Adjusted-Convolutional (PAC) codes, where existing SC decoders are significantly less reliable. To the best of our knowledge, CRISP constructs the first data-driven decoder for PAC codes and attains near-optimal performance on the $\text{PAC}(32,16)$ code. S. Ashwin Hebbar, Viraj Nadkarni, Ashok Vardhan Makkuva, Suma Bhat, Sewoong Oh, Pramod Viswanath |
ICML | 4 |
| 2022 | Idiomatic Expression Paraphrasing without Strong SupervisionabstractIdiomatic expressions (IEs) play an essential role in natural language. In this paper, we study the task of idiomatic sentence paraphrasing (ISP), which aims to paraphrase a sentence with an IE by replacing the IE with its literal paraphrase. The lack of large-scale corpora with idiomatic-literal parallel sentences is a primary challenge for this task, for which we consider two separate solutions. First, we propose an unsupervised approach to ISP, which leverages an IE's contextual information and definition and does not require a parallel sentence training set. Second, we propose a weakly supervised approach using back-translation to jointly perform paraphrasing and generation of sentences with IEs to enlarge the small-scale parallel sentence training dataset. Other significant derivatives of the study include a model that replaces a literal phrase in a sentence with an IE to generate an idiomatic expression and a large scale parallel dataset with idiomatic/literal sentence pairs. The effectiveness of the proposed solutions compared to competitive baselines is seen in the relative gains of over 5.16 points in BLEU, over 8.75 points in METEOR, and over 19.57 points in SARI when the generated sentences are empirically validated on a parallel dataset using automatic and manual evaluations. We demonstrate the practical utility of ISP as a preprocessing step in En-De machine translation. Jianing Zhou, Ziheng Zeng, Hongyu Gong, Suma Bhat |
AAAI | 4 |
| 2022 | Using Machine Learning Explainability Methods to Personalize Interventions for Students
Paul Hur, Haejin Lee, Suma Bhat, Nigel Bosch |
EDM | 3 |
| 2022 | Getting BART to Ride the Idiomatic Train: Learning to Represent Idiomatic ExpressionsabstractAbstract Idiomatic expressions (IEs), characterized by their non-compositionality, are an important part of natural language. They have been a classical challenge to NLP, including pre-trained language models that drive today’s state-of-the-art. Prior work has identified deficiencies in their contextualized representation stemming from the underlying compositional paradigm of representation. In this work, we take a first-principles approach to build idiomaticity into BART using an adapter as a lightweight non-compositional language expert trained on idiomatic sentences. The improved capability over baselines (e.g., BART) is seen via intrinsic and extrinsic methods, where idiom embeddings score 0.19 points higher in homogeneity score for embedding clustering, and up to 25% higher sequence accuracy on the idiom processing tasks of IE sense disambiguation and span detection. Ziheng Zeng, Suma Bhat |
Trans. Assoc. Comput. Linguistics | 2 |
| 2021 | Abusive Language Detection in Heterogeneous Contexts: Dataset Collection and the Role of Supervised AttentionabstractAbusive language is a massive problem in online social platforms. Existing abusive language detection techniques are particularly ill-suited to comments containing heterogeneous abusive language patterns, i.e., both abusive and non-abusive parts. This is due in part to the lack of datasets that explicitly annotate heterogeneity in abusive language. We tackle this challenge by providing an annotated dataset of abusive language in over 11,000 comments from YouTube. We account for heterogeneity in this dataset by separately annotating both the comment as a whole and the individual sentences that comprise each comment. We then propose an algorithm that uses a supervised attention mechanism to detect and categorize abusive content using multi-task learning. We empirically demonstrate the challenges of using traditional techniques on heterogeneous content and the comparative gains in performance of the proposed approach over state-of-the-art methods. Hongyu Gong, Alberto Valido, Katherine M. Ingram, Giulia Fanti, Suma Bhat, Dorothy Espelage |
AAAI | 5 |
| 2021 | Paraphrase Generation: A Survey of the State of the ArtabstractThis paper focuses on paraphrase generation, which is a widely studied natural language generation task in NLP.With the development of neural models, paraphrase generation research has exhibited a gradual shift to neural methods in the recent years.This has provided architectures for contextualized representation of an input text and generating fluent, diverse and human-like paraphrases.This paper surveys various approaches to paraphrase generation with a main focus on neural methods. Jianing Zhou, Suma Bhat |
EMNLP (1) | 2 |
| 2021 | A Social Network Analysis of Online Engagement for College Students Traditionally Underrepresented in STEMabstractLittle is known about the online learning behaviors of students traditionally underrepresented in STEM fields (i.e., UR-STEM students), as well as how those behaviors impact important learning outcomes. The present study examined the relationship between online discussion forum engagement and success for UR-STEM and non-UR-STEM students, using the Community of Inquiry (CoI) model as our theoretical framework. Social network analysis and nested regression models were used to explore how three different measures of forum engagement—1) total number of posts written, 2) number of help-seeking posts written and replied to, and 3) level of connectivity—were related to improvement (i.e., relative performance gains) for 70 undergraduate students enrolled in an online introductory STEM course. We found a significant positive relationship between help-seeking and improvement and nonsignificant effects of general posting and connectivity; these results held for UR-STEM and non-UR-STEM students alike. Our findings suggest that online help-seeking has benefits for course improvement beyond what can be predicted by posting alone and that one need not be well connected in a class network to achieve positive learning outcomes. Finally, UR-STEM students demonstrated greater grade improvement than their non-UR-STEM counterparts, which suggests that the online environment has the potential to combat barriers to success that disproportionately affect underrepresented students. Destiny Williams-Dobosz, Renato Ferreira Leitão Azevedo, Amos Jeng, Vyom Nayan Thakkar, Suma Bhat, Nigel Bosch, Michelle Perry |
LAK | 5 |
| 2021 | Modeling Consistency Using Engagement Patterns in Online CoursesabstractConsistency of learning behaviors is known to play an important role in learners’ engagement in a course and impact their learning outcomes. Despite significant advances in the area of learning analytics (LA) in measuring various self-regulated learning behaviors, using LA to measure consistency of online course engagement patterns remains largely unexplored. This study focuses on modeling consistency of learners in online courses to address this research gap. Toward this, we propose a novel unsupervised algorithm that combines sequence pattern mining and ideas from information retrieval with a clustering algorithm to first extract engagement patterns of learners, represent learners in a vector space of these patterns and finally group them into groups with similar consistency levels. Using clickstream data recorded in a popular learning management system over two offerings of a STEM course, we validate our proposed approach to detect learners that are inconsistent in their behaviors. We find that our method not only groups learners by consistency levels, but also provides reliable instructor support at an early stage in a course. Jianing Zhou, Suma Bhat |
LAK | 2 |
| 2021 | Self-Supervised Euphemism Detection and Identification for Content ModerationabstractFringe groups and organizations have a long history of using euphemisms—ordinary-sounding words with a secret meaning—to conceal what they are discussing. Nowadays, one common use of euphemisms is to evade content moderation policies enforced by social media platforms. Existing tools for enforcing policy automatically rely on keyword searches for words on a "ban list", but these are notoriously imprecise: even when limited to swearwords, they can still cause embarrassing false positives [1]. When a commonly used ordinary word acquires a euphemistic meaning, adding it to a keyword-based ban list is hopeless: consider "pot" (storage container or marijuana?) or "heater" (household appliance or firearm?) The current generation of social media companies instead hire staff to check posts manually, but this is expensive, inhumane, and not much more effective. It is usually apparent to a human moderator that a word is being used euphemistically, but they may not know what the secret meaning is, and therefore whether the message violates policy. Also, when a euphemism is banned, the group that used it need only invent another one, leaving moderators one step behind.This paper will demonstrate unsupervised algorithms that, by analyzing words in their sentence-level context, can both detect words being used euphemistically, and identify the secret meaning of each word. Compared to the existing state of the art, which uses context-free word embeddings, our algorithm for detecting euphemisms achieves 30–400% higher detection accuracies of unlabeled euphemisms in a text corpus. Our algorithm for revealing euphemistic meanings of words is the first of its kind, as far as we are aware. In the arms race between content moderators and policy evaders, our algorithms may help shift the balance in the direction of the moderators. Wanzheng Zhu, Hongyu Gong, Rohan Bansal, Zachary Weinberg, Nicolas Christin, Giulia Fanti, Suma Bhat |
SP | 7 |
| 2021 | Idiomatic Expression Identification using Semantic CompatibilityabstractAbstract Idiomatic expressions are an integral part of natural language and constantly being added to a language. Owing to their non-compositionality and their ability to take on a figurative or literal meaning depending on the sentential context, they have been a classical challenge for NLP systems. To address this challenge, we study the task of detecting whether a sentence has an idiomatic expression and localizing it when it occurs in a figurative sense. Prior research for this task has studied specific classes of idiomatic expressions offering limited views of their generalizability to new idioms. We propose a multi-stage neural architecture with attention flow as a solution. The network effectively fuses contextual and lexical information at different levels using word and sub-word representations. Empirical evaluations on three of the largest benchmark datasets with idiomatic expressions of varied syntactic patterns and degrees of non-compositionality show that our proposed model achieves new state-of-the-art results. A salient feature of the model is its ability to identify idioms unseen during training with gains from 1.4% to 30.8% over competitive baselines on the largest dataset. Ziheng Zeng, Suma Bhat |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | Enriching Word Embeddings with Temporal and Spatial InformationabstractThe meaning of a word is closely linked to sociocultural factors that can change over time and location, resulting in corresponding meaning changes.Taking a global view of words and their meanings in a widely used language, such as English, may require us to capture more refined semantics for use in time-specific or location-aware situations, such as the study of cultural trends or language use.However, popular vector representations for words do not adequately include temporal or spatial information.In this work, we present a model for learning word representation conditioned on time and location.In addition to capturing meaning changes over time and location, we require that the resulting word embeddings retain salient semantic and geometric properties.We train our model on time-and locationstamped corpora, and show using both quantitative and qualitative evaluations that it can capture semantics across time and locations.We note that our model compares favorably with the state-of-the-art for time-specific embedding, and serves as a new benchmark for location-specific embeddings. Hongyu Gong, Suma Bhat, Pramod Viswanath |
CoNLL | 2 |
| 2020 | Effective Forum Curation via Multi-task Learning
Faeze Brahman, Nikhil Varghese, Suma Bhat, Snigdha Chaturvedi |
EDM | 3 |
| 2020 | Rich Syntactic and Semantic Information Helps Unsupervised Text Style TransferabstractText style transfer aims to change an input sentence to an output sentence by changing its text style while preserving the content.Previous efforts on unsupervised text style transfer only use the surface features of words and sentences.As a result, the transferred sentences may either have inaccurate or missing information compared to the inputs.We address this issue by explicitly enriching the inputs via syntactic and semantic structures, from which richer features are then extracted to better capture the original information.Experiments on two text-style-transfer tasks show that our approach improves the content preservation of a strong unsupervised baseline model thereby demonstrating improved transfer performance. Hongyu Gong, Linfeng Song, Suma Bhat |
INLG | 3 |
| 2020 | FUSE: Multi-faceted Set Expansion by Coherent Clustering of Skip-Grams
Wanzheng Zhu, Hongyu Gong, Chao Zhang 0014, Jingbo Shang, Suma Bhat, Jiawei Han 0001 |
ECML/PKDD (3) | 6 |
| 2019 | PaRe: A Paper-Reviewer Matching Approach Using a Common Topic SpaceabstractOmer Anjum, Hongyu Gong, Suma Bhat, Wen-Mei Hwu, JinJun Xiong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Omer Anjum, Hongyu Gong, Suma Bhat, Wen-Mei W. Hwu, Jinjun Xiong |
EMNLP/IJCNLP (1) | 3 |
| 2019 | DiAd: Domain Adaptation for Learning at ScaleabstractMassive online courses occupy an important place in the educational landscape of today. We study an approach to scale predictive analytic models derived from online course discussion fora--specifically that of confusion detection--onto other courses. The primary challenge here is the lack of labeled examples in a new course and this calls for unsupervised domain adaptation (DA). As a first step in exploring DA in the education domain, we propose a simple algorithm, DiAd, which adapts a classifier trained on a course with labeled data by selectively choosing instances from a new course (with no labeled data) that are most dissimilar to the course with labeled data and on which the classifier is very confident of classification. Our algorithm is empirically validated on the confusion detection task across multiple online courses. We find that DiAd outperforms other methods on the target domain, while showing a comparable performance to a popular method that uses labeled data from the target domain. Ziheng Zeng, Snigdha Chaturvedi, Suma Bhat, Dan Roth 0001 |
LAK | 3 |
| 2019 | Context-Sensitive Malicious Spelling Error CorrectionabstractMisspelled words of the malicious kind work by changing specific keywords and are intended to thwart existing automated applications for cyber-environment control such as harassing content detection on the Internet and email spam detection. In this paper, we focus on malicious spelling correction, which requires an approach that relies on the context and the surface forms of targeted keywords. In the context of two applications-profanity detection and email spam detection-we show that malicious misspellings seriously degrade their performance. We then propose a context-sensitive approach for malicious spelling correction using word embeddings and demonstrate its superior performance compared to state-of-the-art spell checkers. Hongyu Gong, Suma Bhat, Pramod Viswanath |
WWW | 3 |
| 2018 | Document Similarity for Texts of Varying Lengths via Hidden TopicsabstractMeasuring similarity between texts is an important task for several applications.Available approaches to measure document similarity are inadequate for document pairs that have non-comparable lengths, such as a long document and its summary.This is because of the lexical, contextual and the abstraction gaps between a long document of rich details and its concise summary of abstract information.In this paper, we present a document matching approach to bridge this gap, by comparing the texts in a common space of hidden topics.We evaluate the matching algorithm on two matching tasks and find that it consistently and widely outperforms strong baselines.We also highlight the benefits of the incorporation of domain knowledge to text matching. Hongyu Gong, Tarek Sakakini, Suma Bhat, Jinjun Xiong |
ACL (1) | 3 |
| 2018 | Using Conversational Agents to Explain Medication Instructions to Older Adults
Renato Ferreira Leitão Azevedo, Daniel G. Morrow, James Graumlich, Ann Willemsen-Dunlap, Mark Hasegawa-Johnson, Thomas S. Huang, Kuangxiao Gu, Suma Bhat, Tarek Sakakini, Victor Sadauskas, Donald Halpin |
AMIA | 8 |
| 2018 | Who they are and what they want: Understanding the reasons for MOOC enrollment
R. Wes Crues, Nigel Bosch, Carolyn J. Anderson, Michelle Perry, Suma Bhat, Najmuddin Shaik |
EDM | 5 |
| 2018 | Preposition Sense Disambiguation and RepresentationabstractPrepositions are highly polysemous, and their variegated senses encode significant semantic information.In this paper we match each preposition's left-and right context, and their interplay to the geometry of the word vectors to the left and right of the preposition.Extracting these features from a large corpus and using them with machine learning models makes for an efficient preposition sense disambiguation (PSD) algorithm, which is comparable to and better than state-of-the-art on two benchmark datasets.Our reliance on no linguistic tool allows us to scale the PSD algorithm to a large corpus and learn sensespecific preposition representations.The crucial abstraction of preposition senses as word representations permits their use in downstream applications-phrasal verb paraphrasing and preposition selection-with new state-ofthe-art results. Hongyu Gong, Jiaqi Mu, Suma Bhat, Pramod Viswanath |
EMNLP | 3 |
| 2018 | Refocusing the lens on engagement in MOOCsabstractMassive open online courses (MOOCs) continue to see increasing enrollment and adoption by universities, although they are still not fully understood and could perhaps be significantly improved. For example, little is known about the relationships between the ways in which students choose to use MOOCs (e.g., sampling lecture videos, discussing topics with fellow students) and their overall level of engagement with the course, although these relationships are likely key to effective course implementation. In this paper we propose a multilevel definition of student engagement with MOOCs and explore the connections between engagement and students' behaviors across five unique courses. We modeled engagement using ordinal penalized logistic regression with the least absolute shrinkage and selection operator (LASSO), and found several predictors of engagement that were consistent across courses. In particular, we found that discussion activities (e.g., viewing forum posts) were positively related to engagement, whereas other types of student behaviors (e.g., attempting quizzes) were consistently related to less engagement with the course. Finally, we discuss implications of unexpected findings that replicated across courses, future work to explore these implications, and relevance of our findings for MOOC course design. R. Wes Crues, Nigel Bosch, Michelle Perry, Lawrence Angrave, Najmuddin Shaik, Suma Bhat |
L@S | 6 |
| 2018 | Embedding Syntax and Semantics of Prepositions via Tensor DecompositionabstractHongyu Gong, Suma Bhat, Pramod Viswanath. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Hongyu Gong, Suma Bhat, Pramod Viswanath |
NAACL-HLT | 2 |
| 2018 | How do Gender, Learning Goals, and Forum Participation Predict Persistence in a Computer Science MOOC?abstractMassive Open Online Courses (MOOCs)—in part, because of their free, flexible, and relatively anonymous nature—may provide a means for helping overcome the large gender gap in Computer Science (CS). This study examines why women and men chose to enroll in a CS MOOC and how this is related to successful behavior in the course by (a) using k-means clustering to explore the reasons why women and men enrolled in this MOOC and then (b) analyzing if these reasons are related to forum participation and, ultimately, persistence in the course. Findings suggest that women and men have different reasons for taking this CS MOOC, and they persist at different rates, an outcome that is moderated by forum participation. R. Wes Crues, Genevieve M. Henricks, Michelle Perry, Suma Bhat, Carolyn J. Anderson, Najmuddin Shaik, Lawrence Angrave |
ACM Trans. Comput. Educ. | 4 |
| 2018 | A comparison of grammatical proficiency measures in the automated assessment of spontaneous speech
Su-Youn Yoon, Suma Bhat |
Speech Commun. | 2 |
| 2017 | Geometry of CompositionalityabstractThis paper proposes a simple test for compositionality (i.e., literal usage) of a word or phrase in a context-specific way. The test is computationally simple, relying on no external resources and only uses a set of trained word vectors. Experiments show that the proposed method is competitive with state of the art and displays high accuracy in context-specific compositionality detection of a variety of natural language phenomena (idiomaticity, sarcasm, metaphor) for different datasets in multiple languages. The key insight is to connect compositionality to a curious geometric property of word embeddings, which is of independent interest. Hongyu Gong, Suma Bhat, Pramod Viswanath |
AAAI | 2 |
| 2017 | MORSE: Semantic-ally Drive-n MORpheme SEgment-erabstractIn this paper we present a novel framework for morpheme segmentation which uses the morpho-syntactic regularities preserved by word representations, in addition to orthographic features, to segment words into morphemes.This framework is the first to consider vocabulary-wide syntactico-semantic information for this task.We also analyze the deficiencies of available benchmarking datasets and introduce our own dataset that was created on the basis of compositionality.We validate our algorithm across different datasets and languages and present new state-of-the-art results. Tarek Sakakini, Suma Bhat, Pramod Viswanath |
ACL (1) | 2 |
| 2017 | Using Computer Agents to Explain Clinical Test Results
Renato Ferreira Leitão Azevedo, Kuangxiao Gu, Yang Zhang 0001, Victor Sadauskas, Tarek Sakakini, Daniel G. Morrow, Mark Hasegawa-Johnson, Thomas S. Huang, Suma Bhat, Ann Willemsen-Dunlap, Donald Halpin, James Graumlich, William Schuh |
AMIA | 9 |
| 2017 | Dr. Babel Fish: A Machine Translator to Simplify Providers' Language
Tarek Sakakini, Renato Ferreira Leitão Azevedo, Victor Sadauskas, Kuangxiao Gu, Yang Zhang 0001, Suma Bhat, Daniel G. Morrow, Mark Hasegawa-Johnson, Thomas S. Huang, Ann Willemsen-Dunlap, Donald Halpin, James Graumlich |
AMIA | 6 |
| 2017 | Learner Affect Through the Looking Glass: Characterization and Detection of Confusion in Online Courses
Ziheng Zeng, Snigdha Chaturvedi, Suma Bhat |
EDM | 3 |
| 2017 | Geometry of Polysemy
Jiaqi Mu, Suma Bhat, Pramod Viswanath |
ICLR (Poster) | 2 |
| 2015 | Seeing the Instructor in Two Video Styles: Preferences and Patterns
Suma Bhat, Phakpoom Chinprutthiwong, Michelle Perry |
EDM | 1 |
| 2015 | Automatic assessment of syntactic complexity for spontaneous speech scoring
Suma Bhat, Su-Youn Yoon |
Speech Commun. | 1 |
| 2014 | Shallow Analysis Based Assessment of Syntactic Complexity for Automated Speech ScoringabstractDesigning measures that capture various aspects of language ability is a central task in the design of systems for automatic scoring of spontaneous speech.In this study, we address a key aspect of language proficiency assessment -syntactic complexity.We propose a novel measure of syntactic complexity for spontaneous speech that shows optimum empirical performance on real world data in multiple ways.First, it is both robust and reliable, producing automatic scores that agree well with human rating compared to the stateof-the-art.Second, the measure makes sense theoretically, both from algorithmic and native language acquisition points of view. Suma Bhat, Huichao Xue, Su-Youn Yoon |
ACL (1) | 1 |
| 2014 | Machine-guided Solution to Mathematical Word Problems
Bussaba Amnueypornsakul, Suma Bhat |
PACLIC | 2 |
| 2013 | Student perceptions of differences in visual communication mode for an online course in engineeringabstractOnline courses have the promise of extending the horizons of today's academic landscape with their cost-effective and convenient model compared to traditional learning environments. Despite the promising nature of the learning model, there continue to be several challenges that hinder learning, one of which is lack of instructor presence. This study aims at understanding the effect of instructor presence on student satisfaction in an online setting of a course in engineering. We conducted a student-centered pilot experiment to assess engineering students' perceptions of two modes of online minilectures: the first, a presentation with the instructor appearing in window, created using an off-the-shelf screen-capture software; the second, a presentation with the instructor overlaid in the slides created using recent visual communication technology that overlays the video of the instructor without any background images or outline boxes. The instructor was the same in both the presentations. Our focus here is on the following factors: 1. Comparing overall student satisfaction after watching the two modes; 2. Comparing the perceived non-verbal immediacy factors of the instructor; and, 3. Comparing the preference of video mode for future online courses. Preliminary results suggest a preference of the video mode with the instructor overlaid over that with the instructor in a box. The effect sizes of the differences in overall satisfaction between the experimental groups and their perceived levels of non-verbal immediacy factors when viewing the online lecture in the two modes are encouraging enough to pursue more longitudinal studies with the set-up. Suma Bhat, Geoffrey L. Herman |
FIE | 1 |
| 2012 | Assessment of ESL Learners' Syntactic Competence Based on Similarity Measures
Su-Youn Yoon, Suma Bhat |
EMNLP-CoNLL | 2 |
| 2010 | A knowledge-rich approach to identifying semantic relations between nominals
Roxana Girju, Brandon Beamer, Alla Rozovskaya, Andrew Fister, Suma Bhat |
Inf. Process. Manag. | 5 |
| 2009 | Knowing the Unseen: Estimating Vocabulary Size over Unseen Samples
Suma Bhat, Richard Sproat |
ACL/IJCNLP | 1 |