VLDB 2026 Research / reviewers in the wild / expert
Natalie Parde
dblp:133/4570 · also Natalie Paige Parde
· DBLP profile ↗
37ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0003-0072-7499ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 7 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Multidisciplinary Summarization of Hospital Stays: Efficient Sentence-Level Clinical Section Categorization
Baris Karacan, Vaibhav Bhargava, Barbara Di Eugenio, Natalie Parde, Mary A. Khetani, Yu-Shan Tseng, Vanessa Barbosa, Julie Vignato, Lindsey Knake, Rajashree Dahal, Emily Spellman, Danielle Hitzel, Janine Petitgout, Kristi Haughey, Amanda Karstens, Brianna Clarahan, Rachel Dawson, Lauren Boyd, Mackenzie Weis, Angie Tipton, Jaewon Bae, Catherine K. Craven, Karen Dunn Lopez, Andrew D. Boyd |
AIME (2) | 4 |
| 2026 | Empathy Speaks in Metaphors: The Empathy-Metaphor Corpus of Figurative Language in Empathetic Text
Gyeongeun Lee, Natalie Parde |
LREC | 2 |
| 2025 | Introduction: Explainability, AI literacy, and language development
Gyu-Ho Shin, Natalie Parde |
Comput. Speech Lang. | 2 |
| 2024 | Pouring Your Heart Out: Investigating the Role of Figurative Language in Online Expressions of EmpathyabstractEmpathy is a social mechanism used to support and strengthen emotional connection with others, including in online communities.However, little is currently known about the nature of these online expressions, nor the particular factors that may lead to their improved detection.In this work, we study the role of a specific and complex subcategory of linguistic phenomena, figurative language, in online expressions of empathy.Our extensive experiments reveal that incorporating features regarding the use of metaphor, idiom, and hyperbole into empathy detection models improves their performance, resulting in impressive maximum F 1 scores of 0.942 and 0.809 for identifying posts without and with empathy, respectively.We hope that the outcomes from this research inspire further work towards improved AI-driven support systems, including empathetic chatbots tailored to specific needs.Additionally, our research may provide avenues for enhancing empathy training and feedback mechanisms for peer supporters in online communities, ultimately elevating the quality of support available to those seeking help.519 Gyeongeun Lee, Christina Wong, Meghan Guo, Natalie Parde |
ACL (1) | 4 |
| 2024 | SLaCAD: A Spoken Language Corpus for Early Alzheimer's Disease DetectionabstractIdentifying early markers of Alzheimer’s disease (AD) trajectory enables intervention in early disease stages when our currently-available interventions are most likely to be beneficial. Research has shown that alterations in speech, as well as linguistic and semantic deviations in spontaneous conversation detected using natural language processing, manifest early in AD prior to some other observed cognitive deficits. Recent studies show that cerebrospinal fluid (CSF) levels serve as useful early biomarkers for identifying early AD, but CSF biomarkers are challenging to collect. A simpler alternative that has seen very rapid development is based on the use of plasma biomarkers as a blood draw is minimally invasive. Associating verbal and nonverbal characteristics from speech data with CSF and plasma biomarkers may open the door to less invasive, more efficient methods for early AD detection. We present SLaCAD, a new dataset to facilitate this process. We describe our data collection procedures, analyze the resulting corpus, and present preliminary findings that relate measures extracted from the audio and transcribed text to clinical diagnoses, CSF levels, and plasma biomarkers. Our findings demonstrate the feasibility of this and indicate that the collected data can be used to improve assessments of early AD. Shahla Farzana, Edoardo Stoppa, Alex Leow, Tamar Gollan, Raeanne Moore, David Salmon, Douglas Galasko, Erin Sundermann, Natalie Parde |
LREC/COLING | 9 |
| 2024 | Towards Comprehensive Language Analysis for Clinically Enriched Spontaneous DialogueabstractContemporary NLP has rapidly progressed from feature-based classification to fine-tuning and prompt-based techniques leveraging large language models. Many of these techniques remain understudied in the context of real-world, clinically enriched spontaneous dialogue. We fill this gap by systematically testing the efficacy and overall performance of a wide variety of NLP techniques ranging from feature-based to in-context learning on transcribed speech collected from patients with bipolar disorder, schizophrenia, and healthy controls taking a focused, clinically-validated language test. We observe impressive utility of a range of feature-based and language modeling techniques, finding that these approaches may provide a plethora of information capable of upholding clinical truths about these subjects. Building upon this, we establish pathways for future research directions in automated detection and understanding of psychiatric conditions. Baris Karacan, Ankit Aich, Avery Quynh, Amy E. Pinkham, Philip D. Harvey, Colin A. Depp, Natalie Parde |
LREC/COLING | 7 |
| 2024 | AcnEmpathize: A Dataset for Understanding Empathy in Dermatology ConversationsabstractEmpathy is critical for effective communication and mental health support, and in many online health communities people anonymously engage in conversations to seek and provide empathetic support. The ability to automatically recognize and detect empathy contributes to the understanding of human emotions expressed in text, therefore advancing natural language understanding across various domains. Existing empathy and mental health-related corpora focus on broader contexts and lack domain specificity, but similarly to other tasks (e.g., learning distinct patterns associated with COVID-19 versus skin allergies in clinical notes), observing empathy within different domains is crucial to providing tailored support. To address this need, we introduce AcnEmpathize, a dataset that captures empathy expressed in acne-related discussions from forum posts focused on its emotional and psychological effects. We find that transformer-based models trained on our dataset demonstrate excellent performance at empathy classification. Our dataset is publicly released to facilitate analysis of domain-specific empathy in online conversations and advance research in this challenging and intriguing domain. Gyeongeun Lee, Natalie Parde |
LREC/COLING | 2 |
| 2024 | CORI: CJKV Benchmark with Romanization Integration - a Step towards Cross-lingual Transfer beyond Textual ScriptsabstractNaively assuming English as a source language may hinder cross-lingual transfer for many languages by failing to consider the importance of language contact. Some languages are more well-connected than others, and target languages can benefit from transferring from closely related languages; for many languages, the set of closely related languages does not include English. In this work, we study the impact of source language for cross-lingual transfer, demonstrating the importance of selecting source languages that have high contact with the target language. We also construct a novel benchmark dataset for close contact Chinese-Japanese-Korean-Vietnamese (CJKV) languages to further encourage in-depth studies of language contact. To comprehensively capture contact between these languages, we propose to integrate Romanized transcription beyond textual scripts via Contrastive Learning objectives, leading to enhanced cross-lingual representations and effective zero-shot cross-lingual transfer. Hoang Nguyen 0006, Ye Liu 0006, Natalie Parde, Eugene Rohrbaugh, Philip S. Yu |
LREC/COLING | 4 |
| 2024 | CareCorpus: A Corpus of Real-World Solution-Focused Caregiver Strategies for Personalized Pediatric Rehabilitation Service DesignabstractIn pediatric rehabilitation services, one intervention approach involves using solution-focused caregiver strategies to support children in their daily life activities. The manual sharing of these strategies is not scalable, warranting need for an automated approach to recognize and select relevant strategies. We introduce CareCorpus, a dataset of 780 real-world strategies written by caregivers. Strategies underwent dual-annotation by three trained annotators according to four established rehabilitation classes (i.e., environment/context, n=325 strategies; a child’s sense of self, n=151 strategies; a child’s preferences, n=104 strategies; and a child’s activity competences, n=62 strategies) and a no-strategy class (n=138 instances) for irrelevant or indeterminate instances. The average percent agreement was 80.18%, with a Cohen’s Kappa of 0.75 across all classes. To validate this dataset, we propose multi-grained classification tasks for detecting and categorizing strategies, and establish new performance benchmarks ranging from F1=0.53-0.79. Our results provide a first step towards a smart option to sort caregiver strategies for use in designing pediatric rehabilitation care plans. This novel, interdisciplinary resource and application is also anticipated to generalize to other pediatric rehabilitation service contexts that target children with developmental need. Mina Valizadeh, Vera C. Kaelin, Mary A. Khetani, Natalie Parde |
LREC/COLING | 4 |
| 2024 | Humanistic Buddhism Corpus: A Challenging Domain-Specific Dataset of English Translations for Classical and Modern ChineseabstractWe introduce the Humanistic Buddhism Corpus (HBC), a dataset containing over 80,000 Chinese-English parallel phrases extracted and translated from publications in the domain of Buddhism. HBC is one of the largest free domain-specific datasets that is publicly available for research, containing text from both classical and modern Chinese. Moreover, since HBC originates from religious texts, many phrases in the dataset contain metaphors and symbolism, and are subject to multiple interpretations. Compared to existing machine translation datasets, HBC presents difficult unique challenges. In this paper, we describe HBC in detail. We evaluate HBC within a machine translation setting, validating its use by establishing performance benchmarks using a Transformer model with different transfer learning setups. Youheng W. Wong, Natalie Parde, Erdem Koyuncu |
LREC/COLING | 2 |
| 2024 | CareCorpus+: Expanding and Augmenting Caregiver Strategy Data to Support Pediatric RehabilitationabstractCaregiver strategy classification in pediatric rehabilitation contexts is strongly motivated by real-world clinical constraints but highly underresourced and seldom studied in natural language processing settings.We introduce a large dataset of 3,062 caregiver strategies in this setting, a five-fold increase over the nearest contemporary dataset.These strategies are manually categorized into clinically established constructs with high agreement (κ=0.68-0.89).We also propose two techniques to further address identified data constraints.First, we manually supplement target task data with relevant public data from online child health forums.Next, we propose a novel data augmentation technique to generate synthetic caregiver strategies with high downstream task utility.Extensive experiments showcase the quality of our dataset.They also establish evidence that both the publicly available data and the synthetic strategies result in large performance gains, with relative F 1 increases of 22.6% and 50.9%, respectively. Shahla Farzana, Ivana Lucero, Vivian Villegas, Vera C. Kaelin, Mary A. Khetani, Natalie Parde |
EMNLP | 6 |
| 2023 | Towards Domain-Agnostic and Domain-Adaptive Dementia Detection from Spoken LanguageabstractHealth-related speech datasets are often small and varied in focus.This makes it difficult to leverage them to effectively support healthcare goals.Robust transfer of linguistic features across different datasets orbiting the same goal carries potential to address this concern.To test this hypothesis, we experiment with domain adaptation (DA) techniques on heterogeneous spoken language data to evaluate generalizability across diverse datasets for a common task: dementia detection.We find that adapted models exhibit better performance across conversational and task-oriented datasets.The feature-augmented DA method achieves a 22% increase in accuracy adapting from a conversational to task-specific dataset compared to a jointly trained baseline.This suggests promising capacity of these techniques to allow for productive use of disparate data for a complex spoken language healthcare task. Shahla Farzana, Natalie Parde |
ACL (1) | 2 |
| 2023 | "Where is history": Toward Designing a Voice Assistant to help Older Adults locate Interface Features quicklyabstractOlder adults often struggle to locate a function quickly in feature-rich user interfaces (UIs). Mobile UIs not only pack a ton of features in a small screen but also get frequent updates to their visual layouts—thereby exacerbating the problem. This paper explores a design solution where users could search for a UI feature using spoken-word queries. We investigated: 1) what type of questions older users ask when facing interaction challenges in unfamiliar scenarios, 2) how those query types compare with younger users’ inquiries, and 3) how older adults use a voice assistant design probe in a Wizard-of-Oz (WoZ) study. Results reveal five query types when verbally articulating interaction issues: validation, directed and undirected informational, navigational, and conceptual. In the WoZ study, older users typically asked for help following a series of non-unique or off-task feature selections (n = 13/15), and in 77% of those instances, they completed the task in the next interaction. Ja Eun Yu, Natalie Parde, Debaleena Chattopadhyay |
CHI | 2 |
| 2023 | Linguistic Cognitive Load Analysis on Dialogues with an Intelligent Virtual Assistant
Mohammad Arvan, Mina Valizadeh, Parian Haghighat, Heejin Jeong, Natalie Parde |
CogSci | 6 |
| 2023 | What Clued the AI Doctor In? On the Influence of Data Source and Quality for Transformer-Based Medical Self-Disclosure DetectionabstractRecognizing medical self-disclosure is important in many healthcare contexts, but it has been under-explored by the NLP community.We conduct a three-pronged investigation of this task.We (1) manually expand and refine the only existing medical self-disclosure corpus, resulting in a new, publicly available dataset of 3,919 social media posts with clinically validated labels and high compatibility with the existing task-specific protocol.We also (2) study the merits of pretraining task domain and text style by comparing Transformer-based models for this task, pretrained from general, medical, and social media sources.Our BERTweet condition outperforms the existing state of the art for this task by a relative F 1 score increase of 16.73%.Finally, we (3) compare data augmentation techniques for this task, to assess the extent to which medical self-disclosure data may be further synthetically expanded.We discover that this task poses many challenges for data augmentation techniques, and we provide an in-depth analysis of identified trends. Mina Valizadeh, Xing Qian, Pardis Ranjbar-Noiey, Cornelia Caragea, Natalie Parde |
EACL | 5 |
| 2023 | A Content-based Skincare Product Recommendation SystemabstractConsumer interest in cosmetics, and particularly skincare products, has surged globally in recent years. Traditional methods of selecting skincare products involve relying on best-sellers or in-store recommendations. However, these approaches are ineffective because they fail to account for individual variations in skin conditions and consumer compatibility. This research aims to design a skincare product recommendation system based on users' skin types and ingredient compositions of products. The proposed method employs content-based filtering to identify chemical components of products and find products with similar ingredient compositions. Unlike many existing sys-tems that require users to input product names, the new system takes into account their desired beauty effects to accommodate those with limited skincare knowledge. The resulting system returns personalized recommendations across multiple product categories. It could contribute to improved compatibility between users and skincare products, enhanced user satisfaction and streamlined selection processes. Gyeongeun Lee, Xunfei Jiang, Natalie Parde |
ICMLA | 3 |
| 2023 | Investigating Reproducibility at Interspeech Conferences: A Longitudinal and Comparative PerspectiveabstractReproducibility is a key aspect for scientific advancement across disciplines, and reducing barriers for open science is a focus area for the theme of Interspeech 2023. Availability of source code is one of the indicators that facilitates reproducibility. However, less is known about the rates of reproducibility at Interspeech conferences in comparison to other conferences in the field. In order to fill this gap, we have surveyed 27,717 papers at seven conferences across speech and language processing disciplines. We find that despite having a close number of accepted papers to the other conferences, Interspeech has up to 40% less source code availability. In addition to reporting the difficulties we have encountered during our research, we also provide recommendations and possible directions to increase reproducibility for further studies. Mohammad Arvan, A. Seza Dogruöz, Natalie Parde |
INTERSPEECH | 3 |
| 2022 | The AI Doctor Is In: A Survey of Task-Oriented Dialogue Systems for Healthcare ApplicationsabstractTask-oriented dialogue systems are increasingly prevalent in healthcare settings, and have been characterized by a diverse range of architectures and objectives.Although these systems have been surveyed in the medical community from a non-technical perspective, a systematic review from a rigorous computational perspective has to date remained noticeably absent.As a result, many important implementation details of healthcare-oriented dialogue systems remain limited or underspecified, slowing the pace of innovation in this area.To fill this gap, we investigated an initial pool of 4070 papers from well-known computer science, natural language processing, and artificial intelligence venues, identifying 70 papers discussing the system-level implementation of task-oriented dialogue systems for healthcare applications.We conducted a comprehensive technical review of these papers, and present our key findings including identified gaps and corresponding recommendations. Mina Valizadeh, Natalie Parde |
ACL (1) | 2 |
| 2022 | Demystifying Neural Fake News via Linguistic Feature-Based InterpretationabstractThe spread of fake news can have devastating ramifications, and recent advancements to neural fake news generators have made it challenging to understand how misinformation generated by these models may best be confronted. We conduct a feature-based study to gain an interpretative understanding of the linguistic attributes that neural fake news generators may most successfully exploit. When comparing models trained on subsets of our features and confronting the models with increasingly advanced neural fake news, we find that stylistic features may be the most robust. We discuss our findings, subsequent analyses, and broader implications in the pages within. Ankit Aich, Souvik Bhattacharya, Natalie Parde |
COLING | 3 |
| 2022 | Reproducibility in Computational Linguistics: Is Source Code Enough?abstractThe availability of source code has been put forward as one of the most critical factors for improving the reproducibility of scientific research.This work studies trends in source code availability at major computational linguistics conferences; namely, ACL, EMNLP, LREC, NAACL, and COLING.We observe positive trends, especially in conferences that actively promote reproducibility.We follow this by conducting a reproducibility study of eight papers published in EMNLP 2021, finding that source code releases leave much to be desired.Moving forward, we suggest all conferences require self-contained artifacts and provide a venue to evaluate such artifacts at the time of publication.Authors can include small-scale experiments and explicit scripts to generate each result to improve the reproducibility of their work. Mohammad Arvan, Luís Pina, Natalie Parde |
EMNLP | 3 |
| 2022 | Telling a Lie: Analyzing the Language of Information and Misinformation during Global Health EventsabstractThe COVID-19 pandemic and other global health events are unfortunately excellent environments for the creation and spread of misinformation, and the language associated with health misinformation may be typified by unique patterns and linguistic markers. Allowing health misinformation to spread unchecked can have devastating ripple effects; however, detecting and stopping its spread requires careful analysis of these linguistic characteristics at scale. We analyze prior investigations focusing on health misinformation, associated datasets, and detection of misinformation during health crises. We also introduce a novel dataset designed for analyzing such phenomena, comprised of 2.8 million news articles and social media posts spanning the early 1900s to the present. Our annotation guidelines result in strong agreement between independent annotators. We describe our methods for collecting this data and follow this with a thorough analysis of the themes and linguistic features that appear in information versus misinformation. Finally, we demonstrate a proof-of-concept misinformation detection task to establish dataset validity, achieving a strong performance benchmark (accuracy = 75%; F1 = 0.7). Ankit Aich, Natalie Parde |
LREC | 2 |
| 2022 | TweetTaglish: A Dataset for Investigating Tagalog-English Code-SwitchingabstractDeploying recent natural language processing innovations to low-resource settings allows for state-of-the-art research findings and applications to be accessed across cultural and linguistic borders. One low-resource setting of increasing interest is code-switching, the phenomenon of combining, swapping, or alternating the use of two or more languages in continuous dialogue. In this paper, we introduce a large dataset (20k+ instances) to facilitate investigation of Tagalog-English code-switching, which has become a popular mode of discourse in Philippine culture. Tagalog is an Austronesian language and former official language of the Philippines spoken by over 23 million people worldwide, but it and Tagalog-English are under-represented in NLP research and practice. We describe our methods for data collection, as well as our labeling procedures. We analyze our resulting dataset, and finally conclude by providing results from a proof-of-concept regression task to establish dataset validity, achieving a strong performance benchmark (R2=0.797-0.909; RMSE=0.068-0.057). Megan Herrera, Ankit Aich, Natalie Parde |
LREC | 3 |
| 2022 | Are Interaction Patterns Helpful for Task-Agnostic Dementia Detection? An Empirical ExplorationabstractDementia often manifests in dialog through specific behaviors such as requesting clarification, communicating repetitive ideas, and stalling, prompting conversational partners to probe or otherwise attempt to elicit information.Dialog act (DA) sequences can have predictive power for dementia detection through their potential to capture these meaningful interaction patterns.However, most existing work in this space relies on content-dependent features, raising questions about their generalizability beyond small reference sets or across different cognitive tasks.In this paper, we adapt an existing DA annotation scheme for two different cognitive tasks present in a popular dementia detection dataset.We show that a DA tagging model leveraging neural sentence embeddings and other information from previous utterances and speaker tags achieves strong performance for both tasks.We also propose content-free interaction features and show that they yield high utility in distinguishing dementia and control subjects across different tasks.Our study provides a step toward better understanding how interaction patterns in spontaneous dialog affect cognitive modeling across different tasks, which carries implications for the design of non-invasive and low-cost cognitive health monitoring tools for use at scale. Shahla Farzana, Natalie Parde |
SIGDIAL | 2 |
| 2021 | Identifying Medical Self-Disclosure in Online CommunitiesabstractMina Valizadeh, Pardis Ranjbar-Noiey, Cornelia Caragea, Natalie Parde. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Mina Valizadeh, Pardis Ranjbar-Noiey, Cornelia Caragea, Natalie Parde |
NAACL-HLT | 4 |
| 2020 | Exploring MMSE Score Prediction Using Verbal and Non-Verbal Cues
Shahla Farzana, Natalie Parde |
INTERSPEECH | 2 |
| 2020 | Modeling Dialogue in Conversational Cognitive Health Screening InterviewsabstractAutomating straightforward clinical tasks can reduce workload for healthcare professionals, increase accessibility for geographically-isolated patients, and alleviate some of the economic burdens associated with healthcare. A variety of preliminary screening procedures are potentially suitable for automation, and one such domain that has remained underexplored to date is that of structured clinical interviews. A task-specific dialogue agent is needed to automate the collection of conversational speech for further (either manual or automated) analysis, and to build such an agent, a dialogue manager must be trained to respond to patient utterances in a manner similar to a human interviewer. To facilitate the development of such an agent, we propose an annotation schema for assigning dialogue act labels to utterances in patient-interviewer conversations collected as part of a clinically-validated cognitive health screening task. We build a labeled corpus using the schema, and show that it is characterized by high inter-annotator agreement. We establish a benchmark dialogue act classification model for the corpus, thereby providing a proof of concept for the proposed annotation schema. The resulting dialogue act corpus is the first such corpus specifically designed to facilitate automated cognitive health screening, and lays the groundwork for future exploration in this area. Shahla Farzana, Mina Valizadeh, Natalie Parde |
LREC | 3 |
| 2019 | AI Meets Austen: Towards Human-Robot Discussions of Literary Metaphor
Natalie Parde, Rodney D. Nielsen |
AIED (2) | 1 |
| 2018 | Reading With Robots: Towards a Human-Robot Book Discussion System for Elderly AdultsabstractAs people age, it is critical that they maintain not only their physical health, but also their cognitive health―for instance, by engaging in cognitive exercise. Recent advancements in AI have uncovered novel ways through which to facilitate such exercise. In this thesis, I propose the first human-robot dialogue system designed specifically to promote cognitive exercise in elderly adults, through discussions about interesting metaphors in books. I describe my work to date, including the development of a new, large corpus and an approach for automatically scoring metaphor novelty. Finally, I outline my plans for incorporating this work into the proposed system. Natalie Parde |
AAAI | 1 |
| 2018 | Exploring the Terrain of Metaphor Novelty: A Regression-Based Approach for Automatically Scoring MetaphorsabstractAutomatically scoring metaphor novelty has been largely unexplored, but could be of benefit to a wide variety of NLP applications. We introduce a large, publicly available metaphor novelty dataset to stimulate research in this area, and propose a regression-based approach to automatically score the novelty of potential metaphors that are expressed as word pairs. We additionally investigate which types of features are most useful for this task, and show that our approach outperforms baseline metaphor novelty scoring and standard metaphor detection approaches on this task. Natalie Parde, Rodney D. Nielsen |
AAAI | 1 |
| 2018 | Automatically Generating Questions about Novel Metaphors in LiteratureabstractThe automatic generation of stimulating questions is crucial to the development of intelligent cognitive exercise applications.We developed an approach that generates appropriate Questioning the Author queries based on novel metaphors in diverse syntactic relations in literature.We show that the generated questions are comparable to human-generated questions in terms of naturalness, sensibility, and depth, and score slightly higher than human-generated questions in terms of clarity.We also show that questions generated about novel metaphors are rated as cognitively deeper than questions generated about non-or conventional metaphors, providing evidence that metaphor novelty can be leveraged to promote cognitive exercise. Natalie Parde, Rodney D. Nielsen |
INLG | 1 |
| 2018 | A Corpus of Metaphor Novelty Scores for Syntactically-Related Word Pairs
Natalie Parde, Rodney D. Nielsen |
LREC | 1 |
| 2017 | Finding Patterns in Noisy Crowds: Regression-based Annotation Aggregation for Crowdsourced DataabstractCrowdsourcing offers a convenient means of obtaining labeled data quickly and inexpensively.However, crowdsourced labels are often noisier than expert-annotated data, making it difficult to aggregate them meaningfully.We present an aggregation approach that learns a regression model from crowdsourced annotations to predict aggregated labels for instances that have no expert adjudications.The predicted labels achieve a correlation of 0.594 with expert labels on our data, outperforming the best alternative aggregation method by 11.9%.Our approach also outperforms the alternatives on third-party datasets. Natalie Parde, Rodney D. Nielsen |
EMNLP | 1 |
| 2017 | Towards a Top-down Policy Engineering Framework for Attribute-based Access ControlabstractAttribute-based access control (ABAC) is a logical access control methodology where authorization to perform a set of operations is based on attributes of the user, the objects being accessed, the environment, and a number of other attribute sources that may be relevant to the current request. Once fully implemented within an enterprise, ABAC promotes information sharing while maintaining control of the information. However, the cost of developing ABAC policies can be a significant obstacle for organizations to migrate from traditional access control models to ABAC. Most organizations have high-level requirement specifications that define security policies and include a set of access control policies. Taking advantage of this rich source of information, we introduce a top-down policy engineering framework for ABAC that aims to automatically extract policies from unrestricted natural language documents and then, we present our methodology to extract policy related information using deep neural networks. We first create an annotated dataset comprised of 2660 sentences from real-world policy documents. We then train a deep recurrent neural network (RNN) to identify sentences containing access control policies (ACP) from irrelevant content. We applied the RNN to our new dataset as well as to five other, smaller datasets that have been employed in prior work on this task, and show that our model outperforms the state-of-the-art and leads to a performance improvement of 5.58% over the previously reported results. Masoud Narouei, Hamed Khanpour, Hassan Takabi, Natalie Parde, Rodney D. Nielsen |
SACMAT | 4 |
| 2015 | "Is It Rectangular?" Using I Spy as an Interactive, Game-Based Approach to Multimodal Robot LearningabstractTraining robots about the objects in their environment requires a multimodal correlation of features extracted from visual and linguistic sources. This work abstracts the task of collecting multimodal training data for object and feature learning by encapsulating it in an interactive game, I Spy, played between human players and robots. It introduces the concept of the game, briefly describes its methodology, and finally presents an evaluation of the game's performance and its appeal to human players. Natalie Parde, Michalis Papakostas, Konstantinos Tsiakas, Rodney D. Nielsen |
AAAI | 1 |
| 2015 | Grounding the Meaning of Words through Vision and Interactive Gameplay
Natalie Parde, Adam Hair, Michalis Papakostas, Konstantinos Tsiakas, Maria Dagioglou, Vangelis Karkaletsis, Rodney D. Nielsen |
IJCAI | 1 |
| 2013 | Data-Driven Mapping Using Local PatternsabstractThe problem of mapping a data flow graph onto a reconfigurable architecture has been difficult to solve quickly and optimally. Anytime algorithms have the potential to meet both goals by generating a good solution quickly and improving that solution over time, but they have not been shown to be practical for mapping. The key insight into this paper is that mapping algorithms based on search trees can be accelerated using a database of examples of high quality mappings. The depth of the search tree is reduced by placing patterns of nodes rather than single nodes at each level. The branching factor is reduced by placing patterns only in arrangements present in a dictionary constructed from examples. We present two anytime algorithms that make use of patterns and dictionaries: Anytime A*and Anytime Multiline Tree Rollup. We compare these algorithms to simulated annealing and to results from human mappers playing the online game UNTANGLED. The anytime algorithms outperform simulated annealing and the best game players in the majority of cases, and the combined results from all algorithms provide an informative comparison between architecture choices. Gayatri Mehta, Krunalkumar Patel, Natalie Parde, Nancy S. Pollard |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2013 | UNTANGLED: A Game Environment for Discovery of Creative Mapping StrategiesabstractThe problem of creating efficient mappings of dataflow graphs onto specific architectures (i.e., solving theplace and routeproblem) is incredibly challenging. The difficulty is especially acute in the area of Coarse-Grained Reconfigurable Architectures (CGRAs) to the extent that solving the mapping problem may remove a significant bottleneck to adoption. We believe that the next generation of mapping algorithms will exhibit pattern recognition, the ability to learn from experience, and identification of creative solutions, all of which are human characteristics. This manuscript describes our game UNTANGLED, developed and fine-tuned over the course of a year to allow us to capture and analyze human mapping strategies. It also describes our results to date. We find that the mapping problem can be crowdsourced very effectively, that players can outperform existing algorithms, and that successful player strategies share many elements in common. Based on our observations and analysis, we make concrete recommendations for future research directions for mapping onto CGRAs. Gayatri Mehta, Carson Crawford, Xiaozhong Luo, Natalie Parde, Krunalkumar Patel, Brandon Rodgers, Anil Kumar Sistla, Anil Yadav, Marc Reisner |
ACM Trans. Reconfigurable Technol. Syst. | 4 |