EDBT 2026 Demo / reviewers in the wild / expert
Yunyao Li 0001
dblp:60/2319
· DBLP profile ↗
84ranked-venue papers
16as first author
32since 2021 · last 2026
0000-0001-8433-8719ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 4 first-author · 24 since 2021Databases, data management, data science and information retrieval · 31 · 13 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 9 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PAGER: Proactive Monitoring Agent for Enterprise AI AssistantabstractWe present a Proactive Monitoring Agent designed for large-scale customer data platforms, such as Adobe Experience Platform (AEP), to predict and prevent workflow disruptions before they impact business operations. Unlike existing reactive solutions that assist engineers only after failures occur, our agent anticipates potential failures across multiple workflow stages, explains its predictions in natural language, and interacts with customer support engineers through a conversational interface. The system integrates a machine learning-based Prediction Module, Knowledge Graph APIs for contextual data access, and a Query Processor that powers an interactive Q&A experience, enabling timely and actionable insights to minimize operational risks and maximize business continuity. Sujan Dutta, Junior Francisco Garcia Ayala, Pranav Umakant Pujar, Sai Sree Harsha, Dan Luo 0004, Nikhil Vasudeva, Bikas Saha, Pritom Baruah, Yunyao Li 0001 |
AAAI | 9 |
| 2026 | Feedback by Design: Understanding and Overcoming User Feedback Barriers in Conversational AgentsabstractHigh-quality feedback is essential for effective human–AI interaction. It bridges knowledge gaps, corrects digressions, and shapes system behavior; both during interaction and throughout model development. Yet despite its importance, human feedback to AI is often infrequent and low quality. This gap motivates a critical examination of human feedback during interactions with AIs. To understand and overcome the challenges preventing users from giving high-quality feedback, we conducted two studies examining feedback dynamics between humans and conversational agents (CAs). Our formative study, through the lens of Grice’s maxims, identified four Feedback Barriers—Common Ground, Verifiability, Communication, and Informativeness—that prevent high-quality feedback by users. Building on these findings, we derive three design desiderata and show that systems incorporating scaffolds aligned with these desiderata enabled users to provide higher-quality feedback. Finally, we detail a call for action to the broader AI community for advances in Large Language Models capabilities to overcome Feedback Barriers. Zheng Zhang 0043, Namita Krishnan, Ziang Xiao, Yunyao Li 0001 |
CHI | 7 |
| 2026 | Enhancing Enterprise Assistant Responses with Rich Multimodal Artifacts from Product Documentation
Sohan Patnaik, Sai Sree Harsha, H. S. V. N. S. Kowndinya Renduchintala, Milan Aggarwal, Sumit Bhatia, Yunyao Li 0001 |
SIGIR | 6 |
| 2025 | Rewind and Render: Towards Factually Accurate Text-to-Video Generation with Distilled Knowledge RetrievalabstractText-to-Video (T2V) models, despite recent advancements, struggle with factual accuracy, especially for knowledge-dense content. We introduce FACT-V (Factual Accuracy in Content Translation to Video), a system integrating multi-source knowledge retrieval into T2V pipelines. FACT-V offers two key benefits: i) improved factual accuracy of generated videos through dynamically retrieved information, and ii) increased interpretability by providing users with the augmented prompt information. A preliminary evaluation demonstrates the potential of knowledge-augmented approaches in improving the accuracy and reliability of T2V systems, particularly for entity-specific or time-sensitive prompts. Arjun Chandra, Yunyao Li 0001, Simone Conia |
AAAI | 4 |
| 2025 | Evaluation and Incident Prevention in an Enterprise AI AssistantabstractEnterprise AI Assistants are increasingly deployed in domains where accuracy is paramount, making each erroneous output a potentially significant incident. This paper presents a comprehensive framework for monitoring, benchmarking, and continuously improving such complex, multi-component systems under active development by multiple teams. Our approach encompasses three key elements: (1) a hierarchical ``severity'' framework for incident detection that identifies and categorizes errors while attributing component-specific error rates, facilitating targeted improvements; (2) a scalable and principled methodology for benchmark construction, evaluation, and deployment, designed to accommodate multiple development teams, mitigate overfitting risks, and assess the downstream impact of system modifications; and (3) a continual improvement strategy leveraging multidimensional evaluation, enabling the identification and implementation of diverse enhancement opportunities. By adopting this holistic framework, organizations can systematically enhance the reliability and performance of their AI Assistants, ensuring their efficacy in critical enterprise environments. We conclude by discussing how this multifaceted evaluation approach opens avenues for various classes of enhancements, paving the way for more robust and trustworthy AI systems. Akash Maharaj, David T. Arbour, Uttaran Bhattacharya, Anup B. Rao, Austin Zane, Avi Feller, Kun Qian 0002, Yunyao Li 0001 |
AAAI | 9 |
| 2025 | ECLAIR: Enhanced Clarification for Interactive Responses in an Enterprise AI AssistantabstractLarge language models (LLMs) have shown remarkable progress in understanding and generating natural language across various applications. However, they often struggle with resolving ambiguities in real-world, enterprise-level interactions, where context and domain-specific knowledge play a crucial role. In this demonstration, we introduce ECLAIR (Enhanced CLArification for Interactive Responses), a multi-agent framework for interactive disambiguation. ECLAIR enhances ambiguous user query clarification through an interactive process where custom agents are defined, ambiguity reasoning is conducted by the agents, clarification questions are generated, and user feedback is leveraged to refine the final response. When tested on real-world customer data, ECLAIR demonstrates significant improvements in clarification question generation compared to standard few-shot methods. John Murzaku, Zifan Liu, Vaishnavi Muppala, Md. Mehrab Tanjim, Xiang Chen 0010, Yunyao Li 0001 |
AAAI | 6 |
| 2025 | ECLAIR: Enhanced Clarification for Interactive ResponsesabstractWe present ECLAIR (Enhanced CLArification for Interactive Responses), a novel unified and end-to-end framework for interactive disambiguation in enterprise AI assistants. ECLAIR generates clarification questions for ambiguous user queries and resolves ambiguity based on the user's response. We introduce a generalized architecture capable of integrating ambiguity information from multiple downstream agents, enhancing context-awareness in resolving ambiguities and allowing enterprise specific definition of agents. We further define agents within our system that provide domain-specific grounding information. We conduct experiments comparing ECLAIR to few-shot prompting techniques and demonstrate ECLAIR's superior performance in clarification question generation and ambiguity resolution. John Murzaku, Zifan Liu, Md. Mehrab Tanjim, Vaishnavi Muppala, Xiang Chen 0010, Yunyao Li 0001 |
AAAI | 6 |
| 2025 | KG-TRICK: Unifying Textual and Relational Information Completion of Knowledge for Multilingual Knowledge GraphsabstractMultilingual knowledge graphs (KGs) provide high-quality relational and textual information for various NLP applications, but they are often incomplete, especially in non-English languages. Previous research has shown that combining information from KGs in different languages aids either Knowledge Graph Completion (KGC), the task of predicting missing relations between entities, or Knowledge Graph Enhancement (KGE), the task of predicting missing textual information for entities. Although previous efforts have considered KGC and KGE as independent tasks, we hypothesize that they are interdependent and mutually beneficial. To this end, we introduce KG-TRICK, a novel sequence-to-sequence framework that unifies the tasks of textual and relational information completion for multilingual KGs. KG-TRICK demonstrates that: i) it is possible to unify the tasks of KGC and KGE into a single framework, and ii) combining textual information from multiple languages is beneficial to improve the completeness of a KG. As part of our contributions, we also introduce WikiKGE10++, the largest manually-curated benchmark for textual information completion of KGs, which features over 25,000 entities across 10 diverse languages. Zelin Zhou, Simone Conia, Shenglei Huang, Umar Farooq Minhas, Saloni Potdar, Henry Xiao, Yunyao Li 0001 |
COLING | 9 |
| 2025 | FISQL: Enhancing Text-to-SQL Systems with Rich Interactive Feedback
Rakesh Menon, Kun Qian 0002, Ishika Joshi, Daniel Pandyan, Yunyao Li 0001 |
EDBT | 7 |
| 2025 | CoPrompter: User-Centric Evaluation of LLM Instruction Alignment for Improved Prompt Engineering
Ishika Joshi, Simra Shahid, Shreeya Venneti, Manushree Vasu, Yantao Zheng, Yunyao Li 0001, Balaji Krishnamurthy, Gromit Yeuk-Yin Chan |
IUI | 6 |
| 2025 | Text-to-SQL Domain Adaptation via Human-LLM Collaborative Data AnnotationabstractText-to-SQL models, which parse natural language (NL) questions to executable SQL queries, are increasingly adopted in real-world applications. However, deploying such models in the real world often requires adapting them to the highly specialized database schemas used in specific applications. We find that existing text-to-SQL models experience significant performance drops when applied to new schemas, primarily due to the lack of domain-specific data for fine-tuning. This data scarcity also limits the ability to effectively evaluate model performance in new domains. Continuously obtaining high-quality text-to-SQL data for evolving schemas is prohibitively expensive in real-world scenarios. To bridge this gap, we propose SQLsynth, a human-in-the-loop text-to-SQL data annotation system. SQLsynth streamlines the creation of high-quality text-to-SQL datasets through human-LLM collaboration in a structured workflow. A within-subjects user study comparing SQLsynth with manual annotation and ChatGPT shows that SQLsynth significantly accelerates text-to-SQL data annotation, reduces cognitive load, and produces datasets that are more accurate, natural, and diverse. Our code is available at https://github.com/magic-YuanTian/SQLsynth. Fei Wu 0029, Tung Mai, Kun Qian 0002, Siddhartha Sahai, Tianyi Zhang 0001, Yunyao Li 0001 |
IUI | 8 |
| 2025 | EVOSCHEMA: TOWARDS TEXT-TO-SQL ROBUSTNESS AGAINST SCHEMA EVOLUTIONabstractNeural text-to-SQL models, which translate natural language questions (NLQs) into SQL queries given a database schema, have achieved remarkable performance. However, database schemas frequently evolve to meet new requirements. Such schema evolution often leads to performance degradation for models trained on static schemas. Existing work either mainly focuses on simply paraphrasing some syntactic or semantic mappings among NLQ, DB and SQL, or lacks a comprehensive and controllable way to investigate the model robustness issue under the schema evolution, which is insufficient when facing the increasingly complex and rich database schema changes in reality, especially in the LLM era. To address the challenges posed by schema evolution, we present EvoSchema, a comprehensive benchmark designed to assess and enhance the robustness of text-to-SQL systems under real-world schema changes. EvoSchema introduces a novel schema evolution taxonomy, encompassing ten perturbation types across column-level and table-level modifications, systematically simulating the dynamic nature of database schemas. Through EvoSchema, we conduct an in-depth evaluation spanning different open-source and closed-source LLMs, revealing that table-level perturbations have a significantly greater impact on model performance compared to column-level changes. Furthermore, EvoSchema inspires the development of more resilient text-to-SQL systems, in terms of both model training and database design. The models trained on EvoSchema's diverse schema designs can force the model to distinguish the schema difference for the same questions to avoid learning spurious patterns, which demonstrate remarkable robustness compared to those trained on unperturbed data on average. This benchmark offers valuable insights into model behavior and a path forward for designing systems capable of thriving in dynamic, real-world environments. Tianshu Zhang 0001, Kun Qian 0002, Siddhartha Sahai, Shaddy Garg, Huan Sun 0001, Yunyao Li 0001 |
Proc. VLDB Endow. | 7 |
| 2025 | Introduction to the Special Issue on Human-Centric Generative AIabstractGenerative AI increasingly reshapes how people engage with interactive systems. It now plays a vital role in designing, studying, and refining human-centered methods that let individuals interact and collaborate with AI, strengthening their agency and control. This special issue highlights the human role in Generative AI and seeks approaches that equip diverse stakeholders across socio-technical contexts to understand, direct, and steer these systems while enabling responsible innovation. We publish in this special issue original research on new interaction techniques that integrate human input into Generative AI’s continual development, studies of interaction paradigms that support more effective human–AI collaboration, and work that deepens understanding of model capabilities. Thus, we aim to build a research community around Human-Centric GenAI that empowers people to actively shape systems in line with their values, needs, and expectations. Yunyao Li 0001, Mary Lou Maher, Ziang Xiao |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2024 | Enhancing Machine Translation Experiences with Multilingual Knowledge GraphsabstractTranslating entity names, especially when a literal translation is not correct, poses a significant challenge. Although Machine Translation (MT) systems have achieved impressive results, they still struggle to translate cultural nuances and language-specific context. In this work, we show that the integration of multilingual knowledge graphs into MT systems can address this problem and bring two significant benefits: i) improving the translation of utterances that contain entities by leveraging their human-curated aliases from a multilingual knowledge graph, and, ii) increasing the interpretability of the translation process by providing the user with information from the knowledge graph. Simone Conia, Umar Farooq Minhas, Yunyao Li 0001 |
AAAI | 5 |
| 2024 | Construction of Paired Knowledge Graph - Text Datasets Informed by Cyclic EvaluationabstractDatasets that pair Knowledge Graphs (KG) and text together (KG-T) can be used to train forward and reverse neural models that generate text from KG and vice versa. However models trained on datasets where KG and text pairs are not equivalent can suffer from more hallucination and poorer recall. In this paper, we verify this empirically by generating datasets with different levels of noise and find that noisier datasets do indeed lead to more hallucination. We argue that the ability of forward and reverse models trained on a dataset to cyclically regenerate source KG or text is a proxy for the equivalence between the KG and the text in the dataset. Using cyclic evaluation we find that manually created WebNLG is much better than automatically created TeKGen and T-REx. Informed by these observations, we construct a new, improved dataset called LAGRANGE using heuristics meant to improve equivalence between KG and text and show the impact of each of the heuristics on cyclic evaluation. We also construct two synthetic datasets using large language models (LLMs), and observe that these are conducive to models that perform significantly well on cyclic generation of text, but less so on cyclic generation of KGs, probably because of a lack of a consistent underlying ontology. Ali Mousavi 0003, Xin Zhan, He Bai 0002, Peng Shi 0010, Theodoros Rekatsinas, Benjamin Han, Yunyao Li 0001, Jeffrey Pound, Joshua M. Susskind, Natalie Schluter, Ihab F. Ilyas, Navdeep Jaitly |
LREC/COLING | 7 |
| 2024 | StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children's Story-Based LearningabstractJiaju Chen, Yuxuan Lu, Shao Zhang, Bingsheng Yao, Yuanzhe Dong, Ying Xu, Yunyao Li, Qianwen Wang, Dakuo Wang, Yuling Sun. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Jiaju Chen, Yuxuan Lu 0003, Shao Zhang, Bingsheng Yao, Yuanzhe Dong, Yunyao Li 0001, Dakuo Wang, Yuling Sun |
EMNLP | 7 |
| 2024 | Towards Cross-Cultural Machine Translation with Retrieval-Augmented Generation from Multilingual Knowledge GraphsabstractTranslating text that contains entity names is a challenging task, as cultural-related references can vary significantly across languages.These variations may also be caused by transcreation, an adaptation process that entails more than transliteration and word-for-word translation.In this paper, we address the problem of cross-cultural translation on two fronts: (i) we introduce XC-Translate, the first large-scale, manually-created benchmark for machine translation that focuses on text that contains potentially culturally-nuanced entity names, and (ii) we propose KG-MT, a novel end-to-end method to integrate information from a multilingual knowledge graph into a neural machine translation model by leveraging a dense retrieval mechanism.Our experiments and analyses show that current machine translation systems and large language models still struggle to translate texts containing entity names, whereas KG-MT outperforms state-of-the-art approaches by a large margin, obtaining a 129% and 62% relative improvement compared to NLLB-200 and GPT-4, respectively. Simone Conia, Umar Farooq Minhas, Saloni Potdar, Yunyao Li 0001 |
EMNLP | 6 |
| 2024 | AGRaME: Any-Granularity Ranking with Multi-Vector EmbeddingsabstractRanking is a fundamental problem in search, however, existing ranking algorithms usually restrict the granularity of ranking to full passages or require a specific dense index for each desired level of granularity.Such lack of flexibility in granularity negatively affects many applications that can benefit from more granular ranking, such as sentence-level ranking for open-domain QA, or proposition-level ranking for attribution.In this work, we introduce the idea of any-granularity ranking 1 which leverages multi-vector embeddings to rank at varying levels of granularity while maintaining encoding at a single (coarser) level of granularity.We propose a multi-granular contrastive loss for training multi-vector approaches and validate its utility with both sentences and propositions as ranking units.Finally, we demonstrate the application of proposition-level ranking to post-hoc citation addition in retrievalaugmented generation, surpassing the performance of prompt-driven citation generation. Revanth Gangi Reddy, Omar Attia, Yunyao Li 0001, Heng Ji 0001, Saloni Potdar |
EMNLP | 3 |
| 2024 | Entity Disambiguation via Fusion Entity DecodingabstractJunxiong Wang, Ali Mousavi, Omar Attia, Ronak Pradeep, Saloni Potdar, Alexander Rush, Umar Farooq Minhas, Yunyao Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Junxiong Wang, Ali Mousavi 0003, Omar Attia, Ronak Pradeep, Saloni Potdar, Alexander M. Rush, Umar Farooq Minhas, Yunyao Li 0001 |
NAACL-HLT | 8 |
| 2023 | When to Use What: An In-Depth Comparative Empirical Analysis of OpenIE Systems for Downstream ApplicationsabstractOpen Information Extraction (OpenIE) has been used in the pipelines of various NLP tasks.Unfortunately, there is no clear consensus on which models to use for which tasks.Muddying things further is the lack of comparisons that take differing training sets into account.In this paper, we present an application-focused empirical survey of neural OpenIE models, training sets, and benchmarks in an effort to help users choose the most suitable OpenIE systems for their applications.We find that the different assumptions made by different models and datasets have a statistically significant effect on performance, making it important to choose the most appropriate model for one's applications.We demonstrate the applicability of our recommendations on a downstream Complex QA application. Kevin Pei, Ishan Jindal, Kevin Chen-Chuan Chang, ChengXiang Zhai, Yunyao Li 0001 |
ACL (1) | 5 |
| 2023 | Increasing Coverage and Precision of Textual Information in Multilingual Knowledge GraphsabstractRecent work in Natural Language Processing and Computer Vision has been using textual information -e.g., entity names and descriptions -available in knowledge graphs to ground neural models to high-quality structured data.However, when it comes to non-English languages, the quantity and quality of textual information are comparatively scarce.To address this issue, we introduce the novel task of automatic Knowledge Graph Enhancement (KGE) and perform a thorough investigation on bridging the gap in both the quantity and quality of textual information between English and non-English languages.More specifically, we: i) bring to light the problem of increasing multilingual coverage and precision of entity names and descriptions in Wikidata; ii) demonstrate that state-of-the-art methods, namely, Machine Translation (MT), Web Search (WS), and Large Language Models (LLMs), struggle with this task; iii) present M-NTA, a novel unsupervised approach that combines MT, WS, and LLMs to generate high-quality textual information; and, iv) study the impact of increasing multilingual coverage and precision of non-English textual information in Entity Linking, Knowledge Graph Completion, and Question Answering.As part of our effort towards better multilingual knowledge graphs, we also introduce WikiKGE-10, the first human-curated benchmark to evaluate KGE approaches in 10 languages across 7 language families. Simone Conia, Umar Farooq Minhas, Ihab F. Ilyas, Yunyao Li 0001 |
EMNLP | 6 |
| 2023 | Powering an AI Chatbot with Expert Sourcing to Support Credible Health Information AccessabstractDuring a public health crisis like the COVID-19 pandemic, a credible and easy-to-access information portal is highly desirable. It helps with disease prevention, public health planning, and misinformation mitigation. However, creating such an information portal is challenging because 1) domain expertise is required to identify and curate credible and intelligible content, 2) the information needs to be updated promptly in response to the fast-changing environment, and 3) the information should be easily accessible by the general public; which is particularly difficult when most people do not have the domain expertise about the crisis. In this paper, we presented an expert-sourcing framework and created Jennifer, an AI chatbot, which serves as a credible and easy-to-access information portal for individuals during the COVID-19 pandemic. Jennifer was created by a team of over 150 scientists and health professionals around the world, deployed in the real world and answered thousands of user questions about COVID-19. We evaluated Jennifer from two key stakeholders’ perspectives, expert volunteers and information seekers. We first interviewed experts who contributed to the collaborative creation of Jennifer to learn about the challenges in the process and opportunities for future improvement. We then conducted an online experiment that examined Jennifer’s effectiveness in supporting information seekers in locating COVID-19 information and gaining their trust. We share the key lessons learned and discuss design implications for building expert-sourced and AI-powered information portals, along with the risks and opportunities of misinformation mitigation and beyond. Ziang Xiao, Qingzi Vera Liao, Michelle X. Zhou, Tyrone Grandison, Yunyao Li 0001 |
IUI | 5 |
| 2022 | A Simulation-Based Evaluation Framework for Interactive AI Systems and Its ApplicationabstractInteractive AI (IAI) systems are increasingly popular as the human-centered AI design paradigm is gaining strong traction. However, evaluating IAI systems, a key step in building such systems, is particularly challenging, as their output highly depends on the performed user actions. Developers often have to rely on limited and mostly qualitative data from ad-hoc user testing to assess and improve their systems. In this paper, we present InteractEva; a systematic evaluation framework for IAI systems. We also describe how we have applied InteractEva to evaluate a commercial IAI system, leading to both quality improvements and better data-driven design decisions. Maeda F. Hanafi, Yannis Katsis, Martín Santillán Cooper, Yunyao Li 0001 |
AAAI | 4 |
| 2022 | InteractEva: A Simulation-Based Evaluation Framework for Interactive AI SystemsabstractEvaluating interactive AI (IAI) systems is a challenging task, as their output highly depends on the performed user actions. As a result, developers often depend on limited and mostly qualitative data derived from user testing to improve their systems. In this paper, we present InteractEva; a systematic evaluation framework for IAI systems. InteractEva employs (a) a user simulation backend to test the system against different use cases and user interactions at scale with (b) an interactive frontend allowing developers to perform important quantitative evaluation tasks, including acquiring a performance overview, performing error analysis, and conducting what-if studies. The framework has supported the evaluation and improvement of an industrial IAI text extraction system, results of which will be presented during our demonstration. Yannis Katsis, Maeda F. Hanafi, Martín Santillán Cooper, Yunyao Li 0001 |
AAAI | 4 |
| 2022 | Universal Proposition Bank 2.0abstractSemantic role labeling (SRL) represents the meaning of a sentence in the form of predicate-argument structures. Such shallow semantic analysis is helpful in a wide range of downstream NLP tasks and real-world applications. As treebanks enabled the development of powerful syntactic parsers, the accurate predicate-argument analysis demands training data in the form of propbanks. Unfortunately, most languages simply do not have corresponding propbanks due to the high cost required to construct such resources. To overcome such challenges, Universal Proposition Bank 1.0 (UP1.0) was released in 2017, with high-quality propbank data generated via a two-stage method exploiting monolingual SRL and multilingual parallel data. In this paper, we introduce Universal Proposition Bank 2.0 (UP2.0), with significant enhancements over UP1.0: (1) propbanks with higher quality by using a state-of-the-art monolingual SRL and improved auto-generation of annotations; (2) expanded language coverage (from 7 to 9 languages); (3) span annotation for the decoupling of syntactic analysis; and (4) Gold data for a subset of the languages. We also share our experimental results that confirm the significant quality improvements of the generated propbanks. In addition, we present a comprehensive experimental evaluation on how different implementation choices impact the quality of the resulting data. We release these resources to the research community and hope to encourage more research on cross-lingual SRL. Ishan Jindal, Alexandre Rademaker, Michal Ulewicz, Ha Linh, Khoi-Nguyen Tran, Huaiyu Zhu 0001, Yunyao Li 0001 |
LREC | 8 |
| 2022 | Label Definitions Improve Semantic Role LabelingabstractArgument classification is at the core of Semantic Role Labeling.Given a sentence and the predicate, a semantic role label is assigned to each argument of the predicate.While semantic roles come with meaningful definitions, existing work has treated them as symbolic.Learning symbolic labels usually requires ample training data, which is frequently unavailable due to the cost of annotation.We instead propose to retrieve and leverage the definitions of these labels from the annotation guidelines.For example, the verb predicate "work" has arguments defined as "worker", "job", "employer", etc.Our model achieves state-of-theart performance on the CoNLL09 English SRL dataset injected with label definitions given the predicate senses.The performance improvement is even more pronounced in low-resource settings when training data is scarce. 1 Li Zhang 0039, Ishan Jindal, Yunyao Li 0001 |
NAACL-HLT | 3 |
| 2021 | Who needs to know what, when?: Broadening the Explainable AI (XAI) Design Space by Looking at Explanations Across the AI LifecycleabstractThe interpretability or explainability of AI systems (XAI) has been a topic gaining renewed attention in recent years across AI and HCI communities. Recent work has drawn attention to the emergent explainability requirements of in situ, applied projects, yet further exploratory work is needed to more fully understand this space. This paper investigates applied AI projects and reports on a qualitative interview study of individuals working on AI projects at a large technology and consulting company. Presenting an empirical understanding of the range of stakeholders in industrial AI projects, this paper also draws out the emergent explainability practices that arise as these projects unfold, highlighting the range of explanation audiences (who), as well as how their explainability needs evolve across the AI project lifecycle (when). We discuss the importance of adopting a sociotechnical lens in designing AI systems, noting how the “AI lifecycle” can serve as a design metaphor to further the XAI design field. Shipi Dhanorkar, Christine T. Wolf, Kun Qian 0002, Anbang Xu, Lucian Popa 0001, Yunyao Li 0001 |
Conference on Designing Interactive Systems | 6 |
| 2021 | AutoText: An End-to-End AutoAI Framework for TextabstractBuilding models for natural language processing (NLP) tasks remains a daunting task for many, requiring significant technical expertise, efforts, and resources. In this demonstration, we present AutoText, an end-to-end AutoAI framework for text, to lower the barrier of entry in building NLP models. AutoText combines state-of-the-art AutoAI optimization techniques and learning algorithms for NLP tasks into a single extensible framework. Through its simple, yet powerful UI, non-AI experts (e.g., domain experts) can quickly generate performant NLP models with support to both control (e.g., via specifying constraints) and understand learned models. Arunima Chaudhary, Alayt Issak, Kiran Kate, Yannis Katsis, Abel N. Valente, Dakuo Wang, Alexandre V. Evfimievski, Sairam Gurajada, Ban Kawas, Cristiano Malossi, Lucian Popa 0001, Tejaswini Pedapati, Horst Samulowitz, Martin Wistuba, Yunyao Li 0001 |
AAAI | 15 |
| 2021 | KAAPA: Knowledge Aware Answers from PDF AnalysisabstractWe present KaaPa (Knowledge Aware Answers from Pdf Analysis), an integrated solution for machine reading comprehension over both text and tables extracted from PDFs. KaaPa enables interactive question refinement using facets generated from an automatically induced Knowledge Graph. In addition it provides a concise summary of the supporting evidence for the provided answers by aggregating information across multiple sources. KaaPa can be applied consistently to any collection of documents in English with zero domain adaptation effort. We showcase the use of KaaPa for QA on scientific literature using the COVID-19 Open Research Dataset. Nicolas R. Fauceglia, Mustafa Canim, Alfio Massimiliano Gliozzo, Jennifer J. Liang, Nancy Xin Ru Wang, Douglas Burdick, Nandana Mihindukulasooriya, Vittorio Castelli, Guy Feigenblat, David Konopnicki, Yannis Katsis, Radu Florian, Yunyao Li 0001, Salim Roukos, Avirup Sil |
AAAI | 13 |
| 2021 | LNN-EL: A Neuro-Symbolic Approach to Short-text Entity LinkingabstractHang Jiang, Sairam Gurajada, Qiuhao Lu, Sumit Neelam, Lucian Popa, Prithviraj Sen, Yunyao Li, Alexander Gray. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Sairam Gurajada, Qiuhao Lu, Sumit Neelam, Lucian Popa 0001, Prithviraj Sen, Yunyao Li 0001, Alexander G. Gray |
ACL/IJCNLP (1) | 7 |
| 2021 | Explainability for Natural Language ProcessingabstractThis lecture-style tutorial, which mixes in an interactive literature browsing component, is intended for the many researchers and practitioners working with text data and on applications of natural language processing (NLP) in data science and knowledge discovery. The focus of the tutorial is on the issues of transparency and interpretability as they relate to building models for text and their applications to knowledge discovery. As black-box models have gained popularity for a broad range of tasks in recent years, both the research and industry communities have begun developing new techniques to render them more transparent and interpretable. Reporting from an interdisciplinary team of social science, human-computer interaction (HCI), and NLP/knowledge management researchers, our tutorial has two components: an introduction to explainable AI (XAI) in the NLP domain and a review of the state-of-the-art research; and findings from a qualitative interview study of individuals working on real-world NLP projects as they are applied to various knowledge extraction and discovery at a large, multinational technology and consulting corporation. The first component will introduce core concepts related to explainability in NLP. Then, we will discuss explainability for NLP tasks and report on a systematic literature review of the state-of-the-art literature in AI, NLP and HCI conferences. The second component reports on our qualitative interview study, which identifies practical challenges and concerns that arise in real-world development projects that require the modeling and understanding of text data. Marina Danilevsky, Shipi Dhanorkar, Yunyao Li 0001, Lucian Popa 0001, Kun Qian 0002, Anbang Xu |
KDD | 3 |
| 2021 | Data Science with Human in the LoopabstractThe aim of this workshop is to stimulate research on human-computer interaction challenges in data science. We invite researchers and practitioners interested in understanding how to optimize the human-computer cooperation and how to minimize human effort along the data science pipeline in a wide range of data science tasks and real-life applications. One over-arching challenge is to raise the level of abstraction of human-computer interaction to more sophisticated interaction models that better reflect a human's conceptual model and understanding. This workshop will bring together the interdisciplinary researchers from academia, research labs and practice to share, exchange, learn, and develop preliminary results, new concepts, ideas, principles, and methodologies on understanding and improving human-computer interaction for cost-effective development of data science models and for knowledge discovery. We expect the workshop to help develop and grow a strong community of researchers who are interested in this topic, and yield future collaborations and scientific exchanges across the relevant areas of data mining, machine learning, data and knowledge management, human-machine interaction, and user interfaces. Eduard C. Dragut, Yunyao Li 0001, Lucian Popa 0001, Slobodan Vucetic |
KDD | 2 |
| 2020 | PARTNER: Human-in-the-Loop Entity Name Understanding with Deep LearningabstractEntity name disambiguation is an important task for many text-based AI tasks. Entity names usually have internal semantic structures that are useful for resolving different variations of the same entity. We present, PARTNER, a deep learning-based interactive system for entity name understanding. Powered by effective active learning and weak supervision, PARTNER can learn deep learning-based models for identifying entity name structure with low human effort. PARTNER also allows the user to design complex normalization and variant generation functions without coding skills. Kun Qian 0002, Poornima Chozhiyath Raman, Yunyao Li 0001, Lucian Popa 0001 |
AAAI | 3 |
| 2020 | Exploiting Node Content for Multiview Graph Convolutional Network and Adversarial RegularizationabstractNetwork representation learning (NRL) is crucial in the area of graph learning.Recently, graph autoencoders and its variants have gained much attention and popularity among various types of node embedding approaches.Most existing graph autoencoder-based methods aim to minimize the reconstruction errors of the input network while not explicitly considering the semantic relatedness between nodes.In this paper, we propose a novel network embedding method which models the consistency across different views of networks.More specifically, we create a second view from the input network which captures the relation between nodes based on node content and enforce the latent representations from the two views to be consistent by incorporating a multiview adversarial regularization module.The experimental studies on benchmark datasets prove the effectiveness of this method, and demonstrate that our method compares favorably with the state-of-the-art algorithms on challenging tasks such as link prediction and node clustering.We also evaluate our method on a real-world application, i.e., 30-day unplanned ICU readmission prediction, and achieve promising results compared with several baseline methods. Qiuhao Lu, Nisansa de Silva, Dejing Dou, Thien Huu Nguyen, Prithviraj Sen, Berthold Reinwald, Yunyao Li 0001 |
COLING | 7 |
| 2020 | Learning Structured Representations of Entity Names using ActiveLearning and Weak SupervisionabstractStructured representations of entity names are useful for many entity-related tasks such as entity normalization and variant generation.Learning the implicit structured representations of entity names without context and external knowledge is particularly challenging.In this paper, we present a novel learning framework that combines active learning and weak supervision to solve this problem.Our experimental evaluation show that this framework enables the learning of high-quality models from merely a dozen or so labeled examples. Kun Qian 0002, Poornima Chozhiyath Raman, Yunyao Li 0001, Lucian Popa 0001 |
EMNLP (1) | 3 |
| 2020 | Learning Explainable Linguistic Expressions with Neural Inductive Logic Programming for Sentence ClassificationabstractPrithviraj Sen, Marina Danilevsky, Yunyao Li, Siddhartha Brahma, Matthias Boehm, Laura Chiticariu, Rajasekar Krishnamurthy. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Prithviraj Sen, Marina Danilevsky, Yunyao Li 0001, Siddhartha Brahma, Matthias Boehm 0001, Laura Chiticariu, Rajasekar Krishnamurthy |
EMNLP (1) | 3 |
| 2020 | Small but Mighty: New Benchmarks for Split and RephraseabstractSplit and Rephrase is a text simplification task of rewriting a complex sentence into simpler ones.As a relatively new task, it is paramount to ensure the soundness of its evaluation benchmark and metric.We find that the widely used benchmark dataset universally contains easily exploitable syntactic cues caused by its automatic generation process.Taking advantage of such cues, we show that even a simple rule-based model can perform on par with the state-of-the-art model.To remedy such limitations, we collect and release two crowdsourced benchmark datasets.We not only make sure that they contain significantly more diverse syntax, but also carefully control for their quality according to a welldefined set of criteria.While no satisfactory automatic metric exists, we apply fine-grained manual evaluation based on these criteria using crowdsourcing, showing that our datasets better represent the task and are significantly more challenging for the models. 1 Li Zhang 0039, Huaiyu Zhu 0001, Siddhartha Brahma, Yunyao Li 0001 |
EMNLP (1) | 4 |
| 2019 | Low-resource Deep Entity Resolution with Transfer and Active LearningabstractEntity resolution (ER) is the task of identifying different representations of the same real-world entities across databases.It is a key step for knowledge base creation and text mining.Recent adaptation of deep learning methods for ER mitigates the need for dataset-specific feature engineering by constructing distributed representations of entity records.While these methods achieve stateof-the-art performance over benchmark data, they require large amounts of labeled data, which are typically unavailable in realistic ER applications.In this paper, we develop a deep learning-based method that targets lowresource settings for ER through a novel combination of transfer learning and active learning.We design an architecture that allows us to learn a transferable model from a highresource setting to a low-resource one.To further adapt to the target dataset, we incorporate active learning that carefully selects a few informative examples to fine-tune the transferred model.Empirical evaluation demonstrates that our method achieves comparable, if not better, performance compared to state-of-the-art learning-based methods while using an order of magnitude fewer labels. Jungo Kasai, Kun Qian 0002, Sairam Gurajada, Yunyao Li 0001, Lucian Popa 0001 |
ACL (1) | 4 |
| 2019 | Bridging the Semantic Gap with SQL Query Logs in Natural Language Interfaces to DatabasesabstractA critical challenge in constructing a natural language interface to database (NLIDB) is bridging the semantic gap between a natural language query (NLQ) and the underlying data. Two specific ways this challenge exhibits itself is through keyword mapping and join path inference. Keyword mapping is the task of mapping individual keywords in the original NLQ to database elements (such as relations, attributes or values). It is challenging due to the ambiguity in mapping the user's mental model and diction to the schema definition and contents of the underlying database. Join path inference is the process of selecting the relations and join conditions in the FROM clause of the final SQL query, and is difficult because NLIDB users lack the knowledge of the database schema or SQL and therefore cannot explicitly specify the intermediate tables and joins needed to construct a final SQL query. In this paper, we propose leveraging information from the SQL query log of a database to enhance the performance of existing NLIDBs with respect to these challenges. We present a system Templar that can be used to augment existing NLIDBs. Our extensive experimental evaluation demonstrates the effectiveness of our approach, leading up to 138% improvement in top-1 accuracy in existing NLIDBs by leveraging SQL query log information. Christopher Baik, H. V. Jagadish, Yunyao Li 0001 |
ICDE | 3 |
| 2018 | Exploiting Structure in Representation of Named Entities using Active LearningabstractFundamental to several knowledge-centric applications is the need to identify named entities from their textual mentions. However, entities lack a unique representation and their mentions can differ greatly. These variations arise in complex ways that cannot be captured using textual similarity metrics. However, entities have underlying structures, typically shared by entities of the same entity type, that can help reason over their name variations. Discovering, learning and manipulating these structures typically requires high manual effort in the form of large amounts of labeled training data and handwritten transformation programs. In this work, we propose an active-learning based framework that drastically reduces the labeled data required to learn the structures of entities. We show that programs for mapping entity mentions to their structures can be automatically generated using human-comprehensible labels. Our experiments show that our framework consistently outperforms both handwritten programs and supervised learning models. We also demonstrate the utility of our framework in relation extraction and entity resolution tasks. Nikita Bhutani, Kun Qian 0002, Yunyao Li 0001, H. V. Jagadish, Mauricio A. Hernández, Mitesh Vasa |
COLING | 3 |
| 2018 | DIMSIM: An Accurate Chinese Phonetic Similarity Algorithm Based on Learned High Dimensional EncodingabstractPhonetic similarity algorithms identify words and phrases with similar pronunciation which are used in many natural language processing tasks.However, existing approaches are designed mainly for Indo-European languages and fail to capture the unique properties of Chinese pronunciation.In this paper, we propose a high dimensional encoded phonetic similarity algorithm for Chinese, DIMSIM.The encodings are learned from annotated data to separately map initial and final phonemes into n-dimensional coordinates.Pinyin phonetic similarities are then calculated by aggregating the similarities of initial, final and tone.DIMSIM demonstrates a 7.5X improvement on mean reciprocal rank over the state-of-theart phonetic similarity approaches. Marina Danilevsky, Sara Noeman, Yunyao Li 0001 |
CoNLL | 4 |
| 2018 | LUSTRE: An Interactive System for Entity Structured Representation and Variant GenerationabstractMany data analysis and data integration applications need to account for multiple representations of entities. The variations in entity mentions arise in complex ways that are hard to capture using a textual similarity function. More sophisticated functions require the knowledge of underlying structure in the representation of entities. People traditionally identify these structures manually and write programs to manipulate them: such work is tedious and cumbersome. We have built LUSTRE, an active learning based system that can learn the structured representations of entities interactively from a few labels. In the background, it automatically generates programs to map entity mentions to their representations and to standardize them to a unique representation. Furthermore, LUSTRE provides a user-friendly interface to allow user declaratively specify normalization and variant generation functions for downstream applications. Kun Qian 0002, Nikita Bhutani, Yunyao Li 0001, H. V. Jagadish, Mauricio A. Hernández |
ICDE | 3 |
| 2017 | SEER: Auto-Generating Information Extraction Rules from User-Specified ExamplesabstractTime-consuming and complicated best describe the current state of the Information Extraction (IE) field. Machine learning approaches to IE require large collections of labeled datasets that are difficult to create and use obscure mathematical models, occasionally returning unwanted results that are unexplainable. Rule-based approaches, while resulting in easy-to-understand IE rules, are still time-consuming and labor-intensive. SEER combines the best of these two approaches: a learning model for IE rules based on a small number of user-specified examples. In this paper, we explain the design behind SEER and present a user study comparing our system against a commercially available tool in which users create IE rules manually. Our results show that SEER helps users complete text extraction tasks more quickly, as well as more accurately. Maeda F. Hanafi, Azza Abouzeid, Laura Chiticariu, Yunyao Li 0001 |
CHI | 4 |
| 2017 | CROWD-IN-THE-LOOP: A Hybrid Approach for Annotating Semantic RolesabstractCrowdsourcing has proven to be an effective method for generating labeled data for a range of NLP tasks.However, multiple recent attempts of using crowdsourcing to generate gold-labeled training data for semantic role labeling (SRL) reported only modest results, indicating that SRL is perhaps too difficult a task to be effectively crowdsourced.In this paper, we postulate that while producing SRL annotation does require expert involvement in general, a large subset of SRL labeling tasks is in fact appropriate for the crowd.We present a novel workflow in which we employ a classifier to identify difficult annotation tasks and route each task either to experts or crowd workers according to their difficulties.Our experimental evaluation shows that the proposed approach reduces the workload for experts by over two-thirds, and thus significantly reduces the cost of producing SRL annotation at little loss in quality. Chenguang Wang 0001, Alan Akbik, Laura Chiticariu, Yunyao Li 0001, Anbang Xu |
EMNLP | 4 |
| 2017 | Active Learning for Black-Box Semantic Role Labeling with Neural FactorsabstractActive learning is a useful technique for tasks for which unlabeled data is abundant but manual labeling is expensive. One example of such a task is semantic role labeling (SRL), which relies heavily on labels from trained linguistic experts. One challenge in applying active learning algorithms for SRL is that the complete knowledge of the SRL model is often unavailable, against the common assumption that active learning methods are aware of the details of the underlying models. In this paper, we present an active learning framework for black-box SRL models (i.e., models whose details are unknown). In lieu of a query strategy based on model details, we propose a neural query strategy model that embeds both language and semantic information to automatically learn the query strategy from predictions of an SRL model alone. Our experimental results demonstrate the effectiveness of both this new active learning framework and the neural query strategy model. Chenguang Wang 0001, Laura Chiticariu, Yunyao Li 0001 |
IJCAI | 3 |
| 2017 | Synthesizing Extraction Rules from User Examples with SEERabstractOur demonstration showcases SEER's end-to-end Information Extraction (IE) workflow where users highlight texts they wish to extract. Given a small set of user-specified example extractions, SEER synthesizes easy-to-understand IE rules and suggests them to the user. In addition to rule suggestions, users can quickly pick the desired rule by filtering the rule suggestion by accepting or rejecting proposed extractions. SEER's workflow allows users to jump start the IE rule development cycle; it is a less time-consuming alternative to machine learning methods that require large labeled datasets or rule-based approaches that are labor-intensive. SEER's design principles and learning algorithm are motivated by how rule developers naturally construct data extraction rules. Maeda F. Hanafi, Azza Abouzeid, Laura Chiticariu, Yunyao Li 0001 |
SIGMOD Conference | 4 |
| 2017 | Natural Language Data Management and Interfaces: Recent Development and Open ChallengesabstractThe volume of natural language text data has been rapidly increasing over the past two decades, due to factors such as the growth of the Web, the low cost associated to publishing and the progress on the digitization of printed texts. This growth combined with the proliferation of natural language systems for search and retrieving information provides tremendous opportunities for studying some of the areas where database systems and natural language processing systems overlap. This tutorial explores two more relevant areas of overlap to the database community: (1) managing natural language text data in a relational database, and (2) developing natural language interfaces to databases. The tutorial presents state-of-the-art methods, related systems, research opportunities and challenges covering both areas. Yunyao Li 0001, Davood Rafiei |
SIGMOD Conference | 1 |
| 2017 | Creation and Interaction with Large-scale Domain-Specific Knowledge BasesabstractThe ability to create and interact with large-scale domain-specific knowledge bases from unstructured/semi-structured data is the foundation for many industry-focused cognitive systems. We will demonstrate the Content Services system that provides cloud services for creating and querying high-quality domain-specific knowledge bases by analyzing and integrating multiple (un/semi)structured content sources. We will showcase an instantiation of the system for a financial domain. We will also demonstrate both cross-lingual natural language queries and programmatic API calls for interacting with this knowledge base. Shreyas Bharadwaj, Laura Chiticariu, Marina Danilevsky, Samarth Dhingra, Samved Divekar, Arnaldo Carreno-Fuentes, Nitin Gupta 0005, Sang-Don Han, Mauricio A. Hernández, C. T. Howard Ho, Parag Jain, Salil Joshi 0001, Hima P. Karanam, Saravanan Krishnan, Rajasekar Krishnamurthy, Yunyao Li 0001, Satishkumaar Manivannan, Ashish R. Mittal, Fatma Özcan 0001, Abdul Quamar, Poornima Chozhiyath Raman, Diptikalyan Saha, Karthik Sankaranarayanan, Jaydeep Sen, Prithviraj Sen, Shivakumar Vaithyanathan, Mitesh Vasa, Huaiyu Zhu 0001 |
Proc. VLDB Endow. | 17 |
| 2016 | Multilingual Aliasing for Auto-Generating Proposition BanksabstractSemantic Role Labeling (SRL) is the task of identifying the predicate-argument structure in sentences with semantic frame and role labels. For the English language, the Proposition Bank provides both a lexicon of all possible semantic frames and large amounts of labeled training data. In order to expand SRL beyond English, previous work investigated automatic approaches based on parallel corpora to automatically generate Proposition Banks for new target languages (TLs). However, this approach heuristically produces the frame lexicon from word alignments, leading to a range of lexicon-level errors and inconsistencies. To address these issues, we propose to manually alias TL verbs to existing English frames. For instance, the German verb drehen may evoke several meanings, including “turn something” and “film something”. Accordingly, we alias the former to the frame TURN.01 and the latter to a group of frames that includes FILM.01 and SHOOT.03. We execute a large-scale manual aliasing effort for three target languages and apply the new lexicons to automatically generate large Proposition Banks for Chinese, French and German with manually curated frames. We present a detailed evaluation in which we find that our proposed approach significantly increases the quality and consistency of the generated Proposition Banks. We release these resources to the research community. Alan Akbik, Yunyao Li 0001 |
COLING | 3 |
| 2016 | K-SRL: Instance-based Learning for Semantic Role LabelingabstractSemantic role labeling (SRL) is the task of identifying and labeling predicate-argument structures in sentences with semantic frame and role labels. A known challenge in SRL is the large number of low-frequency exceptions in training data, which are highly context-specific and difficult to generalize. To overcome this challenge, we propose the use of instance-based learning that performs no explicit generalization, but rather extrapolates predictions from the most similar instances in the training data. We present a variant of k-nearest neighbors (kNN) classification with composite features to identify nearest neighbors for SRL. We show that high-quality predictions can be derived from a very small number of similar instances. In a comparative evaluation we experimentally demonstrate that our instance-based learning approach significantly outperforms current state-of-the-art systems on both in-domain and out-of-domain data, reaching F1-scores of 89,28% and 79.91% respectively. Alan Akbik, Yunyao Li 0001 |
COLING | 2 |
| 2016 | Towards Semi-Automatic Generation of Proposition Banks for Low-Resource LanguagesabstractAnnotation projection based on parallel corpora has shown great promise in inexpensively creating Proposition Banks for languages for which high-quality parallel corpora and syntactic parsers are available.In this paper, we present an experimental study where we apply this approach to three languages that lack such resources: Tamil, Bengali and Malayalam.We find an average quality difference of 6 to 20 absolute F-measure points vis-avis high-resource languages, which indicates that annotation projection alone is insufficient in low-resource scenarios.Based on these results, we explore the possibility of using annotation projection as a starting point for inexpensive data curation involving both experts and non-experts.We give an outline of what such a process may look like and present an initial study to discuss its potential and challenges. Alan Akbik, Vishwajeet Kumar, Yunyao Li 0001 |
EMNLP | 3 |
| 2015 | Generating High Quality Proposition Banks for Multilingual Semantic Role LabelingabstractAlan Akbik, Laura Chiticariu, Marina Danilevsky, Yunyao Li, Shivakumar Vaithyanathan, Huaiyu Zhu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Alan Akbik, Laura Chiticariu, Marina Danilevsky, Yunyao Li 0001, Shivakumar Vaithyanathan, Huaiyu Zhu 0001 |
ACL (1) | 4 |
| 2015 | An In-depth Analysis of the Effect of Text Normalization in Social MediaabstractRecent years have seen increased interest in text normalization in social media, as the informal writing styles found in Twitter and other social media data often cause problems for NLP applications. Unfortunately, most current approaches narrowly regard the normalization task as a “one size fits all” task of replacing non-standard words with their standard counterparts. In this work we build a taxonomy of normalization edits and present a study of normalization to examine its effect on three different downstream applications (dependency parsing, named entity recognition, and text-to-speech synthesis). The results suggest that how the normalization task should be viewed is highly dependent on the targeted application. The results also show that normalization must be thought of as more than word replacement in order to produce results comparable to those seen on clean text. Tyler Baldwin, Yunyao Li 0001 |
HLT-NAACL | 2 |
| 2015 | VINERy: A Visual IDE for Information ExtractionabstractInformation Extraction (IE) is the key technology enabling analytics over unstructured and semi-structured data. Not surprisingly, it is becoming a critical building block for a wide range of emerging applications. To satisfy the rising demands for information extraction in real-world applications, it is crucial to lower the barrier to entry for IE development and enable users with general computer science background to develop higher quality extractors. In this demonstration 1 , we present VINERy, an intuitive yet expressive visual IDE for information extraction. We show how it supports the full cycle of IE development without requiring a single line of code and enables a wide range of users to develop high quality IE extractors with minimal efforts. The extractors visually built in VINERY are automatically translated into semantically equivalent extractors in a state-of-the-art declarative language for IE. We also demonstrate how the auto-generated extractors can then be imported into a conventional Eclipse-based IDE for further enhancement. The results of our user studies indicate that VINERY is a significant step forward in facilitating extractor development for both expert and novice IE developers. Yunyao Li 0001, Elmer Kim, Marc A. Touchette, Ramiya Venkatachalam |
Proc. VLDB Endow. | 1 |
| 2014 | Enterprise Search in the Big Data Era: Recent Developments and Open ChallengesabstractEnterprise search allows users in an enterprise to retrieve desired information through a simple search interface. It is widely viewed as an important productivity tool within an enterprise. While Internet search engines have been highly successful, enterprise search remains notoriously challenging due to a variety of unique challenges, and is being made more so by the increasing heterogeneity and volume of enterprise data. On the other hand, enterprise search also presents opportunities to succeed in ways beyond current Internet search capabilities. This tutorial presents an organized overview of these challenges and opportunities, and reviews the state-of-the-art techniques for building a reliable and high quality enterprise search engine, in the context of the rise of big data. Yunyao Li 0001, Huaiyu Zhu 0001 |
Proc. VLDB Endow. | 1 |
| 2014 | VLDB 2014 Ph.D. Workshop - An OverviewabstractThe VLDB 2014 PhD Workshop is an one-day event to be held in Hangzhou, China on September 1st, 2014, in conjunction with VLDB 2014. The aim of this workshop is to provide helpful feedback, useful information and networking opportunities that can benefit the students' dissertation work as well as their long-term career. The selection process and the workshop program were carefully designed with this specific goal in mind. The accepted submissions are included in the online proceedings for the Workshop at http://www.vldb.org/2014/phd_workshop_proceedings.html Yunyao Li 0001, Erich J. Neuhold |
Proc. VLDB Endow. | 1 |
| 2014 | Understand users' comprehension and preferences for composing information visualizationsabstractWe are developing an automated visualization system that helps users combine two or more existing information graphics to form an integrated view. To establish empirical foundations for building such a system, we designed and conducted two studies on Amazon Mechanical Turk to understand users’ comprehension and preferences of composite visualization under different conditions (e.g., data and tasks). In Study 1, we collected more than 1,500 textual descriptions capturing about 500 participants’ insights of given information graphics, which resulted in a task-oriented taxonomy of visual insights. In Study 2, we asked 240 participants to rank composite visualizations by their suitability for acquiring a given visual insight identified in Study 1, which resulted in ranked user preferences of visual compositions for acquiring each type of insight. In this article, we report the details of our two studies and discuss the broader implications of our crowdsourced research methodology and results to HCI-driven visualization research. Huahai Yang, Yunyao Li 0001, Michelle X. Zhou |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2013 | Adaptive Parser-Centric Text Normalization
Congle Zhang, Tyler Baldwin, C. T. Howard Ho, Benny Kimelfeld, Yunyao Li 0001 |
ACL (1) | 5 |
| 2013 | I can do text analytics!: designing development tools for novice developersabstractText analytics, an increasingly important application domain, is hampered by the high barrier to entry due to the many conceptual difficulties novice developers encounter. This work addresses the problem by developing a tool to guide novice developers to adopt the best practices employed by expert developers in text analytics and to quickly harness the full power of the underlying system. Taking a user centered task analytical approach, the tool development went through multiple design iterations and evaluation cycles. In the latest evaluation, we found that our tool enables novice developers to develop high quality extractors on par with the state of art within a few hours and with minimal training. Finally, we discuss our experience and lessons learned in the context of designing user interfaces to reduce the barriers to entry into complex domains of expertise. Huahai Yang, Daina Pupons Wickham, Laura Chiticariu, Yunyao Li 0001, Benjamin Nguyen, Arnaldo Carreno-Fuentes |
CHI | 4 |
| 2013 | Rule-Based Information Extraction is Dead! Long Live Rule-Based Information Extraction Systems!abstractThe rise of "Big Data" analytics over unstructured text has led to renewed interest in information extraction (IE).We surveyed the landscape of IE technologies and identified a major disconnect between industry and academia: while rule-based IE dominates the commercial world, it is widely regarded as dead-end technology by the academia.We believe the disconnect stems from the way in which the two communities measure the benefits and costs of IE, as well as academia's perception that rulebased IE is devoid of research challenges.We make a case for the importance of rule-based IE to industry practitioners.We then lay out a research agenda in advancing the state-of-theart in rule-based IE systems which we believe has the potential to bridge the gap between academic research and industry practice. Laura Chiticariu, Yunyao Li 0001, Frederick Reiss 0001 |
EMNLP | 2 |
| 2013 | OpinionBlocks: A Crowd-Powered, Self-improving Interactive Visual Analytic System for Understanding Opinion Text
Mengdie Hu, Huahai Yang, Michelle X. Zhou, Liang Gou, Yunyao Li 0001, Eben M. Haber |
INTERACT (2) | 5 |
| 2012 | Gumshoe quality toolkit: administering programmable searchabstractEnterprise search is challenging due to various reasons, notably the dynamic terminology and domain structure that are specific to the enterprise, combined with the fact that search deployments are typically managed by domain experts who are not necessarily search experts. To address that, it has been proposed to design search architectures that feature two principles: comprehensibility of the ranking mechanism and customizability of the search engine by means of intuitive runtime rules. The proposed demonstration operates on top of an engine implementation based on this search philosophy, and provides an administrator toolkit to realize the two principles. In particular, the toolkit provides a complete visualization of the provenance (hence ranking) of search results, embeds an editor for programming runtime rules, facilitates the investigation of (the cause of) missing or low-ranked desired results, and provides suggestions of rewrite rules to handle such results. Zhuowei Bao, Benny Kimelfeld, Yunyao Li 0001, Sriram Raghavan, Huahai Yang |
CIKM | 3 |
| 2012 | Automatic suggestion of query-rewrite rules for enterprise searchabstractEnterprise search is challenging for several reasons, notably the dynamic terminology and jargon that are specific to the enterprise domain. This challenge is partly addressed by having domain experts maintaining the enterprise search engine and adapting it to the domain specifics. Those administrators commonly address user complaints about relevant documents missing from the top matches. For that, it has been proposed to allow administrators to influence search results by crafting query-rewrite rules, each specifying how queries of a certain pattern should be modified or augmented with additional queries. Upon a complaint, the administrator seeks a semantically coherent rule that is capable of pushing the desired documents up to the top matches. However, the creation and maintenance of rewrite rules is highly tedious and time consuming. Our goal in this work is to ease the burden on search administrators by automatically suggesting rewrite rules. This automation entails several challenges. One major challenge is to select, among many options, rules that are ``natural'' from a semantic perspective (e.g., corresponding to closely related and syntactically complete concepts). Towards that, we study a machine-learning classification approach. The second challenge is to accommodate the cross-query effect of rules---a rule introduced in the context of one query can eliminate the desired results for other queries and the desired effects of other rules. We present a formalization of this challenge as a generic computational problem. As we show that this problem is highly intractable in terms of complexity theory, we present heuristic approaches and optimization thereof. In an experimental study within IBM intranet search, those heuristics achieve near-optimal quality and well scale to large data sets. Zhuowei Bao, Benny Kimelfeld, Yunyao Li 0001 |
SIGIR | 3 |
| 2011 | A Graph Approach to Spelling Correction in Domain-Centric Search
Zhuowei Bao, Benny Kimelfeld, Yunyao Li 0001 |
ACL | 3 |
| 2011 | Facilitating pattern discovery for relation extraction with semantic-signature-based clusteringabstractHand-crafted textual patterns have been the mainstay device of practical relation extraction for decades. However, there has been little work on reducing the manual effort involved in the discovery of effective textual patterns for relation extraction. In this paper, we propose a clustering-based approach to facilitate the pattern discovery for relation extraction. Specifically, we define the notion of semantic signature to represent the most salient features of a textual fragment. We then propose a novel clustering algorithm based on semantic signature, S2C, and its enhancement S2C+. Experiments on two real-world data sets show that, when compared with k-means clustering, S2C and S2C+ are at least an order of magnitude faster, while generating high quality clusters that are at least comparable to the best clusters generated by k-means without requiring any manual tuning. Finally, a user study confirms that our clustering-based approach can indeed help users discover effective textual patterns for relation extraction with only a fraction of the manual effort required by the conventional approach. Yunyao Li 0001, Vivian Chu, Sebastian Blohm, Huaiyu Zhu 0001, C. T. Howard Ho |
CIKM | 1 |
| 2011 | Selectivity estimation for extraction operators over text dataabstractRecently, there has been increasing interest in extending relational query processing to efficiently support extraction operators, such as dictionaries and regular expressions, over text data. Many text processing queries are sophisticated in that they involve multiple extraction and join operators, resulting in many possible query plans. However, there has been little research on building the selectivity or cost estimation for these extraction operators, which is crucial for an optimizer to pick a good query plan. In this paper, we define the problem of selectivity estimation for dictionaries and regular expressions, and propose to develop document synopses over a text corpus, from which the selectivity can be estimated. We first adapt the language models in the Natural Language Processing literature to form the top-k n-gram synopsis as the baseline document synopsis. Then we develop two classes of novel document synopses: stratified bloom filter synopsis and roll-up synopsis. We also develop techniques to decompose a complicated regular expression into subparts to achieve more effective and accurate estimation. We conduct experiments over the Enron email corpus using both real-world and synthetic workloads to compare the accuracy of the selectivity estimation over different classes and variations of synopses. The results show that, the top-k stratified bloom filter synopsis and the roll-up synopsis is the most accurate in dictionary and regular expression selectivity estimation respectively. Daisy Zhe Wang, Yunyao Li 0001, Frederick Reiss 0001, Shivakumar Vaithyanathan |
ICDE | 3 |
| 2011 | Rewrite rules for search database systemsabstractThe results of a search engine can be improved by consulting auxiliary data. In a search database system, the association between the user query and the auxiliary data is driven by rewrite rules that augment the user query with a set of alternative queries. This paper develops a framework that formalizes the notion of a rewrite program, which is essentially a collection of hedge-rewriting rules. When applied to a search query, the rewrite program produces a set of alternative queries that constitutes a least fixpoint (lfp). The main focus of the paper is on the lfp-convergence of a rewrite program, where a rewrite program is lfp-convergent if the least fixpoint of every search query is finite. Determining whether a given rewrite program is lfp-convergent is undecidable; to accommodate that, the paper proposes a safety condition, and shows that safety guarantees lfp-convergence, and that safety can be decided in polynomial time. The effectiveness of the safety condition in capturing lfp-convergence is illustrated by an application to a rewrite program in an implemented system that is intended for widespread use. Ronald Fagin, Benny Kimelfeld, Yunyao Li 0001, Sriram Raghavan, Shivakumar Vaithyanathan |
PODS | 3 |
| 2011 | The SystemT IDE: an integrated development environment for information extraction rulesabstractInformation Extraction (IE)-the problem of extracting structured information from unstructured text - has become the key enabler for many enterprise applications such as semantic search, business analytics and regulatory compliance. While rule-based IE systems are widely used in practice due to their well-known "explainability," developing high-quality information extraction rules is known to be a labor-intensive and time-consuming iterative process. Laura Chiticariu, Vivian Chu, Sajib Dasgupta, Thilo W. Goetz, C. T. Howard Ho, Rajasekar Krishnamurthy, Alexander Lang, Yunyao Li 0001, Bin Liu 0002, Sriram Raghavan, Frederick Reiss 0001, Shivakumar Vaithyanathan, Huaiyu Zhu 0001 |
SIGMOD Conference | 8 |
| 2010 | SystemT: An Algebraic Approach to Declarative Information Extraction
Laura Chiticariu, Rajasekar Krishnamurthy, Yunyao Li 0001, Sriram Raghavan, Frederick Reiss 0001, Shivakumar Vaithyanathan |
ACL | 3 |
| 2010 | Domain Adaptation of Rule-Based Annotators for Named-Entity Recognition Tasks
Laura Chiticariu, Rajasekar Krishnamurthy, Yunyao Li 0001, Frederick Reiss 0001, Shivakumar Vaithyanathan |
EMNLP | 3 |
| 2010 | Understanding queries in a search database systemabstractIt is well known that a search engine can significantly benefit from an auxiliary database, which can suggest interpretations of the search query by means of the involved concepts and their interrelationship. The difficulty is to translate abstract notions like concept and interpretation into a concrete search algorithm that operates over the auxiliary database. To surpass existing heuristics, there is a need for a formal basis, which is realized in this paper through the framework of a search database system, where an interpretation is identified as a parse. It is shown that the parses of a query can be generated in polynomial time in the combined size of the input and the output, even if parses are restricted to those having a nonempty evaluation. Identifying that one parse is more specific than another is important for ranking answers, and this framework captures the precise semantics of being more specific; moreover, performing this comparison between parses is tractable. Lastly, the paper studies the problem of finding the most specific parses. Unfortunately, this problem turns out to be intractable in the general case. However, under reasonable assumptions, the parses can be enumerated in an order of decreasing specificity, with polynomial delay and polynomial space. Ronald Fagin, Benny Kimelfeld, Yunyao Li 0001, Sriram Raghavan, Shivakumar Vaithyanathan |
PODS | 3 |
| 2010 | Enterprise information extraction: recent developments and open challengesabstractInformation extraction (IE) - the problem of extracting structured information from unstructured text - has become an increasingly important topic in recent years. A SIGMOD 2006 tutorial [3] outlined challenges and opportunities for the database community to advance the state of the art in information extraction, and posed the following grand challenge: "Can we build a System R for information extraction? Laura Chiticariu, Yunyao Li 0001, Sriram Raghavan, Frederick Reiss 0001 |
SIGMOD Conference | 2 |
| 2009 | Enabling enterprise mashups over unstructured text feeds with InfoSphere MashupHub and SystemTabstractEnterprise mashup scenarios often involve feeds derived from data created primarily for eye consumption, such as email, news, calendars, blogs, and web feeds. These data sources can test the capabilities of current data mashup products, as the attributes needed to perform join, aggregation, and other operations are often buried within unstructured feed text. Information extraction technology is a key enabler in such scenarios, using annotators to convert unstructured text into structured information that can facilitate mashup operations. David E. Simmen, Frederick Reiss 0001, Yunyao Li 0001, Suresh Thalamati |
SIGMOD Conference | 3 |
| 2008 | Regular Expression Learning for Information Extraction
Yunyao Li 0001, Rajasekar Krishnamurthy, Sriram Raghavan, Shivakumar Vaithyanathan, H. V. Jagadish |
EMNLP | 1 |
| 2008 | Enabling Schema-Free XQuery with meaningful query focus
Yunyao Li 0001, Cong Yu 0001, H. V. Jagadish |
VLDB J. | 1 |
| 2007 | Enabling Domain-Awareness for a Generic Natural Language Interface
Yunyao Li 0001, Ishan Chaudhuri, Huahai Yang, Satinder Singh 0001, H. V. Jagadish |
AAAI | 1 |
| 2007 | Making database systems usableabstractDatabase researchers have striven to improve the capability of a database in terms of both performance and functionality. We assert that the usability of a database is as important as its capability. In this paper, we study why database systems today are so difficult to use. We identify a set of five pain points and propose a research agenda to address these. In particular, we introduce a presentation data model and recommend direct data manipulation with a schema later approach. We also stress the importance of provenance and of consistency across presentation models. H. V. Jagadish, Adriane Chapman, Aaron Elkiss, Magesh Jayapandian, Yunyao Li 0001, Arnab Nandi 0001, Cong Yu 0001 |
SIGMOD Conference | 5 |
| 2007 | DaNaLIX: a domain-adaptive natural language interface for querying XMLabstractWe present DaNaLIX, a prototype domain-adaptive natural language interface for querying XML. Our system is an extension of NaLIX, a generic natural language interface for querying XML. While retaining the portability of a purely generic system like NaLIX, DaNaLIX can exploit domain knowledge, whenever available, to its advantage for query translation. More importantly, in DaNaLIX such domain knowledge does not have to be pre-defined; instead it can be automatically obtained from the interactions between a user and the system. In this demonstration, we describe the overall architecture of DaNaLIX. We also demonstrate how a generic system like DaNaLIX can take advantage of domain knowledge to improve its usability and query translation accuracy. In addition, we show DaNaLIX still possesses the portability of a generic system by using data collections from three different domains. Finally, we present how domain knowledge can be obtained through user interactions in an automatic fashion. Yunyao Li 0001, Ishan Chaudhuri, Huahai Yang, Satinder Singh 0001, H. V. Jagadish |
SIGMOD Conference | 1 |
| 2007 | NaLIX: A generic natural language search environment for XML dataabstractWe describe the construction of a generic natural language query interface to an XML database. Our interface can accept a large class of English sentences as a query, which can be quite complex and include aggregation, nesting, and value joins, among other things. This query is translated, potentially after reformulation, into an XQuery expression. The translation is based on mapping grammatical proximity of natural language parsed tokens in the parse tree of the query sentence to proximity of corresponding elements in the XML data to be retrieved. Iterative search in the form of followup queries is also supported. Our experimental assessment, through a user study, demonstrates that this type of natural language interface is good enough to be usable now, with no restrictions on the application domain. Yunyao Li 0001, Huahai Yang, H. V. Jagadish |
ACM Trans. Database Syst. | 1 |
| 2006 | Constructing a Generic Natural Language Interface for an XML Database
Yunyao Li 0001, Huahai Yang, H. V. Jagadish |
EDBT | 1 |
| 2006 | Term Disambiguation in Natural Language Query for XML
Yunyao Li 0001, Huahai Yang, H. V. Jagadish |
FQAS | 1 |
| 2006 | Getting work done on the web: supporting transactional queriesabstractMany searches on the web have a transactional intent. We argue that pages satisfying transactional needs can be distinguished from the more common pages that have some information and links, but cannot be used to execute a transaction. Based on this hypothesis, we provide a recipe for constructing a transaction annotator. By constructing an annotator with one corpus and then demonstrating its classification performance on another,we establish its robustness. Finally, we show experimentally that a search procedure that exploits such pre-annotation greatly outperforms traditional search for retrieving transactional pages. Yunyao Li 0001, Rajasekar Krishnamurthy, Shivakumar Vaithyanathan, H. V. Jagadish |
SIGIR | 1 |
| 2005 | NaLIX: an interactive natural language interface for querying XMLabstractDatabase query languages can be intimidating to the non-expert, leading to the immense recent popularity for keyword based search in spite of its significant limitations. The holy grail has been the development of a natural language query interface. We present NaLIX, a generic interactive natural language query interface to an XML database. Our system can accept an arbitrary English language sentence as query input, which can include aggregation, nesting, and value joins, among other things. This query is translated, potentially after reformulation, into an XQuery expression that can be evaluated against an XML database. The translation is done through mapping grammatical proximity of natural language parsed tokens to proximity of corresponding elements in the result XML. In this demonstration, we show that NaLIX, while far from being able to pass the Turing test, is perfectly usable in practice, and able to handle even quite complex queries in a variety of application domains. In addition, we also demonstrate how carefully designed features in NaLIX facilitate the interactive query process and improve the usability of the interface. Yunyao Li 0001, Huahai Yang, H. V. Jagadish |
SIGMOD Conference | 1 |
| 2004 | Schema-Free XQuery
Yunyao Li 0001, Cong Yu 0001, H. V. Jagadish |
VLDB | 1 |