EDBT 2026 Demo / reviewers in the wild / expert
Hoifung Poon
dblp:78/4609
· DBLP profile ↗
46ranked-venue papers
12as first author
21since 2021 · last 2026
0000-0002-9067-0918ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 11 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring the Future of AI in Clinical Collaboration: A Study on Tumor Board Case PreparationabstractMultidisciplinary tumor boards (MTBs) bring specialists together to identify therapies for complex cancer cases, but preparing for them is time-intensive. Clinicians must extract key details from extensive records and evaluate treatment options. While large language models (LLMs) show promise in medicine for basic tasks like summarizing notes, little is known about their role in high-stakes tasks like MTB preparation. We conducted a mixed-methods study with 16 oncologists using two AI systems to prepare patient cases for MTB: an off-the-shelf assistant (Copilot) and a task-specific multi-agent system (Healthcare Agent Orchestrator, HAO). We analyzed oncologist prompts, AI responses, and oncologists’ perception of AI. Participants showed greater willingness to adopt HAO but were often overconfident in AI summaries and skeptical of AI-recommended therapies. Trust calibration strategies, such as source links and agent-trajectories, failed to align trust with system capabilities. We conclude with how AI systems should be built to support clinicians in high-stakes tasks. Amanda K. Hall, Ruican Rachel Zhong, Selin S. Everett, Alyssa Unell, Matthias Blondeel, Jonathan Carlson, Katie Claveau, Thulasee Jose, Tristan Naumann, David C. Rhew, Naiteek Sangani, Frank Tuan, James Weinstein, Varun Mishra 0001, Elizabeth D. Mynatt, T. Scott Saponas, Leonardo Schettini, J. Samuel Preston, Yu Gu 0017, Naoto Usuyama, Zelalem Gero, Cliff Wong, Noel Codella, Hoifung Poon, Shrey Jain, Matthew P. Lungren, Eric Horvitz |
CHI | 27 |
| 2025 | Boltzmann Attention Sampling for Image Analysis with Small ObjectsabstractDetecting and segmenting small objects, such as lung nodules and tumor lesions, remains a critical challenge in image analysis. These objects often occupy less than 0.1% of an image, making traditional transformer architectures inefficient and prone to performance degradation due to redundant attention computations on irrelevant regions. Existing sparse attention mechanisms rely on rigid hierarchical structures, which are poorly suited for detecting small, variable, and uncertain object locations. In this paper, we propose BoltzFormer, a novel transformer-based architecture designed to address these challenges through dynamic sparse attention. BoltzFormer identifies and focuses attention on relevant areas by modeling uncertainty using a Boltzmann distribution with an annealing schedule. Initially, a higher temperature allows broader area sampling in early layers, when object location uncertainty is greatest. As the temperature decreases in later layers, attention becomes more focused, enhancing efficiency and accuracy. BoltzFormer seamlessly integrates into existing transformer architectures via a modular Boltzmann attention sampling mechanism. Comprehensive evaluations on benchmark datasets demonstrate that BoltzFormer significantly improves segmentation performance for small objects while reducing attention computation by an order of magnitude compared to previous state-of-the-art methods. Theodore Zhao, Sid Kiblawi, Naoto Usuyama, Ho Hin Lee, Sam Preston, Hoifung Poon, Mu Wei |
CVPR | 6 |
| 2025 | MuirBench: A Comprehensive Benchmark for Robust Multi-image UnderstandingabstractWe introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tasks (e.g., scene understanding, ordering) that involve 10 categories of multi-image relations (e.g., multiview, temporal relations). Comprising 11,264 images and 2,600 multiple-choice questions, MuirBench is created in a pairwise manner, where each standard instance is paired with an unanswerable variant that has minimal semantic differences, in order for a reliable assessment. Evaluated upon 20 recent multi-modal LLMs, our results reveal that even the best-performing models like GPT-4o and Gemini Pro find it challenging to solve MuirBench, achieving 68.0% and 49.3% in accuracy. Open-source multimodal LLMs trained on single images can hardly generalize to multi-image questions, hovering below 33.3% in accuracy. These results highlight the importance of MuirBench in encouraging the community to develop multimodal LLMs that can look beyond a single image, suggesting potential pathways for future improvements. Fei Wang 0060, James Y. Huang, Zekun Li 0007, Qin Liu 0010, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu 0014, Wenxuan Zhou 0002, Kai Zhang 0008, Tianyi Lorena Yan, Wenjie Mo 0001, Hsiang-Hui Liu, Pan Lu, Chunyuan Li, Chaowei Xiao, Kai-Wei Chang 0001, Dan Roth 0001, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001 |
ICLR | 20 |
| 2025 | From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context LearningabstractNan Xu, Fei Wang, Sheng Zhang, Hoifung Poon, Muhao Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Nan Xu 0014, Fei Wang 0060, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001 |
NAACL (Long Papers) | 4 |
| 2024 | DocLens: Multi-aspect Fine-grained Medical Text EvaluationabstractYiqing Xie, Sheng Zhang, Hao Cheng, Pengfei Liu, Zelalem Gero, Cliff Wong, Tristan Naumann, Hoifung Poon, Carolyn Rose. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yiqing Xie, Sheng Zhang 0012, Hao Cheng 0002, Zelalem Gero, Cliff Wong, Tristan Naumann, Hoifung Poon, Carolyn P. Rosé |
ACL (1) | 8 |
| 2024 | Pareto Optimal Learning for Estimating Large Language Model ErrorsabstractLarge Language Models (LLMs) have shown impressive abilities in many applications.When a concrete and precise answer is desired, it is important to have a quantitative estimation of the potential error rate.However, this can be challenging due to the text-in-text-out nature of generative models.We present a method based on Pareto optimization that generates a risk score to estimate the probability of error in an LLM response by integrating multiple sources of information.We prove theoretically that the error estimator optimized in our framework aligns with the LLM and the information sources in an Pareto optimal manner.Experimental results show that the risk scores estimated by our method are well correlated with the true LLM error rate, thus facilitating error correction.By dynamically combining with prompting strategies such as self-verification and information retrieval, we demonstrate the proposed method can be utilized to increase the performance of an LLM, surpassing state-ofthe-art task specific models. Theodore Zhao, Mu Wei, J. Samuel Preston, Hoifung Poon |
ACL (1) | 4 |
| 2024 | mDPO: Conditional Preference Optimization for Multimodal Large Language ModelsabstractDirect preference optimization (DPO) has shown to be an effective method for large language model (LLM) alignment.Recent works have attempted to apply DPO to multimodal scenarios but have found it challenging to achieve consistent improvement.Through a comparative experiment, we identify the unconditional preference problem in multimodal preference optimization, where the model overlooks the image condition.To address this problem, we propose MDPO, a multimodal DPO objective that prevents the over-prioritization of language-only preferences by also optimizing image preference.Moreover, we introduce a reward anchor that forces the reward to be positive for chosen responses, thereby avoiding the decrease in their likelihood-an intrinsic problem of relative preference optimization.Experiments on two multimodal LLMs of different sizes and three widely used benchmarks demonstrate that MDPO effectively addresses the unconditional preference problem in multimodal preference optimization and significantly improves model performance, particularly in reducing hallucination. Fei Wang 0060, Wenxuan Zhou 0002, James Y. Huang, Nan Xu 0014, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001 |
EMNLP | 6 |
| 2024 | UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionabstractLarge language models (LLMs) have demonstrated remarkable generalizability, such as understanding arbitrary entities and relations. Instruction tuning has proven effective for distilling LLMs into more cost-efficient models such as Alpaca and Vicuna. Yet such student models still trail the original LLMs by large margins in downstream applications. In this paper, we explore targeted distillation with mission-focused instruction tuning to train student models that can excel in a broad application class such as open information extraction. Using named entity recognition (NER) for case study, we show how ChatGPT can be distilled into much smaller UniversalNER models for open NER. For evaluation, we assemble the largest NER benchmark to date, comprising 43 datasets across 9 diverse domains such as biomedicine, programming, social media, law, finance. Without using any direct supervision, UniversalNER attains remarkable NER accuracy across tens of thousands of entity types, outperforming general instruction-tuned models such as Alpaca and Vicuna by over 30 absolute F1 points in average. With a tiny fraction of parameters, UniversalNER not only acquires ChatGPT's capability in recognizing arbitrary entity types, but also outperforms its NER accuracy by 7-9 absolute F1 points in average. Remarkably, UniversalNER even outperforms by a large margin state-of-the-art multi-task instruction-tuned systems such as InstructUIE, which uses supervised NER examples. We also conduct thorough ablation studies to assess the impact of various components in our distillation approach. We release the distillation recipe, data, and UniversalNER models to facilitate future research on targeted distillation. Wenxuan Zhou 0002, Sheng Zhang 0012, Yu Gu 0017, Muhao Chen 0001, Hoifung Poon |
ICLR | 5 |
| 2023 | Continual Contrastive Finetuning Improves Low-Resource Relation ExtractionabstractRelation extraction (RE), which has relied on structurally annotated corpora for model training, has been particularly challenging in lowresource scenarios and domains.Recent literature has tackled low-resource RE by selfsupervised learning, where the solution involves pretraining the entity pair embedding by RE-based objective and finetuning on labeled data by classification-based objective.However, a critical challenge to this approach is the gap in objectives, which prevents the RE model from fully utilizing the knowledge in pretrained representations.In this paper, we aim at bridging the gap and propose to pretrain and finetune the RE model using consistent objectives of contrastive learning.Since in this kind of representation learning paradigm, one relation may easily form multiple clusters in the representation space, we further propose a multi-center contrastive loss that allows one relation to form multiple clusters to better align with pretraining.Experiments on two document-level RE datasets, BioRED and Re-DocRED, demonstrate the effectiveness of our method.Particularly, when using 1% end-task training data, our method outperforms PLMbased RE classifier by 10.5% and 6.1% on the two datasets, respectively. Wenxuan Zhou 0002, Sheng Zhang 0012, Tristan Naumann, Muhao Chen 0001, Hoifung Poon |
ACL (1) | 5 |
| 2023 | Exploring the Boundaries of GPT-4 in RadiologyabstractQianchu Liu, Stephanie Hyland, Shruthi Bannur, Kenza Bouzid, Daniel Castro, Maria Wetscherek, Robert Tinn, Harshita Sharma, Fernando Pérez-García, Anton Schwaighofer, Pranav Rajpurkar, Sameer Khanna, Hoifung Poon, Naoto Usuyama, Anja Thieme, Aditya Nori, Matthew Lungren, Ozan Oktay, Javier Alvarez-Valle. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Qianchu Liu, Stephanie L. Hyland, Shruthi Bannur, Kenza Bouzid, Daniel C. Castro, Maria Wetscherek, Robert Tinn, Harshita Sharma, Fernando Pérez-García, Anton Schwaighofer, Pranav Rajpurkar, Sameer Tajdin Khanna, Hoifung Poon, Naoto Usuyama, Anja Thieme, Aditya V. Nori, Matthew P. Lungren, Ozan Oktay, Javier Alvarez-Valle |
EMNLP | 13 |
| 2023 | Optimizing Bi-Encoder for Named Entity Recognition via Contrastive Learning
Sheng Zhang 0012, Hao Cheng 0002, Jianfeng Gao 0001, Hoifung Poon |
ICLR | 4 |
| 2023 | Precision Health in the Age of Large Language ModelsabstractMedicine today is imprecise. Among the top 20 drugs in the U.S., up to 80% of patients are non-responders. The goal of precision health is to provide the right intervention for the right people at the right time. The key to realize this dream is to develop a data-driven, learning system that can instantly incorporate new health information to optimize care delivery and accelerate biomedical discovery. In reality, however, the health ecosystem is mired in overwhelming unstructured data and excruciating manual processing. For example, in cancer, standard of care often fails, and clinical trials are the last hope. Yet less than 3% of patients could find a matching trial, whereas 40% of trial failures simply stem from insufficient recruitment. Discovery is painfully slow as a new drug may take billions of dollars and over a decade to develop. Hoifung Poon, Tristan Naumann, Sheng Zhang 0012, Javier González Hernández |
KDD | 1 |
| 2023 | LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One DayabstractConversational generative AI has demonstrated remarkable promise for empowering biomedical practitioners, but current investigations focus on unimodal text. Multimodal conversational AI has seen rapid progress by leveraging billions of image-text pairs from the public web, but such general-domain vision-language models still lack sophistication in understanding and conversing about biomedical images. In this paper, we propose a cost-efficient approach for training a vision-language conversational assistant that can answer open-ended research questions of biomedical images. The key idea is to leverage a large-scale, broad-coverage biomedical figure-caption dataset extracted from PubMed Central, use GPT-4 to self-instruct open-ended instruction-following data from the captions, and then fine-tune a large general-domain vision-language model using a novel curriculum learning method. Specifically, the model first learns to align biomedical vocabulary using the figure-caption pairs as is, then learns to master open-ended conversational semantics using GPT-4 generated instruction-following data, broadly mimicking how a layperson gradually acquires biomedical knowledge. This enables us to train a Large Language and Vision Assistant for BioMedicine (LLaVA-Med) in less than 15 hours (with eight A100s). LLaVA-Med exhibits excellent multimodal conversational capability and can follow open-ended instruction to assist with inquiries about a biomedical image. On three standard biomedical visual question answering datasets, LLaVA-Med outperforms previous supervised state-of-the-art on certain metrics. To facilitate biomedical multimodal research, we will release our instruction-following data and the LLaVA-Med model. Chunyuan Li, Cliff Wong, Sheng Zhang 0012, Naoto Usuyama, Tristan Naumann, Hoifung Poon, Jianfeng Gao 0001 |
NeurIPS | 8 |
| 2023 | Compositional Zero-Shot Domain Transfer with Text-to-Text ModelsabstractAbstract Label scarcity is a bottleneck for improving task performance in specialized domains. We propose a novel compositional transfer learning framework (DoT51) for zero-shot domain transfer. Without access to in-domain labels, DoT5 jointly learns domain knowledge (from masked language modelling of unlabelled in-domain free text) and task knowledge (from task training on more readily available general-domain data) in a multi-task manner. To improve the transferability of task training, we design a strategy named NLGU: We simultaneously train natural language generation (NLG) for in-domain label-to-data generation, which enables data augmentation for self-finetuning and natural language understanding (NLU) for label prediction. We evaluate DoT5 on the biomedical domain and the resource-lean subdomain of radiology, focusing on natural language inference, text summarization, and embedding learning. DoT5 demonstrates the effectiveness of compositional transfer learning through multi-task learning. In particular, DoT5 outperforms the current state-of-the-art in zero-shot transfer by over 7 absolute points in accuracy on RadNLI. We validate DoT5 with ablations and a case study demonstrating its ability to solve challenging NLI examples requiring in-domain expertise. Fangyu Liu 0001, Qianchu Liu, Shruthi Bannur, Fernando Pérez-García, Naoto Usuyama, Sheng Zhang 0012, Tristan Naumann, Aditya V. Nori, Hoifung Poon, Javier Alvarez-Valle, Ozan Oktay, Stephanie L. Hyland |
Trans. Assoc. Comput. Linguistics | 9 |
| 2022 | Making the Most of Text Semantics to Improve Biomedical Vision-Language Processing
Benedikt Boecking, Naoto Usuyama, Shruthi Bannur, Daniel C. Castro, Anton Schwaighofer, Stephanie L. Hyland, Maria Wetscherek, Tristan Naumann, Aditya V. Nori, Javier Alvarez-Valle, Hoifung Poon, Ozan Oktay |
ECCV (36) | 11 |
| 2022 | BioGPT: generative pre-trained transformer for biomedical text generation and miningabstractPre-trained language models have attracted increasing attention in the biomedical domain, inspired by their great success in the general natural language domain. Among the two main branches of pre-trained language models in the general language domain, i.e. BERT (and its variants) and GPT (and its variants), the first one has been extensively studied in the biomedical domain, such as BioBERT and PubMedBERT. While they have achieved great success on a variety of discriminative downstream biomedical tasks, the lack of generation ability constrains their application scope. In this paper, we propose BioGPT, a domain-specific generative Transformer language model pre-trained on large-scale biomedical literature. We evaluate BioGPT on six biomedical natural language processing tasks and demonstrate that our model outperforms previous models on most tasks. Especially, we get 44.98%, 38.42% and 40.76% F1 score on BC5CDR, KD-DTI and DDI end-to-end relation extraction tasks, respectively, and 78.2% accuracy on PubMedQA, creating a new record. Our case study on text generation further demonstrates the advantage of BioGPT on biomedical literature to generate fluent descriptions for biomedical terms. Renqian Luo, Liai Sun, Yingce Xia, Tao Qin 0001, Sheng Zhang 0012, Hoifung Poon, Tie-Yan Liu |
Briefings Bioinform. | 6 |
| 2022 | Domain-Specific Language Model Pretraining for Biomedical Natural Language ProcessingabstractPretraining large neural language models, such as BERT, has led to impressive gains on many natural language processing (NLP) tasks. However, most pretraining efforts focus on general domain corpora, such as newswire and Web. A prevailing assumption is that even domain-specific pretraining can benefit by starting from general-domain language models. In this article, we challenge this assumption by showing that for domains with abundant unlabeled text, such as biomedicine, pretraining language models from scratch results in substantial gains over continual pretraining of general-domain language models. To facilitate this investigation, we compile a comprehensive biomedical NLP benchmark from publicly available datasets. Our experiments show that domain-specific pretraining serves as a solid foundation for a wide range of biomedical NLP tasks, leading to new state-of-the-art results across the board. Further, in conducting a thorough evaluation of modeling choices, both for pretraining and task-specific fine-tuning, we discover that some common practices are unnecessary with BERT models, such as using complex tagging schemes in named entity recognition. To help accelerate research in biomedical NLP, we have released our state-of-the-art pretrained and task-specific models for the community, and created a leaderboard featuring our BLURB benchmark (short for Biomedical Language Understanding & Reasoning Benchmark) at https://aka.ms/BLURB . Yu Gu 0017, Robert Tinn, Hao Cheng 0002, Michael Lucas, Naoto Usuyama, Xiaodong Liu 0003, Tristan Naumann, Jianfeng Gao 0001, Hoifung Poon |
ACM Trans. Comput. Heal. | 9 |
| 2021 | Self-Supervised Self-Supervision by Combining Deep Learning and Probabilistic LogicabstractLabeling training examples at scale is a perennial challenge in machine learning. Self-supervision methods compensate for the lack of direct supervision by leveraging prior knowledge to automatically generate noisy labeled examples. Deep probabilistic logic (DPL) is a unifying framework for self-supervised learning that represents unknown labels as latent variables and incorporates diverse self-supervision using probabilistic logic to train a deep neural network end-to-end using variational EM. While DPL is successful at combining pre-specified self-supervision, manually crafting self-supervision to attain high accuracy may still be tedious and challenging. In this paper, we propose Self-Supervised Self-Supervision (S4), which adds to DPL the capability to learn new self-supervision automatically. Starting from an initial "seed," S4 iteratively uses the deep neural network to propose new self-supervision. These are either added directly (a form of structured self-training) or verified by a human expert (as in feature-based active learning). Experiments show that S4 is able to automatically propose accurate self-supervision and can often nearly match the accuracy of supervised methods with a tiny fraction of the human effort. Hunter Lang, Hoifung Poon |
AAAI | 2 |
| 2021 | Modular Self-Supervision for Document-Level Relation ExtractionabstractExtracting relations across large text spans has been relatively underexplored in NLP, but it is particularly important for high-value domains such as biomedicine, where obtaining high recall of the latest findings is crucial for practical applications.Compared to conventional information extraction confined to short text spans, document-level relation extraction faces additional challenges in both inference and learning.Given longer text spans, state-of-the-art neural architectures are less effective and taskspecific self-supervision such as distant supervision becomes very noisy.In this paper, we propose decomposing document-level relation extraction into relation detection and argument resolution, taking inspiration from Davidsonian semantics.This enables us to incorporate explicit discourse modeling and leverage modular self-supervision for each sub-problem, which is less noise-prone and can be further refined end-to-end via variational EM.We conduct a thorough evaluation in biomedical machine reading for precision oncology, where cross-paragraph relation mentions are prevalent.Our method outperforms prior state of the art, such as multi-scale learning and graph neural networks, by over 20 absolute F1 points.The gain is particularly pronounced among the most challenging relation instances whose arguments never co-occur in a paragraph. Sheng Zhang 0012, Cliff Wong, Naoto Usuyama, Tristan Naumann, Hoifung Poon |
EMNLP (1) | 6 |
| 2021 | Domain-Specific Pretraining for Vertical Search: Case Study on Biomedical LiteratureabstractInformation overload is a prevalent challenge in many high-value domains. A prominent case in point is the explosion of the biomedical literature on COVID-19, which swelled to hundreds of thousands of papers in a matter of months. In general, biomedical literature expands by two papers every minute, totalling over a million new papers every year. Search in the biomedical realm, and many other vertical domains is challenging due to the scarcity of direct supervision from click logs. Self-supervised learning has emerged as a promising direction to overcome the annotation bottleneck. We propose a general approach for vertical search based on domain-specific pretraining and present a case study for the biomedical domain. Despite being substantially simpler and not using any relevance labels for training or development, our method performs comparably or better than the best systems in the official TREC-COVID evaluation, a COVID-related biomedical search competition. Using distributed computing in modern cloud infrastructure, our system can scale to tens of millions of articles on PubMed and has been deployed as Microsoft Biomedical Search, a new search experience for biomedical literature: https://aka.ms/biomedsearch. Yu Wang 0009, Jinchao Li, Tristan Naumann, Chenyan Xiong, Hao Cheng 0002, Robert Tinn, Cliff Wong, Naoto Usuyama, Richard Rogahn, Zhihong Shen, Eric Horvitz, Paul N. Bennett, Jianfeng Gao 0001, Hoifung Poon |
KDD | 15 |
| 2021 | Targeted Adversarial Training for Natural Language UnderstandingabstractLis Pereira, Xiaodong Liu, Hao Cheng, Hoifung Poon, Jianfeng Gao, Ichiro Kobayashi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Lis Pereira, Xiaodong Liu 0003, Hao Cheng 0002, Hoifung Poon, Jianfeng Gao 0001, Ichiro Kobayashi 0001 |
NAACL-HLT | 4 |
| 2019 | Augmenting subnetwork inference with information extracted from the scientific literatureabstractMany biological studies involve either (i) manipulating some aspect of a cell or its environment and then simultaneously measuring the effect on thousands of genes, or (ii) systematically manipulating each gene and then measuring the effect on some response of interest. A common challenge that arises in these studies is to explain how genes identified as relevant in the given experiment are organized into a subnetwork that accounts for the response of interest. The task of inferring a subnetwork is typically dependent on the information available in publicly available, structured databases, which suffer from incompleteness. However, a wealth of potentially relevant information resides in the scientific literature, such as information about genes associated with certain concepts of interest, as well as interactions that occur among various biological entities. We contend that by exploiting this information, we can improve the explanatory power and accuracy of subnetwork inference in multiple applications. Here we propose and investigate several ways in which information extracted from the scientific literature can be used to augment subnetwork inference. We show that we can use literature-extracted information to (i) augment the set of entities identified as being relevant in a subnetwork inference task, (ii) augment the set of interactions used in the process, and (iii) support targeted browsing of a large inferred subnetwork by identifying entities and interactions that are closely related to concepts of interest. We use this approach to uncover the pathways involved in interactions between a virus and a host cell, and the pathways that are regulated by a transcription factor associated with breast cancer. Our experimental results demonstrate that these approaches can provide more accurate and more interpretable subnetworks. Integer program code, background network data, and pathfinding code are available at https://github.com/Craven-Biostat-Lab/subnetwork_inference. Sid Kiblawi, Deborah Chasman, Amanda Henning, Eunju Park, Hoifung Poon, Michael Gould, Paul Ahlquist, Mark Craven |
PLoS Comput. Biol. | 5 |
| 2018 | Deep Probabilistic Logic: A Unifying Framework for Indirect SupervisionabstractDeep learning has emerged as a versatile tool for a wide range of NLP tasks, due to its superior capacity in representation learning.But its applicability is limited by the reliance on annotated examples, which are difficult to produce at scale.Indirect supervision has emerged as a promising direction to address this bottleneck, either by introducing labeling functions to automatically generate noisy examples from unlabeled text, or by imposing constraints over interdependent label decisions.A plethora of methods have been proposed, each with respective strengths and limitations.Probabilistic logic offers a unifying language to represent indirect supervision, but end-to-end modeling with probabilistic logic is often infeasible due to intractable inference and learning.In this paper, we propose deep probabilistic logic (DPL) as a general framework for indirect supervision, by composing probabilistic logic with deep learning.DPL models label decisions as latent variables, represents prior knowledge on their relations using weighted first-order logical formulas, and alternates between learning a deep neural network for the end task and refining uncertain formula weights for indirect supervision, using variational EM.This framework subsumes prior indirect supervision methods as special cases, and enables novel combination via infusion of rich domain and linguistic knowledge.Experiments on biomedical machine reading demonstrate the promise of this approach. Hoifung Poon |
EMNLP | 2 |
| 2018 | EZLearn: Exploiting Organic Supervision in Automated Data AnnotationabstractMany real-world applications require automated data annotation, such as identifying tissue origins based on gene expressions and classifying images into semantic categories. Annotation classes are often numerous and subject to changes over time, and annotating examples has become the major bottleneck for supervised learning methods. In science and other high-value domains, large repositories of data samples are often available, together with two sources of organic supervision: a lexicon for the annotation classes, and text descriptions that accompany some data samples. Distant supervision has emerged as a promising paradigm for exploiting such indirect supervision by automatically annotating examples where the text description contains a class mention in the lexicon. However, due to linguistic variations and ambiguities, such training data is inherently noisy, which limits the accuracy in this approach. In this paper, we introduce an auxiliary natural language processing system for the text modality, and incorporate co-training to reduce noise and augment signal in distant supervision. Without using any manually labeled data, our EZLearn system learned to accurately annotate data samples in functional genomics and scientific figure comprehension, substantially outperforming state-of-the-art supervised methods trained on tens of thousands of annotated examples. Maxim Grechkin, Hoifung Poon, Bill Howe |
IJCAI | 2 |
| 2017 | Distant Supervision for Relation Extraction beyond the Sentence BoundaryabstractThe growing demand for structured knowledge has led to great interest in relation extraction, especially in cases with limited supervision.However, existing distance supervision approaches only extract relations expressed in single sentences.In general, cross-sentence relation extraction is under-explored, even in the supervised-learning setting.In this paper, we propose the first approach for applying distant supervision to crosssentence relation extraction.At the core of our approach is a graph representation that can incorporate both standard dependencies and discourse relations, thus providing a unifying way to model relations within and across sentences.We extract features from multiple paths in this graph, increasing accuracy and robustness when confronted with linguistic variation and analysis error.Experiments on an important extraction task for precision medicine show that our approach can learn an accurate cross-sentence extractor, using only a small existing knowledge base and unlabeled text from biomedical research articles.Compared to the existing distant supervision paradigm, our approach extracted twice as many relations at similar precision, thus demonstrating the prevalence of cross-sentence relations and the promise of our approach. Chris Quirk, Hoifung Poon |
EACL (1) | 2 |
| 2017 | Estimating Accuracy from Unlabeled Data: A Probabilistic Logic ApproachabstractWe propose an efficient method to estimate the accuracy of classifiers using only unlabeled data. We consider a setting with multiple classification problems where the target classes may be tied together through logical constraints. For example, a set of classes may be mutually exclusive, meaning that a data instance can belong to at most one of them. The proposed method is based on the intuition that: (i) when classifiers agree, they are more likely to be correct, and (ii) when the classifiers make a prediction that violates the constraints, at least one classifier must be making an error. Experiments on four real-world data sets produce accuracy estimates within a few percent of the true accuracy, using solely unlabeled data. Our models also outperform existing state-of-the-art solutions in both estimating accuracies, and combining multiple classifier outputs. The results emphasize the utility of logical constraints in estimating accuracy, thus validating our intuition. Emmanouil A. Platanios, Hoifung Poon, Tom M. Mitchell, Eric Horvitz |
NIPS | 2 |
| 2017 | Cross-Sentence N-ary Relation Extraction with Graph LSTMsabstractPast work in relation extraction has focused on binary relations in single sentences. Recent NLP inroads in high-value domains have sparked interest in the more general setting of extracting n-ary relations that span multiple sentences. In this paper, we explore a general relation extraction framework based on graph long short-term memory networks (graph LSTMs) that can be easily extended to cross-sentence n-ary relation extraction. The graph formulation provides a unified way of exploring different LSTM approaches and incorporating various intra-sentential and inter-sentential dependencies, such as sequential, syntactic, and discourse relations. A robust contextual representation is learned for the entities, which serves as input to the relation classifier. This simplifies handling of relations with arbitrary arity, and enables multi-task learning with related relations. We evaluate this framework in two important precision medicine settings, demonstrating its effectiveness with both conventional supervised learning and distant supervision. Cross-sentence extraction produced larger knowledge bases. and multi-task learning significantly improved extraction accuracy. A thorough analysis of various LSTM approaches yielded useful insight the impact of linguistic analysis on extraction accuracy. Nanyun Peng 0001, Hoifung Poon, Chris Quirk, Kristina Toutanova, Scott Yih |
Trans. Assoc. Comput. Linguistics | 2 |
| 2016 | Compositional Learning of Embeddings for Relation Paths in Knowledge Base and TextabstractModeling relation paths has offered significant gains in embedding models for knowledge base (KB) completion.However, enumerating paths between two entities is very expensive, and existing approaches typically resort to approximation with a sampled subset.This problem is particularly acute when text is jointly modeled with KB relations and used to provide direct evidence for facts mentioned in it.In this paper, we propose the first exact dynamic programming algorithm which enables efficient incorporation of all relation paths of bounded length, while modeling both relation types and intermediate nodes in the compositional path representations.We conduct a theoretical analysis of the efficiency gain from the approach.Experiments on two datasets show that it addresses representational limitations in prior approaches and improves accuracy in KB completion. Kristina Toutanova, Xi Victoria Lin, Scott Yih, Hoifung Poon, Chris Quirk |
ACL (1) | 4 |
| 2016 | Unifying Logical and Statistical AIabstractIntelligent agents must be able to handle the complexity and uncertainty of the real world. Logical AI has focused mainly on the former, and statistical AI on the latter. Markov logic combines the two by attaching weights to first-order formulas and viewing them as templates for features of Markov networks. Inference algorithms for Markov logic draw on ideas from satisfiability, Markov chain Monte Carlo and knowledge-based model construction. Learning algorithms are based on the voted perceptron, pseudo-likelihood and inductive logic programming. Markov logic has been successfully applied to a wide variety of problems in natural language understanding, vision, computational biology, social networks and others, and is the basis of the open-source Alchemy system. Pedro M. Domingos, Daniel Lowd, Stanley Kok, Aniruddh Nath, Hoifung Poon, Matthew Richardson, Parag Singla |
LICS | 5 |
| 2015 | Model Selection for Type-Supervised Learning with Application to POS TaggingabstractModel selection (picking, for example, the feature set and the regularization strength) is crucial for building high-accuracy NLP models.In supervised learning, we can estimate the accuracy of a model on a subset of the labeled data and choose the model with the highest accuracy.In contrast, here we focus on type-supervised learning, which uses constraints over the possible labels for word types for supervision, and labeled data is either not available or very small.For the setting where no labeled data is available, we perform a comparative study of previously proposed and one novel model selection criterion on type-supervised POS-tagging in nine languages.For the setting where a small labeled set is available, we show that the set should be used for semi-supervised learning rather than for model selection onlyusing it for model selection reduces the error by less than 5%, whereas using it for semi-supervised learning reduces the error by 44%. Kristina Toutanova, Waleed Ammar, Pallavi Choudhury, Hoifung Poon |
CoNLL | 4 |
| 2015 | Representing Text for Joint Embedding of Text and Knowledge BasesabstractModels that learn to represent textual and knowledge base relations in the same continuous latent space are able to perform joint inferences among the two kinds of relations and obtain high accuracy on knowledge base completion (Riedel et al., 2013).In this paper we propose a model that captures the compositional structure of textual relations, and jointly optimizes entity, knowledge base, and textual relation representations.The proposed model significantly improves performance over a model that does not share parameters among textual relations with common sub-structure. Kristina Toutanova, Danqi Chen 0001, Patrick Pantel, Hoifung Poon, Pallavi Choudhury, Michael Gamon |
EMNLP | 4 |
| 2015 | Grounded Semantic Parsing for Complex Knowledge ExtractionabstractRecently, there has been increasing interest in learning semantic parsers with indirect supervision, but existing work focuses almost exclusively on question answering.Separately, there have been active pursuits in leveraging databases for distant supervision in information extraction, yet such methods are often limited to binary relations and none can handle nested events.In this paper, we generalize distant supervision to complex knowledge extraction, by proposing the first approach to learn a semantic parser for extracting nested event structures without annotated examples, using only a database of such complex events and unannotated text.The key idea is to model the annotations as latent variables, and incorporate a prior that favors semantic parses containing known events.Experiments on the GENIA event extraction dataset show that our approach can learn from and extract complex biological pathway events.Moreover, when supplied with just five example words per event type, it becomes competitive even among supervised systems, outperforming 19 out of 24 teams that participated in the original shared task. Ankur P. Parikh, Hoifung Poon, Kristina Toutanova |
HLT-NAACL | 2 |
| 2014 | Literome: PubMed-scale genomic knowledge base in the cloudabstractMOTIVATION: Advances in sequencing technology have led to an exponential growth of genomics data, yet it remains a formidable challenge to interpret such data for identifying disease genes and drug targets. There has been increasing interest in adopting a systems approach that incorporates prior knowledge such as gene networks and genotype-phenotype associations. The majority of such knowledge resides in text such as journal publications, which has been undergoing its own exponential growth. It has thus become a significant bottleneck to identify relevant knowledge for genomic interpretation as well as to keep up with new genomics findings. RESULTS: In the Literome project, we have developed an automatic curation system to extract genomic knowledge from PubMed articles and made this knowledge available in the cloud with a Web site to facilitate browsing, searching and reasoning. Currently, Literome focuses on two types of knowledge most pertinent to genomic medicine: directed genic interactions such as pathways and genotype-phenotype associations. Users can search for interacting genes and the nature of the interactions, as well as diseases and drugs associated with a single nucleotide polymorphism or gene. Users can also search for indirect connections between two entities, e.g. a gene and a disease might be linked because an interacting gene is associated with a related disease. AVAILABILITY AND IMPLEMENTATION: Literome is freely available at literome.azurewebsites.net. Download for non-commercial use is available via Web services. Hoifung Poon, Chris Quirk, Charlie DeZiel, David Heckerman |
Bioinform. | 1 |
| 2013 | Grounded Unsupervised Semantic Parsing
Hoifung Poon |
ACL (1) | 1 |
| 2013 | Probabilistic Frame Induction
Jackie Chi Kit Cheung, Hoifung Poon, Lucy Vanderwende |
HLT-NAACL | 2 |
| 2011 | Sum-Product Networks: A New Deep Architecture
Hoifung Poon, Pedro M. Domingos |
UAI | 1 |
| 2010 | Unsupervised Ontology Induction from Text
Hoifung Poon, Pedro M. Domingos |
ACL | 1 |
| 2010 | Joint Inference for Knowledge Extraction from Biomedical Literature
Hoifung Poon, Lucy Vanderwende |
HLT-NAACL | 1 |
| 2009 | Language ID in the Context of Harvesting Language Data off the Web
Fei Xia 0004, William D. Lewis, Hoifung Poon |
EACL | 3 |
| 2009 | Unsupervised Semantic Parsing
Hoifung Poon, Pedro M. Domingos |
EMNLP | 1 |
| 2009 | Unsupervised Morphological Segmentation with Log-Linear Models
Hoifung Poon, Colin Cherry, Kristina Toutanova |
HLT-NAACL | 1 |
| 2008 | A General Method for Reducing the Complexity of Relational Inference and its Application to MCMC
Hoifung Poon, Pedro M. Domingos, Marc Sumner |
AAAI | 1 |
| 2008 | Joint Unsupervised Coreference Resolution with Markov Logic
Hoifung Poon, Pedro M. Domingos |
EMNLP | 1 |
| 2007 | Joint Inference in Information Extraction
Hoifung Poon, Pedro M. Domingos |
AAAI | 1 |
| 2006 | Unifying Logical and Statistical AI
Pedro M. Domingos, Stanley Kok, Hoifung Poon, Matthew Richardson, Parag Singla |
AAAI | 3 |
| 2006 | Sound and Efficient Inference with Probabilistic and Deterministic Dependencies
Hoifung Poon, Pedro M. Domingos |
AAAI | 1 |