VLDB 2026 Research / reviewers in the wild / expert
Sheng Zhang 0012
dblp:69/6137-12
· DBLP profile ↗
24ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0003-3672-5436ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 7 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MuirBench: A Comprehensive Benchmark for Robust Multi-image UnderstandingabstractWe introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tasks (e.g., scene understanding, ordering) that involve 10 categories of multi-image relations (e.g., multiview, temporal relations). Comprising 11,264 images and 2,600 multiple-choice questions, MuirBench is created in a pairwise manner, where each standard instance is paired with an unanswerable variant that has minimal semantic differences, in order for a reliable assessment. Evaluated upon 20 recent multi-modal LLMs, our results reveal that even the best-performing models like GPT-4o and Gemini Pro find it challenging to solve MuirBench, achieving 68.0% and 49.3% in accuracy. Open-source multimodal LLMs trained on single images can hardly generalize to multi-image questions, hovering below 33.3% in accuracy. These results highlight the importance of MuirBench in encouraging the community to develop multimodal LLMs that can look beyond a single image, suggesting potential pathways for future improvements. Fei Wang 0060, James Y. Huang, Zekun Li 0007, Qin Liu 0010, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu 0014, Wenxuan Zhou 0002, Kai Zhang 0008, Tianyi Lorena Yan, Wenjie Mo 0001, Hsiang-Hui Liu, Pan Lu, Chunyuan Li, Chaowei Xiao, Kai-Wei Chang 0001, Dan Roth 0001, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001 |
ICLR | 19 |
| 2025 | From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context LearningabstractNan Xu, Fei Wang, Sheng Zhang, Hoifung Poon, Muhao Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Nan Xu 0014, Fei Wang 0060, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001 |
NAACL (Long Papers) | 3 |
| 2024 | DocLens: Multi-aspect Fine-grained Medical Text EvaluationabstractYiqing Xie, Sheng Zhang, Hao Cheng, Pengfei Liu, Zelalem Gero, Cliff Wong, Tristan Naumann, Hoifung Poon, Carolyn Rose. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yiqing Xie, Sheng Zhang 0012, Hao Cheng 0002, Zelalem Gero, Cliff Wong, Tristan Naumann, Hoifung Poon, Carolyn P. Rosé |
ACL (1) | 2 |
| 2024 | mDPO: Conditional Preference Optimization for Multimodal Large Language ModelsabstractDirect preference optimization (DPO) has shown to be an effective method for large language model (LLM) alignment.Recent works have attempted to apply DPO to multimodal scenarios but have found it challenging to achieve consistent improvement.Through a comparative experiment, we identify the unconditional preference problem in multimodal preference optimization, where the model overlooks the image condition.To address this problem, we propose MDPO, a multimodal DPO objective that prevents the over-prioritization of language-only preferences by also optimizing image preference.Moreover, we introduce a reward anchor that forces the reward to be positive for chosen responses, thereby avoiding the decrease in their likelihood-an intrinsic problem of relative preference optimization.Experiments on two multimodal LLMs of different sizes and three widely used benchmarks demonstrate that MDPO effectively addresses the unconditional preference problem in multimodal preference optimization and significantly improves model performance, particularly in reducing hallucination. Fei Wang 0060, Wenxuan Zhou 0002, James Y. Huang, Nan Xu 0014, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001 |
EMNLP | 5 |
| 2024 | UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionabstractLarge language models (LLMs) have demonstrated remarkable generalizability, such as understanding arbitrary entities and relations. Instruction tuning has proven effective for distilling LLMs into more cost-efficient models such as Alpaca and Vicuna. Yet such student models still trail the original LLMs by large margins in downstream applications. In this paper, we explore targeted distillation with mission-focused instruction tuning to train student models that can excel in a broad application class such as open information extraction. Using named entity recognition (NER) for case study, we show how ChatGPT can be distilled into much smaller UniversalNER models for open NER. For evaluation, we assemble the largest NER benchmark to date, comprising 43 datasets across 9 diverse domains such as biomedicine, programming, social media, law, finance. Without using any direct supervision, UniversalNER attains remarkable NER accuracy across tens of thousands of entity types, outperforming general instruction-tuned models such as Alpaca and Vicuna by over 30 absolute F1 points in average. With a tiny fraction of parameters, UniversalNER not only acquires ChatGPT's capability in recognizing arbitrary entity types, but also outperforms its NER accuracy by 7-9 absolute F1 points in average. Remarkably, UniversalNER even outperforms by a large margin state-of-the-art multi-task instruction-tuned systems such as InstructUIE, which uses supervised NER examples. We also conduct thorough ablation studies to assess the impact of various components in our distillation approach. We release the distillation recipe, data, and UniversalNER models to facilitate future research on targeted distillation. Wenxuan Zhou 0002, Sheng Zhang 0012, Yu Gu 0017, Muhao Chen 0001, Hoifung Poon |
ICLR | 2 |
| 2023 | Continual Contrastive Finetuning Improves Low-Resource Relation ExtractionabstractRelation extraction (RE), which has relied on structurally annotated corpora for model training, has been particularly challenging in lowresource scenarios and domains.Recent literature has tackled low-resource RE by selfsupervised learning, where the solution involves pretraining the entity pair embedding by RE-based objective and finetuning on labeled data by classification-based objective.However, a critical challenge to this approach is the gap in objectives, which prevents the RE model from fully utilizing the knowledge in pretrained representations.In this paper, we aim at bridging the gap and propose to pretrain and finetune the RE model using consistent objectives of contrastive learning.Since in this kind of representation learning paradigm, one relation may easily form multiple clusters in the representation space, we further propose a multi-center contrastive loss that allows one relation to form multiple clusters to better align with pretraining.Experiments on two document-level RE datasets, BioRED and Re-DocRED, demonstrate the effectiveness of our method.Particularly, when using 1% end-task training data, our method outperforms PLMbased RE classifier by 10.5% and 6.1% on the two datasets, respectively. Wenxuan Zhou 0002, Sheng Zhang 0012, Tristan Naumann, Muhao Chen 0001, Hoifung Poon |
ACL (1) | 2 |
| 2023 | Optimizing Bi-Encoder for Named Entity Recognition via Contrastive Learning
Sheng Zhang 0012, Hao Cheng 0002, Jianfeng Gao 0001, Hoifung Poon |
ICLR | 1 |
| 2023 | Precision Health in the Age of Large Language ModelsabstractMedicine today is imprecise. Among the top 20 drugs in the U.S., up to 80% of patients are non-responders. The goal of precision health is to provide the right intervention for the right people at the right time. The key to realize this dream is to develop a data-driven, learning system that can instantly incorporate new health information to optimize care delivery and accelerate biomedical discovery. In reality, however, the health ecosystem is mired in overwhelming unstructured data and excruciating manual processing. For example, in cancer, standard of care often fails, and clinical trials are the last hope. Yet less than 3% of patients could find a matching trial, whereas 40% of trial failures simply stem from insufficient recruitment. Discovery is painfully slow as a new drug may take billions of dollars and over a decade to develop. Hoifung Poon, Tristan Naumann, Sheng Zhang 0012, Javier González Hernández |
KDD | 3 |
| 2023 | LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One DayabstractConversational generative AI has demonstrated remarkable promise for empowering biomedical practitioners, but current investigations focus on unimodal text. Multimodal conversational AI has seen rapid progress by leveraging billions of image-text pairs from the public web, but such general-domain vision-language models still lack sophistication in understanding and conversing about biomedical images. In this paper, we propose a cost-efficient approach for training a vision-language conversational assistant that can answer open-ended research questions of biomedical images. The key idea is to leverage a large-scale, broad-coverage biomedical figure-caption dataset extracted from PubMed Central, use GPT-4 to self-instruct open-ended instruction-following data from the captions, and then fine-tune a large general-domain vision-language model using a novel curriculum learning method. Specifically, the model first learns to align biomedical vocabulary using the figure-caption pairs as is, then learns to master open-ended conversational semantics using GPT-4 generated instruction-following data, broadly mimicking how a layperson gradually acquires biomedical knowledge. This enables us to train a Large Language and Vision Assistant for BioMedicine (LLaVA-Med) in less than 15 hours (with eight A100s). LLaVA-Med exhibits excellent multimodal conversational capability and can follow open-ended instruction to assist with inquiries about a biomedical image. On three standard biomedical visual question answering datasets, LLaVA-Med outperforms previous supervised state-of-the-art on certain metrics. To facilitate biomedical multimodal research, we will release our instruction-following data and the LLaVA-Med model. Chunyuan Li, Cliff Wong, Sheng Zhang 0012, Naoto Usuyama, Tristan Naumann, Hoifung Poon, Jianfeng Gao 0001 |
NeurIPS | 3 |
| 2023 | Compositional Zero-Shot Domain Transfer with Text-to-Text ModelsabstractAbstract Label scarcity is a bottleneck for improving task performance in specialized domains. We propose a novel compositional transfer learning framework (DoT51) for zero-shot domain transfer. Without access to in-domain labels, DoT5 jointly learns domain knowledge (from masked language modelling of unlabelled in-domain free text) and task knowledge (from task training on more readily available general-domain data) in a multi-task manner. To improve the transferability of task training, we design a strategy named NLGU: We simultaneously train natural language generation (NLG) for in-domain label-to-data generation, which enables data augmentation for self-finetuning and natural language understanding (NLU) for label prediction. We evaluate DoT5 on the biomedical domain and the resource-lean subdomain of radiology, focusing on natural language inference, text summarization, and embedding learning. DoT5 demonstrates the effectiveness of compositional transfer learning through multi-task learning. In particular, DoT5 outperforms the current state-of-the-art in zero-shot transfer by over 7 absolute points in accuracy on RadNLI. We validate DoT5 with ablations and a case study demonstrating its ability to solve challenging NLI examples requiring in-domain expertise. Fangyu Liu 0001, Qianchu Liu, Shruthi Bannur, Fernando Pérez-García, Naoto Usuyama, Sheng Zhang 0012, Tristan Naumann, Aditya V. Nori, Hoifung Poon, Javier Alvarez-Valle, Ozan Oktay, Stephanie L. Hyland |
Trans. Assoc. Comput. Linguistics | 6 |
| 2022 | BioGPT: generative pre-trained transformer for biomedical text generation and miningabstractPre-trained language models have attracted increasing attention in the biomedical domain, inspired by their great success in the general natural language domain. Among the two main branches of pre-trained language models in the general language domain, i.e. BERT (and its variants) and GPT (and its variants), the first one has been extensively studied in the biomedical domain, such as BioBERT and PubMedBERT. While they have achieved great success on a variety of discriminative downstream biomedical tasks, the lack of generation ability constrains their application scope. In this paper, we propose BioGPT, a domain-specific generative Transformer language model pre-trained on large-scale biomedical literature. We evaluate BioGPT on six biomedical natural language processing tasks and demonstrate that our model outperforms previous models on most tasks. Especially, we get 44.98%, 38.42% and 40.76% F1 score on BC5CDR, KD-DTI and DDI end-to-end relation extraction tasks, respectively, and 78.2% accuracy on PubMedQA, creating a new record. Our case study on text generation further demonstrates the advantage of BioGPT on biomedical literature to generate fluent descriptions for biomedical terms. Renqian Luo, Liai Sun, Yingce Xia, Tao Qin 0001, Sheng Zhang 0012, Hoifung Poon, Tie-Yan Liu |
Briefings Bioinform. | 5 |
| 2021 | Modular Self-Supervision for Document-Level Relation ExtractionabstractExtracting relations across large text spans has been relatively underexplored in NLP, but it is particularly important for high-value domains such as biomedicine, where obtaining high recall of the latest findings is crucial for practical applications.Compared to conventional information extraction confined to short text spans, document-level relation extraction faces additional challenges in both inference and learning.Given longer text spans, state-of-the-art neural architectures are less effective and taskspecific self-supervision such as distant supervision becomes very noisy.In this paper, we propose decomposing document-level relation extraction into relation detection and argument resolution, taking inspiration from Davidsonian semantics.This enables us to incorporate explicit discourse modeling and leverage modular self-supervision for each sub-problem, which is less noise-prone and can be further refined end-to-end via variational EM.We conduct a thorough evaluation in biomedical machine reading for precision oncology, where cross-paragraph relation mentions are prevalent.Our method outperforms prior state of the art, such as multi-scale learning and graph neural networks, by over 20 absolute F1 points.The gain is particularly pronounced among the most challenging relation instances whose arguments never co-occur in a paragraph. Sheng Zhang 0012, Cliff Wong, Naoto Usuyama, Tristan Naumann, Hoifung Poon |
EMNLP (1) | 1 |
| 2021 | Joint Universal Syntactic and Semantic ParsingabstractWhile numerous attempts have been made to jointly parse syntax and semantics, high performance in one domain typically comes at the price of performance in the other. This trade-off contradicts the large body of research focusing on the rich interactions at the syntax–semantics interface. We explore multiple model architectures that allow us to exploit the rich syntactic and semantic annotations contained in the Universal Decompositional Semantics (UDS) dataset, jointly parsing Universal Dependencies and UDS to obtain state-of-the-art results in both formalisms. We analyze the behavior of a joint model of syntax and semantics, finding patterns supported by linguistic theory at the syntax–semantics interface. We then investigate to what degree joint modeling generalizes to a multilingual setting, where we find similar trends across 8 languages. Elias Stengel-Eskin, Kenton Murray, Sheng Zhang 0012, Aaron Steven White, Benjamin Van Durme |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | Universal Decompositional Semantic ParsingabstractWe introduce a transductive model for parsing into Universal Decompositional Semantics (UDS) representations, which jointly learns to map natural language utterances into UDS graph structures and annotate the graph with decompositional semantic attribute scores.We also introduce a strong pipeline model for parsing into the UDS graph structure, and show that our transductive parser performs comparably while additionally performing attribute prediction.By analyzing the attribute prediction errors, we find the model captures natural relationships between attribute groups. Elias Stengel-Eskin, Aaron Steven White, Sheng Zhang 0012, Benjamin Van Durme |
ACL | 3 |
| 2020 | The Universal Decompositional Semantics Dataset and Decomp ToolkitabstractWe present the Universal Decompositional Semantics (UDS) dataset (v1.0), which is bundled with the Decomp toolkit (v0.1). UDS1.0 unifies five high-quality, decompositional semantics-aligned annotation sets within a single semantic graph specification—with graph structures defined by the predicative patterns produced by the PredPatt tool and real-valued node and edge attributes constructed using sophisticated normalization procedures. The Decomp toolkit provides a suite of Python 3 tools for querying UDS graphs using SPARQL. Both UDS1.0 and Decomp0.1 are publicly available at http://decomp.io. Aaron Steven White, Elias Stengel-Eskin, Siddharth Vashishtha, Venkata Subrahmanyan Govindarajan, Dee Ann Reisinger, Tim Vieira, Keisuke Sakaguchi, Sheng Zhang 0012, Francis Ferraro, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme |
LREC | 8 |
| 2019 | AMR Parsing as Sequence-to-Graph TransductionabstractWe propose an attention-based model that treats AMR parsing as sequence-to-graph transduction.Unlike most AMR parsers that rely on pre-trained aligners, external semantic resources, or data augmentation, our proposed parser is aligner-free, and it can be effectively trained with limited amounts of labeled AMR data.Our experimental results outperform all previously reported SMATCH scores, on both AMR 2.0 (76.3% F1 on LDC2017T10) and AMR 1.0 (70.2% F1 on LDC2014T12). Sheng Zhang 0012, Xutai Ma, Kevin Duh, Benjamin Van Durme |
ACL (1) | 1 |
| 2019 | Broad-Coverage Semantic Parsing as TransductionabstractSheng Zhang, Xutai Ma, Kevin Duh, Benjamin Van Durme. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sheng Zhang 0012, Xutai Ma, Kevin Duh, Benjamin Van Durme |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Neural-Davidsonian Semantic Proto-role LabelingabstractWe present a model for semantic proto-role labeling (SPRL) using an adapted bidirectional LSTM encoding strategy that we call Neural-Davidsonian: predicate-argument structure is represented as pairs of hidden states corresponding to predicate and argument head tokens of the input sequence.We demonstrate:(1) state-of-the-art results in SPRL, and (2) that our network naturally shares parameters between attributes, allowing for learning new attribute types with limited added supervision. Rachel Rudinger, Adam R. Teichert, Ryan Culkin, Sheng Zhang 0012, Benjamin Van Durme |
EMNLP | 4 |
| 2018 | Cross-lingual Decompositional Semantic ParsingabstractWe introduce the task of cross-lingual decompositional semantic parsing: mapping content provided in a source language into a decompositional semantic analysis based on a target language.We present: (1) a form of decompositional semantic analysis designed to allow systems to target varying levels of structural complexity (shallow to deep analysis), (2) an evaluation metric to measure the similarity between system output and reference semantic analysis, (3) an end-to-end model with a novel annotating mechanism that supports intra-sentential coreference, and (4) an evaluation dataset on which our model outperforms strong baselines by at least 1.75 F 1 score. Sheng Zhang 0012, Xutai Ma, Rachel Rudinger, Kevin Duh, Benjamin Van Durme |
EMNLP | 1 |
| 2017 | Selective Decoding for Cross-lingual Open Information ExtractionabstractCross-lingual open information extraction is the task of distilling facts from the source language into representations in the target language. We propose a novel encoder-decoder model for this problem. It employs a novel selective decoding mechanism, which explicitly models the sequence labeling process as well as the sequence generation process on the decoder side. Compared to a standard encoder-decoder model, selective decoding significantly increases the performance on a Chinese-English cross-lingual open IE dataset by 3.87-4.49 BLEU and 1.91-5.92 F1. We also extend our approach to low-resource scenarios, and gain promising improvement. Sheng Zhang 0012, Kevin Duh, Benjamin Van Durme |
IJCNLP(1) | 1 |
| 2017 | Ordinal Common-sense InferenceabstractHumans have the capacity to draw common-sense inferences from natural language: various things that are likely but not certain to hold based on established discourse, and are rarely stated explicitly. We propose an evaluation of automated common-sense inference based on an extension of recognizing textual entailment: predicting ordinal human responses on the subjective likelihood of an inference holding in a given context. We describe a framework for extracting common-sense knowledge from corpora, which is then used to construct a dataset for this ordinal entailment task. We train a neural sequence-to-sequence model on this dataset, which we use to score and generate possible inferences. Further, we annotate subsets of previously established datasets via our ordinal annotation protocol in order to then analyze the distinctions between these and what we have constructed. Sheng Zhang 0012, Rachel Rudinger, Kevin Duh, Benjamin Van Durme |
Trans. Assoc. Comput. Linguistics | 1 |
| 2016 | Universal Decompositional Semantics on Universal DependenciesabstractAaron Steven White, Drew Reisinger, Keisuke Sakaguchi, Tim Vieira, Sheng Zhang, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2016. Aaron Steven White, Dee Ann Reisinger, Keisuke Sakaguchi, Tim Vieira, Sheng Zhang 0012, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme |
EMNLP | 5 |
| 2015 | What Is the Longest River in the USA? Semantic Parsing for Aggregation QuestionsabstractAnswering natural language questions against structured knowledge bases (KB) has been attracting increasing attention in both IR and NLP communities. The task involves two main challenges: recognizing the questions' meanings, which are then grounded to a given KB. Targeting simple factoid questions, many existing open domain semantic parsers jointly solve these two subtasks, but are usually expensive in complexity and resources.In this paper, we propose a simple pipeline framework to efficiently answer more complicated questions, especially those implying aggregation operations, e.g., argmax, argmin.We first develop a transition-based parsing model to recognize the KB-independent meaning representation of the user's intention inherent in the question. Secondly, we apply a probabilistic model to map the meaning representation, including those aggregation functions, to a structured query.The experimental results showed that our method can better understand aggregation questions, outperforming the state-of-the-art methods on the Free917 dataset while still maintaining promising performance on a more challenging dataset, WebQuestions, without extra training. Kun Xu 0005, Sheng Zhang 0012, Yansong Feng 0002, Songfang Huang, Dongyan Zhao 0001 |
AAAI | 2 |
| 2014 | Answering Natural Language Questions via Phrasal Semantic Parsing
Kun Xu 0005, Sheng Zhang 0012, Yansong Feng 0002, Dongyan Zhao 0001 |
NLPCC | 2 |