VLDB 2026 Research / reviewers in the wild / expert
Shiyue Zhang 0001
dblp:186/8393-1
· DBLP profile ↗
17ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0001-7027-9076ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Localizing Factual Inconsistencies in Attributable Text GenerationabstractAbstract There has been an increasing interest in detecting hallucinations in model-generated texts, both manually and automatically, at varying levels of granularity. However, most existing methods fail to precisely pinpoint the errors. In this work, we introduce QASemConsistency, a new formalism for localizing factual inconsistencies in attributable text generation, at a fine-grained level. Drawing inspiration from Neo-Davidsonian formal semantics, we propose decomposing the generated text into minimal predicate-argument level propositions, expressed as simple question-answer (QA) pairs, and assess whether each individual QA pair is supported by a trusted reference text. As each QA pair corresponds to a single semantic relation between a predicate and an argument, QASemConsistency effectively localizes the unsupported information. We first demonstrate the effectiveness of the QASemConsistency methodology for human annotation, by collecting crowdsourced annotations of granular consistency errors, while achieving a substantial inter-annotator agreement. This benchmark includes more than 3K instances spanning various tasks of attributable text generation. We also show that QASemConsistency yields factual consistency scores that correlate well with human judgments. Finally, we implement several methods for automatically detecting localized factual inconsistencies, with both supervised entailment models and LLMs.1 Arie Cattan, Paul Roit, Shiyue Zhang 0001, David Wan, Roee Aharoni, Idan Szpektor, Mohit Bansal, Ido Dagan |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | Improving Instruct Models for Free: A Study on Partial AdaptationabstractInstruct models, obtained from various instruction tuning or post-training steps, are commonly deemed superior and more usable than their base counterpart.While the model gains instruction following ability, instruction tuning may lead to forgetting the knowledge from pre-training or it may encourage the model to become overly conversational or verbose.This, in turn, can lead to degradation of in-context few-shot learning performance.In this work, we study the performance trajectory between base and instruct models by scaling down the strength of instruction-tuning via the partial adaption method.We show that, across several model families and model sizes, reducing the strength of instruction-tuning results in material improvement on a few-shot in-context learning benchmark covering a variety of classic natural language tasks.This comes at the cost of losing some degree of instruction following ability as measured by AlpacaEval.Our study shines light on the potential trade-off between in-context learning and instruction following abilities that is worth considering in practice. Ozan Irsoy, Pengxiang Cheng 0001, Jennifer L. Chen, Daniel Preotiuc-Pietro, Shiyue Zhang 0001, Duccio Pappadopulo |
EMNLP | 5 |
| 2025 | RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language ModelsabstractBang An, Shiyue Zhang, Mark Dredze. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shiyue Zhang 0001, Mark Dredze |
NAACL (Long Papers) | 2 |
| 2023 | Extractive is not Faithful: An Investigation of Broad Unfaithfulness Problems in Extractive SummarizationabstractThe problems of unfaithful summaries have been widely discussed under the context of abstractive summarization.Though extractive summarization is less prone to the common unfaithfulness issues of abstractive summaries, does that mean extractive is equal to faithful?Turns out that the answer is no.In this work, we define a typology with five types of broad unfaithfulness problems (including and beyond not-entailment) that can appear in extractive summaries, including incorrect coreference, incomplete coreference, incorrect discourse, incomplete discourse, as well as other misleading information.We ask humans to label these problems out of 1600 English summaries produced by 16 diverse extractive systems.We find that 30% of the summaries have at least one of the five issues.To automatically detect these problems, we find that 5 existing faithfulness evaluation metrics for summarization have poor correlations with human judgment.To remedy this, we propose a new metric, EXTEVAL, that is designed for detecting unfaithful extractive summaries and is shown to have the best performance.We hope our work can increase the awareness of unfaithfulness problems in extractive summarization and help future work to evaluate and resolve these issues.1 * Equal contribution. 1 Our data and code are publicly available at https: //github.com/ZhangShiyue/extractive_is_ not_faithful. Document:(CNN) Most climbers who try don't succeed in summiting the 29,035-foot-high Mount Everest, the world's tallest peak.But they do leave their trash.Thousands of pounds of it.That's why an experienced climbing group from the Indian army plans to trek up the 8,850-meter mountain to pick up at least 4,000 kilograms (more than 8,000 pounds) of waste from the high-altitude camps, according to India Today.The mountain is part of the Himalaya mountain range on the border between Nepal and the Tibet region.The 34-member team plans to depart for Kathmandu on Saturday and start the ascent in mid-May.The upcoming trip marks the 50th anniversary of the first Indian team to scale Mount Everest [...]More than 200 climbers have died attempting to climb the peak, part of a UNESCO World Heritage Site.The Indian expedition isn't the first attempt to clean up the trash left by generations of hikers[...] Summary 1 (incorrect coreference): (CNN) Most climbers who try don't succeed in summiting the 29,035-foot-high Mount Everest, the world's tallest peak.That's why an experienced climbing group from the Indian army plans to trek up the 8,850-meter mountain to pick up at least 4,000 kilograms (more than 8,000 pounds) of waste from the high-altitude camps, according to India Today.[...] Summary 2 (incomplete coreference & incorrect discourse) : That's why an experienced climbing group from the Indian army plans to trek up the 8,850-meter mountain to pick up at least 4,000 kilograms More than 200 climbers have died to clean up the trash [...] Summary 3 (incomplete discourse & incomplete coreference): But they do leave their trash.Thousands of pounds of it.[... Shiyue Zhang 0001, David Wan, Mohit Bansal |
ACL (1) | 1 |
| 2023 | MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-EntropiesabstractShiyue Zhang, Shijie Wu, Ozan Irsoy, Steven Lu, Mohit Bansal, Mark Dredze, David Rosenberg. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shiyue Zhang 0001, Ozan Irsoy, Steven Lu 0003, Mohit Bansal, Mark Dredze, David S. Rosenberg |
ACL (1) | 1 |
| 2023 | HistAlign: Improving Context Dependency in Language Generation by Aligning with HistoryabstractLanguage models (LMs) can generate hallucinations and incoherent outputs, which highlights their weak context dependency.Cache-LMs, which augment LMs with a memory of recent history, can increase context dependency and have shown remarkable performance in diverse language generation tasks.However, we find that even with training, the performance gain stemming from the cache component of current cache-LMs is suboptimal due to the misalignment between the current hidden states and those stored in the memory.In this work, we present HISTALIGN, a new training approach to ensure good cache alignment such that the model receives useful signals from the history.We first prove our concept on a simple and synthetic task where the memory is essential for correct predictions, and we show that the cache component of HISTALIGN is better aligned and improves overall performance.Next, we evaluate HISTALIGN on diverse downstream language generation tasks, including prompt continuation, abstractive summarization, and data-to-text.We demonstrate that HISTALIGN improves text coherence and faithfulness in open-ended and conditional generation settings, respectively.HISTALIGN is also generalizable across different model families, showcasing its strength in improving context dependency of LMs in diverse scenarios.1 David Wan, Shiyue Zhang 0001, Mohit Bansal |
EMNLP | 2 |
| 2023 | Summarization Programs: Interpretable Abstractive Summarization with Neural Modular Trees
Swarnadeep Saha, Shiyue Zhang 0001, Peter Hase, Mohit Bansal |
ICLR | 2 |
| 2022 | How can NLP Help Revitalize Endangered Languages? A Case Study and Roadmap for the Cherokee LanguageabstractMore than 43% of the languages spoken in the world are endangered, and language loss currently occurs at an accelerated rate because of globalization and neocolonialism.Saving and revitalizing endangered languages has become very important for maintaining the cultural diversity on our planet.In this work, we focus on discussing how NLP can help revitalize endangered languages.We first suggest three principles that may help NLP practitioners to foster mutual understanding and collaboration with language communities, and we discuss three ways in which NLP can potentially assist in language education.We then take Cherokee, a severely-endangered Native American language, as a case study.After reviewing the language's history, linguistic features, and existing resources, we (in collaboration with Cherokee community members) arrive at a few meaningful ways NLP practitioners can collaborate with community partners.We suggest two approaches to enrich the Cherokee language's resources with machine-in-the-loop processing, and discuss several NLP tools that people from the Cherokee community have shown interest in.We hope that our work serves not only to inform the NLP community about Cherokee, but also to provide inspiration for future work on endangered languages in general. 1 Shiyue Zhang 0001, Benjamin Frey, Mohit Bansal |
ACL (1) | 1 |
| 2022 | Masked Part-Of-Speech Model: Does Modeling Long Context Help Unsupervised POS-tagging?abstractPrevious Part-Of-Speech (POS) induction models usually assume certain independence assumptions (e.g., Markov, unidirectional, local dependency) that do not hold in real languages.For example, the subject-verb agreement can be both long-term and bidirectional.To facilitate flexible dependency modeling, we propose a Masked Part-of-Speech Model (MPoSM), inspired by the recent success of Masked Language Models (MLM).MPoSM can model arbitrary tag dependency and perform POS induction through the objective of masked POS reconstruction.We achieve competitive results on both the English Penn WSJ dataset as well as the universal treebank containing 10 diverse languages.Though modeling the long-term dependency should ideally help this task, our ablation study shows mixed trends in different languages.To better understand this phenomenon, we design a novel synthetic experiment that can specifically diagnose the model's ability to learn tag agreement.Surprisingly, we find that even strong baselines fail to solve this problem consistently in a very simplified setting: the agreement between adjacent words.Nonetheless, MPoSM achieves overall better performance.Lastly, we conduct a detailed error analysis to shed light on other remaining challenges. 1 Shiyue Zhang 0001, Mohit Bansal |
NAACL-HLT | 2 |
| 2021 | Continuous Language Generative FlowabstractZineng Tang, Shiyue Zhang, Hyounghun Kim, Mohit Bansal. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zineng Tang, Shiyue Zhang 0001, Hyounghun Kim, Mohit Bansal |
ACL/IJCNLP (1) | 2 |
| 2021 | EmailSum: Abstractive Email Thread SummarizationabstractShiyue Zhang, Asli Celikyilmaz, Jianfeng Gao, Mohit Bansal. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Shiyue Zhang 0001, Asli Celikyilmaz, Jianfeng Gao 0001, Mohit Bansal |
ACL/IJCNLP (1) | 1 |
| 2021 | Finding a Balanced Degree of Automation for Summary EvaluationabstractHuman evaluation for summarization tasks is reliable but brings in issues of reproducibility and high costs.Automatic metrics are cheap and reproducible but sometimes poorly correlated with human judgment.In this work, we propose flexible semiautomatic to automatic summary evaluation metrics, following the Pyramid human evaluation method.Semi-automatic Lite 2 Pyramid retains the reusable human-labeled Summary Content Units (SCUs) for reference(s) but replaces the manual work of judging SCUs' presence in system summaries with a natural language inference (NLI) model.Fully automatic Lite 3 Pyramid further substitutes SCUs with automatically extracted Semantic Triplet Units (STUs) via a semantic role labeling (SRL) model.Finally, we propose in-between metrics, Lite 2.x Pyramid, where we use a simple regressor to predict how well the STUs can simulate SCUs and retain SCUs that are more difficult to simulate, which provides a smooth transition and balance between automation and manual evaluation.Comparing to 15 existing metrics, we evaluate human-metric correlations on 3 existing meta-evaluation datasets and our newlycollected PyrXSum (with 100/10 XSum examples/systems).It shows that Lite 2 Pyramid consistently has the best summary-level correlations; Lite 3 Pyramid works better than or comparable to other automatic metrics; Lite 2.x Pyramid trades off small correlation drops for larger manual effort reduction, which can reduce costs for future data collection. 1 Shiyue Zhang 0001, Mohit Bansal |
EMNLP (1) | 1 |
| 2020 | ChrEn: Cherokee-English Machine Translation for Endangered Language RevitalizationabstractCherokee is a highly endangered Native American language spoken by the Cherokee people.The Cherokee culture is deeply embedded in its language.However, there are approximately only 2,000 fluent first language Cherokee speakers remaining in the world, and the number is declining every year.To help save this endangered language, we introduce ChrEn, a Cherokee-English parallel dataset, to facilitate machine translation research between Cherokee and English.Compared to some popular machine translation language pairs, ChrEn is extremely low-resource, only containing 14k sentence pairs in total.We split our parallel data in ways that facilitate both in-domain and out-of-domain evaluation.We also collect 5k Cherokee monolingual data to enable semi-supervised learning.Besides these datasets, we propose several Cherokee-English and English-Cherokee machine translation systems.We compare SMT (phrase-based) versus NMT (RNN-based and Transformer-based) systems; supervised versus semi-supervised (via language model, back-translation, and BERT/Multilingual-BERT) methods; as well as transfer learning versus multilingual joint training with 4 other languages.Our best results are 15.8/12.7 BLEU for in-domain and 6.5/5.0BLEU for out-of-domain Chr-En/En-Chr translations, respectively, and we hope that our dataset and systems will encourage future work by the community for Cherokee language revitalization. 1 Shiyue Zhang 0001, Benjamin Frey, Mohit Bansal |
EMNLP (1) | 1 |
| 2019 | Addressing Semantic Drift in Question Generation for Semi-Supervised Question AnsweringabstractShiyue Zhang, Mohit Bansal. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shiyue Zhang 0001, Mohit Bansal |
EMNLP/IJCNLP (1) | 1 |
| 2017 | Flexible and Creative Chinese Poetry Generation Using Neural MemoryabstractJiyuan Zhang, Yang Feng, Dong Wang, Yang Wang, Andrew Abel, Shiyue Zhang, Andi Zhang. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Jiyuan Zhang 0001, Yang Feng 0004, Dong Wang 0013, Andrew Abel, Shiyue Zhang 0001, Andi Zhang 0002 |
ACL (1) | 6 |
| 2017 | Memory-augmented Neural Machine TranslationabstractNeural machine translation (NMT) has achieved notable success in recent times, however it is also widely recognized that this approach has limitations with handling infrequent words and word pairs.This paper presents a novel memoryaugmented NMT (M-NMT) architecture, which stores knowledge about how words (usually infrequently encountered ones) should be translated in a memory and then utilizes them to assist the neural model.We use this memory mechanism to combine the knowledge learned from a conventional statistical machine translation system and the rules learned by an NMT system, and also propose a solution for out-of-vocabulary (OOV) words based on this framework.Our experiments on two Chinese-English translation tasks demonstrated that the M-NMT architecture outperformed the NMT baseline by 9.0 and 2.7 BLEU points on the two tasks, respectively.Additionally, we found this architecture resulted in a much more effective OOV treatment compared to competitive methods. Yang Feng 0004, Shiyue Zhang 0001, Andi Zhang 0002, Dong Wang 0013, Andrew Abel |
EMNLP | 2 |
| 2017 | Memory visualization for gated recurrent neural networks in speech recognitionabstractRecurrent neural networks (RNNs) have shown clear superiority in sequence modeling, particularly the ones with gated units, such as long short-term memory (LSTM) and gated recurrent unit (GRU). However, the dynamic properties behind the remarkable performance remain unclear in many applications, e.g., automatic speech recognition (ASR). This paper employs visualization techniques to study the behavior of LSTM and GRU when performing speech recognition tasks. Our experiments show some interesting patterns in the gated memory, and some of them have inspired simple yet effective modifications on the network structure. We report two of such modifications: (1) lazy cell update in LSTM, and (2) shortcut connections for residual learning. Both modifications lead to more comprehensible and powerful networks. Zhiyuan Tang, Ying Shi 0001, Dong Wang 0013, Yang Feng 0004, Shiyue Zhang 0001 |
ICASSP | 5 |