EDBT 2026 Demo / reviewers in the wild / expert
Zhengbao Jiang
dblp:174/2536
· DBLP profile ↗
24ranked-venue papers
14as first author
13since 2021 · last 2024
0000-0002-0315-6727ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 11 first-author · 13 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Instruction-tuned Language Models are Better Knowledge LearnersabstractZhengbao Jiang, Zhiqing Sun, Weijia Shi, Pedro Rodriguez, Chunting Zhou, Graham Neubig, Xi Lin, Wen-tau Yih, Srini Iyer. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhengbao Jiang, Zhiqing Sun, Pedro Rodríguez 0001, Chunting Zhou, Graham Neubig, Xi Victoria Lin, Scott Yih, Srinivasan Iyer 0001 |
ACL (1) | 1 |
| 2024 | Beyond Memorization: The Challenge of Random Memory Access in Language ModelsabstractRecent developments in Language Models (LMs) have shown their effectiveness in NLP tasks, particularly in knowledge-intensive tasks.However, the mechanisms underlying knowledge storage and memory access within their parameters remain elusive.In this paper, we investigate whether a generative LM (e.g., GPT-2) is able to access its memory sequentially or randomly.Through carefully-designed synthetic tasks, covering the scenarios of full recitation, selective recitation and grounded question answering, we reveal that LMs manage to sequentially access their memory while encountering challenges in randomly accessing memorized content.We find that techniques including recitation and permutation improve the random memory access capability of LMs.Furthermore, by applying this intervention to realistic scenarios of open-domain question answering, we validate that enhancing random access by recitation leads to notable improvements in question answering.The code to reproduce our experiments can be found at https://github. com/sail-sg/lm-random-memory-access. Tongyao Zhu, Qian Liu 0033, Liang Pang 0001, Zhengbao Jiang, Min-Yen Kan |
ACL (1) | 4 |
| 2024 | GPTScore: Evaluate as You DesireabstractJinlan Fu, See-Kiong Ng, Zhengbao Jiang, Pengfei Liu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, Pengfei Liu 0003 |
NAACL-HLT | 3 |
| 2023 | Active Retrieval Augmented GenerationabstractZhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, Graham Neubig. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Zhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun, Qian Liu 0033, Jane Dwivedi-Yu, Yiming Yang 0002, Jamie Callan, Graham Neubig |
EMNLP | 1 |
| 2023 | PEER: A Collaborative Language Model
Timo Schick, Jane Dwivedi-Yu, Zhengbao Jiang, Fabio Petroni, Patrick S. H. Lewis, Gautier Izacard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, Sebastian Riedel 0001 |
ICLR | 3 |
| 2023 | DocPrompting: Generating Code by Retrieving the Docs
Shuyan Zhou, Uri Alon 0002, Frank F. Xu, Zhengbao Jiang, Graham Neubig |
ICLR | 4 |
| 2022 | Understanding and Improving Zero-shot Multi-hop Reasoning in Generative Question AnsweringabstractGenerative question answering (QA) models generate answers to questions either solely based on the parameters of the model (the closed-book setting) or additionally retrieving relevant evidence (the open-book setting). Generative QA models can answer some relatively complex questions, but the mechanism through which they do so is still poorly understood. We perform several studies aimed at better understanding the multi-hop reasoning capabilities of generative QA models. First, we decompose multi-hop questions into multiple corresponding single-hop questions, and find marked inconsistency in QA models’ answers on these pairs of ostensibly identical question chains. Second, we find that models lack zero-shot multi-hop reasoning ability: when trained only on single-hop questions, models generalize poorly to multi-hop questions. Finally, we demonstrate that it is possible to improve models’ zero-shot multi-hop reasoning capacity through two methods that approximate real multi-hop natural language (NL) questions by training on either concatenation of single-hop questions or logical forms (SPARQL). In sum, these results demonstrate that multi-hop reasoning does not emerge naturally in generative QA models, but can be encouraged by advances in training or modeling techniques. Code is available at https://github.com/jzbjyb/multihop. Zhengbao Jiang, Jun Araki, Haibo Ding, Graham Neubig |
COLING | 1 |
| 2022 | Retrieval as Attention: End-to-end Learning of Retrieval and Reading within a Single TransformerabstractSystems for knowledge-intensive tasks such as open-domain question answering (QA) usually consist of two stages: efficient retrieval of relevant documents from a large corpus and detailed reading of the selected documents to generate answers.Retrievers and readers are usually modeled separately, which necessitates a cumbersome implementation and is hard to train and adapt in an end-to-end fashion.In this paper, we revisit this design and eschew the separate architecture and training in favor of a single Transformer that performs Retrieval as Attention (ReAtt), and end-to-end training solely based on supervision from the end QA task.We demonstrate for the first time that a single model trained end-to-end can achieve both competitive retrieval and QA performance, matching or slightly outperforming state-of-the-art separately trained retrievers and readers.Moreover, end-to-end adaptation significantly boosts its performance on out-of-domain datasets in both supervised and unsupervised settings, making our model a simple and adaptable solution for knowledgeintensive tasks.Code and models are available at https://github.com/jzbjyb/ReAtt. Zhengbao Jiang, Luyu Gao, Zhiruo Wang 0001, Jun Araki, Haibo Ding, Jamie Callan, Graham Neubig |
EMNLP | 1 |
| 2022 | SPE: Symmetrical Prompt Enhancement for Fact ProbingabstractPretrained language models (PLMs) have been shown to accumulate factual knowledge during pretraining (Petroni et al., 2019).Recent works probe PLMs for the extent of this knowledge through prompts either in discrete or continuous forms.However, these methods do not consider symmetry of the task: object prediction and subject prediction.In this work, we propose Symmetrical Prompt Enhancement (SPE), a continuous prompt-based method for factual probing in PLMs that leverages the symmetry of the task by constructing symmetrical prompts for subject and object prediction.Our results on a popular factual probing dataset, LAMA, show significant improvement of SPE over previous probing methods. Yiyuan Li, Tong Che, Yezhen Wang, Zhengbao Jiang, Caiming Xiong, Snigdha Chaturvedi |
EMNLP | 4 |
| 2022 | OmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question AnsweringabstractZhengbao Jiang, Yi Mao, Pengcheng He, Graham Neubig, Weizhu Chen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Zhengbao Jiang, Graham Neubig, Weizhu Chen |
NAACL-HLT | 1 |
| 2021 | CoRI: Collective Relation Integration with Data Augmentation for Open Information ExtractionabstractZhengbao Jiang, Jialong Han, Bunyamin Sisman, Xin Luna Dong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zhengbao Jiang, Jialong Han, Bunyamin Sisman, Xin Dong 0001 |
ACL/IJCNLP (1) | 1 |
| 2021 | GSum: A General Framework for Guided Neural Abstractive SummarizationabstractZi-Yi Dou, Pengfei Liu, Hiroaki Hayashi, Zhengbao Jiang, Graham Neubig. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zi-Yi Dou, Pengfei Liu 0003, Hiroaki Hayashi, Zhengbao Jiang, Graham Neubig |
NAACL-HLT | 4 |
| 2021 | How Can We Know When Language Models Know? On the Calibration of Language Models for Question AnsweringabstractAbstract Recent works have shown that language models (LM) capture different types of knowledge regarding facts or common sense. However, because no model is perfect, they still fail to provide appropriate answers in many cases. In this paper, we ask the question, “How can we know when language models know, with confidence, the answer to a particular query?” We examine this question from the point of view of calibration, the property of a probabilistic model’s predicted probabilities actually being well correlated with the probabilities of correctness. We examine three strong generative models—T5, BART, and GPT-2—and study whether their probabilities on QA tasks are well calibrated, finding the answer is a relatively emphatic no. We then examine methods to calibrate such models to make their confidence scores correlate better with the likelihood of correctness through fine-tuning, post-hoc probability modification, or adjustment of the predicted outputs or inputs. Experiments on a diverse range of datasets demonstrate the effectiveness of our methods. We also perform analysis to study the strengths and limitations of these methods, shedding light on further improvements that may be made in methods for calibrating LMs. We have released the code at https://github.com/jzbjyb/lm-calibration. Zhengbao Jiang, Jun Araki, Haibo Ding, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Generalizing Natural Language Analysis through Span-relation RepresentationsabstractNatural language processing covers a wide variety of tasks predicting syntax, semantics, and information content, and usually each type of output is generated with specially designed architectures.In this paper, we provide the simple insight that a great variety of tasks can be represented in a single unified format consisting of labeling spans and relations between spans, thus a single task-independent model can be used across different tasks.We perform extensive experiments to test this insight on 10 disparate tasks spanning dependency parsing (syntax), semantic role labeling (semantics), relation extraction (information content), aspect based sentiment analysis (sentiment), and many others, achieving performance comparable to state-of-the-art specialized models.We further demonstrate benefits of multi-task learning, and also show that the proposed method makes it easy to analyze differences and similarities in how the model handles different tasks.Finally, we convert these datasets into a unified format to build a benchmark, which provides a holistic testbed for evaluating future models for generalized natural language analysis. Zhengbao Jiang, Wei Xu 0004, Jun Araki, Graham Neubig |
ACL | 1 |
| 2020 | Incorporating External Knowledge through Pre-training for Natural Language to Code GenerationabstractOpen-domain code generation aims to generate code in a general-purpose programming language (such as Python) from natural language (NL) intents.Motivated by the intuition that developers usually retrieve resources on the web when writing code, we explore the effectiveness of incorporating two varieties of external knowledge into NL-to-code generation: automatically mined NL-code pairs from the online programming QA forum StackOverflow and programming language API documentation.Our evaluations show that combining the two sources with data augmentation and retrieval-based data re-sampling improves the current state-of-the-art by up to 2.2% absolute BLEU score on the code generation testbed CoNaLa.The code and resources are available at https://github.com/ neulab/external-knowledge-codegen. Frank F. Xu, Zhengbao Jiang, Bogdan Vasilescu, Graham Neubig |
ACL | 2 |
| 2020 | X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language ModelsabstractLanguage models (LMs) have proven surprisingly successful at capturing factual knowledge by completing cloze-style fill-in-theblank questions such as "Punta Cana is located in _."However, while knowledge is both written and queried in many languages, studies on LMs' factual representation ability have almost invariably been performed on English.To assess factual knowledge retrieval in LMs in different languages, we create a multilingual benchmark of cloze-style probes for 23 typologically diverse languages.To properly handle language variations, we expand probing methods from single-to multi-word entities, and develop several decoding algorithms to generate multi-token predictions.Extensive experimental results provide insights about how well (or poorly) current state-of-theart LMs perform at this task in languages with more or fewer available resources.We further propose a code-switching-based method to improve the ability of multilingual LMs to access knowledge, and verify its effectiveness on several benchmark languages.Benchmark data and code have be released at https: //x-factr.github.io. Zhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding, Graham Neubig |
EMNLP (1) | 1 |
| 2020 | Graph-Revised Convolutional Network
Donghan Yu, Ruohong Zhang, Zhengbao Jiang, Yuexin Wu, Yiming Yang 0002 |
ECML/PKDD (3) | 3 |
| 2020 | How Can We Know What Language Models KnowabstractRecent work has presented intriguing results examining the knowledge contained in language models (LMs) by having the LM fill in the blanks of prompts such as “ Obama is a __ by profession”. These prompts are usually manually created, and quite possibly sub-optimal; another prompt such as “ Obama worked as a __ ” may result in more accurately predicting the correct profession. Because of this, given an inappropriate prompt, we might fail to retrieve facts that the LM does know, and thus any given prompt only provides a lower bound estimate of the knowledge contained in an LM. In this paper, we attempt to more accurately estimate the knowledge contained in LMs by automatically discovering better prompts to use in this querying process. Specifically, we propose mining-based and paraphrasing-based methods to automatically generate high-quality and diverse prompts, as well as ensemble methods to combine answers from different prompts. Extensive experiments on the LAMA benchmark for extracting relational knowledge from LMs demonstrate that our methods can improve accuracy from 31.1% to 39.6%, providing a tighter lower bound on what LMs know. We have released the code and the resulting LM Prompt And Query Archive (LPAQA) at https://github.com/jzbjyb/LPAQA . Zhengbao Jiang, Frank F. Xu, Jun Araki, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 1 |
| 2019 | Improving Open Information Extraction via Iterative Rank-Aware LearningabstractOpen information extraction (IE) is the task of extracting open-domain assertions from natural language sentences.A key step in open IE is confidence modeling, ranking the extractions based on their estimated quality to adjust precision and recall of extracted assertions.We found that the extraction likelihood, a confidence measure used by current supervised open IE systems, is not well calibrated when comparing the quality of assertions extracted from different sentences.We propose an additional binary classification loss to calibrate the likelihood to make it more globally comparable, and an iterative learning process, where extractions generated by the open IE model are incrementally included as training samples to help the model learn from trial and error.Experiments on OIE2016 demonstrate the effectiveness of our method.1 Zhengbao Jiang, Graham Neubig |
ACL (1) | 1 |
| 2018 | Personalizing Search Results Using Hierarchical RNN with Query-aware AttentionabstractSearch results personalization has become an effective way to improve the quality of search engines. Previous studies extracted information such as past clicks, user topical interests, query click entropy and so on to tailor the original ranking. However, few studies have taken into account the sequential information underlying previous queries and sessions. Intuitively, the order of issued queries is important in inferring the real user interests. And more recent sessions should provide more reliable personal signals than older sessions. In addition, the previous search history and user behaviors should influence the personalization of the current query depending on their relatedness. To implement these intuitions, in this paper we employ a hierarchical recurrent neural network to exploit such sequential information and automatically generate user profile from historical data. We propose a query-aware attention model to generate a dynamic user profile based on the input query. Significant improvement is observed in the experiment with data from a commercial search engine when compared with several traditional personalization models. Our analysis reveals that the attention model is able to attribute higher weights to more related past sessions after fine training. Songwei Ge, Zhicheng Dou, Zhengbao Jiang, Jian-Yun Nie, Ji-Rong Wen |
CIKM | 3 |
| 2018 | Supervised Search Result Diversification via Subtopic AttentionabstractSearch result diversification aims to retrieve diverse results to satisfy as many different information needs as possible. Supervised methods have been proposed recently to learn ranking functions and they have been shown to produce superior results to unsupervised methods. However, these methods use implicit approaches based on the principle of Maximal Marginal Relevance (MMR). In this paper, we propose a learning framework for explicit result diversification where subtopics are explicitly modeled. Based on the information contained in the sequence of selected documents, we use the attention mechanism to capture the subtopics to be focused on while selecting the next document, which naturally fits our task of document selection for diversification. As a preliminary attempt, we employ recurrent neural networks and max pooling to instantiate the framework. We use both distributed representations and traditional relevance features to model documents in the implementation. The framework is flexible to model query intent in either a flat list or a hierarchy. Experimental results show that the proposed method significantly outperforms all the existing search result diversification approaches. Zhengbao Jiang, Zhicheng Dou, Wayne Xin Zhao, Jian-Yun Nie, Ji-Rong Wen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2017 | Learning to Diversify Search Results via Subtopic AttentionabstractSearch result diversification aims to retrieve diverse results to satisfy as many different information needs as possible. Supervised methods have been proposed recently to learn ranking functions and they have been shown to produce superior results to unsupervised methods. However, these methods use implicit approaches based on the principle of Maximal Marginal Relevance (MMR). In this paper, we propose a learning framework for explicit result diversification where subtopics are explicitly modeled. Based on the information contained in the sequence of selected documents, we use attention mechanism to capture the subtopics to be focused on while selecting the next document, which naturally fits our task of document selection for diversification. The framework is implemented using recurrent neural networks and max-pooling which combine distributed representations and traditional relevance features. Our experiments show that the proposed method significantly outperforms all the existing methods. Zhengbao Jiang, Ji-Rong Wen, Zhicheng Dou, Wayne Xin Zhao, Jian-Yun Nie |
SIGIR | 1 |
| 2017 | Generating Query Facets Using Knowledge BasesabstractA query facet is a significant list of information nuggets that explains an underlying aspect of a query. Existing algorithms mine facets of a query by extracting frequent lists contained in top search results. The coverage of facets and facet items mined by these kind of methods might be limited, because only a small number of search results are used. In order to solve this problem, we propose mining query facets by using knowledge bases which contain high-quality structured data. Specifically, we first generate facets based on the properties of the entities which are contained in Freebase and correspond to the query. Second, we mine initial query facets from search results, then expanding them by finding similar entities from Freebase. Experimental results show that our proposed method can significantly improve the coverage of facet items over the state-of-the-art algorithms. Zhengbao Jiang, Zhicheng Dou, Ji-Rong Wen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | Automatically Mining Facets for Queries from Their Search ResultsabstractWe address the problem of finding query facets which are multiple groups of words or phrases that explain and summarize the content covered by a query. We assume that the important aspects of a query are usually presented and repeated in the query’s top retrieved documents in the style of lists, and query facets can be mined out by aggregating these significant lists. We propose a systematic solution, which we refer to as QDMiner, to automatically mine query facets by extracting and grouping frequent lists from free text, HTML tags, and repeat regions within top search results. Experimental results show that a large number of lists do exist and useful query facets can be mined by QDMiner. We further analyze the problem of list duplication, and find better query facets can be mined by modeling fine-grained similarities between lists and penalizing the duplicated lists. Zhicheng Dou, Zhengbao Jiang, Sha Hu 0002, Ji-Rong Wen, Ruihua Song |
IEEE Trans. Knowl. Data Eng. | 2 |