VLDB 2026 Research / reviewers in the wild / expert
Peng Shi 0010
dblp:172/7638-10
· DBLP profile ↗
15ranked-venue papers
3as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HRScene: How Far are VLMs from Effective High-Resolution Image Understanding?abstractHigh-resolution image (HRI) understanding aims to process images with a large number of pixels, such as pathological images and agricultural aerial images, both of which can exceed 1 million pixels. Vision Large Language Models (VLMs) can allegedly handle HRIs, however, there is a lack of a comprehensive benchmark for VLMs to evaluate HRI understanding. To address this gap, we introduce HRScene, a novel unified benchmark for HRI understanding with rich scenes. HRScene incorporates 25 real-world datasets and 2 synthetic diagnostic datasets with resolutions ranging from 1,024 $\times$ 1,024 to 35,503 $\times$ 26,627. HRScene is collected and re-annotated by 10 graduate-level annotators, covering 25 scenarios, ranging from microscopic to radiology images, street views, long-range pictures, and telescope images. It includes HRIs of real-world objects, scanned documents, and composite multi-image. The two diagnostic evaluation datasets are synthesized by combining the target image with the gold answer and distracting images in different orders, assessing how well models utilize regions in HRI. We conduct extensive experiments involving 28 VLMs, including Gemini 2.0 Flash and GPT-4o. Experiments on HRScene show that current VLMs achieve an average accuracy of around 50% on real-world tasks, revealing significant gaps in HRI understanding. Results on synthetic datasets reveal that VLMs struggle to effectively utilize HRI regions, showing significant Regional Divergence and lost-in-middle, shedding light on future research. Yusen Zhang 0001, Wenliang Zheng, Aashrith Madasu, Peng Shi 0010, Ryo Kamoi, Zhuoyang Zou, Shu Zhao 0006, Sarkar Snigdha Sarathi Das, Xiaoxin Lu, Ranran Haoran Zhang, Avitej Iyer, Renze Lou, Wenpeng Yin 0001, Rui Zhang 0037 |
ICCV | 4 |
| 2025 | You Only Read Once (YORO): Learning to Internalize Database Knowledge for Text-to-SQLabstractHideo Kobayashi, Wuwei Lan, Peng Shi, Shuaichen Chang, Jiang Guo, Henghui Zhu, Zhiguo Wang, Patrick Ng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hideo Kobayashi, Wuwei Lan, Peng Shi 0010, Shuaichen Chang, Henghui Zhu, Zhiguo Wang 0006, Patrick Ng |
NAACL (Long Papers) | 3 |
| 2024 | Construction of Paired Knowledge Graph - Text Datasets Informed by Cyclic EvaluationabstractDatasets that pair Knowledge Graphs (KG) and text together (KG-T) can be used to train forward and reverse neural models that generate text from KG and vice versa. However models trained on datasets where KG and text pairs are not equivalent can suffer from more hallucination and poorer recall. In this paper, we verify this empirically by generating datasets with different levels of noise and find that noisier datasets do indeed lead to more hallucination. We argue that the ability of forward and reverse models trained on a dataset to cyclically regenerate source KG or text is a proxy for the equivalence between the KG and the text in the dataset. Using cyclic evaluation we find that manually created WebNLG is much better than automatically created TeKGen and T-REx. Informed by these observations, we construct a new, improved dataset called LAGRANGE using heuristics meant to improve equivalence between KG and text and show the impact of each of the heuristics on cyclic evaluation. We also construct two synthetic datasets using large language models (LLMs), and observe that these are conducive to models that perform significantly well on cyclic generation of text, but less so on cyclic generation of KGs, probably because of a lack of a consistent underlying ontology. Ali Mousavi 0003, Xin Zhan, He Bai 0002, Peng Shi 0010, Theodoros Rekatsinas, Benjamin Han, Yunyao Li 0001, Jeffrey Pound, Joshua M. Susskind, Natalie Schluter, Ihab F. Ilyas, Navdeep Jaitly |
LREC/COLING | 4 |
| 2024 | RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented GenerationabstractDespite Retrieval-Augmented Generation (RAG) has shown promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to the modular nature of RAG, evaluation of long-form responses and reliability of measurements. In this paper, we propose a fine-grained evaluation framework, RAGChecker, that incorporates a suite of diagnostic metrics for both the retrieval and generation modules. Meta evaluation verifies that RAGChecker has significantly better correlations with human judgments than other evaluation metrics. Using RAGChecker, we evaluate 8 RAG systems and conduct an in-depth analysis of their performance, revealing insightful patterns and trade-offs in the design choices of RAG architectures. The metrics of RAGChecker can guide researchers and practitioners in developing more effective RAG systems. Dongyu Ru, Xiangkun Hu, Tianhang Zhang, Peng Shi 0010, Shuaichen Chang, Cheng Jiayang, Cunxiang Wang, Shichao Sun, Huanyu Li 0010, Binjie Wang, Jiarong Jiang, Tong He 0002, Zhiguo Wang 0006, Pengfei Liu 0003, Yue Zhang 0004, Zheng Zhang 0001 |
NeurIPS | 5 |
| 2023 | Unified Low-Resource Sequence Labeling by Sample-Aware Dynamic Sparse FinetuningabstractUnified Sequence Labeling articulates different sequence labeling tasks such as Named Entity Recognition, Relation Extraction, Semantic Role Labeling, etc. in a generalized sequence-to-sequence format.Unfortunately, this requires formatting different tasks into specialized augmented formats which are unfamiliar to the base pretrained language model (PLMs).This necessitates model fine-tuning and significantly bounds its usefulness in datalimited settings where fine-tuning large models cannot properly generalize to the target format.To address this challenge and leverage PLM knowledge effectively, we propose FISH-DIP, a sample-aware dynamic sparse finetuning strategy.It selectively finetunes a fraction of parameters informed by highly regressing examples during the fine-tuning process.By leveraging the dynamism of sparsity, our approach mitigates the impact of well-learned samples and prioritizes underperforming instances for improvement in generalization.Across five tasks of sequence labeling, we demonstrate that FISH-DIP can smoothly optimize the model in low-resource settings, offering up to 40% performance improvements over full fine-tuning depending on target evaluation settings.Also, compared to in-context learning and other parameter-efficient fine-tuning (PEFT) approaches, FISH-DIP performs comparably or better, notably in extreme lowresource settings. Sarkar Snigdha Sarathi Das, Ranran Haoran Zhang, Peng Shi 0010, Wenpeng Yin 0001, Rui Zhang 0037 |
EMNLP | 3 |
| 2023 | Binding Language Models in Symbolic Languages
Zhoujun Cheng, Tianbao Xie, Peng Shi 0010, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir R. Radev, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, Tao Yu 0009 |
ICLR | 3 |
| 2022 | Generation-Focused Table-Based Intermediate Pre-training for Free-Form Question AnsweringabstractQuestion answering over semi-structured tables has attracted significant attention in the NLP community. However, most of the existing work focus on questions that can be answered with short-form answer, i.e. the answer is often a table cell or aggregation of multiple cells. This can mismatch with the intents of users who want to ask more complex questions that require free-form answers such as explanations. To bridge the gap, most recently, pre-trained sequence-to-sequence language models such as T5 are used for generating free-form answers based on the question and table inputs. However, these pre-trained language models have weaker encoding abilities over table cells and schema. To mitigate this issue, in this work, we present an intermediate pre-training framework, Generation-focused Table-based Intermediate Pre-training (GENTAP), that jointly learns representations of natural language questions and tables. GENTAP learns to generate via two training objectives to enhance the question understanding and table representation abilities for complex questions. Based on experimental results, models that leverage GENTAP framework outperform the existing baselines on FETAQA benchmark. The pre-trained models are not only useful for free-form question answering, but also for few-shot data-to-text generation task, thus showing good transfer ability by obtaining new state-of-the-art results. Peng Shi 0010, Patrick Ng, Feng Nan, Henghui Zhu, Jun Wang 0122, Jiarong Jiang, Alexander Hanbo Li, Rishav Chakravarti, Donald Weidner, Bing Xiang, Zhiguo Wang 0006 |
AAAI | 1 |
| 2022 | Better Language Model with Hypernym Class PredictionabstractClass-based language models (LMs) have been long devised to address context sparsity in n-gram LMs.In this study, we revisit this approach in the context of neural LMs.We hypothesize that class-based prediction leads to an implicit context aggregation for similar words and thus can improve generalization for rare words.We map words that have a common WordNet hypernym to the same class and train large neural LMs by gradually annealing from predicting the class to token prediction during training.Empirically, this curriculum learning strategy consistently improves perplexity over various large, highly-performant state-of-the-art Transformer-based models on two datasets, WikiText-103 and ARXIV.Our analysis shows that the performance improvement is achieved without sacrificing performance on rare words.Finally, we document other attempts that failed to yield empirical gains, and discuss future directions for the adoption of class-based LMs on a larger scale. He Bai 0002, Tong Wang 0012, Alessandro Sordoni, Peng Shi 0010 |
ACL (1) | 4 |
| 2022 | UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsabstractTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, Tao Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Tianbao Xie, Chen Henry Wu, Peng Shi 0010, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong 0005, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao 0002, Dragomir R. Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang 0037, Noah A. Smith, Luke Zettlemoyer, Tao Yu 0009 |
EMNLP | 3 |
| 2021 | Segatron: Segment-Aware Transformer for Language Modeling and UnderstandingabstractTransformers are powerful for sequence modeling. Nearly all state-of-the-art language models and pre-trained language models are based on the Transformer architecture. However, it distinguishes sequential tokens only with the token position index. We hypothesize that better contextual representations can be generated from the Transformer with richer positional information. To verify this, we propose a segment-aware Transformer (Segatron), by replacing the original token position encoding with a combined position encoding of paragraph, sentence, and token. We first introduce the segment-aware mechanism to Transformer-XL, which is a popular Transformer-based language model with memory extension and relative position encoding. We find that our method can further improve the Transformer-XL base model and large model, achieving 17.1 perplexity on the WikiText-103 dataset. We further investigate the pre-training masked language modeling task with Segatron. Experimental results show that BERT pre-trained with Segatron (SegaBERT) can outperform BERT with vanilla Transformer on various NLP tasks, and outperforms RoBERTa on zero-shot sentence representation learning. Our code is available on GitHub. He Bai 0002, Peng Shi 0010, Jimmy Lin, Yuqing Xie 0001, Luchen Tan, Kun Xiong, Wen Gao 0001, Ming Li 0001 |
AAAI | 2 |
| 2021 | Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-TrainingabstractMost recently, there has been significant interest in learning contextual representations for various NLP tasks, by leveraging large scale text corpora to train powerful language models with self-supervised learning objectives, such as Masked Language Model (MLM). Based on a pilot study, we observe three issues of existing general-purpose language models when they are applied in the text-to-SQL semantic parsers: fail to detect the column mentions in the utterances, to infer the column mentions from the cell values, and to compose target SQL queries when they are complex. To mitigate these issues, we present a model pretraining framework, Generation-Augmented Pre-training (GAP), that jointly learns representations of natural language utterance and table schemas, by leveraging generation models to generate high-quality pre-train data. GAP Model is trained on 2 million utterance-schema pairs and 30K utterance-schema-SQL triples, whose utterances are generated by generation models. Based on experimental results, neural semantic parsers that leverage GAP Model as a representation encoder obtain new state-of-the-art results on both Spider and Criteria-to-SQL benchmarks. Peng Shi 0010, Patrick Ng, Zhiguo Wang 0006, Henghui Zhu, Alexander Hanbo Li, Jun Wang 0122, Cícero Nogueira dos Santos, Bing Xiang |
AAAI | 1 |
| 2019 | Bridging the Gap between Relevance Matching and Semantic Matching for Short Text Similarity ModelingabstractJinfeng Rao, Linqing Liu, Yi Tay, Wei Yang, Peng Shi, Jimmy Lin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jinfeng Rao, Linqing Liu, Yi Tay, Hsiu-Wei Yang, Peng Shi 0010, Jimmy Lin |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Aligning Cross-Lingual Entities with Multi-Aspect InformationabstractHsiu-Wei Yang, Yanyan Zou, Peng Shi, Wei Lu, Jimmy Lin, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Hsiu-Wei Yang, Peng Shi 0010, Wei Lu 0011, Jimmy Lin, Xu Sun 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2018 | Farewell Freebase: Migrating the SimpleQuestions Dataset to DBpediaabstractQuestion answering over knowledge graphs is an important problem of interest both commercially and academically. There is substantial interest in the class of natural language questions that can be answered via the lookup of a single fact, driven by the availability of the popular SimpleQuestions dataset. The problem with this dataset, however, is that answer triples are provided from Freebase, which has been defunct for several years. As a result, it is difficult to build “real-world” question answering systems that are operationally deployable. Furthermore, a defunct knowledge graph means that much of the infrastructure for querying, browsing, and manipulating triples no longer exists. To address this problem, we present SimpleDBpediaQA, a new benchmark dataset for simple question answering over knowledge graphs that was created by mapping SimpleQuestions entities and predicates from Freebase to DBpedia. Although this mapping is conceptually straightforward, there are a number of nuances that make the task non-trivial, owing to the different conceptual organizations of the two knowledge graphs. To lay the foundation for future research using this dataset, we leverage recent work to provide simple yet strong baselines with and without neural networks. Michael Azmy, Peng Shi 0010, Jimmy Lin, Ihab F. Ilyas |
COLING | 2 |
| 2016 | Exploiting Mutual Benefits between Syntax and Semantic Roles using Neural NetworkabstractWe investigate mutual benefits between syntax and semantic roles using neural network models, by studying a parsing→SRL pipeline, a SRL→parsing pipeline, and a simple joint model by embedding sharing.The integration of syntactic and semantic features gives promising results in a Chinese Semantic Treebank, demonstrating large potentials of neural models for joint parsing and semantic role labeling. Peng Shi 0010, Zhiyang Teng, Yue Zhang 0004 |
EMNLP | 1 |