EDBT 2026 Demo / reviewers in the wild / expert
Shangsi Chen
dblp:334/0442
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Question answering and dialogue systems · 67% Deep learning architectures and training · 33% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
1.5 | 2 | 2025 | Self Data Augmentation for Open Domain Question Answering · ACM Trans. Inf. Syst. 2025 A Survey for Efficient Open Domain Question Answering · ACL (1) 2023 |
Machine learning › Deep learning architectures and training
data augmentation |
0.9 | 1 | 2025 | Self Data Augmentation for Open Domain Question Answering · ACM Trans. Inf. Syst. 2025 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
0.9 | 1 | 2025 | Self Data Augmentation for Open Domain Question Answering · ACM Trans. Inf. Syst. 2025 |
Information retrieval › document retrieval
passage retrieval |
0.9 | 1 | 2025 | Self Data Augmentation for Open Domain Question Answering · ACM Trans. Inf. Syst. 2025 |
Methods — techniques the papers use, named apart from their topics
pseudo-labeling · 1.7dense passage retriever · 1.7survey · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Self Data Augmentation for Open Domain Question AnsweringabstractInformation Retrieval (IR) constitutes a vital facet of Open Domain Question Answering (ODQA) systems, focusing on the exploration of pertinent information within extensive collections of passages, such as Wikipedia, to facilitate subsequent reader processing. Historically, IR relied on textual overlaps for relevant context retrieval, employing methods like BM25 and TF-IDF, which, however, lacked natural language understanding. The advent of deep learning ushered in a new era, leading to the introduction of Dense Passage Retrievers (DPR), shows superiority over traditional sparse retrievers. These dense retrievers leverage Pre-Trained Language Models (PLMs) to initialize context encoders, enabling the extraction of natural language representations. They utilize the distance between latent vectors of contexts as a metric for assessing similarity. However, DPR methods are heavily reliant on large volumes of meticulously labeled data, such as Natural Questions. The process of data labeling is both costly and time-intensive. In this article, we propose a novel data augmentation methodology Self Data Augmentation (SDA) that employs DPR models to automatically annotate unanswered questions. Specifically, we initiate the process by retrieving relevant pseudo passages for these unlabeled questions. We subsequently introduce three distinct passage selection methods to annotate these pseudo passages. Ultimately, we amalgamate the pseudo-labeled passages with the unanswered questions to create augmented data. Our experimental evaluations conducted on two extensive datasets (Natural Questions and TriviaQA), alongside a relatively small dataset (WebQuestions), utilizing three diverse base models, illustrate the significant enhancement achieved through the incorporation of freshly augmented data. Moreover, our proposed data augmentation method exhibits remarkable flexibility, which is readily adaptable to various dense retrievers. Additionally, we have conducted a comprehensive human study on the augmented data, which further supports our conclusions. Qin Zhang 0011, Mengqi Zheng, Shangsi Chen, Han Liu 0002 |
ACM Trans. Inf. Syst. | 3 |
| 2023 | A Survey for Efficient Open Domain Question AnsweringabstractQin Zhang, Shangsi Chen, Dongkuan Xu, Qingqing Cao, Xiaojun Chen, Trevor Cohn, Meng Fang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Qin Zhang 0011, Shangsi Chen, Dongkuan Xu, Xiaojun Chen 0006, Trevor Cohn |
ACL (1) | 2 |
| 2023 | Joint reasoning with knowledge subgraphs for Multiple Choice Question Answering
Qin Zhang 0011, Shangsi Chen, Xiaojun Chen 0006 |
Inf. Process. Manag. | 2 |