VLDB 2026 Research / reviewers in the wild / expert
Shengwei Gu
dblp:176/3625
· DBLP profile ↗
8ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0003-1003-0185ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Enhancing Answer Selection via Ad-Hoc Knowledge Extraction from Unstructured Web TextsabstractAnswer selection aims to identify the most relevant answers to a given question from a set of candidates. It is the fundamental component of intelligent question answering system. To improve performance, it gradually becomes an effective strategy to integrate external structured knowledge bases (KBs) into the answer selection model. Due to expensive cost of construction and maintenance of such KBs, these models are suffering from domain barriers and information incompleteness. In this paper, we propose a two-stage extraction–comprehension answer selection model, which can extract ad-hoc knowledge from unstructured web texts to enhance the performance of answer selection. For the extraction, two types of snippets are extracted from unstructured web pages and utilized as the source of ad-hoc knowledge. For the comprehension, a selective attention mechanism is employed to extract and integrate ad-hoc knowledge from multiple text snippets obtained in the first stage, which can enrich the representation of question–answer pairs and more accurately identify the correct answers. By incorporating ad-hoc knowledge extracted from both types of snippets, the proposed model achieves state-of-the-art results on two public available benchmark datasets. In particular, on WikiQA, in terms of the two evaluation metrics (mean average precision and mean reciprocal rank), it achieves 9.9[Formula: see text] and 8.4[Formula: see text] higher than the previous non-pretraining-based models, and 3.4[Formula: see text] and 3.2[Formula: see text] higher than the pretraining-based models. Shengwei Gu, Xiangfeng Luo, Hao Wang 0097 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2023 | Accelerating Multi-Exit BERT Inference via Curriculum Learning and Knowledge DistillationabstractThe real-time deployment of bidirectional encoder representations from transformers (BERT) is limited by its slow inference caused by its large number of parameters. Recently, multi-exit architecture has garnered scholarly attention for its ability to achieve a trade-off between performance and efficiency. However, its early exits suffer from a considerable performance reduction compared to the final classifier. To accelerate inference with minimal compensation of performance, we propose a novel training paradigm for multi-exit BERT performing at two levels: training samples and intermediate features. Specifically, for the training samples level, we leverage curriculum learning to guide the training process and improve the generalization capacity of the model. For the intermediate features level, we employ layer-wise distillation learning from shallow to deep layers to resolve the performance deterioration of early exits. The experimental results obtained on the benchmark datasets of textual entailment and answer selection demonstrate that the proposed training paradigm is effective and achieves state-of-the-art results. Furthermore, the layer-wise distillation can completely replace vanilla distillation and deliver superior performance on text entailment datasets. Shengwei Gu, Xiangfeng Luo, Xinzhi Wang 0001, Yike Guo |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2021 | Improving answer selection with global featuresabstractAbstract Given a question and its answer candidates (named QA corpus), answer selection is the task of identifying the most relevant answers to the question. Answer selection is widely used in question answering, web search, and so on. Current deep neural network models primarily utilize local features extracted from input question‐answer pairs (QA pairs). However, the global features contained in QA corpora are under‐utilized, and we argue that these global features substantially contribute to the answer selection task. To verify this point of view, we propose a novel model that combines local and global features for answer selection. In our model, two different global feature extractors are employed to extract statistical global features and deep global features from a QA corpus, respectively. Furthermore, we investigate the integration of these global features with local features in various experimental settings: statistical global features, deep global features, and a combination of statistical and deep global features. Our experimental results show that the global features are effective for answer selection. Our model obtains new state‐of‐the‐art results on two public answer selection datasets and performs especially well on YahooCQA, where it achieves 9.2 and 6% higher precision@1 (P@1) and mean reciprocal rank (MRR) scores than previously published models. Shengwei Gu, Xiangfeng Luo, Hao Wang 0097, Subin Huang |
Expert Syst. J. Knowl. Eng. | 1 |
| 2020 | Inter-sentence and Implicit Causality Extraction from Chinese Corpus
Xianxian Jin, Xinzhi Wang 0001, Xiangfeng Luo, Subin Huang, Shengwei Gu |
PAKDD (1) | 5 |
| 2020 | Improving taxonomic relation learning via incorporating relation descriptions into word embeddingsabstractSummary Taxonomic relations play an important role in various Natural Language Processing (NLP) tasks (eg, information extraction, question answering and knowledge inference). Existing approaches on embedding‐based taxonomic relation learning mainly rely on the word embeddings trained using co‐occurrence‐based similarity learning. However, the performance of these approaches is not quite satisfactory due to the lack of sufficient taxonomic semantic knowledge within word embeddings. To solve this problem, we propose an improved embedding‐based approach to learn taxonomic relations via incorporating relation descriptions into word embeddings. First, to capture additional taxonomic semantic knowledge, we train special word embeddings using not only co‐occurrence information of words but also relation descriptions (eg, taxonomic seed relations and their contextual triples). Then, using the trained word embeddings as features, we employ two learning models to identify and predict taxonomic relations, namely, offset‐based classification model and offset‐based similarity model. Experimental results on four real‐world domain datasets demonstrate that our proposed approach can capture additional taxonomic semantic knowledge and reduce dependence on the training dataset, outperforming the state‐of‐the‐art compared approaches on the taxonomic relation learning task. Subin Huang, Xiangfeng Luo, Hao Wang 0097, Shengwei Gu, Yike Guo |
Concurr. Comput. Pract. Exp. | 5 |
| 2020 | Abstract Concept Instantiation with Context Relevance MeasurementabstractIn different contexts, one abstract concept (e.g., fruit) may be mapped into different concrete instance sets, which is called abstract concept instantiation. It has been widely applied in many applications, such as web search, intelligent recommendation, etc. However, in most abstract concept instantiation models have the following problems: (1) the neglect of incorrect label and label incompleteness in the category structure on which instance selection relies; (2) the subjective design of instance profile for calculating the relevance between instance and contextual constraint. The above problems lead to false prediction in terms of abstract concept instantiation. To tackle these problems, we proposed a novel model to instantiate the abstract concept. Firstly, to alleviate the incorrect label and remedy label incompleteness in the category structure, an improved random-walk algorithm is proposed, called InstanceRank, which not only utilize the category information, but it also exploits the association information to infer the right instances of an abstract concept. Secondly, for better measuring the relevance between instances and contextual constraint, we learn the proper instance profile from different granularity ones. They are designed based on the surrounding text of the instance. Finally, noise reduction and instance filtering are introduced to further enhance the model performance. Experiments on Chinese food abstract concept set show that the proposed model can effectively reduce false positive and false negative of instantiation results. Shengwei Gu, Xiangfeng Luo, Hao Wang 0097, Subin Huang |
J. Web Eng. | 1 |
| 2019 | An unsupervised approach for learning a Chinese IS-A taxonomy from an unstructured corpus
Subin Huang, Xiangfeng Luo, Yike Guo, Shengwei Gu |
Knowl. Based Syst. | 5 |
| 2018 | Semantic Emotion-Topic Model Based Social Emotion Mining
Ruirong Xue, Xiangfeng Luo, Qichen Ma, Shengwei Gu |
J. Web Eng. | 4 |