Shengli Wu 0001

dblp:w/ShengliWu · DBLP profile ↗
← Back
40ranked-venue papers in the field
19as first author
8since 2021 · last 2026
0000-0003-2008-1736ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 18 (11 first)Database Systems & Data Management · 8 (5 first)Knowledge Engineering, Semantic Web & Information Systems · 8 (2 first)Data Mining & Knowledge Discovery · 4Other / Interdisciplinary · 2 (1 first)
YearPublicationVenuePosition
2026 Egocentric Action Recognition with Retrieval-Augmented Learning
abstract
Egocentric Action Recognition (EAR) aims to identify fine-grained actions and interacted objects from first-person videos, forming a core task in egocentric video understanding. Despite recent progress, EAR remains challenged by limited data scale, annotation quality, and long-tailed class distributions. To address these issues, we propose REAR, a Retrieval-augmented framework for EAR that leverages external third-person (exocentric) videos as auxiliary knowledge—without requiring synchronized ego-exo pairs. REAR adopts a dual-branch architecture: one branch extracts egocentric representations, while the other retrieves semantically relevant exocentric features. These are fused via a cross-view integration module that performs staged refinement and attention-based alignment. To mitigate class imbalance, a class-adaptive selector dynamically adjusts retrieval depth based on class frequency, and independent classifiers are trained with logit-adjusted cross-entropy. Extensive experiments across three benchmarks demonstrate that REAR achieves state-of-the-art performance, with significant gains in object recognition and tail-class accuracy. The source code is publicly available at https://github.com/zou-y23/REAR.
Yishan Zou, Chris D. Nugent, Matthew Burns, Shengli Wu 0001, Meng Liu 0006
ICMR4
2026 Complex query answering on knowledge graphs with Griffin and polarity-weighted message passing
Fanghao Li, Shengli Wu 0001, Hu Lu
Knowl. Inf. Syst.4
2025 Cost-effective data fusion in information retrieval
abstract
Abstract Data fusion has demonstrated its effectiveness in enhancing information retrieval across various studies. However, advanced fusion methods typically require a dataset with extensive relevance judgments to train optimal model weights, necessitating labor-intensive and costly manual efforts. This study explores efficient methods for generating training data to facilitate affordable relevance judgments and improve fusion model quality. Experiments conducted on six datasets from TREC’s Precision Medicine and Deep Learning tracks reveal that with careful sampling design, near-optimal fusion weights can be achieved using only 5% of the documents compared to the full TREC judgments. This translates to a dataset comprising 20 queries and 500 relevance-judged documents in total. The findings highlight the potential for sophisticated fusion techniques to become more accessible to researchers and practitioners, delivering substantial performance improvements with minimal judgment effort and cost.
Shengli Wu 0001, Chris D. Nugent, Adrian Moore 0001
Knowl. Inf. Syst.2
2023 Data Fusion Performance Prophecy: A Random Forest Revelation
Zhongmin Zhang, Shengli Wu 0001
iiWAS2
2023 Federated search techniques: an overview of the trends and state of the art
Adamu Garba, Shengli Wu 0001, Shah Khalid
Knowl. Inf. Syst.2
2022 Data Fusion Methods with Graded Relevance Judgment
Yidong Huang, Qiuyu Xu, Chunlin Xu, Shengli Wu 0001
WISA5
2022 Inexpensive and Effective Data Fusion Methods with Performance Weights
Qiuyu Xu, Yidong Huang, Shengli Wu 0001
iiWAS3
2021 Improving Medical Record Search Performance by Particle Swarm Optimization Based Data Fusion Techniques
Qiuyu Xu, Shengli Wu 0001
WISA2
2020 Word Embedding-Based Reformulation for Long Queries in Information Search
Chunlan Huang, Shengli Wu 0001
WISA4
2019 Multi-Level Matching Networks for Text Matching
abstract
Text matching aims to establish the matching relationship between two texts. It is an important operation in some information retrieval related tasks such as question duplicate detection, question answering, and dialog systems. Bidirectional long short term memory (BiLSTM) coupled with attention mechanism has achieved state-of-the-art performance in text matching. A major limitation of existing works is that only high level contextualized word representations are utilized to obtain word level matching results without considering other levels of word representations, thus resulting in incorrect matching decisions for cases where two words with different meanings are very close in high level contextualized word representation space. Therefore, instead of making decisions utilizing single level word representations, a multi-level matching network (MMN) is proposed in this paper for text matching, which utilizes multiple levels of word representations to obtain multiple word level matching results for final text level matching decision. Experimental results on two widely used benchmarks, SNLI and Scaitail, show that the proposed MMN achieves the state-of-the-art performance.
Chunlin Xu, Zhiwei Lin 0002, Shengli Wu 0001, Hui Wang 0001
SIGIR3
2018 LDA-Based Resource Selection for Results Diversification in Federated Search
Liang Li 0011, Zhongmin Zhang, Shengli Wu 0001
WISA3
2018 Evaluation of Score Standardization Methods for Web Search in Support of Results Diversification
Zhongmin Zhang, Chunlin Xu, Shengli Wu 0001
WISA3
2016 Differential Evolution-Based Fusion for Results Diversification of Web Search
Chunlin Xu, Chunlan Huang, Shengli Wu 0001
WAIM (1)3
2015 Detecting Near-Duplicate Documents Using Sentence Level Features
Jinbo Feng, Shengli Wu 0001
DEXA (2)2
2015 A geometric framework for data fusion in information retrieval
Shengli Wu 0001, Fabio Crestani
Inf. Syst.1
2014 Query Performance Prediction By Considering Score Magnitude and Variance Together
abstract
Query Performance prediction aims to evaluate the effectiveness of the results returned by a search system in response to a query without any relevance information. In this paper, we propose a method that considers both magnitude and variance of scores of the ranked list of results to measure the performance of a query. Using six different TREC test sets, we compare our predictor with three of the state-of-the-art techniques. The experimental results show that our method is very competitive. Pairwise comparisons with each of the three other methods show that our predictor performs better in more data sets.
Yongquan Tao, Shengli Wu 0001
CIKM2
2014 Search result diversification via data fusion
abstract
In recent years, researchers have investigated search result diversification through a variety of approaches. In such situations, information retrieval systems need to consider both aspects of relevance and diversity for those retrieved documents. On the other hand, previous research has demonstrated that data fusion is useful for improving performance when we are only concerned with relevance. However, it is not clear if it helps when both relevance and diversity are both taken into consideration. In this short paper, we propose a few data fusion methods to try to improve performance when both relevance and diversity are concerned. Experiments are carried out with 3 groups of top-ranked results submitted to the TREC web diversity task. We find that data fusion is still a useful approach to performance improvement for diversity as for relevance previously.
Shengli Wu 0001, Chunlan Huang
SIGIR1
2014 Adaptive data fusion methods in information retrieval
abstract
Data fusion is currently used extensively in information retrieval for various tasks. It has proved to be a useful technology because it is able to improve retrieval performance frequently. However, in almost all prior research in data fusion, static search environments have been used, and dynamic search environments have generally not been considered. In this article, we investigate adaptive data fusion methods that can change their behavior when the search environment changes. Three adaptive data fusion methods are proposed and investigated. To test these proposed methods properly, we generate a benchmark from a historic T ext RE trieval Conference data set. Experiments with the benchmark show that 2 of the proposed methods are good and may potentially be used in practice.
Shengli Wu 0001, Xiaoqin Zeng, Yaxin Bi
J. Assoc. Inf. Sci. Technol.1
2014 Clustering-Based Ensembles as an Alternative to Stacking
abstract
One of the most popular techniques of generating classifier ensembles is known as stacking which is based on a meta-learning approach. In this paper, we introduce an alternative method to stacking which is based on cluster analysis. Similar to stacking, instances from a validation set are initially classified by all base classifiers. The output of each classifier is subsequently considered as a new attribute of the instance. Following this, a validation set is divided into clusters according to the new attributes and a small subset of the original attributes of the instances. For each cluster, we find its centroid and calculate its class label. The collection of centroids is considered as a meta-classifier. Experimental results show that the new method outperformed all benchmark methods, namely Majority Voting, Stacking J48, Stacking LR, AdaBoost J48, and Random Forest, in 12 out of 22 data sets. The proposed method has two advantageous properties: it is very robust to relatively small training sets and it can be applied in semi-supervised learning problems. We provide a theoretical investigation regarding the proposed method. This demonstrates that for the method to be successful, the base classifiers applied in the ensemble should have greater than 50% accuracy levels.
Anna Jurek-Loughrey, Yaxin Bi, Shengli Wu 0001, Chris D. Nugent
IEEE Trans. Knowl. Data Eng.3
2013 The weighted Condorcet fusion in information retrieval
Shengli Wu 0001
Inf. Process. Manag.1
2012 Linear combination of component results in information retrieval
Shengli Wu 0001
Data Knowl. Eng.1
2011 The Linear Combination Data Fusion Method in Information Retrieval
Shengli Wu 0001, Yaxin Bi, Xiaoqin Zeng
DEXA (2)1
2010 Measuring Impact of Diversity of Classifiers on the Accuracy of Evidential Ensemble Classifiers
Yaxin Bi, Shengli Wu 0001
IPMU (1)2
2010 Retrieval Result Presentation and Evaluation
Shengli Wu 0001, Yaxin Bi, Xiaoqin Zeng
KSEM1
2009 A Comparative Analysis for Detecting Seismic Anomalies in Data Sequences of Outgoing Longwave Radiation
Yaxin Bi, Shengli Wu 0001, Pan Xiong, Xuhui Shen
KSEM2
2009 Assigning appropriate weights for the linear combination data fusion method in information retrieval
Shengli Wu 0001, Yaxin Bi, Xiaoqin Zeng, Lixin Han
Inf. Process. Manag.1
2008 The Experiments with the Linear Combination Data Fusion Method in Information Retrieval
Shengli Wu 0001, Yaxin Bi, Xiaoqin Zeng, Lixin Han
APWeb1
2008 Combining Classifiers through Triplet-Based Belief Functions
Yaxin Bi, Shengli Wu 0001, Xuhui Shen, Pan Xiong
ECML/PKDD (1)2
2007 A Geometric probabilistic framework for data fusion in information retrieval
abstract
Data fusion in information retrieval has been investigated by many researchers and quite a few data fusion methods have been proposed, but why data fusion can bring improvement in effectiveness is still not very clear. In this paper, we use a geometric probabilistic framework to formally describe data fusion, in which each component result returned from an information retrieval system for a given query is represented as a point in a multiple dimensional space. Then all the component results and data fusion results can be explained using geometrical principles. In such a framework, it becomes clear why quite often data fusion can bring improvement in effectiveness and accordingly what the favourable conditions are for data fusion algorithms to achieve better results. The framework can be used as a guideline to make data fusion techniques be used more effectively.
Shengli Wu 0001
FUSION1
2007 Result merging methods in distributed information retrieval with overlapping databases
Shengli Wu 0001, Sally I. McClean
Inf. Retr.1
2006 Evaluation of System Measures for Incomplete Relevance Judgment in IR
Shengli Wu 0001, Sally I. McClean
FQAS1
2006 Testing the cluster hypothesis in distributed information retrieval
Fabio Crestani, Shengli Wu 0001
Inf. Process. Manag.2
2006 Performance prediction of data fusion for information retrieval
Shengli Wu 0001, Sally I. McClean
Inf. Process. Manag.1
2006 Improving high accuracy retrieval by eliminating the uneven correlation effect in data fusion
abstract
Abstract The aim of this research is twofold. On the one hand, high accuracy retrieval has been a concern of the information retrieval community for some time. We aim to investigate this issue via data fusion. On the other hand, the correlation among component results has been proven harmful to data fusion, but it has not been taken into account in data fusion algorithms. In the hope of achieving better performance, we propose a group of algorithms to eliminate the effect of uneven correlation among component results by assigning different weights to all component results or their combinations. Then the linear combination method or a variation is used for fusion. Extensive experimentation is carried out to evaluate the performances of these algorithms with six groups of component results, which are the top 10 systems submitted to Text REtrieval Conference (TREC) 6, 7, 8, 9, 2001, and 2002. The experimental results show that all eight data fusion methods involved outperform the best component system on average. Therefore, we demonstrate that the data fusion technique in general is effective with accurate retrieval results. The experimental results also demonstrate that all six methods presented in this article are effective for eliminating the effect of uneven correlation among component results. All of them outperform CombSum and five of them outperform CombMNZ on average.
Shengli Wu 0001, Sally I. McClean
J. Assoc. Inf. Sci. Technol.1
2005 Data Fusion with Correlation Weights
Shengli Wu 0001, Sally I. McClean
ECIR1
2004 Effectiveness Evaluation and Comparison of Web Search Engines and Meta-search Engines
Shengli Wu 0001
WAIM1
2003 Experiments with Document Archive Size Detection
Shengli Wu 0001, Forbes Gibb, Fabio Crestani
ECIR1
2003 MIND: resource selection and data fusion in multimedia distributed digital libraries
abstract
No abstract available.
Stefano Berretti, Jamie Callan, Henrik Nottelmann, Xiao Mang Shou, Shengli Wu 0001
SIGIR5
2002 Data fusion with estimated weights
abstract
This paper proposes an adptive approach for data fusion of information retrieval systems, which exploits estimated performances of all component input systems without relevance judgement or training. The estimation is conducted prior to the fusion but uses the same data as fusion applies. The experiment shows that our algorithms are competitive with, and often outperform CombMNZ, one of the most effective algorithms in use.
Shengli Wu 0001, Fabio Crestani
CIKM1
2002 Authorization and Access Control of Application Data in Workflow Systems
Shengli Wu 0001, Amit P. Sheth, John A. Miller 0001, Zongwei Luo
J. Intell. Inf. Syst.1