VLDB 2026 Research / reviewers in the wild / expert
Shengli Wu 0001
dblp:w/ShengliWu
· DBLP profile ↗
81ranked-venue papers
26as first author
28since 2021 · last 2026
0000-0003-2008-1736ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 40 · 19 first-author · 8 since 2021Artificial intelligence and machine learning · 37 · 9 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Egocentric Action Recognition with Retrieval-Augmented LearningabstractEgocentric Action Recognition (EAR) aims to identify fine-grained actions and interacted objects from first-person videos, forming a core task in egocentric video understanding. Despite recent progress, EAR remains challenged by limited data scale, annotation quality, and long-tailed class distributions. To address these issues, we propose REAR, a Retrieval-augmented framework for EAR that leverages external third-person (exocentric) videos as auxiliary knowledge—without requiring synchronized ego-exo pairs. REAR adopts a dual-branch architecture: one branch extracts egocentric representations, while the other retrieves semantically relevant exocentric features. These are fused via a cross-view integration module that performs staged refinement and attention-based alignment. To mitigate class imbalance, a class-adaptive selector dynamically adjusts retrieval depth based on class frequency, and independent classifiers are trained with logit-adjusted cross-entropy. Extensive experiments across three benchmarks demonstrate that REAR achieves state-of-the-art performance, with significant gains in object recognition and tail-class accuracy. The source code is publicly available at https://github.com/zou-y23/REAR. Yishan Zou, Chris D. Nugent, Matthew Burns, Shengli Wu 0001, Meng Liu 0006 |
ICMR | 4 |
| 2026 | Subset selection based fusion for biomedical information retrieval tasksabstractTo improve the effectiveness and efficiency of biomedical information retrieval by proposing ranking-based methods for selecting an optimal subset of retrieval systems for data fusion, we propose three ranking-based subset selection methods SFS (Sequential Forward Search), D&P (Diversity & Performance), and P&D (Performance & Diversity). These methods were applied in combination with the Reciprocal Rank Fusion technique. Experiments were conducted on four medical datasets from TREC, using between 62 and 125 candidate retrieval systems, and selecting up to 15 for fusion. The proposed subset selection methods significantly improved retrieval performance. Fusing the selected systems using RRF yielded improvements ranging from 10% to over 60% compared to the best individual retrieval system across the datasets. They also outperform the state-of-the-art technology by a large margin. In summary, our subset selection approach offers a practical and cost-efficient solution for biomedical information retrieval, achieving substantial performance gains while reducing computational overhead. Shengli Wu 0001, Xiangjun Shen, Chris D. Nugent, Hu Lu |
BMC Bioinform. | 2 |
| 2026 | CLIP-driven fine-grained mining for text-based person search
Xianwen Lin, Xia Geng, Shengli Wu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2026 | DSGNet : A Lightweight Network Integrating Depthwise Separable and Ghost Convolutions for Real-Time Surface Defect SegmentationabstractABSTRACT In industrial product manufacturing, the automated detection and localisation of surface defects are of significant importance for ensuring quality control. However, existing computer vision‐based defect detection methods struggle to achieve both lightweight design and high accuracy on resource‐constrained embedded platforms, which limits their application in practical industrial detection environments. To address this issue, we propose DSGNet, a lightweight surface defect segmentation model, which serves as a core defect detection and localisation method for industrial inspection systems. The proposed model adopts an asymmetric encoder‐decoder structure to simplify the overall architecture. We designed an efficient feature extraction network by using four lightweight feature extraction units based on efficient convolutions. Furthermore, we introduce a hierarchical adaptive upsampling fusion (HAFU) mechanism and a lightweight bidirectional multiscale strip attention (LBMSA) feature refinement module to effectively fuse and refine the multilevel features extracted from the encoder. We conducted comprehensive evaluations of DSGNet on three typical surface defect datasets: Neu‐Seg, MSD and MT. While maintaining an extremely low complexity with only 0.49 M parameters, DSGNet achieved impressive mIoU scores of 83.39%, 91.61% and 80.72% on three datasets, respectively. These results indicate that DSGNet is a promising solution that balances lightweight design and detection accuracy for industrial real‐time detection systems, demonstrating strong potential for practical deployment. Our code is available at https://github.com/young‐zyy/DSGNet . Hu Lu, Guo Yang, Shengli Wu 0001 |
Expert Syst. J. Knowl. Eng. | 4 |
| 2026 | Complex query answering on knowledge graphs with Griffin and polarity-weighted message passing
Fanghao Li, Shengli Wu 0001, Hu Lu |
Knowl. Inf. Syst. | 4 |
| 2026 | FA-CDDL: Contrastive deep dictionary learning with frequency augmentation
Tianrui Huang, John Kingsley Arthur, Conghua Zhou, Shengli Wu 0001, Xiangjun Shen, Sirui Tian, Hongtao Li 0001 |
Knowl. Based Syst. | 5 |
| 2026 | Deep contrastive graph clustering with information preservation
Hu Lu, Haotian Hong, Fuhao Shi, Shengli Wu 0001, Lixin Duan, Shaohua Wan 0001 |
Pattern Recognit. | 4 |
| 2026 | Adaptive Attention-Based Unsupervised Domain Adaptation for Egocentric Action RecognitionabstractIn egocentric videos, collecting and annotating supervised data is more complicated and time-consuming than in exocentric videos, limiting research in this area. As a remedy, Unsupervised Domain Adaptation (UDA) enhances model performance on unlabeled target domains by bridging the distribution gap between source and target domains. However, UDA for egocentric action recognition is under-explored, facing unique challenges such as simultaneous learning of verb and noun representations, focusing on human-object interactions, and managing excessive verb-noun combinations. To tackle these issues, we propose a novel Unsupervised Domain Adaptation for Egocentric Action Recognition (UDA-EAR) approach that adaptively models egocentric actions and facilitates cross-domain knowledge transfer, improving recognition performance in unlabeled target domains. Specifically, our UDA-EAR employs adaptive spatio-temporal and spatio-channel attention in a dual-branch pipeline to focus on motion intervals and interaction regions, respectively, allowing specialized learning of discriminative representations while avoiding negative combination dependencies from domain gaps. Additionally, an adversarial domain alignment mechanism aligns the data distributions between source and target domains, effectively transferring fine-grained verb-noun knowledge of egocentric videos. Extensive experiments demonstrate that our UDA-EAR outperforms state-of-the-art baselines on widely used egocentric datasets, significantly improving egocentric action recognition accuracy. Our source codes and datasets are available at https://github.com/zou-y23/UDA-EAR. Yishan Zou, Chris D. Nugent, Matthew Burns, Shengli Wu 0001, Lei Zhu 0002, Meng Liu 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | DT-VNet: Deep Transformer-Based VNet Framework for 3D Prostate MRI SegmentationabstractMagnetic ResonanceImaging (MRI) is widely used in examining and diagnosing prostate diseases due to its high resolution. However, the diverse morphology of prostate tissue presents a significant challenge for precise gland segmentation. Convolutional Neural Networks have demonstrated effectiveness in segmenting prostate regions. Nevertheless, their limited capability in extracting global long-range semantic features often leads to unstable network segmentation performance. To address these challenges, we propose a Deep Transformer-based Vnet framework (DT-VNet), which consists of a symmetric encoder-decoder architecture that explores global contextual features and retains local feature information. To effectively learn global and local features, We propose the Deep Union Transformer (DU-Trans) as an encoding base module for capturing comprehensive information. Additionally, we introduce a Pool Fusion Attention (PFA) module for decoding, which emphasizes learning context dependencies and interaction relationships. PFA can also facilitate the fusion of deep and shallow features. To our knowledge, this is the first study about deep transformer-based Vnet framework for prostate segmentation. We validate and compare our method on several public datasets against current state-of-the-art methods. The results demonstrate the superior performance of our proposed method in segmenting 3D prostate MRI. Yunyao Cai, Hu Lu, Shengli Wu 0001, Stefano Berretti, Shaohua Wan 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Cost-effective data fusion in information retrievalabstractAbstract Data fusion has demonstrated its effectiveness in enhancing information retrieval across various studies. However, advanced fusion methods typically require a dataset with extensive relevance judgments to train optimal model weights, necessitating labor-intensive and costly manual efforts. This study explores efficient methods for generating training data to facilitate affordable relevance judgments and improve fusion model quality. Experiments conducted on six datasets from TREC’s Precision Medicine and Deep Learning tracks reveal that with careful sampling design, near-optimal fusion weights can be achieved using only 5% of the documents compared to the full TREC judgments. This translates to a dataset comprising 20 queries and 500 relevance-judged documents in total. The findings highlight the potential for sophisticated fusion techniques to become more accessible to researchers and practitioners, delivering substantial performance improvements with minimal judgment effort and cost. Shengli Wu 0001, Chris D. Nugent, Adrian Moore 0001 |
Knowl. Inf. Syst. | 2 |
| 2025 | Deep contrastive coordinated multi-view consistency clustering
Fuhao Shi, Shaohua Wan 0001, Shengli Wu 0001, Hui Wei 0001, Hu Lu |
Mach. Learn. | 3 |
| 2025 | CM-DASN: visible-infrared cross-modality person re-identification via dynamic attention selection network
Hu Lu, Tingting Qin, Juanjuan Tu, Shengli Wu 0001 |
Multim. Syst. | 5 |
| 2023 | ECA-CE: An Evolutionary Clustering Algorithm with Initial Population by Clustering Ensemble
Chunlin Xu, Shengli Wu 0001 |
ICAART (3) | 3 |
| 2023 | Data Fusion Performance Prophecy: A Random Forest Revelation
Zhongmin Zhang, Shengli Wu 0001 |
iiWAS | 2 |
| 2023 | Federated search techniques: an overview of the trends and state of the art
Adamu Garba, Shengli Wu 0001, Shah Khalid |
Knowl. Inf. Syst. | 2 |
| 2023 | A geometric framework for multiclass ensemble classifiersabstractAbstract Ensemble classifiers have been investigated by many in the artificial intelligence and machine learning community. Majority voting and weighted majority voting are two commonly used combination schemes in ensemble learning. However, understanding of them is incomplete at best, with some properties even misunderstood. In this paper, we present a group of properties of these two schemes formally under a geometric framework. Two key factors, every component base classifier’s performance and dissimilarity between each pair of component classifiers are evaluated by the same metric—the Euclidean distance. Consequently, ensembling becomes a deterministic problem and the performance of an ensemble can be calculated directly by a formula. We prove several theorems of interest and explain their implications for ensembles. In particular, we compare and contrast the effect of the number of component classifiers on these two types of ensemble schemes. Some important properties of both combination schemes are discussed. And a method to calculate the optimal weights for the weighted majority voting is presented. Empirical investigation is conducted to verify the theoretical results. We believe that the results from this paper are very useful for us to understand the fundamental properties of these two combination schemes and the principles of ensemble classifiers in general. The results are also helpful for us to investigate some issues in ensemble classifiers, such as ensemble performance prediction, diversity, ensemble pruning, and others. Shengli Wu 0001, Weimin Ding |
Mach. Learn. | 1 |
| 2023 | A correlation analysis framework via joint sample and feature selection
Na Qiang, Xiangjun Shen, Ernest Domanaanmwi Ganaa, Yang Yang 0001, Shengli Wu 0001, Zengmin Zhao, Shu-Cheng Huang |
Multim. Tools Appl. | 5 |
| 2022 | Data Fusion Methods with Graded Relevance Judgment
Yidong Huang, Qiuyu Xu, Chunlin Xu, Shengli Wu 0001 |
WISA | 5 |
| 2022 | Inexpensive and Effective Data Fusion Methods with Performance Weights
Qiuyu Xu, Yidong Huang, Shengli Wu 0001 |
iiWAS | 3 |
| 2022 | Diversified feature representation via deep auto-encoder ensemble through multiple activation functions
Na Qiang, Xiangjun Shen, Chang-Bin Huang, Shengli Wu 0001, Timothy Apasiba Abeo, Ernest Domanaanmwi Ganaa, Shu-Cheng Huang |
Appl. Intell. | 4 |
| 2022 | Feature recommendation strategy for graph convolutional networkabstractGraph Convolutional Network (GCN) is a new method for extracting, learning, and inferencing graph data that builds an embedded representation of the target node by aggregating information from neighbouring nodes. GCN is decisive for node classification and link prediction tasks in recent research. Although the existing GCN performs well, we argue that the current design ignores the potential features of the node. In addition, the presence of features with low correlation to nodes can likewise limit the learning ability of the model. Due to the above two problems, we propose Feature Recommendation Strategy (FRS) for Graph Convolutional Network in this paper. The core of FRS is to employ a principled approach to capture both node-to-node and node-to-feature relationships for encoding, then recommending the maximum possible features of nodes and replacing low-correlation features, and finally using GCN for learning of features. We perform a node clustering task on three citation network datasets and experimentally demonstrate that FRS can improve learning on challenging tasks relative to state-of-the-art (SOTA) baselines. Jisheng Qin, Xiaoqin Zeng, Shengli Wu 0001, Yang Zou 0001 |
Connect. Sci. | 3 |
| 2022 | Context-sensitive graph representation learningabstractGraph Convolutional Network (GCN) is a powerful emerging deep learning technique for learning graph data. However, there are still some challenges for GCN. For example, the model is shallow; the performance is poor when labelled nodes are severely scarce. In this paper, we propose a Multi-Semantic Aligned Graph Convolutional Network (MSAGCN), which contains two fundamental operations: multi-angle aggregation and semantic alignment, to resolve two challenges simultaneously. The core of MSAGCN is the aggregation of nodes that belong to the same class from three perspectives: nodes, features, and graph structure, and expects the obtained node features to be mapped nearby. Specifically, multi-angle aggregation is applied to extract features from three angles of the labelled nodes, and semantic alignment is utilised to align the semantics in the extracted features to enhance the similar content from different angles. In this way, the problem of over-smoothing and over-fitting for GCN can be alleviated. We perform the node clustering task on three citation datasets, and the experimental results demonstrate that our method outperforms the state-of-the-art (SOTA) baselines. Jisheng Qin, Xiaoqin Zeng, Shengli Wu 0001, Yang Zou 0001 |
Connect. Sci. | 3 |
| 2022 | Clustering-based fusion for medical information retrieval
Qiuyu Xu, Yidong Huang, Shengli Wu 0001, Chris D. Nugent |
J. Biomed. Informatics | 3 |
| 2021 | Improving Medical Record Search Performance by Particle Swarm Optimization Based Data Fusion Techniques
Qiuyu Xu, Shengli Wu 0001 |
WISA | 2 |
| 2021 | E-GCN: graph convolution with estimated labels
Jisheng Qin, Xiaoqin Zeng, Shengli Wu 0001, E. Tang |
Appl. Intell. | 3 |
| 2021 | Tag-Enhanced Dynamic Compositional Neural Network over arbitrary tree structure for sentence representation
Chunlin Xu, Hui Wang 0001, Shengli Wu 0001, Zhiwei Lin 0002 |
Expert Syst. Appl. | 3 |
| 2021 | TreeLSTM with tag-aware hypernetwork for sentence representation
Chunlin Xu, Hui Wang 0001, Shengli Wu 0001, Zhiwei Lin 0002 |
Neurocomputing | 3 |
| 2021 | Computation of CNN's Sensitivity to Input Perturbation
Lin Xiang 0003, Xiaoqin Zeng, Shengli Wu 0001, Yanjun Liu 0004, Baohua Yuan |
Neural Process. Lett. | 3 |
| 2020 | Word Embedding-Based Reformulation for Long Queries in Information Search
Chunlan Huang, Shengli Wu 0001 |
WISA | 4 |
| 2019 | Multi-Level Compare-Aggregate Model for Text MatchingabstractText matching is important for a variety of natural language processing tasks, such as paraphrase identification and natural language inference. Recent studies have achieved very promising results under the compare-aggregate framework. A limitation of previous approaches following this framework is that they solely conduct matching at word level. In this paper, we propose a multi-level compare-aggregate model (MLCA), which matches each word in one text against the other text at three different levels, word level (word-by-word matching), phrase level (word-by-phrase matching) and sentence level (word-bysentence matching). Then the results of these levels of matching are aggregated for making final matching decision. We evaluate our model on two different tasks: paraphrase identification and natural language inference. Experimental results show that our model achieves the state-of-the-art performance on both tasks. Chunlin Xu, Zhiwei Lin 0002, Hui Wang 0001, Shengli Wu 0001 |
IJCNN | 4 |
| 2019 | Multi-Level Matching Networks for Text MatchingabstractText matching aims to establish the matching relationship between two texts. It is an important operation in some information retrieval related tasks such as question duplicate detection, question answering, and dialog systems. Bidirectional long short term memory (BiLSTM) coupled with attention mechanism has achieved state-of-the-art performance in text matching. A major limitation of existing works is that only high level contextualized word representations are utilized to obtain word level matching results without considering other levels of word representations, thus resulting in incorrect matching decisions for cases where two words with different meanings are very close in high level contextualized word representation space. Therefore, instead of making decisions utilizing single level word representations, a multi-level matching network (MMN) is proposed in this paper for text matching, which utilizes multiple levels of word representations to obtain multiple word level matching results for final text level matching decision. Experimental results on two widely used benchmarks, SNLI and Scaitail, show that the proposed MMN achieves the state-of-the-art performance. Chunlin Xu, Zhiwei Lin 0002, Shengli Wu 0001, Hui Wang 0001 |
SIGIR | 3 |
| 2018 | LDA-Based Resource Selection for Results Diversification in Federated Search
Liang Li 0011, Zhongmin Zhang, Shengli Wu 0001 |
WISA | 3 |
| 2018 | Evaluation of Score Standardization Methods for Web Search in Support of Results Diversification
Zhongmin Zhang, Chunlin Xu, Shengli Wu 0001 |
WISA | 3 |
| 2017 | Information Retrieval with Implicitly Temporal Queries
Shengli Wu 0001 |
IDEAL | 2 |
| 2016 | Differential Evolution-Based Fusion for Results Diversification of Web Search
Chunlin Xu, Chunlan Huang, Shengli Wu 0001 |
WAIM (1) | 3 |
| 2016 | A Mean-Variance Analysis Based Approach for Search Result Diversification in Federated SearchabstractResource Selection is an important step in a federated search environment. The goal of this work was to improve the collection selection process by selecting collections in terms of relevance and diversity, to best answer a user's query. Sampled documents from the Central Sample Database are first ranked by Indri retrieval algorithm and later re-ranked by a Mean-Standard deviation method that reduces uncertainty and improves diversity of collection sources. A comparative evaluation with the R-based diversification metrics shows that the proposed method significantly outperforms the baseline diversification methods; ReDDE+MMR, ReDDE+MAP-IA and state-of-the-art resource selection methods (ReDDE and CORI) in all metrics. Benjamin Ghansah, Shengli Wu 0001 |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 2 |
| 2015 | Detecting Near-Duplicate Documents Using Sentence Level Features
Jinbo Feng, Shengli Wu 0001 |
DEXA (2) | 2 |
| 2015 | A geometric framework for data fusion in information retrieval
Shengli Wu 0001, Fabio Crestani |
Inf. Syst. | 1 |
| 2014 | Query Performance Prediction By Considering Score Magnitude and Variance TogetherabstractQuery Performance prediction aims to evaluate the effectiveness of the results returned by a search system in response to a query without any relevance information. In this paper, we propose a method that considers both magnitude and variance of scores of the ranked list of results to measure the performance of a query. Using six different TREC test sets, we compare our predictor with three of the state-of-the-art techniques. The experimental results show that our method is very competitive. Pairwise comparisons with each of the three other methods show that our predictor performs better in more data sets. Yongquan Tao, Shengli Wu 0001 |
CIKM | 2 |
| 2014 | Search result diversification via data fusionabstractIn recent years, researchers have investigated search result diversification through a variety of approaches. In such situations, information retrieval systems need to consider both aspects of relevance and diversity for those retrieved documents. On the other hand, previous research has demonstrated that data fusion is useful for improving performance when we are only concerned with relevance. However, it is not clear if it helps when both relevance and diversity are both taken into consideration. In this short paper, we propose a few data fusion methods to try to improve performance when both relevance and diversity are concerned. Experiments are carried out with 3 groups of top-ranked results submitted to the TREC web diversity task. We find that data fusion is still a useful approach to performance improvement for diversity as for relevance previously. Shengli Wu 0001, Chunlan Huang |
SIGIR | 1 |
| 2014 | Adaptive data fusion methods in information retrievalabstractData fusion is currently used extensively in information retrieval for various tasks. It has proved to be a useful technology because it is able to improve retrieval performance frequently. However, in almost all prior research in data fusion, static search environments have been used, and dynamic search environments have generally not been considered. In this article, we investigate adaptive data fusion methods that can change their behavior when the search environment changes. Three adaptive data fusion methods are proposed and investigated. To test these proposed methods properly, we generate a benchmark from a historic T ext RE trieval Conference data set. Experiments with the benchmark show that 2 of the proposed methods are good and may potentially be used in practice. Shengli Wu 0001, Xiaoqin Zeng, Yaxin Bi |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2014 | Clustering-Based Ensembles as an Alternative to StackingabstractOne of the most popular techniques of generating classifier ensembles is known as stacking which is based on a meta-learning approach. In this paper, we introduce an alternative method to stacking which is based on cluster analysis. Similar to stacking, instances from a validation set are initially classified by all base classifiers. The output of each classifier is subsequently considered as a new attribute of the instance. Following this, a validation set is divided into clusters according to the new attributes and a small subset of the original attributes of the instances. For each cluster, we find its centroid and calculate its class label. The collection of centroids is considered as a meta-classifier. Experimental results show that the new method outperformed all benchmark methods, namely Majority Voting, Stacking J48, Stacking LR, AdaBoost J48, and Random Forest, in 12 out of 22 data sets. The proposed method has two advantageous properties: it is very robust to relatively small training sets and it can be applied in semi-supervised learning problems. We provide a theoretical investigation regarding the proposed method. This demonstrates that for the method to be successful, the base classifiers applied in the ensemble should have greater than 50% accuracy levels. Anna Jurek-Loughrey, Yaxin Bi, Shengli Wu 0001, Chris D. Nugent |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2013 | Measuring Stability and Discrimination Power of Metrics in Information Retrieval Evaluation
Huaji Shi, Yanzhi Tan, Shengli Wu 0001 |
IDEAL | 4 |
| 2013 | Merging Results from Overlapping Databases in Distributed Information RetrievalabstractIn this paper, we investigate the problem of results merging in distributed information retrieval when overlapping databases are used. We focus on two issues: score normalization and weights assignment for each of the component results. Empirical study with the TREC data has the following three findings: 1. The cubic regression model and logistic regression model are better than the commonly used zero-one score normalization method, 2. The weighting scheme of uneven similarity is an effective method of weights assignment. 3. Score normalization and weights assignment can be used separately or together in a results merging method to improve effectiveness. The findings obtained in this paper are very useful for effectiveness improvement when implementing a distributed information retrieval system. Shengli Wu 0001 |
PDP | 1 |
| 2013 | The weighted Condorcet fusion in information retrieval
Shengli Wu 0001 |
Inf. Process. Manag. | 1 |
| 2013 | Effective Neural Network Ensemble Approach for Improving Generalization PerformanceabstractThis paper, with an aim at improving neural networks' generalization performance, proposes an effective neural network ensemble approach with two novel ideas. One is to apply neural networks' output sensitivity as a measure to evaluate neural networks' output diversity at the inputs near training samples so as to be able to select diverse individuals from a pool of well-trained neural networks; the other is to employ a learning mechanism to assign complementary weights for the combination of the selected individuals. Experimental results show that the proposed approach could construct a neural network ensemble with better generalization performance than that of each individual in the ensemble combining with all the other individuals, and than that of the ensembles with simply averaged weights. Jing Yang 0028, Xiaoqin Zeng, Shuiming Zhong, Shengli Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2012 | A Cluster-Based Classifier Ensemble as an Alternative to the Nearest Neighbor EnsembleabstractThe combination of multiple classifiers, commonly referred to as an ensemble, has previously demonstrated the ability to improve overall classification accuracy in many application domains. Some ensemble techniques, however, cannot easily improve the performance of stable classification methods. One such example of a stable classification method is the k Nearest Neighbor (kNN) Classifier. In this paper we propose an alternative to the kNN ensemble method through the use of a clustering technique applied for the purpose of selecting the neighborhood of a new instance. In addition, a novel combination function based on exponential support (ExSupp) has been introduced. The proposed approach exhibited improved classification results in 16 out 20 data sets which were considered in comparison with a single kNN and a kNN ensemble based approach. Besides higher classification accuracy the proposed method exhibited higher levels of efficiency in terms of classification time. Anna Jurek-Loughrey, Yaxin Bi, Shengli Wu 0001, Chris D. Nugent |
ICTAI | 3 |
| 2012 | Linear combination of component results in information retrieval
Shengli Wu 0001 |
Data Knowl. Eng. | 1 |
| 2012 | Applying the data fusion technique to blog opinion retrieval
Shengli Wu 0001 |
Expert Syst. Appl. | 1 |
| 2012 | Sensitivity-Based Adaptive Learning Rules for Binary Feedforward Neural NetworksabstractThis paper proposes a set of adaptive learning rules for binary feedforward neural networks (BFNNs) by means of the sensitivity measure that is established to investigate the effect of a BFNN's weight variation on its output. The rules are based on three basic adaptive learning principles: the benefit principle, the minimal disturbance principle, and the burden-sharing principle. In order to follow the benefit principle and the minimal disturbance principle, a neuron selection rule and a weight adaptation rule are developed. Besides, a learning control rule is developed to follow the burden-sharing principle. The advantage of the rules is that they can effectively guide the BFNN's learning to conduct constructive adaptations and avoid destructive ones. With these rules, a sensitivity-based adaptive learning (SBALR) algorithm for BFNNs is presented. Experimental results on a number of benchmark data demonstrate that the SBALR algorithm has better learning performance than the Madaline rule II and backpropagation algorithms. Shuiming Zhong, Xiaoqin Zeng, Shengli Wu 0001, Lixin Han |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2011 | The Linear Combination Data Fusion Method in Information Retrieval
Shengli Wu 0001, Yaxin Bi, Xiaoqin Zeng |
DEXA (2) | 1 |
| 2011 | Combination of Evidence-Based Classifiers for Text CategorizationabstractIn this paper we propose an evidential fusion approach to combining the decisions of text classifiers. These text classifiers are generated by four widely used learning algorithms: Support Vector Machine (SVM), kNN (Nearest Neighbour), kNN model-based approach (kNNM), and Rocchio on two text corpora. We first model each classifier output as a list of prioritized decisions and then divide it into the subsets of 2 and 3 decisions which are subsequently represented by the evidential structures in terms of triplet and quartet. We also develop the general formulae based on the Dempster- Shafer theory of evidence for combining such decisions. To validate our method various experiments have been carried out over the data sets of 20-newsgroup and Reuters-21578, and a comparative analysis with an alternative dichotomous structure and with majority voting have also been conducted to demonstrate the advantage of our approach in combining text classifiers. Yaxin Bi, Shengli Wu 0001, Hui Wang 0001, Gongde Guo |
ICTAI | 2 |
| 2011 | Classification by Clusters Analysis - An Ensemble Technique in a Semi-supervised ClassificationabstractIn this work we adopt a previously introduced meta-learning classification method for semi-supervised learning problems. In our previous work we illustrated that the method is successful when applied in a supervised classification problem. In our current work the results demonstrate that following refinements made to the method it can be successfully applied to semi-supervised classification cases. Anna Jurek-Loughrey, Yaxin Bi, Shengli Wu 0001, Chris D. Nugent |
ICTAI | 3 |
| 2010 | Measuring Impact of Diversity of Classifiers on the Accuracy of Evidential Ensemble Classifiers
Yaxin Bi, Shengli Wu 0001 |
IPMU (1) | 2 |
| 2010 | Retrieval Result Presentation and Evaluation
Shengli Wu 0001, Yaxin Bi, Xiaoqin Zeng |
KSEM | 1 |
| 2009 | A Comparative Analysis for Detecting Seismic Anomalies in Data Sequences of Outgoing Longwave Radiation
Yaxin Bi, Shengli Wu 0001, Pan Xiong, Xuhui Shen |
KSEM | 2 |
| 2009 | A Sensitivity-Based Training Algorithm with Architecture Adjusting for MadalinesabstractHow to design proper architectures of neural networks for solving given problems is an important issue in neural network research. Nowadays, the existing training algorithms of neural networks only focus on adjusting neural networks' weights to improve training accuracy, and few of them adaptively adjust the networks' architecture. However, the architecture is indeed very critical for training neural networks to have high performance and needs to be coped with in the training process. In this paper, we present a new training algorithm of Madalines, which takes not only weight but also architecture adjusting into consideration. The algorithm can thus train Madalines with smaller architecture and higher generalization ability. Experimental results have demonstrated that our algorithm is effective. Yanjun Liu 0004, Xiaoqin Zeng, Shuiming Zhong, Shengli Wu 0001 |
SMC | 4 |
| 2009 | Applying statistical principles to data fusion in information retrieval
Shengli Wu 0001 |
Expert Syst. Appl. | 1 |
| 2009 | Assigning appropriate weights for the linear combination data fusion method in information retrieval
Shengli Wu 0001, Yaxin Bi, Xiaoqin Zeng, Lixin Han |
Inf. Process. Manag. | 1 |
| 2008 | The Experiments with the Linear Combination Data Fusion Method in Information Retrieval
Shengli Wu 0001, Yaxin Bi, Xiaoqin Zeng, Lixin Han |
APWeb | 1 |
| 2008 | Classifier Combination Using a Class-indifferent MethodabstractIn this paper we present a novel approach to combining classifiers in the Dempster-Shafer theory framework. This approach models each output given by classifiers as a list of ranked decisions (classes), which is partitioned into a new evidence structure called a triplet. Resulting triplets are then combined by Dempster's rule. With a triplet, its first subset contains a decision corresponding to the largest numeric value of classes, the second subset corresponds to the second largest numeric value and the third subset represents uncertainty information in determining the support for the former two decisions. We carry out a comparative analysis with the combination methods of majority voting, stacking and boosting on the UCI benchmark data to demonstrate the advantage of our approach. Yaxin Bi, Shengli Wu 0001, Pan Xiong, Xuhui Shen |
ECAI | 2 |
| 2008 | A Novel Ensemble Approach for Improving Generalization Ability of Neural Networks
Xiaoqin Zeng, Shengli Wu 0001, Shuiming Zhong |
IDEAL | 3 |
| 2008 | Performance Weights for the Linear Combination Data Fusion Method in Information Retrieval
Shengli Wu 0001, Qili Zhou, Yaxin Bi, Xiaoqin Zeng |
ISMIS | 1 |
| 2008 | Combining Classifiers through Triplet-Based Belief Functions
Yaxin Bi, Shengli Wu 0001, Xuhui Shen, Pan Xiong |
ECML/PKDD (1) | 2 |
| 2007 | A Geometric probabilistic framework for data fusion in information retrievalabstractData fusion in information retrieval has been investigated by many researchers and quite a few data fusion methods have been proposed, but why data fusion can bring improvement in effectiveness is still not very clear. In this paper, we use a geometric probabilistic framework to formally describe data fusion, in which each component result returned from an information retrieval system for a given query is represented as a point in a multiple dimensional space. Then all the component results and data fusion results can be explained using geometrical principles. In such a framework, it becomes clear why quite often data fusion can bring improvement in effectiveness and accordingly what the favourable conditions are for data fusion algorithms to achieve better results. The framework can be used as a guideline to make data fusion techniques be used more effectively. Shengli Wu 0001 |
FUSION | 1 |
| 2007 | Combining Prioritized Decisions in Classification
Yaxin Bi, Shengli Wu 0001, Gongde Guo |
MDAI | 2 |
| 2007 | Applying statistical principles to data fusion in information retrievalabstractData fusion in information retrieval has been investigated by many researchers and quite a few data fusion methods have been proposed. However, their impact on effectiveness has not been well understood. In this paper, we apply statistical principles to data fusion and present a statistical data fusion model, which specifies the algorithm for fusion and conditions to be satisfied. The statistical model can be used as a guideline for data fusion methods. Based on this analysis, we compare CombSum and CombMNZ, which are the two best-known data fusion methods. We explain why sometimes CombMNZ does outperform Comb- Sum and what can be done to make CombSum more effective. Experimental results with TREC data are reported to support the conclusion that our enhancements to the algorithm improve effectiveness. Shengli Wu 0001, Yaxin Bi, Sally I. McClean |
SMC | 1 |
| 2007 | Result merging methods in distributed information retrieval with overlapping databases
Shengli Wu 0001, Sally I. McClean |
Inf. Retr. | 1 |
| 2006 | Evaluation of System Measures for Incomplete Relevance Judgment in IR
Shengli Wu 0001, Sally I. McClean |
FQAS | 1 |
| 2006 | Combining Multiple Sets of Rules for Improving Classification Via Measuring Their Closenesses
Yaxin Bi, Shengli Wu 0001, Xuming Huang, Gongde Guo |
PRICAI | 2 |
| 2006 | Testing the cluster hypothesis in distributed information retrieval
Fabio Crestani, Shengli Wu 0001 |
Inf. Process. Manag. | 2 |
| 2006 | Performance prediction of data fusion for information retrieval
Shengli Wu 0001, Sally I. McClean |
Inf. Process. Manag. | 1 |
| 2006 | Improving high accuracy retrieval by eliminating the uneven correlation effect in data fusionabstractAbstract The aim of this research is twofold. On the one hand, high accuracy retrieval has been a concern of the information retrieval community for some time. We aim to investigate this issue via data fusion. On the other hand, the correlation among component results has been proven harmful to data fusion, but it has not been taken into account in data fusion algorithms. In the hope of achieving better performance, we propose a group of algorithms to eliminate the effect of uneven correlation among component results by assigning different weights to all component results or their combinations. Then the linear combination method or a variation is used for fusion. Extensive experimentation is carried out to evaluate the performances of these algorithms with six groups of component results, which are the top 10 systems submitted to Text REtrieval Conference (TREC) 6, 7, 8, 9, 2001, and 2002. The experimental results show that all eight data fusion methods involved outperform the best component system on average. Therefore, we demonstrate that the data fusion technique in general is effective with accurate retrieval results. The experimental results also demonstrate that all six methods presented in this article are effective for eliminating the effect of uneven correlation among component results. All of them outperform CombSum and five of them outperform CombMNZ on average. Shengli Wu 0001, Sally I. McClean |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2005 | Data Fusion with Correlation Weights
Shengli Wu 0001, Sally I. McClean |
ECIR | 1 |
| 2004 | Effectiveness Evaluation and Comparison of Web Search Engines and Meta-search Engines
Shengli Wu 0001 |
WAIM | 1 |
| 2003 | Experiments with Document Archive Size Detection
Shengli Wu 0001, Forbes Gibb, Fabio Crestani |
ECIR | 1 |
| 2003 | MIND: resource selection and data fusion in multimedia distributed digital librariesabstractNo abstract available. Stefano Berretti, Jamie Callan, Henrik Nottelmann, Xiao Mang Shou, Shengli Wu 0001 |
SIGIR | 5 |
| 2003 | Distributed Information Retrieval: A Multi-Objective Resource Selection ApproachabstractInformation retrieval is becoming increasingly concerned with resource selection and data fusion for distributed archives. In distributed information retrieval, a user submits a query to a broker, which determines a solution for how to yield a given number of documents from all available resources. In this paper, we present a multi-objective model for resource selection, in which four aspects: a document's relevance to the given query, time, monetary cost, and the chance of getting document duplicates from resources, are considered simultaneously. Some variants of this multi-objective model, aimed at achieving better implementation efficiency, are also proposed. Shengli Wu 0001, Fabio Crestani |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |
| 2002 | Data fusion with estimated weightsabstractThis paper proposes an adptive approach for data fusion of information retrieval systems, which exploits estimated performances of all component input systems without relevance judgement or training. The estimation is conducted prior to the fusion but uses the same data as fusion applies. The experiment shows that our algorithms are competitive with, and often outperform CombMNZ, one of the most effective algorithms in use. Shengli Wu 0001, Fabio Crestani |
CIKM | 1 |
| 2002 | Authorization and Access Control of Application Data in Workflow Systems
Shengli Wu 0001, Amit P. Sheth, John A. Miller 0001, Zongwei Luo |
J. Intell. Inf. Syst. | 1 |
| 2001 | GIMS - A Data Warehouse for Storage and Analysis of Genome Sequence and Functional DataabstractEffective analysis of genome sequences and associated functional data requires access to many different kinds of biological information. For example, when analysing gene expression data, it may be useful to have access to the sequences upstream of the genes, or to the cellular location of their protein products. Such information is currently stored in different formats at different sites in a way that does not readily allow integrated analyses. The Genome Information Management System (GIMS) is an object database that integrates genome sequence data with functional data on the transcriptome and on protein-protein interactions in a single data warehouse. We have used GIMS to store the Saccharomyces cerevisiae (yeast) genome and to demonstrate how the integrated storage of diverse kinds of genomic data can be beneficial for analysing data using context-rich queries and analyses. GIMS allows data to be stored in a way that reflects the underlying mechanisms in the organism, and permits complex questions to be asked of the data. This paper provides an overview of the GIMS system and describes some analyses that illustrate its use for analysing functional data sets for S. cerevisiae. Mike Cornell, Norman W. Paton, Shengli Wu 0001, Carole A. Goble, Crispin J. Miller, Paul Kirby, Karen Eilbeck, Andy Brass, Andrew Hayes, Stephen G. Oliver |
BIBE | 3 |