Xiaoling Zhou

dblp:127/6140 · DBLP profile ↗
← Back
33ranked-venue papers
25as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 19 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 8 since 2021Databases, data management, data science and information retrieval · 9 · 7 first-author · 5 since 2021
YearPublicationVenuePosition
2026 ASKD: Reinforcement Learning-Style Knowledge Distillation with Quality-Adaptive Skewness
abstract
Knowledge distillation (KD) is a widely adopted technique for transferring the capabilities of large teacher models to smaller student models, thereby significantly reducing inference costs and memory consumption. However, existing KD methods are all constrained by an inherent greedy optimization objective, rooted in the assumption of teacher superiority: "Trust all teacher-generated outputs (TGOs)" and "Distrust any student-generated outputs (SGOs) unsupported by the teacher". We propose ASKD, a novel KD method with adaptive skewness determined by sample quality, refining this objective to: "Learn TGOs proportionally to their quality, and distrust only low-quality unsupported SGOs". ASKD comprises three key components: (1) A reinforcement learning-style optimization formulation to mitigate the inherent approximation bias in sample-based Kullback-Leibler (KL) divergence approximations used by previous KD methods; (2) Well-designed quality supervision signals to map and achieve adaptive skewness in skewed KL loss, pioneering the usage of sample quality to adjust learning magnitudes; (3) A gradient-clip function on high-quality SGOs for findings that high-quality SGOs in KL loss fail to yield positive updates and even cause adverse effects on some samples. Extensive experiments indicate that ASKD builds high-performance student models across various tasks, including instruction following, mathematical reasoning, and code generation, outperforming state-of-the-art methods comprehensively and surpassing GRPO-like approaches that use advantages as multiplicative factors. We also provide detailed mathematical proofs demonstrating properties such as Lipschitz continuity of the update coefficient and uniform convergence of the loss function, ensuring theoretical rigor for key components of ASKD.
Xiaoling Zhou, Yiyu Liu, Shikun Zhang, Wei Ye 0004
AAAI2
2026 LeLoRA: Learnable Low-Rank Adaptation of Large Language Models
abstract
Fine-tuning large language models (LLMs) is an effective approach to enhancing their performance on specialized downstream tasks.Among the various techniques, low-rank adaptation has garnered significant attention due to its ability to maintain the full performance of fine-tuning while enhancing computational efficiency.However, existing approaches often rely on manually specified and fixed hyperparameters to identify the trainable components within weight matrices, resulting in suboptimal performance and low parameter efficiency.This paper presents a novel Learnable Low-Rank Adaptation (LeLoRA) framework that utilizes dynamically learned fine-tuning strategies to facilitate the effective adaptation of LLMs.Our framework integrates an LLM with a policy network that automatically and adaptively generates matrix-specific adaptation strategies to identify the trainable components of each weight matrix, taking into account their unique characteristics, such as singular values and matrix norms.A reinforcement learningbased optimization algorithm is then employed to iteratively update the LLM and the policy network, ensuring that the generated strategies adapt in real time to the evolving states of the LLM.Extensive experiments have been conducted across various natural language processing tasks.The results across ten different LLMs, ranging from 125M to 70B parameters, provide compelling evidence that LeLoRA consistently outperforms existing baselines in adapting LLMs.
Xiaoling Zhou, Zhemg Lee, Wei Ye 0004, Shikun Zhang
ACL (1)1
2026 Defects inspection system for building facades using drones and deep learning method
Xiaoling Zhou, Robert L. K. Tiong
Expert Syst. Appl.1
2026 Dynamic Knowledge Transfer for Mitigating Spurious Correlations in Deep Learning
Xiaoling Zhou, Zhemg Lee, Wei Ye 0004, Shikun Zhang
Int. J. Comput. Vis.1
2026 Towards robust and generalizable learning via implicit counterfactual data augmentation
Xiaoling Zhou, Ou Wu 0001, Michael Kwok-Po Ng
Pattern Recognit.1
2025 Class and Attribute-Aware Logit Adjustment for Generalized Long-Tail Learning
abstract
Compared to conventional long-tail learning, which focuses on addressing class-wise imbalances, generalized long-tail (GLT) learning considers that samples within each class still conform to long-tailed distributions due to varying attributes, known as attribute imbalance. In the presence of such imbalance, the assumption of equivalence between the class-conditional probability densities of the training and testing sets is no longer tenable. Existing GLT approaches typically employ regularization techniques to avoid directly modeling the class-conditional probability density (CCPD) ratio between training and test data, leading to suboptimal performance. This study aims to directly estimate this ratio, for which a novel class-attribute aware logit-adjusted (CALA) loss incorporating both the CCPD ratio and the class priors is presented. Two new GLT learning methods, named Heuristic-CALA and Meta-CALA, are then proposed, which estimate the CCPD ratio in the CALA loss by leveraging the neighborhood information of samples. Extensive experiments across diverse scenarios susceptible to class and attribute imbalances showcase the state-of-the-art performance of Meta-CALA. Furthermore, while Heuristic-CALA exhibits inferior performance compared to Meta-CALA, it incurs only negligible additional training time compared to the Cross-Entropy loss, yet surpasses existing methods by a significant margin.
Xiaoling Zhou, Ou Wu 0001
AAAI1
2025 All-Optical Nonlinear Diffractive Deep Network for Ultrafast Image Denoising
abstract
Image denoising poses a significant challenge in image processing, aiming to remove noise and artifacts from input images. However, current denoising algorithms implemented on electronic chips frequently encounter latency issues and demand substantial computational resources. In this paper, we introduce an all-optical Nonlinear Diffractive Denoising Deep Network (N3DNet) for image denoising at the speed of light. Initially, we incorporate an image encoding and pre-denoising module into the Diffractive Deep Neural Network and integrate a nonlinear activation function, termed the phase exponential linear function, after each diffractive layer, thereby boosting the network’s nonlinear modeling and denoising capabilities. Subsequently, we devise a new reinforcement learning algorithm called regularization-assisted deep Q-network to optimize N3DNet. Finally, leveraging 3D printing techniques, we fabricate N3DNet using the trained parameters and construct a physical experimental system for real-world applications. A new benchmark dataset, termed MIDD, is constructed for mode image denoising, comprising 120K pairs of noisy/noise-free images captured from real fiber communication systems across various transmission lengths. Through extensive simulation and real experiments, we validate that N3DNet outperforms both traditional and deep learning-based denoising approaches across various datasets. Remarkably, its processing speed is nearly 3,800 times faster than electronic chip-based methods.
Xiaoling Zhou, Zhemg Lee, Wei Ye 0004, Rui Xie 0003, Guanju Peng, Shikun Zhang
CVPR1
2025 HaDeMiF: Hallucination Detection and Mitigation in Large Language Models
abstract
The phenomenon of knowledge hallucinations has raised substantial concerns about the security and reliability of deployed large language models (LLMs). Current methods for detecting hallucinations primarily depend on manually designed individual metrics, such as prediction uncertainty and consistency, and fall short in effectively calibrating model predictions, thus constraining their detection accuracy and applicability in practical applications. In response, we propose an advanced framework, termed HaDeMiF, for detecting and mitigating hallucinations in LLMs. Specifically, hallucinations within the output and semantic spaces of LLMs are comprehensively captured through two compact networks—a novel, interpretable tree model known as the Deep Dynamic Decision Tree (D3T) and a Multilayer Perceptron (MLP)—which take as input a set of prediction characteristics and the hidden states of tokens, respectively. The predictions of LLMs are subsequently calibrated using the outputs from the D3T and MLP networks, aiming to mitigate hallucinations and enhance model calibration. HaDeMiF can be applied during both the inference and fine-tuning phases of LLMs, introducing less than 2% of the parameters relative to the LLMs through the training of two small-scale networks. Extensive experiments conclusively demonstrate the effectiveness of our framework in hallucination detection and model calibration across text generation tasks with responses of varying lengths.
Xiaoling Zhou, Zhemg Lee, Wei Ye 0004, Shikun Zhang
ICLR1
2025 Robustness to Spurious Correlations via Dynamic Knowledge Transfer
abstract
Spurious correlations pose a significant challenge to the robustness of statistical models, often resulting in unsatisfactory performance when distributional shifts occur between training and testing data. To address this, we propose to transfer knowledge across spuriously correlated categories within the deep feature space. Specifically, samples' deep features are enriched using semantic vectors extracted from both their respective category distributions and those of their spuriously correlated counterparts, enabling the generation of diverse class-specific factual and counterfactual augmented deep features. We then demonstrate the feasibility of optimizing a surrogate robust loss instead of conducting explicit augmentations by considering an infinite number of augmentations. As spurious correlations between samples and classes evolve during training, we develop a reinforcement learning-based training framework called Dynamic Knowledge Transfer (DKT) to facilitate dynamic adjustments in the direction and intensity of knowledge transfer. Within this framework, a target network is trained using the derived robust loss to enhance robustness, while a strategy network generates sample-wise augmentation strategies in a dynamic and automatic way. Extensive experiments validate the effectiveness of the DKT framework in mitigating spurious correlations, achieving state-of-the-art performance across three typical learning scenarios susceptible to such correlations.
Xiaoling Zhou, Wei Ye 0004, Zhemg Lee, Shikun Zhang
IJCAI1
2025 Boosting Resilience of Large Language Models through Causality-Driven Robust Optimization
abstract
Large language models (LLMs) have achieved remarkable achievements across diverse applications; however, they remain plagued by spurious correlations and the generation of hallucinated content. Despite extensive efforts to enhance the resilience of LLMs, existing approaches either rely on indiscriminate fine-tuning of all parameters, resulting in parameter inefficiency and lack of specificity, or depend on post-processing techniques that offer limited adaptability and flexibility. This study introduces a novel Causality-driven Robust Optimization (CdRO) approach that selectively updates model components sensitive to causal reasoning, enhancing model causality while preserving valuable pretrained knowledge to mitigate overfitting. Our method begins by identifying the parameter components within LLMs that capture causal relationships, achieved through comparing the training dynamics of parameter matrices associated with the original samples, as well as augmented counterfactual and paraphrased variants. These comparisons are then fed into a lightweight logistic regression model, optimized in real time to dynamically identify and adapt the causal components within LLMs. The identified parameters are subsequently optimized using an enhanced policy optimization algorithm, where the reward function is designed to jointly promote both model generalization and robustness. Extensive experiments across various tasks using twelve different LLMs demonstrate the superior performance of our framework, underscoring its significant effectiveness in reducing the model’s dependence on spurious associations and mitigating hallucinations.
Xiaoling Zhou, Zhemg Lee, Yuncheng Hua, Chengli Xing, Wei Ye 0004, Flora D. Salim, Shikun Zhang
NeurIPS1
2025 Neo-TKGC: Enhancing Temporal Knowledge Graph Completion with Integrated Node Weights and Future Information
abstract
Temporal Knowledge Graph Completion (TKGC) involves predicting and filling in missing facts within time series data, a crucial task with wide-ranging applications across various domains. The dynamic evolution of Temporal Knowledge Graphs (TKGs) adds complexity to this task, making it inherently challenging. Existing research predominantly relies on historical data to complete the missing facts. However, these approaches often overlook the potential of future information and the significance of node weights.To address these challenges, we propose Neo-TKGC, a novel temporal knowledge graph completion model that integrates a graph structure encoding module and a temporal encoding module. The graph structure encoding module introduces node weights to enhance the capabilities of graph neural networks (GNNs) for entity and relation representation learning, implemented using CompGCN. This module can be easily extended to any GNN models utilizing node and edge aggregation. The temporal encoding module leverages both future and historical information to capture relevant contexts and temporal dependencies among entities and relations.By combining node weights and future information, Neo-TKGC achieves more accurate entity and relation representations, thereby improving the model's ability to infer unknown entities. Extensive experiments on three real-world TKGC datasets demonstrate the superior performance of our model compared to existing approaches, achieving at least a 1.7% relative improvement in Hits@1 across most metrics.
Zihan Qiu, Xiaoling Zhou, Chunyan An, Qiang Yang 0015, Zhixu Li
WSDM2
2025 Mitigating spurious correlations with causal logit perturbation
Xiaoling Zhou, Wei Ye 0004, Rui Xie 0003, Shikun Zhang
Inf. Sci.1
2025 Delving Into the Training Dynamics for Image Classification
abstract
In recent years, there has been an increase in exploring and applying the training dynamics (TD) of deep neural networks (DNNs). Current studies typically rely on quite limited TD quantities and apply their sequences to understand or aid training. This study investigates how to create more effective TD representations, and then apply them to improve the training process of real learning tasks. Specifically, first, an epoch-wise vector comprising 142-dimensional TD quantities, such as loss, is extracted for each sample. Second, a new learning strategy with both self-supervised and supervised learning is designed to learn the deep TD representation of each sample on 200 typical image classification tasks. Third, two novel methods for noisy label detection and imbalance learning, respectively, are presented based on deep TD representations. Our study reveals that neighborhoods and logits are the most important TD quantities, unlike the traditional research that focuses on loss and margin. Moreover, our method based on deep TD representations achieves better performance and demonstrates that high-level TD quantities can facilitate understanding model training, leading to improvements in practical learning tasks, such as noisy label detection and imbalance learning. All the codes are available at https://github.com/limengyang1992/TD_Exploring.
Mengyang Li 0001, Xiaoling Zhou, Ou Wu 0001
IEEE Trans. Image Process.2
2025 Valuing Training Data via Causal Inference for In-Context Learning
abstract
In-context learning (ICL) empowers large pre-trained language models (PLMs) to predict outcomes for unseen inputs without parameter updates. However, the efficacy of ICL heavily relies on the choice of demonstration examples. Randomly selecting from the training set frequently leads to inconsistent performance. Addressing this challenge, this study takes a novel approach by focusing on training data valuation through causal inference. Specifically, we introduce the concept of average marginal effect (AME) to quantify the contribution of individual training samples to ICL performance, encompassing both its generalization and robustness. Drawing inspiration from multiple treatment effects and randomized experiments, we initially sample diverse training subsets to construct prompts and evaluate the ICL performance based on these prompts. Subsequently, we employ Elastic Net regression to collectively estimate the AME values for all training data, considering subset compositions and inference performance. Ultimately, we prioritize samples with the highest values to prompt the inference of the test data. Across various tasks and with seven PLMs ranging in size from 0.8B to 33B, our approach consistently achieves state-of-the-art performance. Particularly, it outperforms Vanilla ICL and the best-performing baseline by an average of 14.1% and 5.2%, respectively. Moreover, prioritizing the most valuable samples for prompting leads to a significant enhancement in performance stability and robustness across various learning scenarios. Impressively, the valuable samples exhibit transferability across diverse PLMs and generalize well to out-of-distribution tasks.
Xiaoling Zhou, Wei Ye 0004, Zhemg Lee, Lei Zou 0001, Shikun Zhang
IEEE Trans. Knowl. Data Eng.1
2024 Enhancing In-Context Learning via Implicit Demonstration Augmentation
abstract
Xiaoling Zhou, Wei Ye, Yidong Wang, Chaoya Jiang, Zhemg Lee, Rui Xie, Shikun Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xiaoling Zhou, Wei Ye 0004, Yidong Wang 0003, Chaoya Jiang, Zhemg Lee, Rui Xie 0003, Shikun Zhang
ACL (1)1
2024 Boosting Model Resilience via Implicit Adversarial Data Augmentation
Xiaoling Zhou, Wei Ye 0004, Zhemg Lee, Rui Xie 0003, Shikun Zhang
IJCAI1
2024 Adversarial Training With Anti-Adversaries
abstract
Adversarial training is effective in improving the robustness of deep neural networks. However, existing studies still exhibit significant drawbacks in terms of the robustness, generalization, and fairness of models. In this study, we validate the importance of different perturbation directions (i.e., adversarial and anti-adversarial) and bounds from both theoretical and practical perspectives. The influence of adversarial training on deep learning models in terms of fairness, robustness, and generalization is theoretically investigated under a more general perturbation scope that different samples can have different perturbation directions and varied perturbation bounds. Our theoretical explorations suggest that combining adversaries and anti-adversaries with varied bounds in training can be more effective in achieving better fairness among classes and a better tradeoff among robustness, accuracy, and fairness in some typical learning scenarios compared with standard adversarial training. Inspired by our theoretical findings, a more general learning objective that combines adversaries and anti-adversaries with varied bounds on each training sample is presented. To solve this objective, two adversarial training frameworks based on meta-learning and reinforcement learning are proposed, in which the perturbation direction and bound for each sample are determined by its training characteristics. Furthermore, the role of the combination strategy with varied bounds is explained from a regularization perspective. Extensive experiments under different learning scenarios verify our theoretical findings and the effectiveness of the proposed methodology.
Xiaoling Zhou, Ou Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Investigating the Sample Weighting Mechanism Using an Interpretable Weighting Framework
abstract
Training deep learning models with unequal sample weights has been shown to enhance model performance in various typical learning scenarios, particularly for imbalanced and noisy-label learning scenarios. A deep understanding of the weighting mechanism facilitates the application of existing weighting strategies and illuminates the design of new weighting strategies for real learning tasks. Scholars have focused on exploring existing weighting methods. However, their studies mainly establish how the weights of samples influence the model training. Little headway is made on the weighting mechanism, i.e., which and how the characteristics of a sample influence its weight. In this study, we adopt a data-driven approach to investigate the weighting mechanism by utilizing an interpretable weighting framework. First, a wide range of sample characteristics is extracted from the classifier network during training. Second, the extracted characteristics are fed into a new neural regression tree (NRT), which is a tree model implemented by a neural network, and its output is the weight of the input sample. Third, the NRT is trained using meta-learning within the whole training process. Once the NRT is learned, the weighting mechanism, including the importance of weighting characteristics, prior modes, and specific weighting rules, can be obtained. We conduct extensive experiments on benchmark noisy and imbalanced data corpora. A package of weighting mechanisms is derived from the learned NRT. Furthermore, our proposed interpretable weighting framework exhibits superior performance in comparison to existing weighting strategies.
Xiaoling Zhou, Ou Wu 0001, Mengyang Li 0001
IEEE Trans. Knowl. Data Eng.1
2024 Which Samples Should Be Learned First: Easy or Hard?
abstract
Treating each training sample unequally is prevalent in many machine-learning tasks. Numerous weighting schemes have been proposed. Some schemes take the easy-first mode, whereas others take the hard-first one. Naturally, an interesting yet realistic question is raised. Given a new learning task, which samples should be learned first, easy or hard? To answer this question, both theoretical analysis and experimental verification are conducted. First, a general objective function is proposed and the optimal weight can be derived from it, which reveals the relationship between the difficulty distribution of the training set and the priority mode. Two novel findings are subsequently obtained: besides the easy-first and hard-first modes, there are two other typical modes, namely, medium-first and two-ends-first; the priority mode may be varied if the difficulty distribution of the training set changes greatly. Second, inspired by the findings, a flexible weighting scheme (FlexW) is proposed for selecting the optimal priority mode when there is no prior knowledge or theoretical clues. The four priority modes can be flexibly switched in the proposed solution, thus suitable for various scenarios. Third, a wide range of experiments is conducted to verify the effectiveness of our proposed FlexW and further compare the weighting schemes in different modes under various learning scenarios. On the basis of these works, reasonable and comprehensive answers are obtained for the easy-or-hard question.
Xiaoling Zhou, Ou Wu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Combining Adversaries with Anti-adversaries in Training
abstract
Adversarial training is an effective learning technique to improve the robustness of deep neural networks. In this study, the influence of adversarial training on deep learning models in terms of fairness, robustness, and generalization is theoretically investigated under more general perturbation scope that different samples can have different perturbation directions (the adversarial and anti-adversarial directions) and varied perturbation bounds. Our theoretical explorations suggest that the combination of adversaries and anti-adversaries (samples with anti-adversarial perturbations) in training can be more effective in achieving better fairness between classes and a better tradeoff between robustness and generalization in some typical learning scenarios (e.g., noisy label learning and imbalance learning) compared with standard adversarial training. On the basis of our theoretical findings, a more general learning objective that combines adversaries and anti-adversaries with varied bounds on each training sample is presented. Meta learning is utilized to optimize the combination weights. Experiments on benchmark datasets under different learning scenarios verify our theoretical findings and the effectiveness of the proposed methodology.
Xiaoling Zhou, Ou Wu 0001
AAAI1
2023 CAFNET: Cross-Attention Fusion Network for Infrared and Low Illumination Visible-Light Image
Xiaoling Zhou, Zetao Jiang, Idowu Paul Okuwobi
Neural Process. Lett.1
2023 Retinex-MPCNN: A Retinex and Modified Pulse coupled Neural Network based method for low-illumination visible and infrared image fusion
Xiaoling Zhou, Zetao Jiang, Idowu Paul Okuwobi
Signal Process. Image Commun.1
2022 Understanding Difficulty-Based Sample Weighting with a Universal Difficulty Measure
Xiaoling Zhou, Ou Wu 0001, Weiyao Zhu, Ziyang Liang
ECML/PKDD (3)1
2022 Increasing naturalness of human-machine dialogue: The users' choices inference of options in machine-raised questions
Xiaoling Zhou, Ou Wu 0001
Knowl. Based Syst.1
2021 An MPC-based Controller Framework for Agile Maneuvering of Autonomous Vehicles
abstract
In a rally competition, professional drivers usually adopt a very aggressive strategy. Agile maneuvering such as ‘drifting’ often occurs during the cornering process. In this paper, a controller framework is present for the vehicle's agile maneuver based on Model Predictive Control (MPC). We introduce the 3-state vehicle model and the brush tire model and analyze the trajectory characteristics of drift equilibrium. The proposed control system can track the variable drift state and the reference path simultaneously and is applicable for both regular driving and drift cornering. Different from the previous studies on drift stability or drift cornering, the system can not only realize lane keeping in complex scenarios but also achieve autonomous drift maneuvering during the cornering process. The effectiveness of the system is validated via simulations on the Matlab-Carsim platform.
Xiaoling Zhou, Lei Xie 0007
IV4
2020 A Recurrent Model for Collective Entity Linking with Adaptive Features
abstract
The vast amount of web data enables us to build knowledge bases with unprecedented quality and coverage. Named Entity Disambiguation (NED) is an important task that automatically resolves ambiguous mentions in free text to correct target entries in the knowledge base. Traditional machine learning based methods for NED were outperformed and made obsolete by the state-of-the-art deep learning based models. However, deep learning models are more complex, requiring large amount of training data and lengthy training and parameter tuning time. In this paper, we revisit traditional machine learning techniques and propose a light-weight, tuneable and time-efficient method without using deep learning or deep learning generated features. We propose novel adaptive features that focus on extracting discriminative features to better model similarities between candidate entities and the mention's context. We learn a local ranking model based on traditional and the new adaptive features based on the learning-to-rank framework. While arriving at linking decisions individually via the local model, our method also takes into consideration the correlation between decisions by running multiple recurrent global models, which can be deemed as a learned local search method. Our method attains performances comparable to the state-of-the-art deep learning-based methods on NED benchmark datasets while being significantly faster to train.
Xiaoling Zhou, Yukai Miao, Wei Wang 0011, Jianbin Qin
AAAI1
2019 Antonym-Synonym Classification Based on New Sub-Space Embeddings
abstract
Distinguishing antonyms from synonyms is a key challenge for many NLP applications focused on the lexical-semantic relation extraction. Existing solutions relying on large-scale corpora yield low performance because of huge contextual overlap of antonym and synonym pairs. We propose a novel approach entirely based on pre-trained embeddings. We hypothesize that the pre-trained embeddings comprehend a blend of lexical-semantic information and we may distill the task-specific information using Distiller, a model proposed in this paper. Later, a classifier is trained based on features constructed from the distilled sub-spaces along with some word level features to distinguish antonyms from synonyms. Experimental results show that the proposed model outperforms existing research on antonym synonym distinction in both speed and performance.
Muhammad Asif Ali, Yifang Sun, Xiaoling Zhou, Wei Wang 0011, Xiang Zhao 0002
AAAI3
2018 A CRF-Based Stacking Model with Meta-features for Named Entity Recognition
Shifeng Liu 0002, Yifang Sun, Wei Wang 0011, Xiaoling Zhou
PAKDD (2)4
2017 Interval-valued quintuple implication principle of fuzzy reasoning
Minxia Luo, Xiaoling Zhou
Int. J. Approx. Reason.2
2016 General Purpose Index-Based Method for Efficient MaxRS Query
Xiaoling Zhou, Wei Wang 0011, Jianliang Xu
DEXA (1)1
2016 BEVA: An Efficient Query Processing Algorithm for Error-Tolerant Autocompletion
abstract
Query autocompletion has become a standard feature in many search applications, especially for search engines. A recent trend is to support theerror-tolerant autocompletion, which increases the usability significantly by matching prefixes of database strings and allowing a small number of errors. In this article, we systematically study the query processing problem for error-tolerant autocompletion with a given edit distance threshold. We propose a general framework that encompasses existing methods and characterizes different classes of algorithms and the minimum amount of information they need to maintain under different constraints. We then propose a novel evaluation strategy that achieves the minimum active node size by eliminating ancestor-descendant relationships among active nodes entirely. In addition, we characterize the essence of edit distance computation by a novel data structure namededit vector automaton(EVA). It enables us to compute new active nodes and their associated states efficiently by table lookups. In order to support large distance thresholds, we devise a partitioning scheme to reduce the size and construction cost of the automaton, which results in theuniversal partitioned EVA(UPEVA) to handle arbitrarily large thresholds. Our extensive evaluation demonstrates that our proposed method outperforms existing approaches in both space and time efficiencies.
Xiaoling Zhou, Jianbin Qin, Chuan Xiao 0001, Wei Wang 0011, Xuemin Lin 0001, Yoshiharu Ishikawa
ACM Trans. Database Syst.1
2015 Robustness of reverse triple I algorithms based on interval-valued fuzzy inference
Minxia Luo, Xiaoling Zhou
Int. J. Approx. Reason.2
2013 Improved Spatial Keyword Search Based on IDF Approximation
Xiaoling Zhou, Yifang Sun, Muhammad Aamir Cheema
APWeb1