VLDB 2026 Research / reviewers in the wild / expert
Xiaofei Xu 0002
dblp:95/3275-2
· DBLP profile ↗
29ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0001-6988-7541ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 2 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Wise recommender: LLMs refined by iterative critics
Zhisheng Yang, Xiaofei Xu 0002, Li Li 0006 |
Inf. Softw. Technol. | 2 |
| 2025 | FineRR-ZNS: Enabling Fine-Granularity Read Refreshing for ZNS SSDsabstractZoned namespace (ZNS) SSDs are emerging storage devices offering low cost, high performance, and software definability. By adopting host-managed zone-based sequential programming, ZNS SSDs effectively eliminate the space overhead associated with on-board DRAM memory and garbage collection. However, while background read refreshing serves as a data protection mechanism in conventional block-interface SSDs, state-of-the-art ZNS SSDs lack read refreshing functionality to guarantee data reliability. Moreover, implementing zonelevel read refreshing in ZNS SSDs incurs significant overhead due to the large volume of valid data movements in a zone, leading to degraded I/O performance. To efficiently enable read refreshing for ZNS SSDs, this paper proposes FineRR-ZNS, a fine-granularity read refreshing mechanism for ZNS SSDs. FineRR-ZNS employs a host-controlled fine-granularity read refreshing scheme that selectively determines block-level read refreshing via metadata remapping. A zone reconstruction method is also designed to retrieve remapped data forming complete data during zone-level RR. Specifically, the remapped data after zone reconstruction are still available and prioritized for read access until their respective blocks need the next RR. Evaluation results show that FineRR-ZNS significantly enhances read refreshing efficiency and I/O throughput compared to zone-level read refreshing implemented in the state-of-the-art ZenFS file system. Jun Li 0062, Zhibing Sha, Fan Yang 0110, Xiaofei Xu 0002, Xiaobai Chen, Jieming Yin, Jianwei Liao 0001 |
DAC | 4 |
| 2025 | CoupledCB: Eliminating Wasted Pages in Copyback-based Garbage Collection for SSDsabstractThe management of garbage collection poses significant challenges in high-density NAND flash-based SSDs. The introduction of the copyback command aims to expedite the migration of valid data. However, its odd/even constraint causes wasted pages during migrations, limiting the efficiency of garbage collection. Additionally, while full-sequence programming en-hances write performance in high-density SSDs, it increases write granularity and exacerbates the issue of wasted pages. To address the problem of wasted pages, we propose a novel method called CoupledCB, which utilizes coupled blocks to fill up the wasted space in copyback-based garbage collection. By taking into account the access characteristics of the candidate coupled blocks and workloads, we develop a coupled block selection model assisted by logistic regression. Experimental results show that our proposal significantly enhances garbage collection efficiency and 1/O performance compared to state-of-the-art schemes. Jun Li 0062, Xiaofei Xu 0002, Zhibing Sha, Xiaobai Chen, Jieming Yin, Jianwei Liao 0001 |
DATE | 2 |
| 2025 | Generating Grounded Responses to Counter Misinformation via Learning Efficient Fine-Grained CritiquesabstractFake news and misinformation poses a significant threat to society, making efficient mitigation essential. However, manual fact-checking is costly and lacks scalability. Large Language Models (LLMs) offer promise in automating counter-response generation to mitigate misinformation, but a critical challenge lies in their tendency to hallucinate non-factual information. Existing models mainly rely on LLM self-feedback to reduce hallucination, but this approach is computationally expensive. In this paper, we propose MisMitiFact, Misinformation Mitigation grounded in Facts, an efficient framework for generating fact-grounded counter-responses at scale. MisMitiFact generates simple critique feedback to refine LLM outputs, ensuring responses are grounded in evidence. We develop lightweight, fine-grained critique models trained on data sourced from readily available fact-checking sites to identify and correct errors in key elements such as numerals, entities, and topics in LLM generations. Experiments show that MisMitiFact generates counter-responses of comparable quality to LLMs' self-feedback while using significantly smaller critique models. Importantly, it achieves ~5x increase in feedback generation throughput, making it highly suitable for cost-effective, large-scale misinformation mitigation. Code and additional results are available at https://github.com/xxfwin/MisMitiFact. Xiaofei Xu 0002, Xiuzhen Zhang 0001 |
IJCAI | 1 |
| 2025 | Leveraging Language Model and Knowledge Tracing for Personalized Question Generation
Zhongwei Yin, Li Li 0006, Xiaofei Xu 0002 |
KSEM (2) | 3 |
| 2024 | Harnessing Network Effect for Fake News Mitigation: Selecting Debunkers via Self-Imitation LearningabstractThis study aims to minimize the influence of fake news on social networks by deploying debunkers to propagate true news. This is framed as a reinforcement learning problem, where, at each stage, one user is selected to propagate true news. A challenging issue is episodic reward where the "net" effect of selecting individual debunkers cannot be discerned from the interleaving information propagation on social networks, and only the collective effect from mitigation efforts can be observed. Existing Self-Imitation Learning (SIL) methods have shown promise in learning from episodic rewards, but are ill-suited to the real-world application of fake news mitigation because of their poor sample efficiency. To learn a more effective debunker selection policy for fake news mitigation, this study proposes NAGASIL - Negative sampling and state Augmented Generative Adversarial Self-Imitation Learning, which consists of two improvements geared towards fake news mitigation: learning from negative samples, and an augmented state representation to capture the "real" environment state by integrating the current observed state with the previous state-action pairs from the same campaign. Experiments on two social networks show that NAGASIL yields superior performance to standard GASIL and state-of-the-art fake news mitigation models. Xiaofei Xu 0002, Michael Dann, Xiuzhen Zhang 0001 |
AAAI | 1 |
| 2022 | Identifying Cost-effective Debunkers for Multi-stage Fake News Mitigation CampaignsabstractOnline social networks have become a fertile ground for spreading fake news. Methods to automatically mitigate fake news propagation have been proposed. Some studies focus on selecting top k influential users on social networks as debunkers, but the social influence of debunkers may not translate to wide mitigation information propagation as expected. Other studies assume a given set of debunkers and focus on optimizing intensity for debunkers to publish true news, but as debunkers are fixed, even if with high social influence and/or high intensity to post true news, the true news may not reach users exposed to fake news and therefore mitigation effect may be limited. In this paper, we propose the multi-stage fake news mitigation campaign where debunkers are dynamically selected within budget at each stage. We formulate it as a reinforcement learning problem and propose a greedy algorithm optimized by predicting future states so that the debunkers can be selected in a way that maximizes the overall mitigation effect. We conducted extensive experiments on synthetic and real-world social networks and show that our solution outperforms state-of-the-art baselines in terms of mitigation effect. Xiaofei Xu 0002, Xiuzhen Zhang 0001 |
WSDM | 1 |
| 2022 | Veracity-aware and Event-driven Personalized News Recommendation for Fake News MitigationabstractDespite the tremendous efforts by social media platforms and fact-check services for fake news detection, fake news and misinformation still spread wildly on social media platforms (e.g., Twitter). Consequently, fake news mitigation strategies are urgently needed. Most of the existing work on fake news mitigation focuses on the overall mitigation on a whole social network while ignoring developing concrete mitigation strategies to deter individual users from sharing fake news. In this paper, we propose a novel veracity-aware and event-driven recommendation model to recommend personalised corrective true news to individual users for effectively debunking fake news. Our proposed model Rec4Mit (Recommendation for Mitigation) not only effectively captures a user’s current reading preference with a focus on which event, e.g., US election, from her/his recent reading history containing true and/or fake news, but also accurately predicts the veracity (true or fake) of candidate news. As a result, Rec4Mit can recommend the most suitable true news to best match the user’s preference as well as to mitigate fake news. In particular, for those users who have read fake news of a certain event, Rec4Mit is able to recommend the corresponding true news of the same event. Extensive experiments on real-world datasets show Rec4Mit significantly outperforms the state-of-the-art news recommendation methods in terms of the capability to recommend personalized true news for fake news mitigation. Shoujin Wang, Xiaofei Xu 0002, Xiuzhen Zhang 0001, Yan Wang 0002, Wenzhuo Song |
WWW | 2 |
| 2022 | Pattern-Based Prefetching with Adaptive Cache Management Inside of Solid-State DrivesabstractThis article proposes a pattern-based prefetching scheme with the support of adaptive cache management, at the flash translation layer of solid-state drives ( SSDs ). It works inside of SSDs and has features of OS dependence and uses transparency. Specifically, it first mines frequent block access patterns that reflect the correlation among the occurred I/O requests. Then, it compares the requests in the current time window with the identified patterns to direct prefetching data into the cache of SSDs. More importantly, to maximize the cache use efficiency, we build a mathematical model to adaptively determine the cache partition on the basis of I/O workload characteristics, for separately buffering the prefetched data and the written data. Experimental results show that our proposal can yield improvements on average read latency by 1.8 %– 36.5 % without noticeably increasing the write latency, in contrast to conventional SSD-inside prefetching schemes. Jun Li 0062, Xiaofei Xu 0002, Zhigang Cai, Jianwei Liao 0001, Kenli Li 0001, Balazs Gerofi, Yutaka Ishikawa |
ACM Trans. Storage | 2 |
| 2021 | Topic Enhanced Multi-head Co-Attention: Generating Distractors for Reading ComprehensionabstractIn the construction of multiple-choice Machine Reading Comprehension(MRC), in addition to questions and answers, it is necessary to generate distractors corresponding to them. Recent models based on Seq2Seq have shown good results in text generation, while previous work has only succeeded in producing a few interfering words or phrases per question. Our goal is to generate more meaningful distractors based on reading comprehension of articles that are closer to the semantics of the question to help better diagnose gaps in text understanding. A major drawback of recent studies is that they do not take the relationship between the distractors and background text into account when generating distractors. This often results in the distractors being either too general or too close to the correct answer. We propose a Topic Enhanced Multi-head Co-Attention model (TMCA) based on hierarchical networks to better capture the interactions between sentences. By adding query-relevance loss, our model enables the distractors to be as semantically relevant to the question as possible based on the reading comprehension of the article, while trying to ensure that they are false answers. The results show that the proposed approach achieves superior performance over the baselines in terms of automatic metrics on two multiple-choice datasets (RACE and DREAM). For further evaluation, we use another advanced reading comprehension model and human evaluation to demonstrate that our model outperforms several strong baselines in generating high-quality and educationally meaningful distracters. Pengju Shuai, Zixi Wei, Sishun Liu, Xiaofei Xu 0002, Li Li 0006 |
IJCNN | 4 |
| 2021 | Transformer-based Relation Detect Model for Aspect-based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) aims to detect the sentiment polarities of a sentence with a given aspect. Aspect-term sentiment analysis (ATSA) is a subtask of ABSA, in ATSA, the aspect is given by aspect term: a word or a phrase in sentence. In previous work, most models apply attention mechanism or gating mechanism to capture the key part of the sentence and detect the sentiment polarity by classifying the weighted sum vector, which are regardless of the aspect information during the classification. However, same contexts may show different sentiment polarity with differnet aspects. For example, in sentence “The scene hunky waiters dub diner darling and it sounds like they mean it.” the context “dub diner dearling” shows negative polarity towards the aspect term “waiters”. But in sentence “The diner's husband dub diner darling.”, the same context would show neural polarity towards “husband”. The absent of aspect information would make model get a wrong result in some sentence. To solve this problem, we propose a model called Trasfomer-based context-aspect Relation Detect Model (TRDM), and add two special tokens “[CLSA]” and “[SEPA]” to specify the aspect term in the sentence. TRDM uses the embedding of two tyes of tokens (i.e. “[CLS]” and “[CLSA]”) for classification where “[CLS]” is used to represent the entire sentence in BERT. This scheme enable TRDM to combine the sentence information and aspect information in the step of sentiment polarity classification. We evaluate the performance of our model on two datasets: Restaurant dataset from SemEval2014 and MAMS from NLPCC 2020. Experiment results show that our model obtain noticeable improvement compared with state-of-art transfromer-based models. Zixi Wei, Xiaofei Xu 0002, Lijian Li 0003, Kaixin Qin, Li Li 0006 |
IJCNN | 2 |
| 2020 | Frequent Access Pattern-based Prefetching Inside of Solid-State DrivesabstractThis paper proposes an SSD-inside data prefetching scheme, which has features of OS-dependence and use transparency. To be specific, it first mines frequent block access patterns that reflect the correlation among the occurred requests. Then it compares the requests in the current time window with the identified patterns, to direct fetching data in advance. Furthermore, to maximize the cache use efficiency, we construct a method to adaptively determine the cache partition on the basis of I/O workload characteristics, for separately buffering the prefetched data and the write data. Experimental results demonstrate that our proposal can yield improvements on average read latency by 6.3% to 9.3% without noticeably increasing write latency, in contrast to conventional SSD-inside prefetching schemes. Xiaofei Xu 0002, Zhigang Cai, Jianwei Liao 0001, Yutaka Ishikawa |
DATE | 1 |
| 2019 | Exploring Semantic Change of Chinese Word Using Crawled Web Data
Xiaofei Xu 0002, Yukun Cao, Li Li 0006 |
ICWE | 1 |
| 2019 | Pattern-based Write Scheduling and Read Balance-oriented Wear-Leveling for Solid State DriversabstractThis paper proposes a pattern-based I/O scheduling mechanism, which identifies frequently written data with patterns and dispatches them to the same SSD blocks having a small erase count. The data on the same block are mostly like to be invalided together, so that the overhead of garbage collection can be greatly reduced. Moreover, a read balance-oriented wear-leveling scheme is introduced to extend the lifetime of SSDs. Specifically, it distributes the hot read data in the blocks with a small erase count, to heavily erased blocks in different chips of the same SSD channel, while carrying out wear-leveling. As a result, internal parallelism at the chip level of SSD can be fully exploited for achieving better read data throughput. We conduct a series of simulation tests with a number of disk traces of real-world applications under the SSDsim platform. The experimental results show that the newly proposed mechanism can reduce garbage collection overhead by 11.3%, and the read response time by 12.8% in average, comparing to existing approaches of scheduling and wear-leveling for SSDs. Jun Li 0062, Xiaofei Xu 0002, Xiaoning Peng, Jianwei Liao 0001 |
MSST | 2 |
| 2019 | Opinion extraction by distinguishing term dependencies and digging deep text features
Fei Hu 0004, Li Li 0006, Xiaofei Xu 0002, Jinjing Zhang |
Neural Comput. Appl. | 3 |
| 2019 | An adaptive mechanism to achieve learning rate dynamically
Jinjing Zhang, Fei Hu 0004, Li Li 0006, Xiaofei Xu 0002, Zhanbo Yang, Yanbin Chen |
Neural Comput. Appl. | 4 |
| 2018 | A Deep Prediction Model of Traffic Flow Considering Precipitation ImpactabstractTraffic flow prediction is an important part of intelligent transportation systems (ITS). However, the performance of current traffic flow prediction methods does not meet the expectation. Weather factors such as precipitation in residential areas and tourist destinations affect traffic flow on the surrounding roads. In this paper, we attempt to take precipitation impact into consideration when predicting traffic flow. To realize this idea, we propose a deep traffic flow prediction architecture by introducing a deep bi-directional long short-term memory model, precipitation information, residual connection, regression layer and dropout training method. The proposed model has good ability to capture the deep features of traffic flow. Besides, it can take full advantage of time-aware traffic flow data and additional precipitation data. We evaluate the prediction architecture on the dataset from Caltrans Performance Measurement System (PeMS) with precipitation data from California Data Exchange Center (CDEC) and the dataset from KDD Cup 2017. The experiment results demonstrate that the proposed model for traffic flow prediction obtains high accuracy and generalizes well compared with other models. Fei Hu 0004, Xiaofei Xu 0002, Dengbao Wang, Li Li 0006 |
IJCNN | 3 |
| 2018 | P-DBL: A Deep Traffic Flow Prediction Architecture Based on Trajectory Data
Xiaofei Xu 0002, Jun He 0012, Li Li 0006 |
KSEM (2) | 2 |
| 2018 | A fast pseudo-stochastic sequential cipher generator based on RBMs
Fei Hu 0004, Xiaofei Xu 0002, Changjiu Pu, Li Li 0006 |
Neural Comput. Appl. | 2 |
| 2018 | Content tracking by leveraging hashtag and time information in Twitter social mediaabstractThe content in social media is difficult to analyze because of its informal and unstructured features. Luckily, some social media data like tweets have rich hashtags information, which can be helpful to identify meaningful content and topic information. More importantly, the hashtag usually express the context information of a tweet best. To this end, this paper introduces a context-aware topic model to detect and track the evolution of content by integrating hashtag and time information in text-based social media. Specifically, we develop two methods to cope with different functions of hashtags separately. The first one is named hashtag-generated Topic over Time (hgToT), in which a document is generated jointly by the existing words and hashtags. To enhance the significant effect of hashtags via topic variables, we further develop the second model named hashtag-supervised Topic over Time (hsToT), in which hashtags are treated as useful topic indicators of the tweet. Time information is modeled similarly in both hgToT and hsToT. The proposed two methods are able to capture the hashtags distribution over topics and topic changing over time simultaneously. Experiments on the dataset obtained from Twitter show that both hgToT and hsToT could detect the important information and track the meaningful content and topics successfully. Xiaofei Xu 0002, Li Li 0006, Jinjing Zhang, Shuo He 0001 |
Web Intell. | 1 |
| 2018 | Stochastic gradient descent with variance reduction techniqueabstractGradient descent is prevalent for large scale optimization problems in machine learning, especially its major role is computing and correcting the connection strength of neural network in deep learning. However, choosing a proper learning rate for SGD can be difficult. A too small rate may lead to painfully slow convergence, while too large one would hinder convergence. In this paper, we present a novel variance reduction technique which applies the moving average of gradient termed SMVRG. SMVRG can take a large learning rate by using variance reduction technique. And, we only need to preserve current gradient and the previous average gradient. Our method is employed to Long Short-Term Memory (LSTM). The experiment on two data sets, the IMDB (movie reviews) and SemEval-2016 (sentiment analysis in twitter) shows our method can improve the results significantly. Jinjing Zhang, Fei Hu 0004, Xiaofei Xu 0002, Li Li 0006 |
Web Intell. | 3 |
| 2017 | Memory-Enhanced Latent Semantic Model: Short Text Understanding for Sentiment Analysis
Fei Hu 0004, Xiaofei Xu 0002, Zhanbo Yang, Li Li 0006 |
DASFAA (1) | 2 |
| 2017 | A Piecewise Hybrid of ARIMA and SVMs for Short-Term Traffic Flow Prediction
Li Li 0006, Xiaofei Xu 0002 |
ICONIP (5) | 3 |
| 2017 | AEDR: An Adaptive Mechanism to Achieve Online Learning Rate DynamicallyabstractThe distribution of connection strength between neural network units contains all the information of neural network. However, learning algorithms, as the key to the weight correction of neural network, retain more sensitive hyper-parameters which require endless ways of configuring. In this paper, we present a novel adaptive mechanism called Adaptive Exponential Decay Rate(AEDR). AEDR allows to adaptively calculate exponential decay rate using moving average of gradients and squared gradients over time, otherwise it need tune exponential decay rate. The mechanism reduces the ways of configuring hyper-parameters. Moreover, it enables varying learning rates for different parameters. The mechanism is then applied to Adadelta by Long Short-Term Memory(LSTM) to demonstrate how learning rate adapt dynamically under AEDR. Our mechanism improves the results significantly and achieve state-of-the-art performance on two real data sets, the IMDB(movie reviews) and SemEval-2016(sentiment analysis in twitter). Jinjing Zhang, Fei Hu 0004, Li Li 0006, Xiaofei Xu 0002, Zhanbo Yang |
ICTAI | 4 |
| 2017 | ProductRec: Product Bundle Recommendation Based on User's Sequential Patterns in Social Networking Service EnvironmentabstractWith the overload of information on the Web, Recommender Systems (RSs) are becoming increasingly popular and have been employed to provide suggestions to meet different requirements. RSs are utilized in a variety of areas including movies, music, social tags, user group and products as Web services evoked on the Internet either as mobile Apps or PC-based applications. However, it is challenging to achieve personalized recommendations instead of offering up too many lowest common denominator recommendations. Understanding how products relate to each other is important because it has great impact on the performance. Furthermore, the personalized sequential behavior, which is closely related to a particular product, is essential for recommender systems. Most models simply integrate features from users and items without considering potential product bundle relationships between products exposed by users' personalized sequential behaviors. In this paper, a novel method based on Factorizing Personalized Markov Chain (FPMC) is proposed to comprehensively explore the latent bundle relations from users perspective, along with the hidden correlative semantics between products obtained from logic regression method, which provides a unified view to describe the user preferences, product/item features, and the user sequential patterns in timely manner. The involved semantic features are extracted using deep learning models. We evaluate our method on real-world Amazon datasets and our framework significantly outperforms other baseline models, especially on sparse datasets. The experimental results show that our approach qualitatively captures personalized behaviors with superior recommendation performance. Wenli Yu 0002, Li Li 0006, Xiaofei Xu 0002, Dengbao Wang, Shiping Chen 0001 |
ICWS | 3 |
| 2017 | Weakly Supervised Feature Compression Based Topic Model for Sentiment Classification
Xiaofei Xu 0002, Li Li 0006 |
KSEM | 2 |
| 2017 | Emphasizing Essential Words for Sentiment Classification Based on Recurrent Neural Networks
Fei Hu 0004, Li Li 0006, Zili Zhang 0001, Xiaofei Xu 0002 |
J. Comput. Sci. Technol. | 5 |
| 2016 | Analyzing Topic-Sentiment and Topic Evolution over Time from Social Media
Xiaofei Xu 0002, Li Li 0006 |
KSEM | 2 |
| 2016 | A hybrid method for bilingual text sentiment classification based on deep learningabstractText sentiment classification has occupied a pivotal position in sentiment analysis research, it offers important opinion mining functions. Nowadays, with explosion of information, many researchers are focusing on sentiment classification research on massive amounts of data. However, the traditional machine learning methods cannot acquire text semantic information and most research achievements are about single language, in this paper, a hybrid method which integrates the deep learning features and shallow learning features is proposed. The hybrid method can not only realize single language text sentiment classification but realize bilingual text sentiment classification as well. Models such as recurrent neural networks (RNNs) with long short term memory(LSTM), Naïve Bayes Support Vector Machine (NB-SVM), word vectors and bag-of-words are explored. Firstly, these models are studied separately in sentiment classification task. The paper then integrates the above methods as a whole to complete the task. Different combination strategies are discussed regarding the contribution of each method. The experiments show that the accuracy can reach 89% and the hybrid method performs much better than any other method individually. The proposed method achieves a performance close to the state-of-the-art methods based on the had-engineered features. What's more, the hybrid model can learn more linguistic phenomena with the growth of the accuracy of emotional tendency discrimination when more background knowledge is available. Guolong Liu, Xiaofei Xu 0002, Bailong Deng, Siding Chen, Li Li 0006 |
SNPD | 2 |