VLDB 2026 Research / reviewers in the wild / expert
Xiaofeng Zhu 0004
dblp:60/4671-4
· DBLP profile ↗
7ranked-venue papers
4as first author
3since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Integrating Heterogeneous Sources for Learned Prediction of Vehicular Data ConsumptionabstractIn addition to the multiple sensors to measure parameters that can be used to improve both safety and efficiency, modern vehicles also gather information about external data (e.g., traffic conditions, weather) which, if properly used, could further improve the overall trip experience. Specifically, when it comes to navigation, one source that can provide increased context awareness, especially for autonomous driving, are the High Definition (HD) maps, which have recently witnessed a tremendous growth of popularity in vehicular technology and use. As they are limited to a particular geographic area, different portions need to be downloaded (and processed) on multiple occasions throughout a given trip, along with the other data from other internal and external sources. In this paper, we provide an effective deep learning approach for the recently introduced problem of Predicting Map Data Consumption (PMDC) in the future time instants for a given trip. We propose a novel methodology that integrates multiple data sources (road network, traffic, historic trips, HD maps) and, for a given trip, enables prediction of the map data consumption. Our experimental observations demonstrate the benefits of the proposed approach over the candidate baselines. Andi Zang, Xiaofeng Zhu 0004, Ce Li 0003, Fan Zhou 0002, Goce Trajcevski |
MDM | 2 |
| 2021 | Continual Neural Network Model RetrainingabstractWe propose incremental (re)training of a neural network model to cope with a continuous flow of new data. As such, this is a life-long learning process. We address two challenges of life-long retraining: catastrophic forgetting and efficient retraining. If we combine all past and new data it can easily become intractable to retrain the neural network model. On the other hand, if the model is retrained using only new data, it can easily suffer catastrophic forgetting and thus it is paramount to strike the right balance. Moreover, if we retrain all weights of the model every time new data is collected, retraining tends to require too many computing resources. To solve these two issues, we propose a novel retraining model that can select important samples and important weights utilizing multi-armed bandits. To further address forgetting, we propose a new regularization term focusing on synapse and neuron importance. We analyze multiple datasets to document the outcome of the proposed retraining methods. Various experiments demonstrate that our retraining methodologies mitigate the catastrophic forgetting problem while boosting model performance. Xiaofeng Zhu 0004, Diego Klabjan |
IEEE BigData | 1 |
| 2021 | Towards Predicting Vehicular Data ConsumptionabstractCombining in-car multiple sensors measuring parameters that can be used to improve both safety and efficiency with a plethora of external data sources (e.g., traffic conditions, weather) which, if properly used, can significantly improve the overall trip experience. One source that can help the navigation and provide "context awareness", especially for autonomous driving, are the High Definition (HD) maps, which have recently witnessed a tremendous growth of popularity in vehicular technology and use. As they are limited to a particular geographic area with respect to a given point along a trip, different portions need to be downloaded (and processed) on multiple occasions throughout a given trip, along with the other data from internal and external sources. We take a first step towards formalizing the problem of Predicting Map Data Consumption (PMDC) in the future time instants for a given trip, based on a (time) window from its history, and investigate the use of Long Short-Term Memory (LSTM) networks - a special type of Recurrent Neural Networks (RNN). Significant efforts were focused on generating an appropriate dataset for this study, towards which we fused the information available in multiple heterogeneous data sources. We conducted experimental observations demonstrating the benefits of the proposed approach. Andi Zang, Xiaofeng Zhu 0004, Yuxiang Guo 0001, Fan Zhou 0002, Goce Trajcevski |
MDM | 2 |
| 2020 | Listwise Learning to Rank by Exploring Unique RatingsabstractIn this paper, we propose new listwise learning-to-rank models that mitigate the shortcomings of existing ones. Existing listwise learning-to-rank models are generally derived from the classical Plackett-Luce model, which has three major limitations. (1) Its permutation probabilities overlook ties, i.e., a situation when more than one document has the same rating with respect to a query. This can lead to imprecise permutation probabilities and inefficient training because of selecting documents one by one. (2) It does not favor documents having high relevance. (3) It has a loose assumption that sampling documents at different steps is independent. To overcome the first two limitations, we model ranking as selecting documents from a candidate set based on unique rating levels in decreasing order. The number of steps in training is determined by the number of unique rating levels. More specifically, in each step, we apply multiple multi-class classification tasks to a document candidate set and choose all documents that have the highest rating from the document set. This is in contrast to taking one document step by step in the classical Plackett-Luce model. Afterward, we remove all of the selected documents from the document set and repeat until the remaining documents all have the lowest rating. We propose a new loss function and associated four models for the entire sequence of weighted classification tasks by assigning high weights to the selected documents with high ratings for optimizing Normalized Discounted Cumulative Gain (NDCG). To overcome the final limitation, we further propose a novel and efficient way of refining prediction scores by combining an adapted Vanilla Recurrent Neural Network (RNN) model with pooling given selected documents at previous steps. We encode all of the documents already selected by an RNN model. In a single step, we rank all of the documents with the same ratings using the last cell of the RNN multiple times. We have implemented our models using three settings: neural networks, neural networks with gradient boosting, and regression trees with gradient boosting. We have conducted experiments on four public datasets. The experiments demonstrate that the models notably outperform state-of-the-art learning-to-rank models. Xiaofeng Zhu 0004, Diego Klabjan |
WSDM | 1 |
| 2019 | Suggestion Mining from Online Reviews usingRandom Multimodel Deep LearningabstractIn this paper, we propose a new deep learning based method for suggestion mining. The major challenges of suggesion mining include cross domain issue and the issues caused by unstructured and highly imbalanced data structure. To overcome these challenges, we propose to apply Random Multimodel Deep Learning (RMDL) which combines three different deep learning architectures (DNNs, RNNs and CNNs) and automatically selects the optimal hyper parameter to improve the robustness and flexibility of the model. Our experimental results on the SemEval-2019 competition Task 9 data sets demonstrate that our proposed RMDL outperforms most of the existing suggestion mining methods. Liangji Wang, Xiaofeng Zhu 0004, Dingding Wang 0001 |
ICMLA | 3 |
| 2019 | Frosting Weights for Better Continual TrainingabstractTraining a neural network model can be a lifelong learning process and is a computationally intensive one. A severe adverse effect that may occur in deep neural network models is that they can suffer from catastrophic forgetting during retraining on new data. To avoid such disruptions in the continuous learning, one appealing property is the additive nature of ensemble models. In this paper, we propose two generic ensemble approaches, gradient boosting and meta-learning, to solve the catastrophic forgetting problem in tuning pre-trained neural network models. Xiaofeng Zhu 0004, Goce Trajcevski, Dingding Wang 0001 |
ICMLA | 1 |
| 2017 | Semantic Document Distance Measures and Unsupervised Document Revision DetectionabstractIn this paper, we model the document revision detection problem as a minimum cost branching problem that relies on computing document distances. Furthermore, we propose two new document distance measures, word vector-based Dynamic Time Warping (wDTW) and word vector-based Tree Edit Distance (wTED). Our revision detection system is designed for a large scale corpus and implemented in Apache Spark. We demonstrate that our system can more precisely detect revisions than state-of-the-art methods by utilizing the Wikipedia revision dumps and simulated data sets. Xiaofeng Zhu 0004, Diego Klabjan, Patrick N. Bless |
IJCNLP(1) | 1 |