VLDB 2026 Research / reviewers in the wild / expert
Xiufeng Liu 0001
dblp:43/8119-1
· DBLP profile ↗
27ranked-venue papers in the field
18as first author
10since 2021 · last 2026
0000-0001-5133-6688ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 18 (14 first)Data Mining & Knowledge Discovery · 6 (4 first)Information Retrieval & Web Search · 2Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WassDRO-AD: Distributionally Robust Anomaly Detection for Multivariate IoT Data Streams
Xiufeng Liu 0001, Ruyu Liu, Yanyan Yang 0002 |
DEXA (2) | 1 |
| 2026 | TwinDB: Interactive What-If Analysis for Digital Twins
Xiufeng Liu 0001, Ruyu Liu, Per Sieverts Nielsen, Hua Lu 0001 |
EDBT | 1 |
| 2026 | NavLLM: Interactive LLM-Assisted Navigation over Multidimensional Data Cubes
Xiufeng Liu 0001, Ruyu Liu, Yanyan Yang 0002 |
EDBT | 1 |
| 2025 | Causal Discovery with Inverted Self-attention for Multivariate Time Series
Yusen Liu 0001, Tianqing Zhu, Xiufeng Liu 0001, Huan Huo |
PAKDD (4) | 5 |
| 2025 | FRAME: Feature Rectification for Class Imbalance LearningabstractClass imbalance learning is a challenging task in machine learning applications. To balance training data, traditional class imbalance learning approaches, such as class resampling or reweighting, are commonly applied in the literature. However, these methods can have significant limitations, particularly in the presence of noisy data, missing values, or when applied to advanced learning paradigms like semi-supervised or federated learning. To address these limitations, this paper proposes a novel and theoretically-ensured latentFeatureRectification method for clAss iMbalance lEarning (FRAME). The proposed FRAME can automatically learn multiple centroids for each class in the latent space and then perform class balancing. Unlike data-level methods, FRAME balances feature in the latent space rather than the original space. Compared to algorithm-level methods, FRAME can distinguish different classes based on distance without the need to adjust the learning algorithms. Through latent feature rectification, FRAME can effectively mitigate contaminated noises/missing values without worrying about structural variations in the data. In order to accommodate a wider range of applications, this paper extends FRAME to the following three main learning paradigms: fully-supervised learning, semi-supervised learning, and federated learning. Extensive experiments on 10 binary-class datasets demonstrate that our FRAME can achieve competitive performance than the state-of-the-art methods and its robustness to noises/missing values. Xu Cheng 0003, Fan Shi 0001, Yao Zhang 0021, Huan Li 0003, Xiufeng Liu 0001, Shengyong Chen |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Online Machine Learning for Adaptive Ballast Water ManagementabstractThe paper proposes an innovative solution that employs online machine learning to continuously train and update models using sensor data from ships and ports. The proposed solution enhances the efficiency of ballast water management systems (BWMS), which are automated systems that utilize ultraviolet light and filters to purify and disinfect the ballast water that ships carry for maintaining their stability and balance. The solution allows it to grasp the complex and evolving patterns of ballast water quality and flow rate, as well as the diverse conditions of ships and ports. The solution also offers probabilistic forecasts that consider the uncertainty of future events that could impact the performance of ballast water management systems. An online machine learning architecture is proposed that can accommodate probabilistic based machine learning models and algorithms designed for specific training objectives and strategies. Three training methodologies are introduced: continuous training, scheduled training and threshold-triggered training. The effectiveness and reliability of the solution are demonstrated using actual data from ship and port performances. The results are visualized using time-based line charts and maps. Nadeem Iftikhar, Yi-Chen Lin 0001, Xiufeng Liu 0001, Finn Ebertsen Nordbjerg |
DATA | 3 |
| 2023 | Understanding crowd energy consumption behaviors
Xiufeng Liu 0001, Xu Cheng 0003, Yanyan Yang 0002, Huan Huo, Yongping Liu, Per Sieverts Nielsen |
EDBT | 1 |
| 2023 | Multi-level Correlation Matching for Legal Text Similarity Modeling with Multiple Examples
Xike Xie, Xiufeng Liu 0001 |
WISE | 3 |
| 2023 | A late-mover genetic algorithm for resource-constrained project-scheduling problemsabstractThe Resource-Constrained Project Scheduling Problem (RCPSP) plays a critical role in various management applications. Despite its importance, research efforts are still ongoing to improve lower bounds and reduce deviation values. This study aims to develop an innovative and straightforward algorithm for RCPSPs by integrating the “1+1” evolution strategy into a genetic algorithm framework. Unlike most existing studies, the proposed algorithm eliminates the need for parameter tuning and utilizes real-valued numbers and path representation as chromosomes. Consequently, it does not require priority rules to construct a feasible schedule. The algorithm's performance is evaluated using the RCPSP benchmark and compared to alternative algorithms, such as cWSA, Hybrid PSO, and EESHHO. The experimental results demonstrate that the proposed algorithm is competitive, while the exploration capability remains a challenge for further investigation. Yongping Liu, Xiufeng Liu 0001, Guomin Ji, Xu Cheng 0003, Erling Onstein |
Inf. Sci. | 3 |
| 2021 | A Workload-Aware Change Data Capture Framework for Data Warehousing
Weiping Qu, Xiufeng Liu 0001, Stefan Deßloch |
DaWaK | 2 |
| 2020 | VAP: A Visual Analysis Tool for Energy Consumption Spatio-temporal Pattern DiscoveryabstractIn the context of urbanization and the rapid growth of energy demand, understanding the spatial and temporal dynamics of urban energy use is crucial for identifying energy-saving potentials. In this demo, we present a visual analysis tool, VAP, that allows users to explore the dynamics of urban energy use at different spatial and temporal scales. In contrast to traditional statistical and machine learning methods, the visual analysis based tool focuses on analytical thinking, user interactions and answering business questions by examining different visual analysis views. In the demonstration, conference attendees will interact with VAP and learn its capabilities in discovering typical consumption patterns and spatio-temporal shift patterns from a real-world case study of electricity. Xiufeng Liu 0001, Zhibin Niu, Yanyan Yang 0002, Junqi Wu 0003, Dawei Cheng, Xin Wang 0064 |
EDBT | 1 |
| 2018 | Recognizing Textual Entailment with Attentive Reading and Writing Operations
Liang Liu 0015, Huan Huo, Xiufeng Liu 0001, Vasile Palade, Dunlu Peng, Qingkui Chen |
DASFAA (1) | 3 |
| 2018 | Analysis and Visualization of Urban Emission Measurements in Smart CitiesabstractCities worldwide aim to reduce their greenhouse gas emissions and improve air quality for their citizens. Therefore, there is a need to implement smart city approaches to monitor, model, and understand local emissions to better guide these actions. We present our approach that deploys a number of low-cost sensors through a wireless Internet of Things (IoT) backbone and is thus capable of collecting high-granular data. Based on a flexible architecture, we built an ecosystem of data management and data analytics including processing, integration, analysis, and visualization as well as decision-support systems for cities to better understand their emissions. Our prototype system has so far been tested in two Scandinavian cities. We present this system and demonstrate how to collect, integrate, analyze, and visualize real-time air quality data. Dirk Ahlers, Frank Alexander Kraemer, Anders Eivind Braten, Xiufeng Liu 0001, Fredrik Anthonisen, Patrick Driscoll, John Krogstie |
EDBT | 4 |
| 2018 | Scalable prediction-based online anomaly detection for smart meter data
Xiufeng Liu 0001, Per Sieverts Nielsen |
Inf. Syst. | 1 |
| 2017 | Air Quality Monitoring System and Benchmarking
Xiufeng Liu 0001, Per Sieverts Nielsen |
DaWaK | 1 |
| 2017 | A survey on scholarly data: From big data perspective
Samiya Khan, Xiufeng Liu 0001, Kashish Ara Shakil, Mansaf Alam |
Inf. Process. Manag. | 2 |
| 2017 | CITIESData: a smart city data management framework
Xiufeng Liu 0001, Alfred Heller, Per Sieverts Nielsen |
Knowl. Inf. Syst. | 1 |
| 2017 | Smart Meter Data Analytics: Systems, Algorithms, and BenchmarkingabstractSmart electricity meters have been replacing conventional meters worldwide, enabling automated collection of fine-grained (e.g., every 15 minutes or hourly) consumption data. A variety of smart meter analytics algorithms and applications have been proposed, mainly in the smart grid literature. However, the focus has been on what can be done with the data rather than how to do it efficiently. In this article, we examine smart meter analytics from a software performance perspective. First, we design a performance benchmark that includes common smart meter analytics tasks. These include offline feature extraction and model building as well as a framework for online anomaly detection that we propose. Second, since obtaining real smart meter data is difficult due to privacy issues, we present an algorithm for generating large realistic datasets from a small seed of real data. Third, we implement the proposed benchmark using five representative platforms: a traditional numeric computing platform (Matlab), a relational DBMS with a built-in machine learning toolkit (PostgreSQL/MADlib), a main-memory column store (“System C”), and two distributed data processing platforms (Hive and Spark/Spark Streaming). We compare the five platforms in terms of application development effort and performance on a multicore machine as well as a cluster of 16 commodity servers. Xiufeng Liu 0001, Lukasz Golab, Wojciech M. Golab, Ihab F. Ilyas, Shichao Jin |
ACM Trans. Database Syst. | 1 |
| 2016 | Online Anomaly Energy Consumption Detection Using Lambda Architecture
Xiufeng Liu 0001, Nadeem Iftikhar, Per Sieverts Nielsen, Alfred Heller |
DaWaK | 1 |
| 2015 | Benchmarking Smart Meter Data AnalyticsabstractSmart electricity meters have been replacing conventional meters worldwide, enabling automated collection of fine-grained (every 15 minutes or hourly) consumption data. A variety of smart meter analytics algorithms and applications have been proposed, mainly in the smart grid literature, but the focus thus far has been on what can be done with the data rather than how to do it efficiently. In this paper, we examine smart meter analytics from a software per-formance perspective. First, we propose a performance benchmark that includes common data analysis tasks on smart meter data. Sec-ond, since obtaining large amounts of smart meter data is diffi-cult due to privacy issues, we present an algorithm for generat-ing large realistic data sets from a small seed of real data. Third, we implement the proposed benchmark using five representative platforms: a traditional numeric computing platform (Matlab), a relational DBMS with a built-in machine learning toolkit (Post-greSQL/MADLib), a main-memory column store (“System C”), and two distributed data processing platforms (Hive and Spark). We compare the five platforms in terms of application development effort and performance on a multi-core machine as well as a cluster of 16 commodity servers. We have made the proposed benchmark and data generator freely available online. 1. Xiufeng Liu 0001, Lukasz Golab, Wojciech M. Golab, Ihab F. Ilyas |
EDBT | 1 |
| 2015 | SMAS: A smart meter data analytics systemabstractSmart electricity meters are replacing conventional meters worldwide and have enabled a new application domain: smart meter data analytics. In this paper, we introduce SMAS, our smart meter analytics system, which demonstrates the actionable insight that consumers and utilities can obtain from smart meter data. Notably, we implemented SMAS inside a relational database management system using open source tools: PostgreSQL and the MADLib machine learning toolkit. In the proposed demonstration, conference attendees will interact with SMAS as electricity providers, consultants and consumers, and will perform various analyses on real data sets. Xiufeng Liu 0001, Lukasz Golab, Ihab F. Ilyas |
ICDE | 1 |
| 2014 | Survey of real-time processing systems for big dataabstractIn recent years, real-time processing and analytics systems for big data--in the context of Business Intelligence (BI)--have received a growing attention. The traditional BI platforms that perform regular updates on daily, weekly or monthly basis are no longer adequate to satisfy the fast-changing business environments. However, due to the nature of big data, it has become a challenge to achieve the real-time capability using the traditional technologies. The recent distributed computing technology, MapReduce, provides off-the-shelf high scalability that can significantly shorten the processing time for big data; Its open-source implementation such as Hadoop has become the de-facto standard for processing big data, however, Hadoop has the limitation of supporting real-time updates. The improvements in Hadoop for the real-time capability, and the other alternative real-time frameworks have been emerging in recent years. This paper presents a survey of the open source technologies that support big data processing in a real-time/near real-time fashion, including their system architectures and platforms. Xiufeng Liu 0001, Nadeem Iftikhar, Xike Xie |
IDEAS | 1 |
| 2014 | CloudETL: scalable dimensional ETL for hiveabstractExtract-Transform-Load (ETL) programs process data into data warehouses (DWs). Rapidly growing data volumes demand systems that scale out. Recently, much attention has been given to MapReduce for parallel handling of massive data sets in cloud environments. Hive is the most widely used RDBMS-like system for DWs on MapReduce and provides scalable analytics. It is, however, challenging to do proper dimensional ETL processing with Hive; e.g., the concept of slowly changing dimensions (SCDs) is not supported (and due to lacking support for UPDATEs, SCDs are complex to handle manually). Also the powerful Pig platform for data processing on MapReduce does not support such dimensional ETL processing. To remedy this, we present the ETL framework CloudETL which uses Hadoop to parallelize ETL execution and to process data into Hive. The user defines the ETL process by means of high-level constructs and transformations and does not have to worry about technical MapReduce details. CloudETL supports different dimensional concepts such as star schemas and SCDs. We present how CloudETL works and uses different performance optimizations including a purpose-specific data placement policy to co-locate data. Further, we present a performance study and compare with other cloud-enabled systems. The results show that CloudETL scales very well and outperforms the dimensional ETL capabilities of Hive both with respect to performance and programmer productivity. For example, Hive uses 3.9 times as long to load an SCD in an experiment and needs 112 statements while CloudETL only needs 4. Xiufeng Liu 0001, Christian Thomsen 0001, Torben Bach Pedersen |
IDEAS | 1 |
| 2012 | MapReduce-based Dimensional ETL Made EasyabstractThis paper demonstrates ETLMR , a novel dimensional Extract--Transform--Load (ETL) programming framework that uses Map-Reduce to achieve scalability. ETLMR has built-in native support of data warehouse (DW) specific constructs such as star schemas, snowflake schemas, and slowly changing dimensions (SCDs). This makes it possible to build MapReduce-based dimensional ETL flows very easily. The ETL process can be configured with only few lines of code. We will demonstrate the concrete steps in using ETLMR to load data into a (partly snowflaked) DW schema. This includes configuration of data sources and targets, dimension processing schemes, fact processing, and deployment. In addition, we also present the scalability on large data sets. Xiufeng Liu 0001, Christian Thomsen 0001, Torben Bach Pedersen |
Proc. VLDB Endow. | 1 |
| 2011 | ETLMR: A Highly Scalable Dimensional ETL Framework Based on MapReduce
Xiufeng Liu 0001, Christian Thomsen 0001, Torben Bach Pedersen |
DaWaK | 1 |
| 2011 | The ETLMR MapReduce-Based ETL Framework
Xiufeng Liu 0001, Christian Thomsen 0001, Torben Bach Pedersen |
SSDBM | 1 |
| 2011 | 3XL: Supporting efficient operations on very large OWL Lite triple-stores
Xiufeng Liu 0001, Christian Thomsen 0001, Torben Bach Pedersen |
Inf. Syst. | 1 |