Xiongpai Qin

dblp:50/53 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2023 Flowris: Managing Data Analysis Workflows for Conversational Agent
Jiajia Sun, Yueguo Chen, Xiongpai Qin
DASFAA (4)4
2023 DiCausal: Exploiting Domain Knowledge for Interactive Causal Discovery
Yueguo Chen, Shengwei Huang, Xiongpai Qin, Li Chong
DASFAA (4)4
2021 TS-Benchmark: A Benchmark for Time Series Databases
abstract
Time series data is widely used in scenarios such as supply chain, stock data analysis, and smart manufacturing. A number of time series database systems have been invented to manage and query large volumes of time series data. We observe that the existing benchmarks of time series databases are focused on workloads of complex analysis such as pattern matching and trend prediction whose performance may be highly affected by the data analysis algorithms, instead of the back-end databases. However, in many real applications of time series databases, people are more interested in the performance metrics such as data injection throughput and query processing time. A benchmark is still required to extensively compare the performance of time series databases in such metrics. We introduce such a benchmark called TS-Benchmark which majorly applies a scenario of device monitoring for wind turbines. A DCGAN-based data generation model is proposed to generate large volumes of time series data from some real time series data. The workloads are categorized into three folds: data loading (in batch), streaming data injection, and historical data access (for typical queries). We implement the benchmark and compare four representative time series databases: InfluxDB, TimescaleDB, Druid and OpenTSDB. The results are reported and analyzed.
Yuanzhe Hao, Xiongpai Qin, Yueguo Chen, Xiaoguang Sun, Xiao Zhang 0001, Xiaoyong Du 0001
ICDE2
2018 Rainbow: Adaptive Layout Optimization for Wide Tables
abstract
Popular column stores such as ORC and Parquet have been widely used in many Hadoop-oriented data analysis systems. With the effective column skipping and data compression functionalities provided by column stores, wide tables with hundreds or even thousands of columns are applied by many big data analysis applications to avoid the expensive distributed joins. We found that the performance of such systems can be further improved by optimizing the physical data layout to fit certain workloads and system settings. However, it is nontrivial to perform such optimization manually. In this demo, we present a data layout optimization tool called Rainbow, which leverages workload-driven layout optimization algorithms to adjust data layouts adaptively without intervening the previous data blocks that have been stored. We also provide a Web UI for users to interact with the layout optimization process. Furthermore, Rainbow is open sourced with an accompanying benchmark for performance evaluation of wide tables.
Haoqiong Bian, Youxian Tao, Guodong Jin, Yueguo Chen, Xiongpai Qin, Xiaoyong Du 0001
ICDE5
2016 Entity Fiber Based Partitioning, No Loss Staging and Fast Loading of Log Data
abstract
Real time analysis of fine granularity of log data can help people gain personalized insights on business. For example, real time analysis of e-commerce log data will help us learn recent changes of browsing and shopping behavior of specific customers, which enables us to provide personalized recommendations. To accomplish such analysis, log data should have been loaded quickly into data warehouse without loss. This paper proposes a no loss staging and fast loading solution for log data. Based on open sourced tools such as Kafka, HDFS, and Spark, we have designed and implemented an entity fiber based log data partitioning and staging method, as well as a parallel loading algorithm. Our scheme achieves a data staging performance of around 390,000 records/s, and a data loading performance of around 160,000 records/s.
Xiongpai Qin, Yueguo Chen, Guodong Jin, Yiming Cong, Xiaoyong Du 0001
PDCAT1
2015 A Fast Data Ingestion and Indexing Scheme for Real-Time Log Analytics
Haoqiong Bian, Yueguo Chen, Xiongpai Qin, Xiaoyong Du 0001
APWeb3
2015 Efficient query processing framework for big data warehouse: an almost join-free approach
Huiju Wang, Xiongpai Qin, Xuan Zhou 0001, Zuoyan Qin, Qing Zhu 0010, Shan Wang 0001
Frontiers Comput. Sci.2
2014 HC-Store: putting MapReduce's foot in two camps
Huijui Wang, Xuan Zhou 0001, Yu Cao 0004, Xiongpai Qin, Jidong Chen, Shan Wang 0001
Frontiers Comput. Sci.5
2013 A Survey on Benchmarks for Big Data and Some More Considerations
Xiongpai Qin
IDEAL1
2011 LinearDB: A Relational Approach to Make Data Warehouse Scale Like MapReduce
Huijui Wang, Xiongpai Qin, Shan Wang 0001, Zhanwei Wang
DASFAA (2)2
2011 Parallel Aggregation Queries over Star Schema: A Hierarchical Encoding Scheme and Efficient Percentile Computing as a Case
abstract
Big data analysis is a main challenge we meet recently. Cloud computing is attracting more and more big data analysis applications, due to its well scalability and fault-tolerance. Some aggregation functions, like SUM, can be computed in parallel, because they satisfy distributive law of addition. Unfortunately, some of statistical functions are not naturally parallelizable. That means they do not satisfy distributive law of addition. In this paper, we focus on percentile computing problem. We proposed an iterative-style prediction-based parallel algorithm in a distributed system. Prediction is done through a sampling technique. Experiment results verify the efficiency of our algorithm.
Xiongpai Qin, Huijui Wang, Xiaoyong Du 0001, Shan Wang 0001
ISPA1
2008 A Parallel Recovery Scheme for Update Intensive Main Memory Database Systems
abstract
In update intensive applications, main memory database systems produce large volume of log records, it is critical to write out the log records efficiently to speedup transaction processing. We propose a parallel recovery scheme based on XOR differential logging for main memory database systems in such environments. Some NVRAM is used to temporarily hold log records and decouple transaction committing from disk writes, inherited parallelism properties of differential logging are exploited to accelerate log flushing by using multiple log disks. During recovery, log records are loaded from multiple log disks and applied to data partition in time without the need of reordering according to serialization order, total recovery time is cut down. The scheme employs a data partition based consistent checkpointing method. The log records are classified according to IDs of data partitions accessed. Data partitions are recovered according to loading priorities computed from update frequencies and transaction waiting times, data access demands of new transactions coming after failure recovery are given attention immediately, thus the scheme provides system availability during recovery, which is of importance for large scale main memory database systems.
Xiongpai Qin, Yanqin Xiao, Shan Wang 0001
PDCAT1
2008 COCA: More Accurate Multidimensional Histograms out of More Accurate Correlations Detection
abstract
Detecting and exploiting correlations among columns in relational databases are of great value for query optimizers to generate better query execution plans (QEPs). We propose a more robust and informative metric, namely, entropy correlation coefficients, other than chi-square test to detect correlations among columns in large datasets. We introduce a novel yet simple kind of multi-dimensional synopses named COCA-Hist to cope with different correlations in databases. With the aid of the precise metric of entropy correlation coefficients, correlations of various degrees can be detected effectively; when correlation coefficients testify to mutual independence among columns, the AVI (attribute value independence) assumption can be adopted undoubtedly. COCA can also serve as a data-mining tool with superior qualities as CORDS does. We demonstrate the effectiveness and accuracy of our approach by several experiments.
Xiongpai Qin, Shan Wang 0001
WAIM2