Xiaokun Fang

dblp:337/5204 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Graph data management · 53% Data stream processing · 27% Query processing and optimization · 20%
Computer networks
1 paper
Edge and fog computing · 50% Internet of things and sensor networks · 50%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Graph data management
dynamic graph
0.812024
Enabling Window-Based Monotonic Graph Analytics with Reusable Transitional Results for Pattern-Consistent Queries · Proc. VLDB Endow. 2024
Memory systems
non-volatile memory
0.812024
Enabling Efficient NVM-Based Text Analytics without Decompression · ICDE 2024
Query processing and optimization › parallel query processing
heterogeneous hardware query processing
0.612022
Exploring Query Processing on CPU-GPU Integrated Edge Device · IEEE Trans. Parallel Distributed Syst. 2022
Edge and fog computing
edge analytics
0.612022
Exploring Query Processing on CPU-GPU Integrated Edge Device · IEEE Trans. Parallel Distributed Syst. 2022
Internet of things and sensor networks
query processing
0.612022
Exploring Query Processing on CPU-GPU Integrated Edge Device · IEEE Trans. Parallel Distributed Syst. 2022
Memory systems
DRAM
0.212024
Enabling Efficient NVM-Based Text Analytics without Decompression · ICDE 2024

Methods — techniques the papers use, named apart from their topics

upper bound estimation · 1.5pruning · 1.5fine-grained workload scheduling · 1.1transitional result reuse · 0.8slice merging · 0.8persistence strategy · 0.8persistence strategies · 0.8
YearPublicationVenuePosition
2024 Enabling Efficient NVM-Based Text Analytics without Decompression
abstract
Text analytics directly on compression (TADOC) is a promising technology designed for handling big data analytics. However, a substantial amount of DRAM is required for high performance, which limits its usage in many important scenarios where the capacity of DRAM is limited, such as memory-constrained systems. Non-volatile memory (NVM) is a novel storage technology that combines the advantage of reading per-formance and byte addressability of DRAM with the durability of traditional storage devices like SSD and HDD. Unfortunately, no research demonstrates how to use NVM to reduce DRAM utilization in compressed data analytics. In this paper, we propose N-TADOC, which substitutes DRAM with NVM while maintaining TADOC's analytics performance and space savings. Utilizing an NVM block device to reduce DRAM utilization presents two challenges, including poor data locality in traversing datasets and auxiliary data structure reconstruction on NVM. We develop novel designs to solve these challenges, including a pruning method with NVM pool management, bottom-up upper bound estimation, correspondent data structures, and persistence strategy at different levels of cost. Experimental results show that on four real-world datasets, N-TADOC achieves 2.04× performance speedup compared to the processing directly on the uncompressed data and 70.7% DRAM space saving compared to the original TADOC.
Xiaokun Fang, Feng Zhang 0007, Junxiang Nong, Puyun Hu, Yunpeng Chai, Xiaoyong Du 0001
ICDE1
2024 Enabling Window-Based Monotonic Graph Analytics with Reusable Transitional Results for Pattern-Consistent Queries
abstract
Evolving graphs consisting of slices are large and constantly changing. For example, in Alipay, the graph generates hundreds of millions of new transaction records every day. Analyzing the graph within a temporary window is time-consuming due to the heavy merging of slices. Fortunately, we have discovered that most queries exhibit consistent patterns and possess monotonic properties. As a result, transitional results can be computed within slice generation for reuse. Accordingly, we develop MergeGraph enabling window-based monotonic graph analytics with reusable transitional results for pattern-consistent queries. MergeGraph has three advantages over previous works. First, it is the first system specifically tailored for window-based monotonic graph analytics with pattern-consistent queries. Second, it effectively utilizes transitional results from different slices concurrently. Third, MergeGraph boasts a high degree of expressiveness, supporting a broad spectrum of monotonic graph queries. Experimental results demonstrate that MergeGraph delivers significant performance benefits. In evaluating four typical graph applications, MergeGraph achieves an average speedup of 11.30× compared to state-of-the-art methods.
Zheng Chen 0023, Feng Zhang 0007, Xiaokun Fang, Guanyu Feng, Xiaowei Zhu 0001, Xiaoyong Du 0001
Proc. VLDB Endow.4
2022 Exploring Query Processing on CPU-GPU Integrated Edge Device
abstract
Huge amounts of data have been generated on edge devices every day, which requires efficient data analytics and management. However, due to the limited computing capacity of these edge devices, query processing at the edge faces tremendous pressure. Fortunately, in recent years, hardware vendors have integrated heterogeneous coprocessors, such as GPUs, into the edge device, which can provide much more computing power. Furthermore, the CPU-GPU integrated edge device has shown significant benefits in a variety of situations. Therefore, the exploration of query processing on such CPU-GPU integrated edge devices becomes an urgent need. In this article, we develop a fine-grained query processing engine, called FineQuery, which can perform efficient query processing on CPU-GPU integrated edge devices. Particularly, FineQuery can take advantage of both architectural features of edge devices and query characteristics by performing fine-grained workload scheduling between the CPU and the GPU. Experiments show that on TPC-H workloads, FineQuery reduces 42.81% latency and improves 2.39× bandwidth utilization on average compared to the implementation of using only GPU or CPU. Furthermore, query processing at the edge can bring significant performance-per-cost benefits and energy efficiency. On average, FineQuery at the edge brings 21× performance-per-cost ratio and 4× energy efficiency compared with processing the data on a discrete GPU platform.
Jiesong Liu, Feng Zhang 0007, Hourun Li, Dalin Wang, Weitao Wan, Xiaokun Fang, Jidong Zhai, Xiaoyong Du 0001
IEEE Trans. Parallel Distributed Syst.6