EDBT 2026 Demo / reviewers in the wild / expert
Xingjun Hao
dblp:181/9119
· DBLP profile ↗
6ranked-venue papers
1as first author
2since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Storage systems · 35% Embedded and real-time systems · 18% Parallel and multicore computing · 18% | |
| Databases, data mining, and information retrieval
3 papers |
Query processing and optimization · 75% Database theory · 14% Data stream processing · 6% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
OLAP |
0.7 | 2 | 2020 | OLAP over Probabilistic Data Cubes II: Parallel Materialization and Extended Aggregates · IEEE Trans. Knowl. Data Eng. 2020 OLAP over probabilistic data cubes I: Aggregating, materializing, and querying · ICDE 2016 |
Query processing and optimization › OLAP › data cube
cube materialization |
0.4 | 1 | 2020 | OLAP over Probabilistic Data Cubes II: Parallel Materialization and Extended Aggregates · IEEE Trans. Knowl. Data Eng. 2020 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.4 | 1 | 2019 | Energy-Efficient Task Scheduling for CPU-Intensive Streaming Jobs on Hadoop · IEEE Trans. Parallel Distributed Syst. 2019 |
Embedded and real-time systems › energy-efficient embedded systems
energy-efficient scheduling |
0.4 | 1 | 2019 | Energy-Efficient Task Scheduling for CPU-Intensive Streaming Jobs on Hadoop · IEEE Trans. Parallel Distributed Syst. 2019 |
Parallel and multicore computing
task scheduling |
0.4 | 1 | 2019 | Energy-Efficient Task Scheduling for CPU-Intensive Streaming Jobs on Hadoop · IEEE Trans. Parallel Distributed Syst. 2019 |
Query processing and optimization
aggregate query processing |
0.2 | 1 | 2016 | OLAP over probabilistic data cubes I: Aggregating, materializing, and querying · ICDE 2016 |
Distributed systems
distributed caching |
0.2 | 1 | 2016 | Efficient Storage of Multi-Sensor Object-Tracking Data · IEEE Trans. Parallel Distributed Syst. 2016 |
Storage systems
file systems |
0.2 | 1 | 2016 | Efficient Storage of Multi-Sensor Object-Tracking Data · IEEE Trans. Parallel Distributed Syst. 2016 |
Storage systems › file systems › distributed file system
HDFS |
0.2 | 1 | 2016 | Efficient Storage of Multi-Sensor Object-Tracking Data · IEEE Trans. Parallel Distributed Syst. 2016 |
Storage systems
storage reliability |
0.2 | 1 | 2016 | Efficient Storage of Multi-Sensor Object-Tracking Data · IEEE Trans. Parallel Distributed Syst. 2016 |
Database theory › probabilistic databases
possible world semantics |
0.1 | 1 | 2020 | OLAP over Probabilistic Data Cubes II: Parallel Materialization and Extended Aggregates · IEEE Trans. Knowl. Data Eng. 2020 |
Database theory
probabilistic databases |
0.1 | 1 | 2020 | OLAP over Probabilistic Data Cubes II: Parallel Materialization and Extended Aggregates · IEEE Trans. Knowl. Data Eng. 2020 |
Data models and query languages › uncertain data
probabilistic data model |
0.1 | 1 | 2016 | OLAP over probabilistic data cubes I: Aggregating, materializing, and querying · ICDE 2016 |
Internet of things and sensor networks › wireless sensor network
target tracking |
0.1 | 1 | 2016 | Efficient Storage of Multi-Sensor Object-Tracking Data · IEEE Trans. Parallel Distributed Syst. 2016 |
Methods — techniques the papers use, named apart from their topics
performance modeling · 0.8energy modeling · 0.8binning algorithms · 0.8sketch-based aggregation · 0.7sensor-dependence graph · 0.5clustering · 0.5parallelization · 0.4convolution aggregation · 0.4cost model · 0.2convolution · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Real-time semantic segmentation with local spatial pixel adjustment
Cunjun Xiao, Xingjun Hao, Wenming Zhang |
Image Vis. Comput. | 2 |
| 2021 | Real-time semantic segmentation with weighted factorized-depthwise convolution
Xiaochen Hao, Xingjun Hao, Yaru Zhang |
Image Vis. Comput. | 2 |
| 2020 | OLAP over Probabilistic Data Cubes II: Parallel Materialization and Extended AggregatesabstractOn-Line Analytical Processing (OLAP) enables powerful analytics by quickly computing aggregate values of numerical measures over multiple hierarchical dimensions for massive datasets. However, many types of source data, e.g., from GPS, sensors, and other measurement devices, are intrinsically inaccurate (imprecise and/or uncertain) and thus OLAP cannot be readily applied. In this paper, we address the resulting data veracityproblem in OLAP by proposing the concept of probabilistic data cubes. Such a cube is comprised of a set of probabilistic cuboids which summarize the aggregated values in the form of probability mass functions (pmfs in short) and thus offer insights into the underlying data quality and enable confidence-aware query evaluation and analysis. However, the probabilistic nature of data poses computational challenges, since a probabilistic database can have exponential number of possible worlds under the possible world semantics. Even worse, it is hard to share computations among different cuboids, as aggregation functions that are distributive for traditional data cubes, e.g., SUM, become holistic in probabilistic settings. In this paper, we propose a complete set of techniques for probabilistic data cubes, from cuboid aggregation, over cube materialization, to query evaluation. We study two types of aggregation: convolution and sketch-based, which take polynomial time complexities for aggregation and jointly enable efficient query processing. Also, our proposal is versatile in terms of: 1) its capability of supporting common aggregation functions, i.e., SUM, COUNT, MAX, and AVG; 2) its adaptivity to different materialization strategies, e.g., full versus partial materialization, with support of our devised cost models and parallelization framework; 3) its coverage of common OLAP operations, i.e., probabilistic slicing and dicing queries. Extensive experiments over real and synthetic datasets show that our techniques are effective and scalable. Xike Xie, Xingjun Hao, Torben Bach Pedersen, Peiquan Jin, Wei Yang 0011 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2019 | Energy-Efficient Task Scheduling for CPU-Intensive Streaming Jobs on HadoopabstractHadoop, especially Hadoop 2.0, has been a dominant framework for real-time big data processing. However, Hadoop is not optimized for energy efficiency. Aiming to solve this problem, in this paper, we propose a new framework to improve the energy efficiency of Hadoop 2.0. We focus on the resource manager in Hadoop 2.0, namely YARN, and propose energy-efficient task scheduling mechanisms on YARN. Particularly, we focus on CPU-intensive streaming jobs and classify streaming jobs into two types, namely batch streaming jobs (i.e., a set of jobs are submitted simultaneously) and online streaming jobs (i.e., jobs are continuously submitted one by one). We devise different energy-efficient task scheduling algorithms for each kind of streaming jobs. Specially, we first propose to abstractly model performance and energy consumption by considering the characteristics of tasks as well as the computational resources in YARN. Based on this model, we study the energy efficiency of streaming tasks which consist of the performance model and energy consumption model of task. We propose two key principles for improving energy efficiency: 1) CPU usage aware task allocation, partitions tasks to NMs based on the task characteristic in term of CPU usage; and 2) resource efficient task allocation, reduce idle resource. Then, we propose a D-based binning algorithm for the batch task scheduling and K-based binning algorithm for the online task scheduling that can adapt to continuously arriving tasks. We conduct extensive experiments on a real Hadoop 2.0 cluster and use two kinds of workloads to evaluate the performance and energy efficiency of our proposal. Compared with Storm (the streaming data processing tool in Hadoop 2.0) and other approaches including TAPA and DVFS-MR, our proposal is more energy efficient. The batch task scheduling algorithm reduces up to 10 percent of energy consumption and keeps comparable performance. In addition, the online task scheduling algorithm reduces up to 7 percent over the existing algorithms. Peiquan Jin, Xingjun Hao, Lihua Yue |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | OLAP over probabilistic data cubes I: Aggregating, materializing, and queryingabstractOn-Line Analytical Processing (OLAP) enables powerful analytics by quickly computing aggregate values of numerical measures over multiple hierarchical dimensions for massive datasets. However, many types of source data, e.g., from GPS, sensors, and other measurement devices, are intrinsically inaccurate (imprecise and/or uncertain) and thus OLAP cannot be readily applied. In this paper, we address the resulting data veracity problem in OLAP by proposing the concept of probabilistic data cubes. Such a cube is comprised of a set of probabilistic cuboids which summarize the aggregated values in the form of probability mass functions (pmfs in short) and thus offer insights into the underlying data quality and enable confidence-aware query evaluation and analysis. However, the probabilistic nature of data poses computational challenges as even simple operations are #P-hard under the possible world semantics. Even worse, it is hard to share computations among different cuboids, as aggregation functions that are distributive for traditional data cubes, e.g., SUM and COUNT, become holistic in probabilistic settings. In this paper, we propose a complete set of techniques for probabilistic data cubes, from cuboid aggregation, over cube materialization, to query evaluation. For aggregation, we focus on how to maximize the sharing of computation among cells and cuboids. We present two aggregation methods: convolution and sketch-based. The two methods scale down the time complexities of building a probabilistic cuboid to polynomial and linear, respectively. Each of the two supports both full and partial data cube materialization. Then, we devise a cost model which guides the aggregation methods to be deployed and combined during the cube materialization. We further provide algorithms for probabilistic slicing and dicing queries on the data cube. Extensive experiments over real and synthetic datasets are conducted to show that the techniques are effective and scalable. Xike Xie, Xingjun Hao, Torben Bach Pedersen, Peiquan Jin, Jinchuan Chen |
ICDE | 2 |
| 2016 | Efficient Storage of Multi-Sensor Object-Tracking DataabstractThe rapid development of Internet of Things (IoT) enables people to track objects by deploying multiple sensors, e.g., to track people in indoor spaces using RFID sensors. Multi-sensor object-tracking data are usually produced as records, which are thereby organized into small files and written to servers. However, the small-size property and high arriving-rate of multi-sensor object-tracking data will result in poor I/O performance of file systems such as HDFS. In this paper, we propose the first read/write-optimized solution for storing multi-sensor object-tracking data on HDFS. In particular, we exploit a distributed caching mechanism and a parallel file-merging policy to improve the I/O performance of HDFS. With our design, object-tracking data are first cached by a Distributed Memory File System (DMFS) on top of HDFS. These data are further merged into large files, which are then flushed to HDFS in parallel. We demonstrate that this mechanism is able to improve the write throughput of HDFS and outperforms existing centralized-cache-based approaches. In addition, in order to improve the search performance of object queries over multi-sensor object-tracking data, we propose a Sensor-Dependence Graph (SDG) to model sensor dependence and further present an SDG-based algorithm to efficiently cluster sensors. The object-tracking data from the sensors in the same cluster are merged into the same large files, which can reduce file scans during query processing and therefore improve search performance. We conduct extensive experiments to evaluate the performance of our proposal. The results suggest the efficiency of our proposal with respect to disk-write throughput, memory-write throughput, search performance, and sensor clustering. Xingjun Hao, Peiquan Jin, Lihua Yue |
IEEE Trans. Parallel Distributed Syst. | 1 |