Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xiaogang Shi

dblp:42/9082 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 2017
0000-0002-0245-4749ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 first-authorSystems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Query processing and optimization · 43% Data integration and cleaning · 30% Data stream processing · 16%

Topics — the 3 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
ad hoc data processing
0.522017
UniAD: A Unified Ad Hoc Data Processing System · ACM Trans. Database Syst. 2017
Towards unified ad-hoc data processing · SIGMOD Conference 2014
Query processing and optimization
intermediate representation
0.322017
Towards unified ad-hoc data processing · SIGMOD Conference 2014
UniAD: A Unified Ad Hoc Data Processing System · ACM Trans. Database Syst. 2017
Database theory
higher-order query
0.212014
Towards unified ad-hoc data processing · SIGMOD Conference 2014

Methods — techniques the papers use, named apart from their topics

program transformation · 0.3declarative-procedural optimization · 0.3fixed-point iteration · 0.2asynchronous iteration · 0.2program optimization · 0.2intermediate representation · 0.2
YearPublicationVenuePosition
2017 UniAD: A Unified Ad Hoc Data Processing System
abstract
Instead of constructing complex declarative queries, many users prefer to write their programs using procedural code embedded with simple queries. Since many users are not expert programmers or the programs are written in a rush, these programs usually exhibit poor performance in practice and it is a challenge to automatically and efficiently optimize these programs. In this article, we present UniAD, which stands for Uni fied execution for Ad hoc Data processing, a system designed to simplify the programming of data processing tasks and provide efficient execution for user programs. We provide the background of program semantics and propose a novel intermediate representation, called Unified Intermediate Representation (UniIR), which utilizes a simple and expressive mechanism HOQ to describe the operations performed in programs. By combining both procedural and declarative logics with the proposed intermediate representation, we can perform various optimizations across the boundary between procedural and declarative code. We propose a transformation-based optimizer to automatically optimize programs and implement the UniAD system. The extensive experimental results on various benchmarks demonstrate that our techniques can significantly improve the performance of a wide range of data processing programs.
Xiaogang Shi, Bin Cui 0001, Gillian Dobbie, Beng Chin Ooi
ACM Trans. Database Syst.1
2016 Tornado: A System For Real-Time Iterative Analysis Over Evolving Data
abstract
There is an increasing demand for real-time iterative analysis over evolving data. In this paper, we propose a novel execution model to obtain timely results at given instants. We notice that a loop starting from a good initial guess usually converges fast. Hence we organize the execution of iterative methods over evolving data into a main loop and several branch loops. The main loop is responsible for the gathering of inputs and maintains the approximation to the timely results. When the results are requested by a user, a branch loop is forked from the main loop and iterates until convergence to produce the results. Using the approximation of the main loop, the branch loops can start from a place near the fixed-point and converge quickly. Since the inputs not reflected in the approximation is concerned with the approximation error, we develop a novel bounded asynchronous iteration model to enhance the timeliness. The bounded asynchronous iteration model can achieve fine-grained updates while ensuring correctness for general iterative methods.
Xiaogang Shi, Bin Cui 0001, Yingxia Shao, Yunhai Tong
SIGMOD Conference1
2014 Towards unified ad-hoc data processing
abstract
It is important to provide efficient execution for ad-hoc data processing programs. In contrast to constructing complex declarative queries, many users prefer to write their programs using procedural code with simple queries. As many users are not expert programmers, their programs usually exhibit poor performance in practice and it is a challenge to automatically optimize these programs and efficiently execute the programs. In this paper, we present UniAD, a system designed to simplify the programming of data processing tasks and provide efficient execution for user programs. We propose a novel intermediate representation named UniQL which utilizes HOQs to describe the operations performed in programs. By combining both procedural and declarative logics, we can perform various optimizations across the boundary between procedural and declarative codes. We describe optimizations and conduct extensive empirical studies using UniAD. The experimental results on four benchmarks demonstrate that our techniques can significantly improve the performance of a wide range of data processing programs.
Xiaogang Shi, Bin Cui 0001, Gillian Dobbie, Beng Chin Ooi
SIGMOD Conference1
2013 bCATE: A Balanced Contention-Aware Transaction Execution Model for Highly Concurrent OLTP Systems
Xiaogang Shi, Yanfei Lv, Yingxia Shao, Bin Cui 0001
WAIM1
2010 Improving the performance of hypervisor-based fault tolerance
abstract
Hypervisor-based fault tolerance (HBFT), a checkpoint-recovery mechanism, is an emerging approach to sustaining mission-critical applications. Based on virtualization technology, HBFT provides an economic and transparent solution. However, the advantages currently come at the cost of substantial overhead during failure-free, especially for memory intensive applications. This paper presents an in-depth examination of HBFT and options to improve its performance. Based on the behavior of memory accesses among checkpointing epochs, we introduce two optimizations, read fault reduction and write fault prediction, for the memory tracking mechanism. These two optimizations improve the mechanism by 31.1% and 21.4% respectively for some application. Then, we present software-superpage which efficiently maps large memory regions between virtual machines (VM). By the above optimizations, HBFT is improved by a factor of 1.4 to 2.2 and it achieves a performance which is about 60% of that of the native VM.
Zhefu Jiang, Xiaogang Shi
IPDPS4