VLDB 2026 Research / reviewers in the wild / expert
Jinhan Xin
dblp:305/8889
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0003-1900-7774ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Performance modeling and evaluation · 51% Cloud and datacenter computing · 34% High-performance computing · 11% | |
| Databases, data mining, and information retrieval
2 papers |
Database system architecture and tuning · 42% Query processing and optimization · 42% Machine learning and data management · 16% |
Topics — the 10 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
benchmarking |
0.9 | 1 | 2025 | Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters · ACM Trans. Archit. Code Optim. 2025 |
Performance modeling and evaluation › benchmarking › benchmark design
benchmark suite design |
0.9 | 1 | 2025 | Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters · ACM Trans. Archit. Code Optim. 2025 |
Performance modeling and evaluation › performance monitoring
hardware performance counters |
0.9 | 1 | 2025 | Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters · ACM Trans. Archit. Code Optim. 2025 |
Performance modeling and evaluation
workload characterization |
0.9 | 1 | 2025 | Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters · ACM Trans. Archit. Code Optim. 2025 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.8 | 1 | 2024 | TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data Analytics · IEEE Trans. Computers 2024 |
Cloud and datacenter computing
configuration tuning |
0.8 | 1 | 2024 | TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data Analytics · IEEE Trans. Computers 2024 |
Cloud and datacenter computing › big data analytics
in-memory data analytics |
0.8 | 1 | 2024 | TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data Analytics · IEEE Trans. Computers 2024 |
High-performance computing
performance optimization |
0.8 | 1 | 2024 | TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data Analytics · IEEE Trans. Computers 2024 |
Database system architecture and tuning
configuration tuning |
0.6 | 1 | 2022 | LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL Applications · SIGMOD Conference 2022 |
Machine learning and data management
machine learning for systems |
0.2 | 1 | 2024 | TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data Analytics · IEEE Trans. Computers 2024 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.5early termination · 1.5cherrypick · 1.5stochastic gradient boosted regression tree · 0.9k-means clustering · 0.9independent component analysis · 0.9machine learning · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance CountersabstractWe find existing benchmark suites for smartphone CPU micro-architecture design such as Geekbench 5.0 fail to authentically represent the micro-architecture-level performance behavior of widely used real Android applications with interactive operations such as screen sliding. It is therefore crucial to systematically construct a benchmark suite as a supplementary to Geekbench to represent the user interaction behavior of Android applications for CPU micro-architecture design. The key is to identify a small number of representative programs from a large number of real applications. To this end, a set of features used to represent a program need to be constructed, and these features should be fair for different micro-architectures and can be collected efficiently. However, this is extremely difficult for Android applications. For example, the feature collection tools for Android applications are unavailable for benchmark selection. In this article, we propose a novel benchmark suite construction approach dubbed BEMAP to efficiently build a supplementary benchmark suite from real-world Android applications to represent their user interaction behavior. 1 BEMAP innovates four techniques. The first technique, called two-stage RFC (representative feature construction), constructs program features from performance counters (events) to represent a program for selecting benchmarks from a large number of real Android applications in two stages. The first stage identifies a set of important performance events in terms of IPC (instructions per cycle) by employing a machine learning algorithm named SGBRT (Stochastic Gradient Boosted Regression Tree). The second stage constructs representative features based on the important performance events by using ICA (independent component analysis). The second technique, named SPC-MMA (source performance counters from multiple micro-architectures), collects the performance events from multiple mobile CPUs with different micro-architectures and mixes them as the source of RFC. The goal of these two innovations is to make the program features fair to different mobile CPU micro-architectures. The third technique, called ES (Elbow-Silhouette) approach, artfully leverages the synergy between the elbow method and the silhouette method to determine an optimal K when we use K-Means to group Android applications. The fourth technique is that we design a new tool named AutoProfiler to automatically profile the micro-architecture events (e.g., IPC, L1 Icache misses) of Android applications with interactive operations. Using the proposed BEMAP methodology, 2 we constructed SPBench, a novel benchmark suite supplementary to traditional mobile benchmark suites like Geekbench, for mobile CPU micro-architecture design. It consists of 15 benchmarks selected from 100 real Android applications with 3 common user interaction operations, which is the fifth innovation of this article. The experimental results on four significantly different micro-architectures show that SPBench can represent the micro-architecture performance behaviors of the 100 real-world applications with 3 common user interactive operations on each micro-architecture with significantly higher accuracy than benchmark suites produced by the state-of-the-art approaches. Chenghao Ouyang, Jinhan Xin, Siqi Zeng 0002, Guohui Li 0001, Jianjun Li 0010, Zhibin Yu 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data AnalyticsabstractRecently, experiment-driven machine-learning (ML) based configuration tuning for in-memory data analytics such as Apache Spark become popular because they can achieve high speedups. However, experiment-driven ML-based approaches naturally need alargenumber of iterations and each iteration generates a configuration with a probabilistic strategy and executes the program on a real cluster with the configuration. It therefore takes a long time to optimize the performance of an in-memory data analytics program, and thereby hinders these approaches from being widely used in practice.To address this issue, we propose a novel as well as simple approach dubbedTerminating-It-Early (TIE)to reduce the time needed to perform the experiment executions but to achieve speedups similar to those obtained by experiment-driven ML-based approaches. The key idea is that, during the process of searching for the optimal configuration which produces the shortest execution time for a program, weterminatean experiment program execution with a trial configuration as soon as possible when we find its execution time islonger than a predefined threshold(e.g., the shortest execution time thus far). In contrast, traditional experiment-driven ML-based approaches always run all experiment executions completely.We employ 19 Apache Spark programs running on a physical cluster as well as a virtual cluster to evaluate TIE. We compare thetuning timeused to find the optimal configuration of a program and theoptimized execution timeof a program obtained by TIE against those obtained byCherryPickand a reinforcement learning (RL) based approach. The experimental results show that on physical machines, TIE reduces the tuning time used byCherryPickand the RL-based approach by factors of 2.39× and 1.68× on average, respectively. On virtual machines, the corresponding factors are 2.79× and 1.71×. Moreover, the average optimized execution time of the 19 programs tuned by TIE is slightly shorter than those tuned byCherryPickand the RL-based approach. Chao Chen 0022, Jinhan Xin, Zhibin Yu 0001 |
IEEE Trans. Computers | 2 |
| 2022 | LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL ApplicationsabstractSpark SQL has been widely deployed in industry but it is challenging to tune its performance. Recent studies try to employ machine learning (ML) to solve this problem, but suffer from two drawbacks. First, it takes a long time (high overhead) to collect training samples. Second, the optimal configuration for one input data size of the same application might not be optimal for others. Jinhan Xin, Kai Hwang 0001, Zhibin Yu 0001 |
SIGMOD Conference | 1 |