Jinhan Xin

dblp:305/8889 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0003-1900-7774ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Performance modeling and evaluation · 51% Cloud and datacenter computing · 34% High-performance computing · 11%
Databases, data mining, and information retrieval
2 papers
Database system architecture and tuning · 42% Query processing and optimization · 42% Machine learning and data management · 16%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
benchmarking
0.912025
Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters · ACM Trans. Archit. Code Optim. 2025
Performance modeling and evaluation › benchmarking › benchmark design
benchmark suite design
0.912025
Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters · ACM Trans. Archit. Code Optim. 2025
Performance modeling and evaluation › performance monitoring
hardware performance counters
0.912025
Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters · ACM Trans. Archit. Code Optim. 2025
Performance modeling and evaluation
workload characterization
0.912025
Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters · ACM Trans. Archit. Code Optim. 2025
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.812024
TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data Analytics · IEEE Trans. Computers 2024
Cloud and datacenter computing
configuration tuning
0.812024
TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data Analytics · IEEE Trans. Computers 2024
Cloud and datacenter computing › big data analytics
in-memory data analytics
0.812024
TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data Analytics · IEEE Trans. Computers 2024
High-performance computing
performance optimization
0.812024
TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data Analytics · IEEE Trans. Computers 2024
Database system architecture and tuning
configuration tuning
0.612022
LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL Applications · SIGMOD Conference 2022
Machine learning and data management
machine learning for systems
0.212024
TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data Analytics · IEEE Trans. Computers 2024

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.5early termination · 1.5cherrypick · 1.5stochastic gradient boosted regression tree · 0.9k-means clustering · 0.9independent component analysis · 0.9machine learning · 0.6
YearPublicationVenuePosition
2025 Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters
abstract
We find existing benchmark suites for smartphone CPU micro-architecture design such as Geekbench 5.0 fail to authentically represent the micro-architecture-level performance behavior of widely used real Android applications with interactive operations such as screen sliding. It is therefore crucial to systematically construct a benchmark suite as a supplementary to Geekbench to represent the user interaction behavior of Android applications for CPU micro-architecture design. The key is to identify a small number of representative programs from a large number of real applications. To this end, a set of features used to represent a program need to be constructed, and these features should be fair for different micro-architectures and can be collected efficiently. However, this is extremely difficult for Android applications. For example, the feature collection tools for Android applications are unavailable for benchmark selection. In this article, we propose a novel benchmark suite construction approach dubbed BEMAP to efficiently build a supplementary benchmark suite from real-world Android applications to represent their user interaction behavior. 1 BEMAP innovates four techniques. The first technique, called two-stage RFC (representative feature construction), constructs program features from performance counters (events) to represent a program for selecting benchmarks from a large number of real Android applications in two stages. The first stage identifies a set of important performance events in terms of IPC (instructions per cycle) by employing a machine learning algorithm named SGBRT (Stochastic Gradient Boosted Regression Tree). The second stage constructs representative features based on the important performance events by using ICA (independent component analysis). The second technique, named SPC-MMA (source performance counters from multiple micro-architectures), collects the performance events from multiple mobile CPUs with different micro-architectures and mixes them as the source of RFC. The goal of these two innovations is to make the program features fair to different mobile CPU micro-architectures. The third technique, called ES (Elbow-Silhouette) approach, artfully leverages the synergy between the elbow method and the silhouette method to determine an optimal K when we use K-Means to group Android applications. The fourth technique is that we design a new tool named AutoProfiler to automatically profile the micro-architecture events (e.g., IPC, L1 Icache misses) of Android applications with interactive operations. Using the proposed BEMAP methodology, 2 we constructed SPBench, a novel benchmark suite supplementary to traditional mobile benchmark suites like Geekbench, for mobile CPU micro-architecture design. It consists of 15 benchmarks selected from 100 real Android applications with 3 common user interaction operations, which is the fifth innovation of this article. The experimental results on four significantly different micro-architectures show that SPBench can represent the micro-architecture performance behaviors of the 100 real-world applications with 3 common user interactive operations on each micro-architecture with significantly higher accuracy than benchmark suites produced by the state-of-the-art approaches.
Chenghao Ouyang, Jinhan Xin, Siqi Zeng 0002, Guohui Li 0001, Jianjun Li 0010, Zhibin Yu 0001
ACM Trans. Archit. Code Optim.2
2024 TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data Analytics
abstract
Recently, experiment-driven machine-learning (ML) based configuration tuning for in-memory data analytics such as Apache Spark become popular because they can achieve high speedups. However, experiment-driven ML-based approaches naturally need alargenumber of iterations and each iteration generates a configuration with a probabilistic strategy and executes the program on a real cluster with the configuration. It therefore takes a long time to optimize the performance of an in-memory data analytics program, and thereby hinders these approaches from being widely used in practice.To address this issue, we propose a novel as well as simple approach dubbedTerminating-It-Early (TIE)to reduce the time needed to perform the experiment executions but to achieve speedups similar to those obtained by experiment-driven ML-based approaches. The key idea is that, during the process of searching for the optimal configuration which produces the shortest execution time for a program, weterminatean experiment program execution with a trial configuration as soon as possible when we find its execution time islonger than a predefined threshold(e.g., the shortest execution time thus far). In contrast, traditional experiment-driven ML-based approaches always run all experiment executions completely.We employ 19 Apache Spark programs running on a physical cluster as well as a virtual cluster to evaluate TIE. We compare thetuning timeused to find the optimal configuration of a program and theoptimized execution timeof a program obtained by TIE against those obtained byCherryPickand a reinforcement learning (RL) based approach. The experimental results show that on physical machines, TIE reduces the tuning time used byCherryPickand the RL-based approach by factors of 2.39× and 1.68× on average, respectively. On virtual machines, the corresponding factors are 2.79× and 1.71×. Moreover, the average optimized execution time of the 19 programs tuned by TIE is slightly shorter than those tuned byCherryPickand the RL-based approach.
Chao Chen 0022, Jinhan Xin, Zhibin Yu 0001
IEEE Trans. Computers2
2022 LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL Applications
abstract
Spark SQL has been widely deployed in industry but it is challenging to tune its performance. Recent studies try to employ machine learning (ML) to solve this problem, but suffer from two drawbacks. First, it takes a long time (high overhead) to collect training samples. Second, the optimal configuration for one input data size of the same application might not be optimal for others.
Jinhan Xin, Kai Hwang 0001, Zhibin Yu 0001
SIGMOD Conference1