Chenghao Ouyang

dblp:356/8081 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0001-8698-2356ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 93% Processor architecture and microarchitecture · 7%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
benchmarking
0.912025
Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters · ACM Trans. Archit. Code Optim. 2025
Performance modeling and evaluation › benchmarking › benchmark design
benchmark suite design
0.912025
Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters · ACM Trans. Archit. Code Optim. 2025
Performance modeling and evaluation › performance monitoring
hardware performance counters
0.912025
Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters · ACM Trans. Archit. Code Optim. 2025
Performance modeling and evaluation
workload characterization
0.912025
Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters · ACM Trans. Archit. Code Optim. 2025

Methods — techniques the papers use, named apart from their topics

stochastic gradient boosted regression tree · 0.9k-means clustering · 0.9independent component analysis · 0.9
YearPublicationVenuePosition
2025 Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters
abstract
We find existing benchmark suites for smartphone CPU micro-architecture design such as Geekbench 5.0 fail to authentically represent the micro-architecture-level performance behavior of widely used real Android applications with interactive operations such as screen sliding. It is therefore crucial to systematically construct a benchmark suite as a supplementary to Geekbench to represent the user interaction behavior of Android applications for CPU micro-architecture design. The key is to identify a small number of representative programs from a large number of real applications. To this end, a set of features used to represent a program need to be constructed, and these features should be fair for different micro-architectures and can be collected efficiently. However, this is extremely difficult for Android applications. For example, the feature collection tools for Android applications are unavailable for benchmark selection. In this article, we propose a novel benchmark suite construction approach dubbed BEMAP to efficiently build a supplementary benchmark suite from real-world Android applications to represent their user interaction behavior. 1 BEMAP innovates four techniques. The first technique, called two-stage RFC (representative feature construction), constructs program features from performance counters (events) to represent a program for selecting benchmarks from a large number of real Android applications in two stages. The first stage identifies a set of important performance events in terms of IPC (instructions per cycle) by employing a machine learning algorithm named SGBRT (Stochastic Gradient Boosted Regression Tree). The second stage constructs representative features based on the important performance events by using ICA (independent component analysis). The second technique, named SPC-MMA (source performance counters from multiple micro-architectures), collects the performance events from multiple mobile CPUs with different micro-architectures and mixes them as the source of RFC. The goal of these two innovations is to make the program features fair to different mobile CPU micro-architectures. The third technique, called ES (Elbow-Silhouette) approach, artfully leverages the synergy between the elbow method and the silhouette method to determine an optimal K when we use K-Means to group Android applications. The fourth technique is that we design a new tool named AutoProfiler to automatically profile the micro-architecture events (e.g., IPC, L1 Icache misses) of Android applications with interactive operations. Using the proposed BEMAP methodology, 2 we constructed SPBench, a novel benchmark suite supplementary to traditional mobile benchmark suites like Geekbench, for mobile CPU micro-architecture design. It consists of 15 benchmarks selected from 100 real Android applications with 3 common user interaction operations, which is the fifth innovation of this article. The experimental results on four significantly different micro-architectures show that SPBench can represent the micro-architecture performance behaviors of the 100 real-world applications with 3 common user interactive operations on each micro-architecture with significantly higher accuracy than benchmark suites produced by the state-of-the-art approaches.
Chenghao Ouyang, Jinhan Xin, Siqi Zeng 0002, Guohui Li 0001, Jianjun Li 0010, Zhibin Yu 0001
ACM Trans. Archit. Code Optim.1
2023 DAG-Aware Optimization for Geo-Distributed Data Analytics
abstract
Geo-distributed data analytics has been proposed to analyze geographically distributed data. Existing studies have achieved significant reductions in execution time and data transfer cost ($) of data analytics jobs by optimizing task placement. Given a directed acyclic graph (DAG)-style job, however, they mainly optimize each stage independently, and they tend to distribute tasks and intermediate data across all locations, potentially inflating execution time and data transfer cost of descendent stages and the whole job.
Qingyuan Wang 0005, Bin Gao 0013, Zhi Zhou 0006, Fei Xu 0009, Chenghao Ouyang
ICPP5