Scott Cheng

dblp:226/4169 · also YuHsuan Cheng · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0001-9954-7986ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 63% Planning, search and constraint satisfaction · 37%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 53% Processor architecture and microarchitecture · 20% Hardware accelerators and domain-specific architectures · 20%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
exploration
0.912025
Uncertainty-Guided Exploration for Efficient AlphaZero Training · NeurIPS 2025
Machine learning › Reinforcement learning › exploration
uncertainty-guided exploration
0.912025
Uncertainty-Guided Exploration for Efficient AlphaZero Training · NeurIPS 2025
Machine learning › Reinforcement learning
value function estimation
0.912025
Uncertainty-Guided Exploration for Efficient AlphaZero Training · NeurIPS 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.812024
Speculative Monte-Carlo Tree Search · NeurIPS 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › plan execution
speculative execution
0.812024
Speculative Monte-Carlo Tree Search · NeurIPS 2024
Processor architecture and microarchitecture
CPU optimization
0.712023
Optimizing CPU Performance for Recommendation Systems At-Scale · ISCA 2023
Memory systems
memory access latency
0.712023
Optimizing CPU Performance for Recommendation Systems At-Scale · ISCA 2023
Hardware accelerators and domain-specific architectures › machine learning accelerator
recommendation model inference
0.712023
Optimizing CPU Performance for Recommendation Systems At-Scale · ISCA 2023
Memory systems
software prefetching
0.712023
Optimizing CPU Performance for Recommendation Systems At-Scale · ISCA 2023
Parallel and multicore computing
parallelization strategies
0.212024
Speculative Monte-Carlo Tree Search · NeurIPS 2024
Memory systems › cache
cache performance
0.212023
Optimizing CPU Performance for Recommendation Systems At-Scale · ISCA 2023
Memory systems › memory access patterns
irregular memory access
0.212023
Optimizing CPU Performance for Recommendation Systems At-Scale · ISCA 2023

Methods — techniques the papers use, named apart from their topics

speculative execution · 1.5neural network caching · 1.5monte carlo tree search · 0.9bayesian inference · 0.9software prefetching · 0.7hyperthreading · 0.7
YearPublicationVenuePosition
2025 Uncertainty-Guided Exploration for Efficient AlphaZero Training
abstract
AlphaZero has achieved remarkable success in complex decision-making problems through self-play and neural network training. However, its self-play process remains inefficient due to limited exploration of high-uncertainty positions, the overlooked runner-up decisions in Monte Carlo Tree Search (MCTS), and high variance in value labels. To address these challenges, we propose and evaluate uncertainty-guided exploration by branching from high-uncertainty positions using our proposed Label Change Rate (LCR) metric, which is further refined by a Bayesian inference framework. Our proposed approach leverages runner-up MCTS decisions to create multiple variations, and ensembles value labels across these variations to reduce variance. We investigate three key design parameters for our branching strategy: where to branch, how many variations to branch, and which move to play in the new branch. Our empirical findings indicate that branching with 10 variations per game provides the best performance-exploration balance. Overall, our end-to-end results show an improved sample efficiency over the baseline by 58.5\% on 9x9 Go in the early stage of training and by 47.3\% on 19x19 Go in the late stage of training.
Scott Cheng, Meng-Yu Tsai 0001, Ding-Yong Hong, Mahmut T. Kandemir
NeurIPS1
2024 Speculative Monte-Carlo Tree Search
abstract
Monte-Carlo tree search (MCTS) is an influential sequential decision-making algorithm notably employed in AlphaZero. Despite its success, the primary challenge in AlphaZero training lies in its prolonged time-to-solution due to the high latency imposed by the sequential MCTS process. To address this challenge, this paper proposes and evaluates an inter-decision parallelization strategy called speculative MCTS, a new type of parallelism in AlphaZero which implements speculative execution. This approach allows for the parallel execution of future moves before the current MCTS computations are completed, thus reducing the latency. Additionally, we analyze factors contributing to the overall speedup by studying the synergistic effects of speculation and neural network caching in MCTS. We also provide an analytical model that can be used to evaluate the potential of different speculation strategies before they are implemented and deployed. Our empirical findings indicate that the proposed speculative MCTS can reduce training latency by 5.81$\times$ in 9x9 Go games. Moreover, our study shows that speculative execution can enhance the NN cache hit rate by 26\% during midgame. Overall, our end-to-end evaluation indicates 1.91$\times$ speedup in 19x19 Go training time, compared to the state-of-the-art KataGo program.
Scott Cheng, Mahmut T. Kandemir, Ding-Yong Hong
NeurIPS1
2023 Optimizing CPU Performance for Recommendation Systems At-Scale
abstract
Deep Learning Recommendation Models (DLRMs) are very popular in personalized recommendation systems and are a major contributor to the data-center AI cycles. Due to the high computational and memory bandwidth needs of DLRMs, specifically the embedding stage in DLRM inferences, both CPUs and GPUs are used for hosting such workloads. This is primarily because of the heavy irregular memory accesses in the embedding stage of computation that leads to significant stalls in the CPU pipeline. As the model and parameter sizes keep increasing with newer recommendation models, the computational dominance of the embedding stage also grows, thereby, bringing into question the suitability of CPUs for inference. In this paper, we first quantify the cause of irregular accesses and their impact on caches and observe that off-chip memory access is the main contributor to high latency. Therefore, we exploit two well-known techniques: (1) Software prefetching, to hide the memory access latency suffered by the demand loads and (2) Overlapping computation and memory accesses, to reduce CPU stalls via hyperthreading to minimize the overall execution time. We evaluate our work on a single-core and 24-core configuration with the latest recommendation models and recently released production traces. Our integrated techniques speed up the inference by up to 1.59x, and on average by 1.4x.
Scott Cheng, Vishwas Kalagi, Vrushabh Sanghavi, Samvit Kaul, Meena Arunachalam, Kiwan Maeng, Adwait Jog, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das
ISCA2
2019 Student Cluster Competition 2018, team NTHU: Reproducing performance of multi-physics simulations of the tsunamigenic 2004 sumatra megathrust earthquake on the Intel Skylake architecture
ShaoFu Lin, ChiChen Yang, Scott Cheng, KengJui Hsu, Hung-Hsin Chen, YuanChing Lin, Jerry Chou 0001
Parallel Comput.3
2018 Student cluster competition 2017, team NTHU: Reproducing vectorization of the tersoff multi-body potential on the Intel Skylake and Nvidia P100 architecture
ChanJung Chang, YungChing Lin, Scott Cheng, YuCheng Wang, LiYu Yu, TienChi Yang, Jerry Chou 0001
Parallel Comput.3