James Yen

dblp:68/486 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0004-1415-0832ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 51% Cloud and datacenter computing · 44% Processor architecture and microarchitecture · 4%
Software engineering, system software, and programming languages
1 paper
Runtime systems and virtual machines · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
1.012026
SpiderSense: Lightweight Last-Level Cache Management via Time Period Tagging for LLC-Critical Workloads · ACM Trans. Archit. Code Optim. 2026
Memory systems › cache management
cache partitioning
1.012026
SpiderSense: Lightweight Last-Level Cache Management via Time Period Tagging for LLC-Critical Workloads · ACM Trans. Archit. Code Optim. 2026
Memory systems › memory hierarchy › cache hierarchy
last-level cache
1.012026
SpiderSense: Lightweight Last-Level Cache Management via Time Period Tagging for LLC-Critical Workloads · ACM Trans. Archit. Code Optim. 2026
Cloud and datacenter computing › multi-tenancy
multi-tenant cloud
1.012026
SpiderSense: Lightweight Last-Level Cache Management via Time Period Tagging for LLC-Critical Workloads · ACM Trans. Archit. Code Optim. 2026
Cloud and datacenter computing
performance isolation
1.012026
SpiderSense: Lightweight Last-Level Cache Management via Time Period Tagging for LLC-Critical Workloads · ACM Trans. Archit. Code Optim. 2026
Runtime systems and virtual machines
binary translation
0.912025
ARMing x86 Games: Accelerating Binary Translation Using Software-Only Validated Flag Speculation · MobiSys 2025
Cloud and datacenter computing
virtualization
0.312026
SpiderSense: Lightweight Last-Level Cache Management via Time Period Tagging for LLC-Critical Workloads · ACM Trans. Archit. Code Optim. 2026
Cloud and datacenter computing › virtualization
virtual machine monitor
0.312026
SpiderSense: Lightweight Last-Level Cache Management via Time Period Tagging for LLC-Critical Workloads · ACM Trans. Archit. Code Optim. 2026

Methods — techniques the papers use, named apart from their topics

software-only validated flag speculation · 1.7time period tagging · 1.0dynamic sampling · 1.0
YearPublicationVenuePosition
2026 The Optimal, The Fast, and The Hybrid: Automatic Placement and Routing for AIE Arrays
abstract
Most of the widely deployed deep learning (DL) workloads, such as large language models and convolutional neural networks, are relatively regular compute graphs that exhibit a high degree of compute parallelism. Therefore, they are a natural fit for spatial dataflow accelerator architectures that map computations to an array of many compute cores communicating via shared memory buffers and/or some form of flexible interconnect between them, such as circuit or packet switched networks-on-chip (NoCs). AMD’s adaptive intelligent engine (AIE) arrays in both the Versal FPGAs and Ryzen NPU devices are exemplars of such architectures. Despite their high peak performance, efficiently mapping workloads to these architectures to maximize compute utilization is a challenging task. The application’s compute kernels are first partitioned into logical cores that are then placed at specific physical core locations. Finally, the different types of inter-core communication resources are configured to realize efficient data movement between cores. Each of these steps is a complex optimization problem that determines the ability to find a feasible mapping and directly impacts performance results. The current programming model for AIE arrays relies on manual placement or uses greedy 2D tiling algorithms. These approaches either require significant designer effort or work only for regular 2D-structured computations, but produce poor-quality or unroutable solutions for other cases. To this end, this work presents a versatile automatic placement and routing (PnR) framework for AIE arrays. We evaluate a variety of placement algorithms that guarantee optimality or trade optimality for scalability. We also formulate routing as a modified multi-commodity flow problem that is solved using mixed-integer linear programming. To demonstrate our PnR framework, we integrate it into AMD’s open-source MLIR-AIE toolchain and develop an entire benchmark suite using their programming API for end-to-end performance evaluation. Across 202 synthetic and real-world benchmarks, our PnR framework finds a legal mapping for 200 out of 202 benchmarks (a 99% success rate) compared to the 62% success rate of AMD’s greedy sequential placer in the MLIR-AIE toolchain. It also reduces routing resource usage by 15% compared to manual placement followed by AMD’s router. On-device end-to-end runtime measurements show that our PnR produces solutions that have a 30% speedup over AMD’s placer and are only 7% slower than expert manual placement.
Hang Yan 0014, James Yen, Rongbo Zhang, Andrew Boutros, Vaughn Betz
FCCM2
2026 SpiderSense: Lightweight Last-Level Cache Management via Time Period Tagging for LLC-Critical Workloads
abstract
Multi-tenant clouds enhance resource sharing among Virtual Machines (VMs) to boost overall utilization and reduce power consumption. However, this also introduces interference among workloads from different tenants and impedes VM performance isolation. In this article, we first demonstrate that the last-level cache (LLC) in CPUs, which is inherently shared by all VMs on the same physical machine, becomes a significant contending resource for LLC-critical workloads, leading to notable performance imbalances under the default hardware caching strategy. Although recent studies on LLC scheduling have progressed, they often require detailed profiling of user workloads or rely on hyperparameter tuning, limiting their applicability to private clusters or specific scenarios. We propose SpiderSense, a software-initiated LLC partitioner for managing Virtual Machine Monitors (VMM), to address these limitations. SpiderSense leverages modern yet off-the-shelf server CPU features to adaptively orchestrate LLC allocation among running black-boxed user VMs. SpiderSense dynamically samples VMs and calculates their fair share of LLC to allocate them while fully improving performance isolation among VMs. We experiment with SpiderSense using typical LLC-critical workloads, representative of the types of applications that stress LLC performance, such as Memcached and Llama. Our results show that SpiderSense improves performance by up to 40% in numerous colocation scenarios compared to current solutions.
Zhixiang Wei, Zhibai Huang, James Yen, Tianlei Xiong, Kailiang Xu, Yucheng Zheng, Xingzi Yu, Yun Wang 0039, Zhengwei Qi
ACM Trans. Archit. Code Optim.3
2025 ARMing x86 Games: Accelerating Binary Translation Using Software-Only Validated Flag Speculation
James Yen, Zhibai Huang, Zhixiang Wei, Chen Chen 0067, Senhao Yu, Yun Wang 0039, Hao Wang 0022, Zhengwei Qi
MobiSys1
2003 Multimodal biometrics: issues in design and testing
abstract
Experimental studies show that multimodal biometric systems for small-scale populations perform better than single-mode biometric systems. We examine if such techniques scale to larger populations, introduce a methodology to test the performance of such systems, and assess the feasibility of using commercial off-the-shelf (COTS) products to construct deployable multimodal biometric systems. A key aspect of our approach is to leverage confidence level scores from preexisting single-mode data. An example presents a multimodal biometrics system analysis that explores various normalization and fusion techniques for face and fingerprint classifiers. This multimodal analysis uses a population of about 1000 subjects, a number ten-times larger than seen in any previously reported study. Experimental results combining face and fingerprint biometric classifiers reveal significant performance improvement over single-mode biometric systems.
Robert Snelick, Mike Indovina, James Yen, Alan Mink
ICMI3
1998 Reliability of Conformance Tests
abstract
A conformance test is a software assurance test that is applied in order to determine if specification requirements of the software are being met. It is a time-independent model, where the software object is subjected to an a priori known test suite. The reliability of the software is the probability that it will function properly for values in the input space. Because the input space is usually very large, it is impossible to sample all input values, so in order to provide better sampling coverage, the input space is partitioned into homogeneous subspaces. Samples are drawn from each subspace for testing the software. The conformance tests based on these samples are required to pass all tests in the test suite. Based on these data, the classical statistical estimate of reliability is one. Such an estimate may be unrealistic if the sample sizes are not large. Even in such a scenario a nontrivial confidence interval is provided for the reliability.
Charles Hagwood, Raghu Kacker, James Yen, David Banks, Lynne Rosenthal, Leonard Gallagher, Paul E. Black
COMPSAC3