Thomas Y. Yeh

dblp:58/3980 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
4since 2021 · last 2026
0009-0009-5217-8234ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Hardware accelerators and domain-specific architectures · 87% Integrated circuit design · 5% Emerging computing paradigms · 5%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 5 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
0.612022
Be Like Water: Adaptive Floating Point for Machine Learning · ICML 2022
Hardware accelerators and domain-specific architectures › machine learning accelerator
low-precision arithmetic
0.612022
Be Like Water: Adaptive Floating Point for Machine Learning · ICML 2022
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.612022
Be Like Water: Adaptive Floating Point for Machine Learning · ICML 2022
Integrated circuit design › digital circuit design › arithmetic circuit design
floating-point unit
0.112007
The Art of Deception: Adaptive Precision Reduction for Area Efficient Physics Acceleration · MICRO 2007
Emerging computing paradigms › approximate computing
precision reduction
0.112007
The Art of Deception: Adaptive Precision Reduction for Area Efficient Physics Acceleration · MICRO 2007

Methods — techniques the papers use, named apart from their topics

mixed precision · 1.1adaptive floating point · 1.1precision reduction · 0.1level-of-detail · 0.1LCP solver · 0.1physical simulation · 0.1dynamic precision reduction · 0.1FPU sharing · 0.1
YearPublicationVenuePosition
2026 Fighting Fire with Fire: LLM-Assisted Grading of Handwritten CS Assessments
abstract
Widespread student adoption of large language models (LLMs) has prompted many CS instructors to assign greater weight to handwritten, proctored assessments. However, this approach struggles to scale as class sizes outpace course staff resources. To address this challenge, our study explores LLM-assisted grading to reduce required grading time. While prior work has emphasized tool accuracy, we evaluate both time and accuracy by comparing outcomes when course staff use an LLM-assisted grader versus Gradescope. We also incorporate a mixed-methods analysis of student and staff perceptions. In a CS1 course of 166 students supported by four teaching assistants (TAs), we observed that LLM-assisted grading reduced overall grading time by 40% compared to Gradescope, with time savings of 48% for exams and 25% for quizzes. Across all assessments, short answer questions showed a 46% time improvement, and free response questions showed a 37% time improvement. In terms of accuracy, accepted regrade requests increased negligibly from 0.1% to 0.5% across three exams and six quizzes. Students were generally neutral about LLM-assisted grading, but stressed the value of TA feedback and oversight. Meanwhile, TAs expressed positive sentiments towards the tool, tempered by concerns of skewed perceptions of students caused by the tool. Overall, these findings indicate that LLM-assisted grading can greatly reduce grading time, with only minor accuracy trade-offs that can be mitigated. As a result, LLM-assisted grading emerges as a promising approach for enhancing grading efficiency in CS courses, meriting further exploration for broader adoption.
Jared Apillanes, Jason Lee Weber, Sergio Gago Masagué, Jennifer Wong-Ma, Thomas Y. Yeh
SIGCSE (1)5
2026 Pacing for Mastery: Optimizing LLM Interactions for Learning
abstract
Large Language Models (LLMs) hold significant potential for transforming computer science education, yet concerns over their possible negative effects on student learning and retention have slowed broader instructor adoption. Evidence on LLM use is mixed. While novices may benefit from the generative capabilities of AI, they also risk developing overreliance. To address these concerns, we investigate how the pace of interaction with AI assistants affects learning in introductory CS courses by deploying three AI assistants (Fast, Medium, and Slow) in a classroom setting. Our results show that the slower-paced, Socratic-style AI assistant significantly increases learning, especially for students with less prior knowledge. Although faster-paced interaction benefits more advanced students initially, learning retention degrades enough to negate those gains. Surprisingly, the medium-paced assistant with typical instructor preprompt elements shows no statistically significant improvements. Given that students may use fast-paced commercial AI tools for coursework regardless of policy, offering a slower-paced, Socratic-style AI alternative could meaningfully improve overall student learning outcomes.
Karena Tran, Angela Lombard, Tyler Yu, Haoning Jiang, Thomas Y. Yeh
SIGCSE (1)6
2025 Bridging Novice Programmers and LLMs with Interactivity
Thomas Y. Yeh, Karena Tran, Tyler Yu, Wai On Fong, Tzu-Yi Chen
SIGCSE (1)1
2022 Be Like Water: Adaptive Floating Point for Machine Learning
abstract
In the pursuit of optimizing memory and compute density to accelerate machine learning applications, reduced precision training and inference has been an active area of research. While some approaches selectively apply low precision computations, this may require costly off-chip data transfers or mixed precision support. In this paper, we propose a novel numerical representation, Adaptive Floating Point (AFP), that dynamically adjusts to the characteristics of deep learning data. AFP requires no changes to the model topology, requires no additional training, and applies to all layers of DNN models. We evaluate AFP on a spectrum of representative models in computer vision and NLP, and show that our technique enables ultra-low precision inference of deep learning models while providing accuracy comparable to full precision inference. By dynamically adjusting to ML data, AFP increases memory density by 1.6x, 1.6x, and 3.2x and compute density by 4x, 1.3x, and 12x when compared to BFP, BFloat16, and FP32.
Thomas Y. Yeh, Max Sterner, Zerlina Lai, Brandon Chuang, Alexander Ihler
ICML1
2009 Fool me twice: Exploring and exploiting error tolerance in physics-based animation
abstract
The error tolerance of human perception offers a range of opportunities to trade numerical accuracy for performance in physics-based simulation. However, most prior work on perceptual error tolerance either focus exclusively on understanding the tolerance of the human visual system or burden the application developer with case-specific implementations such as Level-of-Detail (LOD) techniques. In this article, based on a detailed set of perceptual metrics, we propose a methodology to identify the maximum error tolerance of physics simulation. Then, we apply this methodology in the evaluation of four case studies. First, we utilize the methodology in the tuning of the simulation timestep. The second study deals with tuning the iteration count for the LCP solver. Then, we evaluate the perceptual quality of Fast Estimation with Error Control (FEEC) [Yeh et al. 2006]. Finally, we explore the hardware optimization technique of precision reduction.
Thomas Y. Yeh, Glenn Reinman, Sanjay J. Patel, Petros Faloutsos
ACM Trans. Graph.1
2007 ParallAX: an architecture for real-time physics
abstract
Future interactive entertainment applications will featurethe physical simulation of thousands of interacting objectsusing explosions, breakable objects, and cloth effects. Whilethese applications require a tremendous amount of performanceto satisfy the minimum frame rate of 30 FPS, there is a dramatic amount of parallelism in future physics workloads.How will future physics architectures leverage parallelismto achieve the real-time constraint?.
Thomas Y. Yeh, Petros Faloutsos, Sanjay J. Patel, Glenn Reinman
ISCA1
2007 The Art of Deception: Adaptive Precision Reduction for Area Efficient Physics Acceleration
abstract
Physics-based animation has enormous potential to improve the realism of interactive entertainment through dynamic, immersive content creation. Despite the massively parallel nature of physics simulation, fully exploiting this parallelism to reach interactive frame rates will require significant area to place the large number of cores. Fortunately, interactive entertainment requires believability rather than accuracy. Recent work shows that real-time physics has a remarkable tolerance for reduced precision of the significant in floating-point (FP) operations. In this paper, we describe an architecture with a hierarchical floating-point unit (FPU) that leverages dynamic precision reduction to enable efficient FPU sharing among multiple cores. This sharing reduces the area required by these cores, thereby allowing more cores to be packed into a given area and exploiting more parallelism.
Thomas Y. Yeh, Petros Faloutsos, Milos D. Ercegovac, Sanjay J. Patel, Glenn Reinman
MICRO1
2005 Fast and fair: data-stream quality of service
abstract
Chip multiprocessors have the potential to exploit thread level parallelism, particularly in the context of embedded server farms where the available number of threads can be quite high. Unfortunately, both per-core and overall throughput are significantly impacted by the organization of the lowest level on-chip cache. On-chip caches for CMPs must be able to handle the increased demand and contention of multiple cores. To complicate the problem, cache demand changes dynamically with phases changes, context switches, power saving features, and assignments to asymmetric cores.We propose PDAS, a distributed NUCA L2 cache design with an adaptive sharing mechanism. Each core independently measures its dynamic need, and all cache resources are managed to increase utilization, reduce migrations, and lower interference. Per-core performance degradation is bounded while overall throughput is optimized, thus qualitatively improving performance of embedded systems where quality-of-service is an important characteristic.In single thread mode, PDAS, on average, improves by 26%, 27%, and 13% over Private, Shared, and NUCA caches respectively. This improvement is achieved while reducing internal migrations on average by 82% as compared to the NUCA. With thread contention, PDAS increases its performance and power advantage over prior work. The average migration reduction over NUCA increases to over 90%, and average IPC improvements over NUCA are 30%, 14%, and 35% for 2T, 3T, and 4T scenarios.
Thomas Y. Yeh, Glenn Reinman
CASES1
2000 Redundant Arithmetic Optimizations (Research Note)
Thomas Y. Yeh, Hong Wang 0003
Euro-Par1