Teresa Zhang

dblp:420/9628 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Energy-efficient computing · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › inference efficiency
LLM inference optimization
1.012026
Algorithms for Context Engineering in LLM Inference: Optimization of Placement, Compression, and Scheduling · AAAI 2026
Machine learning › Efficient and distributed learning › inference efficiency
memory-efficient inference
1.012026
Algorithms for Context Engineering in LLM Inference: Optimization of Placement, Compression, and Scheduling · AAAI 2026
Energy-efficient computing › energy-quality tradeoff
energy-accuracy tradeoff
1.012026
Algorithms for Context Engineering in LLM Inference: Optimization of Placement, Compression, and Scheduling · AAAI 2026

Methods — techniques the papers use, named apart from their topics

prefetch scheduling · 2.0compression · 2.0approximation algorithm · 2.0
YearPublicationVenuePosition
2026 Algorithms for Context Engineering in LLM Inference: Optimization of Placement, Compression, and Scheduling
abstract
Scaling long-context and agentic LLMs is increasingly limited by memory capacity and bandwidth rather than FLOPs. I propose an algorithmic framework for context engineering that models placement, compression, and scheduling as coupled optimization problems with explicit accuracy-efficiency trade-offs. Concretely, I aim to develop (1) salience-aware retention/eviction policies with provable approximation guarantees relative to an ideal oracle; (2) tier-dependent compression schemes that bound error propagation across memory levels; and (3) probabilistic prefetch/scheduling that controls tail latency. I will evaluate on long-context language modeling and reasoning benchmarks, isolating each component via ablations and comparing against heuristic baselines under controlled bandwidth/capacity regimes. Results target improved throughput and energy metrics at near-baseline quality, advancing principled, hardware-aware inference without requiring custom hardware.
Teresa Zhang
AAAI1
2026 Five-Minute Rule 40 Years Later: A First-Principles Revisit for Modern Memory Hierarchy
abstract
In 1987, Jim Gray and Gianfranco Putzolu introduced the five-minute rule, a simple, storage-memory-economics-based heuristic for deciding when data should live in DRAM rather than on storage. Subsequent revisits to the rule largely retained that economics-only view, leaving host costs, feasibility limits, and workload behavior out of scope. This paper revisits the rule from first principles, integrating host costs, DRAM bandwidth/capacity, and physics-grounded models of SSD performance and cost, and then embedding these elements in a constraint- and workload-aware framework that yields actionable provisioning guidance. We show that, for modern AI platforms, especially GPU-centric hosts paired with ultra-high-IOPS SSDs engineered for fine-grained random access, the DRAM$\leftrightarrow$flash caching threshold collapses from minutes to a few seconds. This shift reframes NAND flash memory as an \emph{active data tier} and exposes a broad research space across the hardware-software stack. We further introduce MQSim-Next, a calibrated SSD simulator that supports validation and sensitivity analysis and facilitates future architectural and system research. Finally, we present two concrete case studies that showcase the software system design space opened by such memory hierarchy paradigm shift. Overall, we turn a classical heuristic into an actionable, feasibility-aware analysis and provisioning framework and set the stage for further research on AI-era memory hierarchy.
Tong Zhang 0002, Vikram S. Mailthody, Linsen Ma, Chris J. Newburn, Teresa Zhang, Jiangpeng Li, Hao Zhong 0006, Wen-Mei W. Hwu
ISCA6