David Tarjan

dblp:61/2174 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 4 first-authorSoftware engineering, systems software and programming languages · 5Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Processor architecture and microarchitecture · 39% GPUs and heterogeneous computing · 25% Memory systems · 14%
Artificial intelligence
2 papers
Efficient and distributed learning · 73% Video understanding and tracking · 16% Language models and text generation · 11%

Topics — the 30 heaviest of 44, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression
0.812024
Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference · ICML 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference · ICML 2024
GPUs and heterogeneous computing
GPU architecture
0.432012
A Hierarchical Thread Scheduler and Register File for Energy-Efficient Throughput Processors · ACM Trans. Comput. Syst. 2012
The Sharing Tracker: Using Ideas from Cache Coherence Hardware to Reduce Off-Chip Memory Traffic with Non-Coherent Caches · SC 2010
Dynamic warp subdivision for integrated branch and memory divergence tolerance · ISCA 2010
Computer vision › Video understanding and tracking
video prediction
0.312018
SDC-Net: Video Prediction Using Spatially-Displaced Convolution · ECCV (7) 2018
Processor architecture and microarchitecture
register file
0.322012
A Hierarchical Thread Scheduler and Register File for Energy-Efficient Throughput Processors · ACM Trans. Comput. Syst. 2012
Energy-efficient mechanisms for managing thread context in throughput processors · ISCA 2011
Natural language and speech › Language models and text generation
large language model inference
0.212024
Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference · ICML 2024
Processor architecture and microarchitecture
out-of-order execution
0.222010
Federation: Boosting per-thread performance of throughput-oriented manycore architectures · ACM Trans. Archit. Code Optim. 2010
Federation: repurposing scalar cores for out-of-order instruction issue · DAC 2008
Processor architecture and microarchitecture › register file
hierarchical register file
0.112012
A Hierarchical Thread Scheduler and Register File for Energy-Efficient Throughput Processors · ACM Trans. Comput. Syst. 2012
Processor architecture and microarchitecture
throughput processor
0.112012
A Hierarchical Thread Scheduler and Register File for Energy-Efficient Throughput Processors · ACM Trans. Comput. Syst. 2012
GPUs and heterogeneous computing
GPU microarchitecture
0.112011
Energy-efficient mechanisms for managing thread context in throughput processors · ISCA 2011
Processor architecture and microarchitecture › register file
register caching
0.112011
Energy-efficient mechanisms for managing thread context in throughput processors · ISCA 2011
Parallel and multicore computing › parallel scheduling
thread scheduling
0.112011
Energy-efficient mechanisms for managing thread context in throughput processors · ISCA 2011
GPUs and heterogeneous computing › control flow divergence
branch divergence
0.112010
Dynamic warp subdivision for integrated branch and memory divergence tolerance · ISCA 2010
Memory systems
cache coherence
0.112010
The Sharing Tracker: Using Ideas from Cache Coherence Hardware to Reduce Off-Chip Memory Traffic with Non-Coherent Caches · SC 2010
Memory systems
cache design
0.112010
The Sharing Tracker: Using Ideas from Cache Coherence Hardware to Reduce Off-Chip Memory Traffic with Non-Coherent Caches · SC 2010
Processor architecture and microarchitecture
latency hiding
0.112010
Dynamic warp subdivision for integrated branch and memory divergence tolerance · ISCA 2010
Processor architecture and microarchitecture
many-core architecture
0.112010
Federation: Boosting per-thread performance of throughput-oriented manycore architectures · ACM Trans. Archit. Code Optim. 2010
GPUs and heterogeneous computing › GPU memory access
memory divergence
0.112010
Dynamic warp subdivision for integrated branch and memory divergence tolerance · ISCA 2010
GPUs and heterogeneous computing › GPU scheduling
warp scheduling
0.112010
Dynamic warp subdivision for integrated branch and memory divergence tolerance · ISCA 2010
Memory systems
cache
0.112009
Increasing memory miss tolerance for SIMD cores · SC 2009
Memory systems
memory access latency
0.112009
Increasing memory miss tolerance for SIMD cores · SC 2009
Processor architecture and microarchitecture
memory latency tolerance
0.112009
Increasing memory miss tolerance for SIMD cores · SC 2009
Processor architecture and microarchitecture
SIMD
0.112009
Increasing memory miss tolerance for SIMD cores · SC 2009
Energy-efficient computing › thermal management
dynamic thermal management
0.122004
Temperature-aware microarchitecture: Modeling and implementation · ACM Trans. Archit. Code Optim. 2004
Temperature-Aware Microarchitecture · ISCA 2003
Energy-efficient computing
thermal management
0.122004
Temperature-aware microarchitecture: Modeling and implementation · ACM Trans. Archit. Code Optim. 2004
Temperature-Aware Microarchitecture · ISCA 2003
Energy-efficient computing
thermal modeling
0.122004
Temperature-aware microarchitecture: Modeling and implementation · ACM Trans. Archit. Code Optim. 2004
Temperature-Aware Microarchitecture · ISCA 2003
Processor architecture and microarchitecture
multicore design
0.112008
Federation: repurposing scalar cores for out-of-order instruction issue · DAC 2008
Processor architecture and microarchitecture
superscalar processor
0.112008
Federation: repurposing scalar cores for out-of-order instruction issue · DAC 2008
Energy-efficient computing
datacenter energy efficiency
0.112014
Rhythm: harnessing data parallel hardware for server workloads · ASPLOS 2014
Processor architecture and microarchitecture
branch prediction
0.112005
Merging path and gshare indexing in perceptron branch prediction · ACM Trans. Archit. Code Optim. 2005

Methods — techniques the papers use, named apart from their topics

key-value eviction · 0.8grouped-query attention · 0.8continued pretraining · 0.8spatially-displaced convolution · 0.3simulation · 0.3data parallel hardware · 0.2execution-driven simulation · 0.1register file reuse · 0.1lookup table conversion · 0.1finite element simulation · 0.1path-based indexing · 0.1hashed perceptron · 0.1equivalent circuit modeling · 0.0equivalent circuit thermal model · 0.0
YearPublicationVenuePosition
2026 SmoothDiffusion-VE: Real-time Generative Video Editing Using Adaptive Feature Cache
abstract
Video editing with diffusion models presents significant challenges, especially under real-time constraints. Current methods either enhance temporal consistency at the cost of slow processing or rely on frame-by-frame editing, leading to flickering and temporal artifacts. To address both challenges, we propose SmoothDiffusion-VE, a streaming-based editing approach that improves temporal consistency and processing speed through our proposed Adaptive Feature Cache (AFC) and motion-guided attention. The AFC dynamically adjusts the caching behavior based on perceptual similarity (LPIPS) between frames, i.e., shifting to a mini-cache mode for similar frames to reduce computational load. Conversely, significant frame changes trigger deeper caching to maintain robust temporal coherence. Our motion-guided attention selectively focuses on dynamic regions using optical flow, reducing unnecessary computations in static areas and accelerating processing. SmoothDiffusion-VE can run 28 FPS on one RTX 4090 GPU, achieving a 1564× speedup over Plug-and-Play Diffusion (PNP) and a 1916× speedup over Diffusion Motion Transfer (DMT), delivering a powerful solution for fast and consistent video editing.
Mustafa Munir, Sophia Zalewski, Shiqiu Liu, David Tarjan, Sushmitha Belede, Anjul Patney, Radu Marculescu
WACV4
2024 Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference
abstract
Transformers have emerged as the backbone of large language models (LLMs). However, generation remains inefficient due to the need to store in memory a cache of key–value representations for past tokens, whose size scales linearly with the input sequence length and batch size. As a solution, we propose Dynamic Memory Compression (DMC), a method for on-line key–value cache compression at inference time. Most importantly, the model learns to apply different compression ratios in different heads and layers. We retrofit pre-trained LLMs such as Llama 2 (7B, 13B and 70B) into DMC Transformers, achieving up to $\sim 3.7 \times$ throughput increase during auto-regressive inference on an NVIDIA H100 GPU. DMC is applied via continued pre-training on a negligible percentage of the original data without adding any extra parameters. We find that DMC preserves the original downstream performance with up to 4$\times$ cache compression, outperforming up-trained grouped-query attention (GQA) and key–value eviction policies (H$_2$O, TOVA). GQA and DMC can be even combined to obtain compounded gains. As a result DMC fits longer contexts and larger batches within any given memory budget. We release the DMC code and models at https://github.com/NVIDIA/Megatron-LM/tree/DMC.
Piotr Nawrot, Adrian Lancucki, Marcin Chochowski, David Tarjan, Edoardo Maria Ponti
ICML4
2018 SDC-Net: Video Prediction Using Spatially-Displaced Convolution
Fitsum A. Reda, Guilin Liu, Kevin J. Shih, Robert Kirby 0001, Jon Barker, David Tarjan, Andrew Tao, Bryan Catanzaro
ECCV (7)6
2014 Rhythm: harnessing data parallel hardware for server workloads
abstract
Trends in increasing web traffic demand an increase in server throughput while preserving energy efficiency and total cost of ownership. Present work in optimizing data center efficiency primarily focuses on the data center as a whole, using off-the-shelf hardware for individual servers. Server capacity is typically increased by adding more machines, which is cheap, though inefficient in the long run in terms of energy and area.
Sandeep R. Agrawal, Valentin Pistol, Jun Pang 0001, David Tarjan, Alvin R. Lebeck
ASPLOS5
2012 A Hierarchical Thread Scheduler and Register File for Energy-Efficient Throughput Processors
abstract
Modern graphics processing units (GPUs) employ a large number of hardware threads to hide both function unit and memory access latency. Extreme multithreading requires a complex thread scheduler as well as a large register file, which is expensive to access both in terms of energy and latency. We present two complementary techniques for reducing energy on massively-threaded processors such as GPUs. First, we investigate a two-level thread scheduler that maintains a small set of active threads to hide ALU and local memory access latency and a larger set of pending threads to hide main memory latency. Reducing the number of threads that the scheduler must consider each cycle improves the scheduler’s energy efficiency. Second, we propose replacing the monolithic register file found on modern designs with a hierarchical register file. We explore various trade-offs for the hierarchy including the number of levels in the hierarchy and the number of entries at each level. We consider both a hardware-managed caching scheme and a software-managed scheme, where the compiler is responsible for orchestrating all data movement within the register file hierarchy. Combined with a hierarchical register file, our two-level thread scheduler provides a further reduction in energy by only allocating entries in the upper levels of the register file hierarchy for active threads. Averaging across a variety of real world graphics and compute workloads, the active thread count can be reduced by a factor of 4 with minimal impact on performance and our most efficient three-level software-managed register file hierarchy reduces register file energy by 54%.
Mark Gebhart, Daniel R. Johnson, David Tarjan, Stephen W. Keckler, William J. Dally, Erik Lindholm, Kevin Skadron
ACM Trans. Comput. Syst.3
2011 Energy-efficient mechanisms for managing thread context in throughput processors
abstract
Modern graphics processing units (GPUs) use a large number of hardware threads to hide both function unit and memory access latency. Extreme multithreading requires a complicated thread scheduler as well as a large register file, which is expensive to access both in terms of energy and latency. We present two complementary techniques for reducing energy on massively-threaded processors such as GPUs. First, we examine register file caching to replace accesses to the large main register file with accesses to a smaller structure containing the immediate register working set of active threads. Second, we investigate a two-level thread scheduler that maintains a small set of active threads to hide ALU and local memory access latency and a larger set of pending threads to hide main memory latency. Combined with register file caching, a two-level thread scheduler provides a further reduction in energy by limiting the allocation of temporary register cache resources to only the currently active subset of threads. We show that on average, across a variety of real world graphics and compute workloads, a 6-entry per-thread register file cache reduces the number of reads and writes to the main register file by 50% and 59% respectively. We further show that the active thread count can be reduced by a factor of 4 with minimal impact on performance, resulting in a 36% reduction of register file energy.
Mark Gebhart, Daniel R. Johnson, David Tarjan, Stephen W. Keckler, William J. Dally, Erik Lindholm, Kevin Skadron
ISCA3
2010 Dynamic warp subdivision for integrated branch and memory divergence tolerance
abstract
SIMD organizations amortize the area and power of fetch, decode, and issue logic across multiple processing units in order to maximize throughput for a given area and power budget. However, throughput is reduced when a set of threads operating in lockstep (a warp) are stalled due to long latency memory accesses. The resulting idle cycles are extremely costly. Multi-threading can hide latencies by interleaving the execution of multiple warps, but deep multi-threading using many warps dramatically increases the cost of the register files (multi-threading depth x SIMD width), and cache contention can make performance worse. Instead, intra-warp latency hiding should first be exploited. This allows threads that are ready but stalled by SIMD restrictions to use these idle cycles and reduces the need for multi-threading among warps. This paper introduces dynamic warp subdivision (DWS), which allows a single warp to occupy more than one slot in the scheduler without requiring extra register file space. Independent scheduling entities allow divergent branch paths to interleave their execution, and allow threads that hit to run ahead. The result is improved latency hiding and memory level parallelism (MLP). We evaluate the technique on a coherent cache hierarchy with private L1 caches and a shared L2 cache. With an area overhead of less than 1%, experiments with eight data-parallel benchmarks show our technique improves performance on average by 1.7X.
Jiayuan Meng, David Tarjan, Kevin Skadron
ISCA2
2010 The Sharing Tracker: Using Ideas from Cache Coherence Hardware to Reduce Off-Chip Memory Traffic with Non-Coherent Caches
abstract
Graphics Processing Units (GPUs) have recently emerged as a new platform for high performance, general-purpose computing. Because current GPUs employ deep multithreading to hide latency, they only have small, per-core caches to capture reuse and eliminate unnecessary off-chip accesses. This paper shows that for general-purpose workloads, the ability to copy cache lines between private caches captures inter-core temporal locality and provides substantial reductions in off-chip bandwidth requirements. Unlike hardware cache coherence, a sharing tracker only needs to track cache lines in the private caches imprecisely, because it is only a performance hint. This simplifies the implementation and is so effective at capturing inter-core reuse that the L2 can be eliminated entirely. The sharing tracker is motivated by but not specific to the GPU and could be used in other manycore organizations.
David Tarjan, Kevin Skadron
SC1
2010 Federation: Boosting per-thread performance of throughput-oriented manycore architectures
abstract
Manycore architectures designed for parallel workloads are likely to use simple, highly multithreaded, in-order cores. This maximizes throughput, but only with enough threads to keep hardware utilized. For applications or phases with more limited parallelism, we describe creating an out-of-order processor on-the-fly, by federating two neighboring in-order cores. We reuse the large register file in the multithreaded cores to implement some out-of-order structures and reengineer other large, associative structures into simpler lookup tables. The resulting federated core provides twice the single-thread performance of the underlying in-order core, allowing the architecture to efficiently support a wider range of parallelism.
Michael Boyer, David Tarjan, Kevin Skadron
ACM Trans. Archit. Code Optim.2
2009 Accelerating leukocyte tracking using CUDA: A case study in leveraging manycore coprocessors
abstract
The availability of easily programmable manycore CPUs and GPUs has motivated investigations into how to best exploit their tremendous computational power for scientific computing. Here we demonstrate how a systems biology application - detection and tracking of white blood cells in video microscopy - can be accelerated by 200times using a CUDA-capable GPU. Because the algorithms and implementation challenges are common to a wide range of applications, we discuss general techniques that allow programmers to make efficient use of a manycore GPU.
Michael Boyer, David Tarjan, Scott T. Acton, Kevin Skadron
IPDPS2
2009 Increasing memory miss tolerance for SIMD cores
abstract
Manycore processors with wide SIMD cores are becoming a popular choice for the next generation of throughput oriented architectures. We introduce a hardware technique called "diverge on miss" that allows SIMD cores to better tolerate memory latency for workloads with non-contiguous memory access patterns. Individual threads within a SIMD "warp" are allowed to slip behind other threads in the same warp, letting the warp continue execution even if a subset of threads are waiting on memory. Diverge on miss can either increase the performance of a given design by up to a factor of 3.14 for a single warp per core, or reduce the number of warps per core needed to sustain a given level of performance from 16 to 2 warps, reducing the area per core by 35%.
David Tarjan, Jiayuan Meng, Kevin Skadron
SC1
2008 Federation: repurposing scalar cores for out-of-order instruction issue
abstract
Future SoCs will contain multiple cores. For workloads with significant parallelism, prior work has shown the benefit of many small, multi-threaded, scalar cores. For workloads that require better single-thread performance, a dedicated, larger core can help but comes at a large opportunity cost in the number of scalar cores that could be provisioned instead. This paper proposes a way to repurpose a pair of scalar cores into a 2-way out-of-order issue core with minimal area overhead. "Federating" scalar cores in this way nevertheless achieves comparable performance to a dedicated out-of-order core and dissipates less power as well.
David Tarjan, Michael Boyer, Kevin Skadron
DAC1
2008 A performance study of general-purpose applications on graphics processors using CUDA
Shuai Che, Michael Boyer, Jiayuan Meng, David Tarjan, Jeremy W. Sheaffer, Kevin Skadron
J. Parallel Distributed Comput.4
2007 Impact of process variations on multicore performance symmetry
Eric Humenay, David Tarjan, Kevin Skadron
DATE2
2005 Merging path and gshare indexing in perceptron branch prediction
abstract
We introduce the hashed perceptron predictor, which merges the concepts behind the gshare, path-based and perceptron branch predictors. This predictor can achieve superior accuracy to a path-based and a global perceptron predictor, previously the most accurate dynamic branch predictors known in the literature. We also show how such a predictor can be ahead pipelined to yield one cycle effective latency. On the SPECint2000 set of benchmarks, the hashed perceptron predictor improves accuracy by up to 15.6% over a MAC-RHSP and 27.2% over a path-based neural predictor.
David Tarjan, Kevin Skadron
ACM Trans. Archit. Code Optim.1
2004 Temperature-aware microarchitecture: Modeling and implementation
abstract
With cooling costs rising exponentially, designing cooling solutions for worst-case power dissipation is prohibitively expensive. Chips that can autonomously modify their execution and power-dissipation characteristics permit the use of lower-cost cooling solutions while still guaranteeing safe temperature regulation. Evaluating techniques for thisdynamic thermal management(DTM), however, requires a thermal model that is practical for architectural studies.This paper describesHotSpot, an accurate yet fast and practical model based on an equivalent circuit of thermal resistances and capacitances that correspond to microarchitecture blocks and essential aspects of the thermal package. Validation was performed using finite-element simulation. The paper also introduces several effective methods for DTM: "temperature-tracking" frequency scaling, "migrating computation" to spare hardware units, and a "hybrid" policy that combines fetch gating with dynamic voltage scaling. The latter two achieve their performance advantage by exploiting instruction-level parallelism, showing the importance of microarchitecture research in helping control the growth of cooling costs.Modeling temperature at the microarchitecture level also shows that power metrics are poor predictors of temperature, that sensor imprecision has a substantial impact on the performance of DTM, and that the inclusion of lateral resistances for thermal diffusion is important for accuracy.
Kevin Skadron, Mircea R. Stan, Karthik Sankaranarayanan, Wei Huang 0004, Sivakumar Velusamy, David Tarjan
ACM Trans. Archit. Code Optim.6
2003 Temperature-Aware Microarchitecture
abstract
With power density and hence cooling costs rising exponentially, processor packaging can no longer be designed for the worst case, and there is an urgent need for runtime processor-level techniques that can regulate operating temperature when the package's capacity is exceeded. Evaluating such techniques, however, requires a thermal model that is practical for architectural studies.This paper describes HotSpot, an accurate yet fast model based on an equivalent circuit of thermal resistances and capacitances that correspond to microarchitecture blocks and essential aspects of the thermal package. Validation was performed using finite-element simulation. The paper also introduces several effective methods for dynamic thermal management (DTM): "temperature-tracking" frequency scaling, localized toggling, and migrating computation to spare hardware units. Modeling temperature at the microarchitecture level also shows that power metrics are poor predictors of temperature, and that sensor imprecision has a substantial impact on the performance of DTM.
Kevin Skadron, Mircea R. Stan, Wei Huang 0004, Sivakumar Velusamy, Karthik Sankaranarayanan, David Tarjan
ISCA6