EDBT 2026 Demo / reviewers in the wild / expert
Andrew D. Hilton
dblp:98/1014
· DBLP profile ↗
12ranked-venue papers
6as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-authorSoftware engineering, systems software and programming languages · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Processor architecture and microarchitecture · 56% Memory systems · 29% Energy-efficient computing · 12% | |
| Network and information security
1 paper |
Hardware security and side channels · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 30 heaviest of 37, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing
thermal management |
0.4 | 1 | 2019 | DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal Management · MICRO 2019 |
Processor architecture and microarchitecture
out-of-order execution |
0.3 | 3 | 2010 | BOLT: Energy-efficient Out-of-Order Latency-Tolerant execution · HPCA 2010 Decoupled store completion/silent deterministic replay: enabling scalable data memory for CPR/CFP processors · ISCA 2009 Ginger: control independence using tag rewriting · ISCA 2007 |
Hardware security and side channels › memory integrity
memory integrity verification |
0.2 | 1 | 2016 | PoisonIvy: Safe speculation for secure memory · MICRO 2016 |
Hardware security and side channels › trusted execution environments
secure processor |
0.2 | 1 | 2016 | PoisonIvy: Safe speculation for secure memory · MICRO 2016 |
Compilers and program optimization › instruction scheduling
compile-time scheduling |
0.2 | 1 | 2016 | Decoupling Loads for Nano-Instruction Set Computers · ISCA 2016 |
Compilers and program optimization
instruction scheduling |
0.2 | 1 | 2016 | Decoupling Loads for Nano-Instruction Set Computers · ISCA 2016 |
Processor architecture and microarchitecture
instruction set architecture |
0.2 | 1 | 2016 | Decoupling Loads for Nano-Instruction Set Computers · ISCA 2016 |
Memory systems
integrity tree |
0.2 | 1 | 2016 | PoisonIvy: Safe speculation for secure memory · MICRO 2016 |
Processor architecture and microarchitecture › instruction set architecture
ISA extension |
0.2 | 1 | 2016 | Decoupling Loads for Nano-Instruction Set Computers · ISCA 2016 |
Memory systems
memory encryption |
0.2 | 1 | 2016 | PoisonIvy: Safe speculation for secure memory · MICRO 2016 |
Processor architecture and microarchitecture › out-of-order execution
out-of-order processor |
0.1 | 1 | 2012 | Flexible register management using reference counting · HPCA 2012 |
Processor architecture and microarchitecture › register file
physical register file |
0.1 | 1 | 2012 | Flexible register management using reference counting · HPCA 2012 |
Memory systems › memory management
reference counting |
0.1 | 1 | 2012 | Flexible register management using reference counting · HPCA 2012 |
Processor architecture and microarchitecture
register management |
0.1 | 1 | 2012 | Flexible register management using reference counting · HPCA 2012 |
Memory systems
cache |
0.1 | 1 | 2019 | DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal Management · MICRO 2019 |
Memory systems › cache management
cache capacity management |
0.1 | 1 | 2019 | DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal Management · MICRO 2019 |
Memory systems › memory hierarchy › cache hierarchy
last-level cache |
0.1 | 1 | 2019 | DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal Management · MICRO 2019 |
Memory systems › cache management
cache miss handling |
0.1 | 1 | 2009 | iCFP: Tolerating all-level cache misses in in-order processors · HPCA 2009 |
Memory systems › cache
cache miss tolerance |
0.1 | 1 | 2009 | iCFP: Tolerating all-level cache misses in in-order processors · HPCA 2009 |
Processor architecture and microarchitecture › checkpoint-based microarchitecture
checkpoint processing and recovery |
0.1 | 1 | 2009 | Decoupled store completion/silent deterministic replay: enabling scalable data memory for CPR/CFP processors · ISCA 2009 |
Processor architecture and microarchitecture › pipelining
continual flow pipeline |
0.1 | 1 | 2009 | iCFP: Tolerating all-level cache misses in in-order processors · HPCA 2009 |
Parallel and multicore computing › parallel computing › parallel program debugging
deterministic replay |
0.1 | 1 | 2009 | Decoupled store completion/silent deterministic replay: enabling scalable data memory for CPR/CFP processors · ISCA 2009 |
Processor architecture and microarchitecture › microprocessor design › processor core design
in-order core |
0.1 | 1 | 2009 | iCFP: Tolerating all-level cache misses in in-order processors · HPCA 2009 |
Processor architecture and microarchitecture › out-of-order execution
instruction window |
0.1 | 1 | 2009 | Decoupled store completion/silent deterministic replay: enabling scalable data memory for CPR/CFP processors · ISCA 2009 |
Processor architecture and microarchitecture › load/store queue
store buffer |
0.1 | 1 | 2009 | Decoupled store completion/silent deterministic replay: enabling scalable data memory for CPR/CFP processors · ISCA 2009 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.1 | 1 | 2016 | Decoupling Loads for Nano-Instruction Set Computers · ISCA 2016 |
Processor architecture and microarchitecture
speculation |
0.1 | 1 | 2016 | PoisonIvy: Safe speculation for secure memory · MICRO 2016 |
Processor architecture and microarchitecture › branch prediction
branch misprediction recovery |
0.1 | 1 | 2007 | Ginger: control independence using tag rewriting · ISCA 2007 |
Processor architecture and microarchitecture
branch prediction |
0.1 | 1 | 2007 | Ginger: control independence using tag rewriting · ISCA 2007 |
Processor architecture and microarchitecture › instruction-level parallelism
control independence |
0.1 | 1 | 2007 | Ginger: control independence using tag rewriting · ISCA 2007 |
Methods — techniques the papers use, named apart from their topics
decoupled loads · 0.5compiler scheduling · 0.5thermal headroom modeling · 0.4dynamic utility prediction · 0.4priority encoder · 0.1bit-matrix reference counting · 0.1microarchitecture simulation · 0.1simulation · 0.1cycle-level simulation · 0.1out-of-order renaming · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal ManagementabstractSprinting is a class of mechanisms that provides a short but significant performance boost while temporarily exceeding the thermal design point. We propose DynaSprint, a software runtime that manages sprints by dynamically predicting utility and modeling thermal headroom. Moreover, we propose a new sprint mechanism for caches, increasing capacity briefly for enhanced performance. For a system that extends last-level cache capacity from 2MB to 4MB per core and can absorb 10J of heat, DynaSprint-guided cache sprints improve performance by 17% on average and by up to 40% over a non-sprinting system. These performance outcomes, within 95% of an oracular policy, are possible because DynaSprint accurately predicts phase behavior and sprint utility. Ziqiang Huang, José A. Joao, Alejandro Rico, Andrew D. Hilton, Benjamin C. Lee |
MICRO | 4 |
| 2018 | MAPS: Understanding Metadata Access Patterns in Secure MemoryabstractSecure memory increases both the latency and energy required for memory accesses. To reduce these overheads, computer architects have sought to cache metadata on the processor chip, but placing metadata in a simple cache has not been as effective as expected. With a detailed analysis of metadata access patterns, we clarify myths in metadata caching and provide insight into more efficient caching strategies. We provide three observations that can help architects design future metadata caches. First, caching all metadata types improves efficiency. Second, the size of the metadata cache should match the reuse distance of the metadata. Third, when designing a better eviction policy, the traditional Belady's MIN algorithm cannot be used as the optimal replacement policy. Tamara Silbergleit Lehman, Andrew D. Hilton, Benjamin C. Lee |
ISPASS | 2 |
| 2018 | A technique for translation from problem to codeabstractStudents in introductory programming courses struggle with how to turn a problem statement into code. We introduce a technique, ``The Seven Steps,'' that provides structure and guidance on how to approach a problem. The first four steps focus on devising an algorithm in words, then the remaining steps are to translate that algorithm to code, test the algorithm, and debug failed test cases. This approach not only gives students a way to solve problems, but also ideas for what to do if they get stuck during the process. Furthermore, it provides a way for instructors to work examples in class that focus on the process of devising the code---instructors can show how to come up with the code, rather than just showing an example. We have used this technique in several introductory programming courses---both in the classroom and online. We describe this technique and results from its use in fall 2017 courses. Andrew D. Hilton, Genevieve M. Lipp, Susan H. Rodger |
ITiCSE | 1 |
| 2016 | Decoupling Loads for Nano-Instruction Set ComputersabstractWe propose an ISA extension that decouples the data access and register write operations in a load instruction. We describe system and hardware support for decoupled loads. Furthermore, we show how compilers can generate better static instruction schedules by hoisting a decoupled load's data access above may-alias stores and branches. We find that decoupled loads improve performance with geometric mean speedups of 8.4%. Ziqiang Huang, Andrew D. Hilton, Benjamin C. Lee |
ISCA | 2 |
| 2016 | PoisonIvy: Safe speculation for secure memoryabstractEncryption and integrity trees guard against physical attacks, but harm performance. Prior academic work has speculated around the latency of integrity verification, but has done so in an insecure manner. No industrial implementations of secure processors have included speculation. This work presents PoisonIvy, a mechanism which speculatively uses data before its integrity has been verified while preserving security and closing address-based side-channels. PoisonIvy reduces performance overheads from 40% to 20% for memory intensive workloads and down to 1.8%, on average. Tamara Silbergleit Lehman, Andrew D. Hilton, Benjamin C. Lee |
MICRO | 2 |
| 2015 | Multi-program benchmark definitionabstractAlthough definition of single-program benchmarks is relatively straight-forward-a benchmark is a program plus a specific input-definition of multi-program benchmarks is more complex. Each program may have a different runtime and they may have different interactions depending on how they align with each other. While prior work has focused on sampling multiprogram benchmarks, little attention has been paid to defining the benchmarks in their entirety. In this work, we propose a four-tuple that formally defines multi-program benchmarks in a well-defined way. We then examine how four different classes of benchmarks created by varying the elements of this tuple align with real-world use-cases. We evaluate the impact of these variations on real hardware, and see drastic variations in results between different benchmarks constructed from the same programs. Notable differences include significant speedups versus slowdowns (e.g., +57% vs -5% or +26% vs -18%), and large differences in magnitude even when the results are in the same direction (e.g., 67% versus 11%). Adam N. Jacobvitz, Andrew D. Hilton, Daniel J. Sorin |
ISPASS | 2 |
| 2012 | Flexible register management using reference countingabstractConventional out-of-order processors that use a unified physical register file allocate and reclaim registers explicitly using a free list that operates as a circular queue. We describe and evaluate a more flexible register management scheme - reference counting. We implement reference counting using a bit-matrix with a column for every physical register and a row for every entity that can hold a physical register, e.g., an in-flight instruction. Columns are NOR'ed together to create a bitvector free list from which registers are allocated using priority encoders. We describe reference counting designs that support micro-architectural techniques including register file power gating, dynamic register move elimination, register file checkpointing, and latency tolerant execution. Performance and circuit simulation show that the energy cost of reference counting is low and is easily recouped by the savings of the techniques it enables. Steven J. Battle, Andrew D. Hilton, Mark Hempstead, Amir Roth |
HPCA | 2 |
| 2010 | BOLT: Energy-efficient Out-of-Order Latency-Tolerant executionabstractLT (latency tolerant) execution is an attractive candidate technique for future out-of-order cores. LT defers the forward slices of LLC (last-level cache) misses to a slice buffer and re-executes them when the misses return. An LT core increases ILP without physically scaling the issue queue and register file and increases MLP without additional software threads that can reduce cache performance. Unfortunately, proposed LT designs are not energy efficient. They require too many additional structures and they defer and re-execute too many instructions to justify their performance gains. In this paper, we address these inefficiencies. We introduce a microarchitecture called BOLT (Better Out-of-Order Latency-Tolerance) that implements LT as an alternative use of SMT (Simultaneous Multi-Threading). We also present a new slice buffer organization and traversal scheme that increases performance and reduces overhead by pruning instances of useless and redundant LT. Collectively, these modifications turn out-of-order LT into a technique that improves performance in an energy-efficient way. Andrew D. Hilton, Amir Roth |
HPCA | 1 |
| 2009 | CPROB: Checkpoint Processing with Opportunistic Minimal RecoveryabstractCPR (Checkpoint Processing and Recovery) is a physical register management scheme that supports a larger instruction window and higher average IPC than conventional ROB-style register management. It does so by restricting mis-speculation recovery to checkpoints created at rename, and leveraging this restriction to aggressively reclaim registers that don't appear in checkpoints. The cost of CPR is checkpoint overhead, which is incurred when a mis-speculation occurs on an instruction for which a checkpoint was not created a priori. Here, CPR must recover to the immediately older checkpoint, squashing instructions older than the mis-speculation itself. In contrast, a ROB processor performs minimal recovery and only squashes instructions younger than the mis-speculation. CPROB is a hybrid register management scheme that preserves CPR's aggressive reclamation while opportunistically minimizing checkpoint overhead. CPROB extends CPR to track and hold the registers needed to perform minimal recovery to un-executed branches within each checkpoint. Recovery registers are held on a best-effort basis only. A checkpoint's recovery registers can be freed spontaneously when all branches in the checkpoint execute. They can also be aggressively victimized if dispatch needs registers to proceed. CPROB naturally adapts the register reclamation policy to dynamic branch behavior. When branch mis-predictions are infrequent and registers are needed to support a large window, CPROB victimizes registers and behaves like CPR. When mis-predictions are frequent and the window is small, CPROB holds on to registers and behaves like ROB. As a result, it out-performs both CPR and ROB for a given program. This performance improvement, combined with reduced checkpoint overhead, makes CPROB more energy-efficient than either ROB or CPR. Andrew D. Hilton, Neeraj Eswaran, Amir Roth |
PACT | 1 |
| 2009 | iCFP: Tolerating all-level cache misses in in-order processorsabstractGrowing concerns about power have revived interest in in-order pipelines. In-order pipelines sacrifice single-thread performance. Specifically, they do not allow execution to flow freely around data cache misses. As a result, they have difficulties overlapping independent misses with one another. Previously proposed techniques like Runahead execution and Multipass pipelining have attacked this problem. In this paper, we go a step further and introduce iCFP (in-order Continual Flow Pipeline), an adaptation of the CFP concept to an in-order processor. When iCFP encounters a primary data cache or 12 miss, it checkpoints the register file and transitions into an "advance " execution mode. Miss-independent instructions execute as usual and even update register state. Miss- dependent instructions are diverted into a slice buffer, un-blocking the pipeline latches. When the miss returns, iCFP "rallies" and executes the contents of the slice buffer, merging miss-dependent state with miss- independent state along the way. An enhanced register dependence tracking scheme and a novel store buffer design facilitate the merging process. Cycle-level simulations show that iCFP out-performs Runahead, Multipass, and SLTP, another non-blocking in-order pipeline design. Andrew D. Hilton, Santosh Nagarakatte, Amir Roth |
HPCA | 1 |
| 2009 | Decoupled store completion/silent deterministic replay: enabling scalable data memory for CPR/CFP processorsabstractCPR/CFP (Checkpoint Processing and Recovery/Continual Flow Pipeline) support an adaptive instruction window that scales to tolerate last-level cache misses. CPR/CFP scale the register file by aggressively reclaiming the destination registers of many in-flight instructions. However, an analogous mechanism does not exist for stores and loads. As the window expands, CPR/CFP processors must track all in-flight stores and loads to support forwarding and detect memory ordering violations. Andrew D. Hilton, Amir Roth |
ISCA | 1 |
| 2007 | Ginger: control independence using tag rewritingabstractThe negative performance impact of branch mis-predictions can be reduced by exploiting control independence (CI). When a branch mis-predicts, the wrong-path instructions up to the point where control converges with the correct path are selectively squashed and replaced with correct-path instructions. Instructions beyond the convergence-point-the branch's control-independent (CI) instructions-are spared from squashing. Exploiting CI requires updating the input data dependences of CI instructions to reflect the selective removal and insertion of logically older instructions and transitively re-dispatching those CI instructions whose inputs have changed. This capability is generally called out-of-order renaming. Previously proposed CI designs use out-of-order renaming schemes that either consume excessive rename/dispatch bandwidth, can only be applied in limited cases, or incur a cost even when the branch would be correctly predicted. Andrew D. Hilton, Amir Roth |
ISCA | 1 |