EDBT 2026 Demo / reviewers in the wild / expert
Shiliang Hu
dblp:62/4126
· DBLP profile ↗
11ranked-venue papers
3as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-authorSoftware engineering, systems software and programming languages · 6 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Memory systems · 60% Processor architecture and microarchitecture · 28% Parallel and multicore computing · 9% | |
| Software engineering, system software, and programming languages
4 papers |
Concurrent programming · 44% Runtime systems and virtual machines · 35% Operating systems · 21% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache coherence |
0.8 | 3 | 2017 | TMI: thread memory isolation for false sharing repair · MICRO 2017 Remix: online detection and repair of cache contention for the JVM · PLDI 2016 LASER: Light, Accurate Sharing dEtection and Repair · HPCA 2016 |
Memory systems › cache coherence
false sharing |
0.5 | 2 | 2017 | TMI: thread memory isolation for false sharing repair · MICRO 2017 LASER: Light, Accurate Sharing dEtection and Repair · HPCA 2016 |
Memory systems › memory interference
cache contention |
0.3 | 2 | 2016 | Remix: online detection and repair of cache contention for the JVM · PLDI 2016 LASER: Light, Accurate Sharing dEtection and Repair · HPCA 2016 |
Concurrent programming
concurrency bugs |
0.2 | 1 | 2016 | Remix: online detection and repair of cache contention for the JVM · PLDI 2016 |
Runtime systems and virtual machines › virtual machine implementation
java virtual machine |
0.2 | 1 | 2016 | Remix: online detection and repair of cache contention for the JVM · PLDI 2016 |
Parallel and multicore computing › parallel computing › parallel program debugging
deterministic replay |
0.2 | 2 | 2013 | QuickRec: prototyping an intel architecture extension for record and replay of multithreaded programs · ISCA 2013 CoreRacer: a practical memory race recorder for multicore x86 TSO processors · MICRO 2011 |
Processor architecture and microarchitecture › debugging support
hardware-assisted deterministic replay |
0.2 | 1 | 2013 | QuickRec: prototyping an intel architecture extension for record and replay of multithreaded programs · ISCA 2013 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2011 | CoreRacer: a practical memory race recorder for multicore x86 TSO processors · MICRO 2011 |
Processor architecture and microarchitecture
instruction set architecture |
0.1 | 2 | 2006 | Reducing Startup Time in Co-Designed Virtual Machines · ISCA 2006 An approach for implementing efficient superscalar CISC processors · HPCA 2006 |
Processor architecture and microarchitecture › multicore design
memory race recording |
0.1 | 1 | 2011 | CoreRacer: a practical memory race recorder for multicore x86 TSO processors · MICRO 2011 |
Processor architecture and microarchitecture
multicore design |
0.1 | 1 | 2011 | CoreRacer: a practical memory race recorder for multicore x86 TSO processors · MICRO 2011 |
Performance modeling and evaluation › performance monitoring
hardware performance counters |
0.1 | 1 | 2016 | Remix: online detection and repair of cache contention for the JVM · PLDI 2016 |
Processor architecture and microarchitecture › binary translation
dynamic binary translation |
0.1 | 1 | 2006 | Reducing Startup Time in Co-Designed Virtual Machines · ISCA 2006 |
Processor architecture and microarchitecture
instruction scheduling |
0.1 | 1 | 2006 | An approach for implementing efficient superscalar CISC processors · HPCA 2006 |
Parallel and multicore computing › thread-level parallelism
multithreaded applications |
0.0 | 1 | 2013 | QuickRec: prototyping an intel architecture extension for record and replay of multithreaded programs · ISCA 2013 |
Memory systems › memory consistency
memory consistency model |
0.0 | 1 | 2011 | CoreRacer: a practical memory race recorder for multicore x86 TSO processors · MICRO 2011 |
Memory systems › memory consistency › memory consistency model
total store order |
0.0 | 1 | 2011 | CoreRacer: a practical memory race recorder for multicore x86 TSO processors · MICRO 2011 |
Runtime systems and virtual machines › binary translation
dynamic binary translation |
0.0 | 1 | 2006 | An approach for implementing efficient superscalar CISC processors · HPCA 2006 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 2006 | Reducing Startup Time in Co-Designed Virtual Machines · ISCA 2006 |
Methods — techniques the papers use, named apart from their topics
runtime detection · 0.5JIT compilation · 0.5hardware prototyping · 0.3hardware performance counters · 0.2hardware race recording · 0.1hardware-software co-design · 0.1SPEC2000 benchmarking · 0.1superblock translation · 0.1hardware assists · 0.1basic block translation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | TMI: thread memory isolation for false sharing repairabstractCache contention in the form of false sharing and true sharing arises when threads overshare cache lines at high frequency. Such oversharing can reduce or negate the performance benefits of parallel execution. Prior systems for detecting and repairing cache contention lack efficiency in detection or repair, contain subtle memory consistency flaws, or require invasive changes to the program environment. Christian DeLozier, Ariel Eizenberg, Shiliang Hu, Gilles Pokam, Joseph Devietti |
MICRO | 3 |
| 2016 | LASER: Light, Accurate Sharing dEtection and RepairabstractContention for shared memory, in the forms of true sharing and false sharing, is a challenging performance bug to discover and to repair. Understanding cache contention requires global knowledge of the program's actual sharing behavior, and can even arise invisibly in the program due to the opaque decisions of the memory allocator. Previous schemes have focused only on false sharing, and impose significant performance penalties or require non-trivial alterations to the operating system or runtime system environment. This paper presents the Light, Accurate Sharing dEtection and Repair (LASER) system, which leverages new performance counter capabilities available on Intel's Haswell architecture that identify the source of expensive cache coherence events. Using records of these events generated by the hardware, we build a system for online contention detection and repair that operates with low performance overhead and does not require any invasive program, compiler or operating system changes. Our experiments show that LASER imposes just 2% average runtime overhead on the Phoenix, Parsec and Splash2x benchmarks. LASER can automatically improve the performance of programs by up to 19% on commodity hardware. Akshitha Sriraman, Brooke Fugate, Shiliang Hu, Gilles Pokam, Chris J. Newburn, Joseph Devietti |
HPCA | 4 |
| 2016 | Remix: online detection and repair of cache contention for the JVMabstractAs ever more computation shifts onto multicore architectures, it is increasingly critical to find effective ways of dealing with multithreaded performance bugs like true and false sharing. Previous approaches to fixing false sharing in unmanaged languages have employed highly-invasive runtime program modifications. We observe that managed language runtimes, with garbage collection and JIT code compilation, present unique opportunities to repair such bugs directly, mirroring the techniques used in manual repairs. We present Remix, a modified version of the Oracle HotSpot JVM which can detect cache contention bugs and repair false sharing at runtime. Remix's detection mechanism leverages recent performance counter improvements on Intel platforms, which allow for precise, unobtrusive monitoring of cache contention at the hardware level. Remix can detect and repair known false sharing issues in the LMAX Disruptor high-performance inter-thread messaging library and the Spring Reactor event-processing framework, automatically providing 1.5-2x speedups over unoptimized code and matching the performance of hand-optimization. Remix also finds a new false sharing bug in SPECjvm2008, and uncovers a true sharing bug in the HotSpot JVM that, when fixed, improves the performance of three NAS Parallel Benchmarks by 7-25x. Remix incurs no statistically-significant performance overhead on other benchmarks that do not exhibit cache contention, making Remix practical for always-on use. Ariel Eizenberg, Shiliang Hu, Gilles Pokam, Joseph Devietti |
PLDI | 2 |
| 2013 | QuickRec: prototyping an intel architecture extension for record and replay of multithreaded programsabstractThere has been significant interest in hardware-assisted deterministic Record and Replay (RnR) systems for multithreaded programs on multiprocessors. However, no proposal has implemented this technique in a hardware prototype with full operating system support. Such an implementation is needed to assess RnR practicality. Gilles Pokam, Klaus Danne, Cristiano Pereira, Rolf Kassa, Tim Kranich, Shiliang Hu, Justin Emile Gottschlich, Nima Honarmand, Nathan Dautenhahn, Samuel T. King, Josep Torrellas |
ISCA | 6 |
| 2011 | A HW/SW co-designed heterogeneous multi-core virtual machine for energy-efficient general purpose computingabstractIt is increasingly challenging to improve single thread performance because power/energy consumption becomes a major barrier to achieve significantly higher performance for general purpose cores. General purpose processors are designed to perform well in a wide variety of market segments, at the cost of having significantly lower performance-per-watt than special purpose processors targeting limited applications or market segments. In this paper, we propose a HW/SW co-designed heterogeneous multi-core virtual machine, called TwinPeaks, which integrates a set of less general but power efficient cores and uses dynamic binary optimization to schedule code regions to run on the most efficient cores. Our experiment and analysis indicate that TwinPeaks with a wide in-order core and a narrow out-of-order core may achieve 108% performance at ~71% energy of a big 4-wide out-of-order core. Youfeng Wu, Shiliang Hu, Edson Borin, Cheng Wang 0013 |
CGO | 2 |
| 2011 | CoreRacer: a practical memory race recorder for multicore x86 TSO processorsabstractShared memory multiprocessors are difficult to program because of the non-deterministic ways in which the memory operations from different threads interleave. To address this issue, many hardware-based memory race recorders have been proposed that efficiently log an ordering of the shared memory interleavings between threads for deterministic replay. These approaches are challenging to integrate into current processors because they change the cache subsystem or the coherence protocol, and they mostly support a sequentially consistent memory model. Gilles Pokam, Cristiano Pereira, Shiliang Hu, Ali-Reza Adl-Tabatabai, Justin Emile Gottschlich, Youfeng Wu |
MICRO | 3 |
| 2010 | TAO: two-level atomicity for dynamic binary optimizationsabstractDynamic binary translation is a key component of Hardware/Software (HW/SW) co-design, which is an enabling technology for processor microarchitecture innovation. There are two well-known dynamic binary optimization techniques based on atomic execution support. Frame-based optimizations leverage processor pipeline support to enable atomic execution of hot traces. Region level optimizations employ transactional-memory-like atomicity support to aggressively optimize large regions of code. In this paper we propose a two-level atomic optimization scheme which not only overcomes the limitations of the two approaches, but also boosts the benefits of the two approaches effectively. Our experiment shows that the combined approach can achieve a total of 21.5% performance improvement over an aggressive out-of-order baseline machine and improve the performance over the frame-based approach by an additional 5.3%. Edson Borin, Youfeng Wu, Cheng Wang 0013, Wei Liu 0014, Maurício Breternitz, Shiliang Hu, Esfir Natanzon, Shai Rotem, Roni Rosner |
CGO | 6 |
| 2009 | Dynamic parallelization of single-threaded binary programs using speculative slicingabstractThe performance of single-threaded programs and legacy binary code is of critical importance in many everyday applications. However, neither can hardware multi-core processors directly speed up single-threaded programs, nor can software automatic parallelizing compilers effectively parallelize legacy binary code and irregular applications. In this paper, we propose a framework and a set of algorithms to dynamically parallelize single-threaded binary programs. Our parallelization is based on program slicing and explores both instruction-level parallelism (ILP) and thread-level parallelism (TLP). To significantly reduce the critical path of the parallel slices, our slicing algorithms exploit speculation to cut rare dependences, and use well-designed program transformations to expose parallelism. Furthermore, because we transparently parallelize binary code at runtime, we perform slicing only on program hot regions. Our experiments demonstrate that the proposed speculative slicing approach extracts more parallelism than any known slicing based parallelization schemes. For the SPEC2000 benchmarks, we can achieve 3x parallelism with infinite number of threads, and 1.8x parallelism with 4 threads. Cheng Wang 0013, Youfeng Wu, Edson Borin, Shiliang Hu, Wei Liu 0014, Dave Sager, Tin-Fook Ngai, Jesse Fang |
ICS | 4 |
| 2006 | An approach for implementing efficient superscalar CISC processorsabstractAn integrated, hardware/software co-designed CISC processor is proposed and analyzed. The objectives are high performance and reduced complexity. Although the x86 ISA is targeted, the overall approach is applicable to other CISC ISAs. To provide high performance on frequently executed code sequences, fully transparent dynamic translation software decomposes CISC superblocks into RISC-style micro-ops. Then, pairs of dependent micro-ops are reordered and fused into macro-ops held in a large, concealed code cache. The macro-ops are fetched from the code cache and processed throughout the pipeline as single units. Consequently, instruction level communication and management are reduced, and processor resources such as the issue buffer and register file ports are better utilized. Moreover, fused instructions lead naturally to pipelined instruction scheduling (issue) logic, and collapsed 3-1 ALUs can be used, resulting in much simplified result forwarding logic. Steady state performance is evaluated for the SPEC2000 benchmarks, and a proposed x86 implementation with complexity similar to a two-wide superscalar processor is shown to provide performance (instructions per cycle) that is equivalent to a conventional four-wide superscalar processor. Shiliang Hu, Ilhyun Kim, Mikko H. Lipasti, James E. Smith 0001 |
HPCA | 1 |
| 2006 | Reducing Startup Time in Co-Designed Virtual MachinesabstractA Co-Designed Virtual Machine allows designers to implement a processor via a combination of hardware and software. Dynamic binary translation converts code written for a conventional (legacy) ISA into optimized code for an underlying implementation-specific ISA. Because translation is done dynamically, an important consideration in such systems is the startup time for performing the initial translations. Beginning with a previously proposed co-designed VM that implements the x86 ISA, we study runtime binary translation overhead effects. The co-designed x86 virtual machine is based on an adaptive translation system that uses a basic block translator for initial emulation and a superblock translator for hotspot optimization. We analyze and model VM startup performance via simulation. We observe that non-hotspot emulation via basic block translation is the major part of the startup overhead. To reduce startup translation overhead, we follow the co-designed hardware / software philosophy and propose hardware assists to dramatically accelerate basic block translations. By combining hardware assists with balanced translation strategies, the co-designed translation system reduces runtime overhead significantly and demonstrates very competitive startup performance when compared with conventional processors running a set of Windows application benchmarks. Shiliang Hu, James E. Smith 0001 |
ISCA | 1 |
| 2004 | Using Dynamic Binary Translation to Fuse Dependent InstructionsabstractInstruction scheduling hardware can be simplified and easily pipelined if pairs of dependent instructions are fused so they share a single instruction scheduling slot. We study an implementation of the x86 ISA that dynamically translates x86 code to an underlying ISA that supports instruction fusing. A microarchitecture that is codesigned with the fused instruction set completes the implementation. We focus on the dynamic binary translator for such a codesigned x86 virtual machine. The dynamic binary translator first cracks x86 instructions belonging to hot superblocks into RISC-style microoperations, and then uses heuristics to fuse together pairs of dependent microoperations. Experimental results with SPEC2000 integer benchmarks demonstrate that: (1) the fused ISA with dynamic binary translation reduces the number of scheduling decisions by about 30% versus a conventional implementation that uses hardware cracking into RISC microoperations; (2) an instruction scheduling slot needs only hold two source register fields even though it may hold two instructions; (3) translations generated in the proposed ISA consume about 30% less storage than a corresponding fixed-length RISC-style ISA. Shiliang Hu, James E. Smith 0001 |
CGO | 1 |