Yunhe Shi

dblp:46/254 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
1since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Runtime systems and virtual machines · 61% Programming languages and type systems · 30% Compilers and program optimization · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Runtime systems and virtual machines
interpreter design
0.112008
Virtual machine showdown: Stack versus registers · ACM Trans. Archit. Code Optim. 2008
Programming languages and type systems
method dispatch
0.112008
Virtual machine showdown: Stack versus registers · ACM Trans. Archit. Code Optim. 2008
Runtime systems and virtual machines
virtual machine architecture
0.112008
Virtual machine showdown: Stack versus registers · ACM Trans. Archit. Code Optim. 2008
Compilers and program optimization
register allocation
0.012008
Virtual machine showdown: Stack versus registers · ACM Trans. Archit. Code Optim. 2008

Methods — techniques the papers use, named apart from their topics

switch dispatch · 0.1inline threading · 0.1direct threading · 0.1bytecode translation · 0.1
YearPublicationVenuePosition
2025 Vehicle self-positioning via Kalman filter using multi-station non-circular signals
Yunhe Shi, Zhongtian Yang
Signal Process.2
2008 Virtual machine showdown: Stack versus registers
abstract
Virtual machines (VMs) enable the distribution of programs in an architecture-neutral format, which can easily be interpreted or compiled. A long-running question in the design of VMs is whether a stack architecture or register architecture can be implemented more efficiently with an interpreter. We extend existing work on comparing virtual stack and virtual register architectures in three ways. First, our translation from stack to register code and optimization are much more sophisticated. The result is that we eliminate an average of more than 46% of executed VM instructions, with the bytecode size of the register machine being only 26% larger than that of the corresponding stack one. Second, we present a fully functional virtual-register implementation of the Java virtual machine (JVM), which supports Intel, AMD64, PowerPC and Alpha processors. This register VM supports inline-threaded, direct-threaded, token-threaded, and switch dispatch. Third, we present experimental results on a range of additional optimizations such as register allocation and elimination of redundant heap loads. On the AMD64 architecture the register machine using switch dispatch achieves an average speedup of 1.48 over the corresponding stack machine. Even using the more efficient inline-threaded dispatch, the register VM achieves a speedup of 1.15 over the equivalent stack-based VM.
Yunhe Shi, Kevin Casey, M. Anton Ertl, David Gregg
ACM Trans. Archit. Code Optim.1
2005 Virtual machine showdown: stack versus registers
abstract
Virtual machines (VMs) are commonly used to distribute programs in an architecture-neutral format, which can easily be interpreted or compiled. A long-running question in the design of VMs is whether stack architecture or register architecture can be implemented more efficiently with an interpreter. We extend existing work on comparing virtual stack and virtual register architectures in two ways. Firstly, our translation from stack to register code is much more sophisticated. The result is that we eliminate an average of more than 47% of executed VM instructions, with the register machine bytecode size only 25% larger than that of the corresponding stack bytecode. Secondly we present an implementation of a register machine in a fully standard-compliant implementation of the Java VM. We find that, on the Pentium 4, the register architecture requires an average of 32.3% less time to execute standard benchmarks if dispatch is performed using a C switch statement. Even if more efficient threaded dispatch is available (which requires labels as first class values), the reduction in running time is still approximately 26.5% for the register architecture.
Yunhe Shi, David Gregg, Andrew Beatty, M. Anton Ertl
VEE1