Xiameng Hu

dblp:164/7904 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Memory systems · 68% Performance modeling and evaluation · 11% Storage systems · 8%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache management
0.932018
Fast Miss Ratio Curve Modeling for Storage Cache · ACM Trans. Storage 2018
Optimal Symbiosis and Fair Scheduling in Shared Cache · IEEE Trans. Parallel Distributed Syst. 2017
Kinetic Modeling of Data Eviction in Cache · USENIX ATC 2016
Memory systems › memory management
memory allocation
0.522017
Optimizing Locality-Aware Memory Management of Key-Value Caches · IEEE Trans. Computers 2017
LAMA: Optimized Locality-aware Memory Allocation for Key-value Cache · USENIX ATC 2015
Performance modeling and evaluation
cache performance modeling
0.422018
Fast Miss Ratio Curve Modeling for Storage Cache · ACM Trans. Storage 2018
Kinetic Modeling of Data Eviction in Cache · USENIX ATC 2016
Memory systems › cache › cache performance
miss ratio curve
0.312018
Fast Miss Ratio Curve Modeling for Storage Cache · ACM Trans. Storage 2018
Memory systems › memory referencing behavior
reuse distance
0.312018
Fast Miss Ratio Curve Modeling for Storage Cache · ACM Trans. Storage 2018
Distributed systems › resource sharing
application co-location
0.312017
Optimal Symbiosis and Fair Scheduling in Shared Cache · IEEE Trans. Parallel Distributed Syst. 2017
Memory systems › cache management
cache interference
0.312017
Optimal Symbiosis and Fair Scheduling in Shared Cache · IEEE Trans. Parallel Distributed Syst. 2017
Storage systems
key-value storage
0.312017
Optimizing Locality-Aware Memory Management of Key-Value Caches · IEEE Trans. Computers 2017
Memory systems › cache management › storage caching
storage cache
0.112018
Fast Miss Ratio Curve Modeling for Storage Cache · ACM Trans. Storage 2018
Processor architecture and microarchitecture
chip multiprocessor
0.112017
Optimal Symbiosis and Fair Scheduling in Shared Cache · IEEE Trans. Parallel Distributed Syst. 2017
Cloud and datacenter computing
quality of service
0.112017
Optimizing Locality-Aware Memory Management of Key-Value Caches · IEEE Trans. Computers 2017
Memory systems › cache
key-value cache
0.112015
LAMA: Optimized Locality-aware Memory Allocation for Key-value Cache · USENIX ATC 2015

Methods — techniques the papers use, named apart from their topics

sampling · 0.6kinetic modeling · 0.6locality analysis · 0.6optimization theory · 0.3
YearPublicationVenuePosition
2018 Fast Miss Ratio Curve Modeling for Storage Cache
abstract
The reuse distance (least recently used (LRU) stack distance) is an essential metric for performance prediction and optimization of storage cache. Over the past four decades, there have been steady improvements in the algorithmic efficiency of reuse distance measurement. This progress is accelerating in recent years, both in theory and practical implementation. In this article, we present a kinetic model of LRU cache memory, based on the average eviction time (AET) of the cached data. The AET model enables fast measurement and use of low-cost sampling. It can produce the miss ratio curve in linear time with extremely low space costs. On storage trace benchmarks, AET reduces the time and space costs compared to former techniques. Furthermore, AET is a composable model that can characterize shared cache behavior through sampling and modeling individual programs or traces.
Xiameng Hu, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003, Chen Ding 0001, Chencheng Ye 0001
ACM Trans. Storage1
2017 Optimizing Locality-Aware Memory Management of Key-Value Caches
abstract
The in-memory cache system is a performance-critical layer in today's web server architectures. Memcached is one of the most effective, representative, and prevalent among such systems. An important problem is on its memory allocation. The default design does not make the best use of the memory. It is unable to adapt when the demand changes, a problem known as slab calcification. This paper introduces locality-aware memory allocation (LAMA), which addresses the problem by first analyzing the locality of Memcached's requests and then reassigning slabs to minimize the miss ratio or the average response time. By evaluating LAMA using various industry and academic workloads, the paper shows that LAMA outperforms existing techniques in the steady-state performance, the speed of convergence, and the ability to adapt to request pattern changes, and overcome slab calcification. The new solution is close to optimal, achieving over 98 percent of the theoretical potential. Furthermore, LAMA can also be adopted in resource partitioning to guarantee quality-of-service (QoS).
Xiameng Hu, Xiaolin Wang 0001, Yingwei Luo, Chen Ding 0001, Song Jiang 0001, Zhenlin Wang 0003
IEEE Trans. Computers1
2017 Optimal Symbiosis and Fair Scheduling in Shared Cache
abstract
On multi-core processors, applications are run sharing the cache. This paper presents optimization theory to co-locate applications to minimize cache interference and maximize performance. The theory precisely specifies MRC-based composition, optimization, and correctness conditions. The paper also presents a new technique called footprint symbiosis to obtain the best shared cache performance underfair CPU allocation as well as a new sampling technique which reduces the cost of locality analysis. When sampling and optimization are combined, the paper shows that it takes less than 0.1 second analysis per program to obtain a co-run that is within 1.5 percent of the best possible performance. In an exhaustive evaluation with 12,870 tests, the best prior work improves co-run performance by 56 percent on average. The new optimization improves it by another 29 percent. Without single co-run test, footprint symbiosis is able to choose co-run choices that are just 8 percent slower than the best co-run solutions found with exhaustive testing.
Xiameng Hu, Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Chen Ding 0001, Zhenlin Wang 0003
IEEE Trans. Parallel Distributed Syst.1
2016 Kinetic Modeling of Data Eviction in Cache
Xiameng Hu, Xiaolin Wang 0001, Yingwei Luo, Chen Ding 0001, Zhenlin Wang 0003
USENIX ATC1
2015 Optimal Footprint Symbiosis in Shared Cache
abstract
On multicore processors, applications are run sharing the cache. This paper presents online optimization to collocate applications to minimize cache interference to maximize performance. The paper formulates the optimization problem and solution, presents a new sampling technique for locality analysis and evaluates it in an exhaustive test of 12,870 cases. For locality analysis, previous sampling was two orders of magnitude faster than full-trace analysis. The new sampling reduces the cost by another two orders of magnitude. The best prior work improves co-run performance by 56% on average. The new optimization improves it by another 29%. When sampling and optimization are combined, the paper shows that it takes less than 0.1 second analysis per program to obtain a co-run that is within 1.5% of the best possible performance.
Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Xiameng Hu, Jacob Brock, Chen Ding 0001, Zhenlin Wang 0003
CCGRID4
2015 LAMA: Optimized Locality-aware Memory Allocation for Key-value Cache
Xiameng Hu, Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Chen Ding 0001, Song Jiang 0001, Zhenlin Wang 0003
USENIX ATC1