Ziqu Yu

dblp:396/7936 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
0009-0004-2023-5672ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Concurrent programming · 91% Operating systems · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Concurrent programming › synchronization
lock contention
0.912025
HTLL: Latency-Aware Scalable Blocking Mutex · IEEE Trans. Parallel Distributed Syst. 2025
Concurrent programming › synchronization
mutex lock
0.912025
HTLL: Latency-Aware Scalable Blocking Mutex · IEEE Trans. Parallel Distributed Syst. 2025
Concurrent programming
synchronization
0.912025
HTLL: Latency-Aware Scalable Blocking Mutex · IEEE Trans. Parallel Distributed Syst. 2025
Operating systems › resource management › process management › CPU scheduling
thread scheduling
0.312025
HTLL: Latency-Aware Scalable Blocking Mutex · IEEE Trans. Parallel Distributed Syst. 2025

Methods — techniques the papers use, named apart from their topics

quota-based scheduling · 0.9latency-aware reordering · 0.9
YearPublicationVenuePosition
2025 HTLL: Latency-Aware Scalable Blocking Mutex
abstract
This paper finds that existing mutex locks suffer from throughput collapses or latency collapses, or both, in the oversubscribed scenarios where applications create more threads than the CPU core number, e.g., database applications like mysql use per thread per connection. We make an in-depth performance analysis on existing locks and then identify three design rules for the lock primitive to achieve scalable performance in oversubscribed scenarios. First, to achieve ideal throughput, the lock design should keep adequate number of active competitors. Second, the active competitors should be arranged carefully to avoid the lock-holder preemption problem. Third, to meet latency requirements, the lock design should track the latency of each competitor and reorder the competitors according to the latency requirement. We propose a new lock library called HTLL that satisfies these rules and achieves both high throughput and low latency even when the cores are oversubscribed. HTLL only requires minimal human effort (e.g., add several lines of code) to annotate the latency requirement. Evaluation results show that HTLL achieves scalable performance in the oversubscribed scenarios. Specifically, for the real-world database, LMDB, HTLL can reduce the tail latency by up to 97% with only an average 5% degradation in throughput, compared with state-of-the-art alternatives such as Malthusian, CST, and Mutexee locks; In comparison to the widely used pthread mutex lock, it can increase the throughput by up to 22% and decrease the latency by up to 80%. Meanwhile, for the under-subscribed scenarios, it also shows comparable performance than state-of-the-art blocking locks.
Ziqu Yu, Jinyu Gu 0001
IEEE Trans. Parallel Distributed Syst.1