EDBT 2026 Demo / reviewers in the wild / expert
Li Wan 0008
dblp:35/5678-8
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0002-6836-2740ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PCMT: Prioritizing Coherence Message Types for NoC Protocol-Level Deadlock FreedomabstractThe Network-on-Chip (NoC) has emerged as a vital interconnect fabric in multi-core processors. However, the implementation of virtual networks requires multiple message queues at each router’s input port to prevent protocol-level deadlocks, resulting in substantial area and power overheads. This complexity poses challenges in maintaining performance within a constrained area budget. In this paper, we introduce PCMT, a virtual-network-free mechanism that leverages the inherent priority of directory-based coherence messages to resolve protocol-level deadlocks within the NoC. Experimental results show that PCMT outperforms both baseline and state-of-the-art solutions across systems with varying numbers of virtual networks. In a MOESI-hammer protocol system with 6 virtual networks, PCMT achieves comparable performance while reducing input buffer overhead by up to 63% compared to the baseline. In a MOESI directory system with 3 virtual networks, PCMT improves overall execution time by up to 4.0% with the same input buffer resources. Additionally, PCMT demonstrates superior deadlock resolution speed and bandwidth consumption compared to the state-of-the-art work. Yufan Jia, Li Wan 0008, Jun Han 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | BoostTM: Best-effort performance guarantees in best-effort hardware transactional memory for distributed manycore architectures
Li Wan 0008, Jun Han 0003 |
J. Syst. Archit. | 1 |
| 2024 | LockillerTM: Enhancing Performance Lower Bounds in Best-Effort Hardware Transactional MemoryabstractConcurrent access to shared data has always been a challenge for developing multi-threaded programs and a bottleneck in the performance of Chip-Multiprocessor (CMP) systems. The challenge has been exacerbated by the need to augment processor cores and network bandwidth to fulfill the low-latency demands of ever-expanding data processing. Existing commercial best-effort Hardware Transactional Memory (HTM) is a common and effective solution. However, its architectural constraints prevent transactions from surviving in exceptions, cache overflow, and coexisting with a non-speculation fallback path, leading to unstable performance and diminishing favor. In this paper, we propose three lightweight mechanisms designed to mitigate the limitations of the best-effort HTM architecture to enhance performance stability. One is the recovery mechanism that supports the dynamic revocation of toxic conflicting requests, dramatically reducing the potential of livelocks. The second is the HTMLock mechanism with hardware and software co-design, which allows transactions using HTM and locks to run concurrently except when encountering actual conflict. Lastly, the switchingMode mechanism enables a running transaction to proactively attempt to switch to HTMLock mode in the event of a non-conflict-induced abort. Gem5 infrastructure is extended to validate and evaluate our mechanisms in a 32-core tiled CMP system. Experimental studies show that LockillerTM outperforms the coarse-grained locking scheme under STAMP benchmarks except for the yada workload, irrespective of thread number and cache size. Furthermore, our approach achieves an average of 1.86x and 1.57x speedup in all benchmarks and different threads under a typical cache size and a maximum of 7.79x and 6.73x speedup in high-contention benchmarks under extreme scenarios with only 8KB L1 cache and 32 threads, compared to best-effort HTM and state-of-the-art HTM respectively. Li Wan 0008, Jun Han 0003 |
IPDPS | 1 |
| 2022 | LosaTM: A Hardware Transactional Memory Integrated With a Low-Overhead Scenario-Awareness Conflict ManagerabstractThe vigorous development of high compute-intensive applications has led to the demand for maximizing the concurrency of multicore processors. The best-effort hardware transactional memory(HTM) is an important technology adopted by vendors to improve the potential concurrency of multicore processors, but the HTM implementations on commercial products have some drawbacks for its simplicity and need some further optimizations to enable more exploitation of concurrency. In this article, we propose and evaluate a novel design of HTM, called LosaTM, which can provide a scenario-awareness conflict management strategy. By leveraging the proposed feature of multiple-grained coherency maintenance in the coherence protocol, LosaTM resolves most false conflicts at a half-cache-line granularity. Furthermore, we design a winner/aborter vector conflict management algorithm to improve the efficiency of LosaTM in handling friendly-fire and unfairness competition that we have newly defined. In order to coordinate these integrated conflict management strategies, a scheduling strategy is also proposed to adaptively select the appropriate management according to the specific conflict scenario. We use gem5 to simulate LosaTM in detail on an 8-core tiled CMP system, and the simulation result shows that it only causes 0.7% of the L1 cache size hardware overhead while achieving a 38% average execution time reduction on the native STAMP. The speedup also demonstrates that LosaTM outperforms the state-of-the-art designs in previous works. Li Wan 0008, Jun Han 0003 |
IEEE Trans. Parallel Distributed Syst. | 2 |