VLDB 2026 Research / reviewers in the wild / expert
Hanmei Yang
dblp:236/7694
· DBLP profile ↗
9ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SNIP: An Adaptive Mixed Precision Framework for Subbyte Large Language Model TrainingabstractTraining large language models (LLMs) efficiently while preserving model quality poses significant challenges, particularly with subbyte precision supported by state-of-the-art GPUs. Current mixed-precision training approaches either apply uniform precision to all GEMM operations or rely on heuristic-based methods that fail to generalize during training, leading to suboptimal convergence and instability. Yunjie Pan, Yongyi Yang, Hanmei Yang, Scott Mahlke |
ASPLOS (2) | 3 |
| 2025 | An Empirical Study of Microscaling Formats for Low-Precision LLM TrainingabstractThis paper presents a comprehensive evaluation of microscaling (MX) quantization in the pre-training of large language models (LLMs), investigating its potential to enhance the computation and memory efficiencies. We systematically examine the effects of key design parameters - including data types, rounding modes, scaling strategies, granularity, and organization - on numerical accuracy and training stability. Our extensive experimental study on Llama3 models reveals critical insights into the challenges of 4-bit training for LLMs and identifies optimal configurations with mixed precisions of 4-bit and 6-bit MX formats that significantly enhance training quality, bridging the gap with higher-precision formats. This research provides valuable guidance on the benefits and limitations of MX quantization, laying the groundwork for future innovations in low-precision LLM training. Hanmei Yang, Summer Deng, Amit Nagpal, Maxim Naumov, Mohammad Janani, Tongping Liu, Hui Guan 0001 |
ARITH | 1 |
| 2025 | Implicit Neural Attention for Removing Blur in Remote Sensing ImagesabstractDeblurring in remote sensing images is a challenging task due to the long-range imaging capabilities of remote sensing sensors, which often results in image blur. Factors contributing to image blur include atmospheric disturbances during long-range imaging or the orbital motion of remote sensing platforms. The existing methods remove blur in remote sensing images using the traditional attention mechanism, which focuses on a limited number of features. However, they often overlook the features among neighboring positions in blurry areas, and these areas contain more relevant features. Leveraging these features can effectively assist in restoring the complex object textures of remote sensing blurry images. To achieve this, we propose a novel implicit neural attention mechanism for assembling more relevant features implied by surrounding coordinates. Specifically, we use the features and their corresponding coordinates to learn the enhanced feature representation with more relevant features, and this representation can be used to derive the deblurred images. Extensive experiments demonstrate that our proposed method, INA-RSDeblur, outperforms the state-of-the-art deblurring methods in remote sensing blurry images. Hanmei Yang, Xiaoxuan Chen, Hang An, Bo Jiang 0014 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | NUMAlloc: A Faster NUMA Memory AllocatorabstractThe NUMA architecture accommodates the hardware trend of an increasing number of CPU cores. It requires the cooperation of memory allocators to achieve good performance for multithreaded applications. Unfortunately, existing allocators do not support NUMA architecture well. This paper presents a novel memory allocator – NUMAlloc, that is designed for the NUMA architecture. is centered on a binding-based memory management. On top of it, proposes an “origin-aware memory management” to ensure the locality of memory allocations and deallocations, as well as a method called “incremental sharing” to balance the performance benefits and memory overhead of using transparent huge pages. According to our extensive evaluation, NUMAlloc has the best performance among all evaluated allocators, running 15.7% faster than the second-best allocator (mimalloc), and 20.9% faster than the default Linux allocator with reasonable memory overhead. NUMAlloc is also scalable to 128 threads and is ready for deployment. Hanmei Yang, Wei Wang 0054, Sandip Kundu, Bo Wu 0002, Hui Guan 0001, Tongping Liu |
ISMM | 1 |
| 2023 | Nonlocal ultrasound image despeckling via improved statistics and rank constraint
Hanmei Yang, Jian Lu 0002, Ye Luo 0004, Heng Zhang 0014 |
Pattern Anal. Appl. | 1 |
| 2023 | MemPerf: Profiling Allocator-Induced Performance SlowdownsabstractThe memory allocator plays a key role in the performance of applications, but none of the existing profilers can pinpoint performance slowdowns caused by a memory allocator. Consequently, programmers may spend time improving application code incorrectly or unnecessarily, achieving low or no performance improvement. This paper designs the first profiler—MemPerf—to identify allocator-induced performance slowdowns without comparing against another allocator. Based on the key observation that an allocator may impact the whole life-cycle of heap objects, including the accesses (or uses) of these objects, MemPerf proposes a life-cycle based detection to identify slowdowns caused by slow memory management operations and slow accesses separately. For the prior one, MemPerf proposes a thread-aware and type-aware performance modeling to identify slow management operations. For slow memory accesses, MemPerf utilizes a top-down approach to identify all possible reasons for slow memory accesses introduced by the allocator, mainly due to cache and TLB misses, and further proposes a unified method to identify them correctly and efficiently. Based on our extensive evaluation, MemPerf reports 98% medium and large allocator-reduced slowdowns (larger than 5%) correctly without reporting any false positives. MemPerf also pinpoints multiple known and unknown design issues in widely-used allocators. Sam Silvestro, Steven (Jiaxun) Tang, Hanmei Yang, Hongyu Liu 0005, Guangming Zeng, Bo Wu 0002, Cong Liu 0005, Tongping Liu |
Proc. ACM Program. Lang. | 4 |
| 2022 | Deadlock prediction via generalized dependencyabstractDeadlocks are notorious bugs in multithreaded programs, causing serious reliability issues. However, they are difficult to be fully expunged before deployment, as their appearances typically depend on specific inputs and thread schedules, which require the assistance of dynamic tools. However, existing deadlock detection tools mainly focus on locks, but cannot detect deadlocks related to condition variables. This paper presents a novel approach to fill this gap. It extends the classic lock dependency to generalized dependency by abstracting the signal for the condition variable as a special resource so that communication deadlocks can be modeled as hold-and-wait cycles as well. It further designs multiple practical mechanisms to record and analyze generalized dependencies. In the end, this paper presents the implementation of the tool, called UnHang. Experimental results on real applications show that UnHang is able to find all known deadlocks and uncover two new deadlocks. Overall, UnHang only imposes around 3% performance overhead and 8% memory overhead, making it a practical tool for the deployment environment. Jinpeng Zhou, Hanmei Yang, Jack Lange, Tongping Liu |
ISSTA | 2 |
| 2022 | Cross-Modal Prostate Cancer Segmentation via Self-Attention DistillationabstractThe automatic and accurate segmentation of the prostate cancer from the multi-modal magnetic resonance images is of prime importance for the disease assessment and follow-up treatment plan. However, how to use the multi-modal image features more efficiently is still a challenging problem in the field of medical image segmentation. In this paper, we develop a cross-modal self-attention distillation network by fully exploiting the encoded information of the intermediate layers from different modalities, and the generated attention maps of different modalities enable the model to transfer significant and discriminative information that contains more details. Moreover, a novel spatial correlated feature fusion module is further employed for learning more complementary correlation and non-linear information of different modality images. We evaluate our model in five-fold cross-validation on 358 MRI images with biopsy confirmed. Without bells and whistles, our proposed network achieves state-of-the-art performance on extensive experiments. Xiaoang Shen, Yudong Zhang 0001, Ye Luo 0004, Jihao Luo, Dandan Zhu 0001, Hanmei Yang, Binghui Zhao |
IEEE J. Biomed. Health Informatics | 7 |
| 2020 | Ultrasound Image Restoration Using Weighted Nuclear Norm MinimizationabstractUltrasound images are often contaminated by speckle noise during the acquisition process, which influences the performance of subsequent applications. The paper introduces a nonconvex low-rank matrix approximation model for ultrasound images restoration, which integrates the weighted nuclear norm minimization (WNNM) and data fidelity term. WNNM can adaptively assign weights on different singular values to preserve more details in restored images. The fidelity term about ultrasound images do not be utilized in existing low-rank ultrasound denoising methods. This optimization question can effectively solved by alternating direction method of multipliers (ADMM). The experimental results on simulated images and real medical ultrasound images demonstrate the excellent performance of the proposed method compared with other four state-of-the-art methods. Hanmei Yang, Heng Zhang 0014, Ye Luo 0004, Jian Lu 0002 |
ICPR | 1 |