VLDB 2026 Research / reviewers in the wild / expert
Huailiang Tan
dblp:59/8452
· DBLP profile ↗
20ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0001-9980-8015ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Innovative Bloom Filter-Based Multilevel Block Classification: A Practical Approach to Enhancing SSD Performance and EnduranceabstractABSTRACT In modern solid‐state drives (SSDs), managing hot and cold data is key to improving performance and extending lifespan. However, previous Bloom filter‐based methods were often complex and lacked solid empirical validation under high‐performance conditions. To address this, we propose a more efficient multilevel Bloom filter classification strategy tailored to SSDs. Our approach optimizes classification accuracy while reducing computational and storage overhead through careful parameter tuning. By utilizing SSD block‐level parallelism to group similar data access patterns, we enhance garbage collection efficiency and extend block life. Unlike previous studies relying on simulations, we validate our method on real SSD hardware. Experimental results show that our strategy improves SSD performance and endurance, offering valuable insights for future firmware optimization. Huailiang Tan, Zaihong He, Jinyou Li, Keqin Li 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2025 | State-driven fairness control for efficient I/O queue scheduling in NVMe virtualization
Yifu Zhu, Xin Kuang, Yanjie Tan, Huailiang Tan, Keqin Li 0001 |
Future Gener. Comput. Syst. | 5 |
| 2025 | CDA-GC: An effective cache data allocation for garbage collection in flash-based solid-state drives
Huailiang Tan, Zaihong He, Jinyou Li, Keqin Li 0001 |
Integr. | 2 |
| 2025 | Dynamic DPU Offloading and Computational Resource Management in Heterogeneous SystemsabstractDPU offloading has emerged as a promising way to enhance data processing efficiency and free up host CPU resources. However, unsuitable offloading may overwhelm the hardware and hurt overall system performance. It is still unclear how to make full use of the shared hardware resources and select optimal execution units for each tenant application. In this paper, we propose DORM, a dynamic DPU offloading and resource management architecture for multi-tenant cloud environments with CPU-DPU heterogeneous platforms. The primary goal of DORM is to minimize host resource consumption and maximize request processing efficiency. By establishing a joint optimization model for offloading decision and resource allocation, we abstract the problem into a mixed integer programming mathematical model. To simplify the complexity of model-solving, we decompose the model into two subproblems: a 0-1 integer programming model for offloading decision-making and a convex optimization problem for fine-grained resource allocation. Besides, DORM presents an orchestrator agent to detect load changes and dynamically adjust the scheduling strategy. Experimental results demonstrate that DORM significantly improves resource efficiency, reducing host CPU core usage by up to 83.3%, increasing per-core throughput by up to 4.61x, and lowering the latency by up to 58.5% compared to baseline systems. Yanjie Tan, Yifu Zhu, Huailiang Tan, Keqin Li 0001 |
IEEE Trans. Computers | 4 |
| 2025 | Simplicity as the Ultimate Principle: The Art of Garbage Collection Management in SSDs Inspired by Natural Data BehaviorabstractAs solid-state drives (SSDs) are increasingly used in various computing environments, effective garbage collection (GC) management is crucial for enhancing performance and extending lifespan. Existing GC strategies rely on complex data categorization techniques to distinguish between hot and cold data. This process is not only computationally expensive but also inefficient. This article introduces a revolutionary simplified GC management method, which we call SUP-GC, based on the core design principle that simplicity is the ultimate principle. SUP-GC requires almost no computation and naturally redefines the data storage pattern. All newly written data is defaulted as hot data and stored directly in hot data blocks; data that needs to be moved during the GC process is considered cold data and uniformly migrated to cold data blocks. Our strategy eliminates the need for precise but resource-intensive real-time analysis of data states and instead adopts a storage strategy that is highly consistent with the natural properties of data. This intuitive partitioning significantly reduces the need for complex judgments about data states, thereby optimizing the storage management process. Experiments and designs on real SSDs have shown that our SUP-GC strategy significantly outperforms existing mainstream GC methods. Huailiang Tan, Keqin Li 0001 |
ACM Trans. Storage | 2 |
| 2024 | A HW/SW Co-Design of Video Dehazing Accelerator Using Decoupled Local Atmospheric Light PriorabstractIn this paper, we introduce DLAPID, a novel decoupled parallel hardware-software co-design architecture for real-time video dehazing. From a software point of view, DLAPID isolates the atmospheric light operation from the initial transmission estimation to take full advantage of the hardware accelerators' parallelization features. For the hardware implementation, we deploy DLAPID both on FPGA and GPU platforms and validate its effectiveness. Using both real-world driving scenario testing sets and ground-truth datasets, we quantitatively and qualitatively assess the proposed method against several SOTA (state-of-the-art) video dehazing models. The outcomes of our experiments demonstrate that our approach achieves better dehazing performance with lower power consumption and has real-time processing capabilities, thereby preventing potential accidents of autonomous vehicles. Yanjie Tan, Yifu Zhu, Feiteng Nie, Huailiang Tan |
DAC | 5 |
| 2024 | DDPM-MoCo: Enhancing the Generation and Detection of Industrial Surface Defects Through Generative and Contrastive Learning
Huailiang Tan |
ICANN (9) | 2 |
| 2024 | MTDA: Efficient and Fair DPU Offloading Method for Multiple TenantsabstractIn modern cloud computing environment, the offloading potential of DPU must be fully exploited for multiple tenants. Existing DPU offloading techniques lack the capability to perform the fair allocation of a DPU domain's internal resources among tenants with various performance requirements. In this article, we propose a virtual multi-channel DPU offloading architecture for multiple tenants (MTDA) and implement it on a BlueField-2 DPU platform to achieve stability and fairness in resource allocation for generic datacenter tasks. MTDA provides an independent virtual channel for each tenant before their requests are submitted to avoid competition among tenants. Considering the diverse requirements of tenants, MTDA constructs a credit-based resource allocation model and a traffic-aware scheduling algorithm to fully utilize the rich computing resources of DPU and improve the fairness of DPU resource allocation. Experimental results show that MTDA increases the throughput by up to 101.2%, 143.2%, 36.1%, and 41.7%, lowers the latency by up to 50.3%, 58.9%, 26.6%, and 29.4%, improves the fairness by up to 98.8%, 99.0%, 98.3%, and 98.4%, and provides more stable performance for multi-tenants, compared with DPDK, iPipe, FairNIC, and LogNIC. Yanjie Tan, Yifu Zhu, Huailiang Tan, Keqin Li 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | Brief Industry Paper: Real-Time Image Dehazing for Automated VehiclesabstractAutonomous vehicles at L2 and above are increasingly relying on stereo vision systems, where haze removal is critical to detect obstacles hidden in fog. Existing image dehazing techniques have low processing speed and high resource consumption, restricting their application scope in practice. In this work, we propose a hardware-software co-design solution for haze removal. It fully decouples calculation of the two main parameters, i.e., atmospheric light and transmittance, in the dehazing process. By eliminating the data dependency, parallelism in hardware acceleration is enhanced. Furthermore, in replacement of the conventional global homogeneous atmospheric light computation, we report a chunk-based heterogeneous method to reduce cache overhead. Our approach is implemented on FPGA, compared against five state-of-the-art (SOTA) works for image haze removal. Evaluation using test sets of real-world foggy driving scenarios shows that our object detection accuracy is over 88%, 9.5%-47.4% better than the SOTA works with neural networks (NN) on GPU, and 25.9%-52.2 % better than the SOTA works on FPGA. The processing speed varies with image resolution and our improvement is generally even more at higher resolution, which is 29.7% faster than the fastest SOTA. We have the lowest overall resource consumption, where the bottleneck BRAM usage is reduced by over 70%. The FPGA solution has circuit-level timing determinism at nanosecond, hence suitable for hard real-time applications. Yanjie Tan, Yifu Zhu, Huailiang Tan, Wanli Chang 0001 |
RTSS | 4 |
| 2023 | MAPD: An FPGA-Based Real-Time Video Haze Removal Accelerator Using Mixed Atmosphere PriorabstractReal-time video dehazing plays a key role in helping autonomous driving detect pedestrians or obstacles in severe foggy weather to prevent potential hazards. Existing video dehazing methods achieve good restoration performance but still suffer from oversaturation and low dehazing speed, especially for high-definition (HD, high-resolution) videos. In this article, we propose a mixed atmosphere prior information video dehazing accelerator (MAPD) and implement it on field programmable gate array (FPGA) to achieve real-time haze removal for HD video. MAPD provides a mixed atmospheric light model by applying heterogeneous atmospheric light in the foreground area to balance brightness deviation, and maintaining the global atmospheric light in the background region. Considering the parallel characteristics of FPGA, MAPD leverages the redundant information between adjacent frames to accelerate the dehazing process and designs an indirect transmission estimation to decrease resource consumption. For comparison, we also implement six dehazing solutions (DCP, color ellipsoid prior (CEP), RDCP, FFVD, MHVD, and REFD) on FPGA, and deploy a graphics processing unit (GPU)-based method$(D^{4})$on the platform with Nvidia 3080 GPU. Experiments using two widely used benchmarks show that MAPD increases the performance by up to 36.5%, 53.5%, 36.3%, 33.3%, 11.9%, and 23.3%, decreases resource consumption by up to 79.7%, 75.0%, 74.8%, 25.6%, 22.6%, and 73.9% and enhances FPS for HD videos by up to 241.6%, 145.9%, 151.7%, 68.6%, 50.6%, and 62.4%, compared with DCP, CEP, RDCP, FFVD, MHVD, and REFD. Compared to$D^{4}$, MAPD also promotes the dehazing performance by up to 21.8%, and increases FPS by up to 487.0%. Yanjie Tan, Yifu Zhu, Huailiang Tan, Keqin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Embedded Transaction Support Inside SSD With Small-Capacity Non-Volatile Disk CacheabstractFlash-based Solid State Drives (SSDs) have proved to be ideal devices that support embedded transaction protocols inside SSDs. Existing embedded transaction protocols in SSDs effectively improve transaction throughput, but still incur high transaction overhead and long recovery time. While it is reasonable to provide a small-capacity non-volatile (NVM-based) disk cache in the SSDs, in this paper, we propose a new embedded transaction protocol called Non-volatile Cache Transaction (NVCTX). NVCTX reduces transaction overhead and provides fast recovery by leveraging the small-capacity NVM-based disk cache from two aspects. First, we store transactional metadata, which is of small amount but is frequently accessed, in the NVM-based disk cache rather than in the flash memory. Second, we introduce two techniques, i.e., a dynamic allocation algorithm and a hybrid storing method, to improve the performance when the capacity of the NVM-based disk cache is very limited. We have implemented NVCTX on a real hardware board called Cosmos+ FPGA platform, and modified ext4 file system and NVMe (Non-Volatile Memory express) driver to be compatible with the transactional interfaces provided by NVCTX. For comparison, we also implement SCC, BPCC, WAL, and X-FTL protocols in the firmware of Cosmos+ FPGA platform. Evaluations using DBMS (Database Management System) and file system workloads show that, compared to four typical transaction protocols (SCC, BPCC, WAL, and X-FTL), NVCTX improves transaction throughput by up to 136.5, 9.4, 131.6 and 29.9 percent, reduces write traffic to flash memory by up to 42.8, 4.1, 62.4, 31.2 percent, lowers garbage collection overhead by up to 93.2, 63, 66.5, 22.1 percent, and shortens recovery time to 1/2574, 1/2559, 1/95 and 1/2 respectively compared with SCC, BPCC, WAL, and X-FTL. Yanjie Tan, Huailiang Tan, Youyou Lu, Zaihong He |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2019 | LPCMsim: A Lightweight Phase Change Memory Simulator
Zaihong He, Jishun Kuang, Yanjie Tan, Shihui Peng, Huailiang Tan |
Future Gener. Comput. Syst. | 5 |
| 2019 | A Virtual Multi-Channel GPU Fair Scheduling Method for Virtual MachinesabstractIn modern virtual computing environment, the 2D/3D rendering performance and parallel computing potential of GPU (graphics processing unit) must be fully exploited for multiple virtual machines (VMs). Existing GPU virtualization techniques are unable to take full advantage of a GPU's powerful 2D/3D hardware-accelerated graphics rendering performance or parallel computing potential, or it has not been considered that the internal resources of a GPU domain are fairly allocated between VMs with different performance requirements. Therefore, we propose a multi-channel GPU virtualization architecture (VMCG), model the corresponding credit allocating and transferring mechanisms, and redesign the virtual multi-channel GPU fair-scheduling algorithm. VMCG provides a separate V-Channel for each guest VM (DomU) that competes with other VMs for the same physical GPU resources, and each DomU submits command request blocks to its respective V-Channel according to the corresponding DomU ID. Through the virtual multi-channel GPU fair-scheduling algorithm, not only do multiple DomUs make full use of native GPU hardware acceleration, but the fairness of GPU resource allocation is significantly improved during GPU-intensive workloads from multiple DomUs running on the same host. Experimental results show that, for 2D/3D graphics applications, performance is close to 96 percent of that of the native GPU, performance is improved by approximately 500 percent for parallel computing applications, and GPU resource-allocation fairness is improved by approximately 60-80 percent. Huailiang Tan, Yanjie Tan, Xiaofei He 0007, Kenli Li 0001, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Parallel implementation and optimization of high definition video real-time dehazing
Huailiang Tan, Xiaofei He 0007, Gaoming Liu |
Multim. Tools Appl. | 1 |
| 2016 | Redundant Network Traffic Elimination with GPU Accelerated Rabin FingerprintingabstractRecently, redundant network traffic elimination has attracted a lot of attention from both the academia and the industry. A core challenge and enabling technique in implementing redundancy elimination is to perform content-based chunking, which typically involves the computationally heavy Rabin fingerprinting algorithm. In this paper, we propose a GPU-based implementation of Rabin fingerprinting to address this issue. To maximize performance gains, a diverse set of optimization strategies, such as efficient buffer management, GPU memory hierarchy optimization, and balanced load distribution, is proposed by either exploiting the intrinsic hardware features or addressing domain-specific challenges. Extensive evaluations on both the overall and microscopic performance reveal the effectiveness of the GPU-accelerated Rabin fingerprinting algorithm, and we can achieve up to 40 Gpbs throughput on a GTX 780 card. The throughput shows 1.87× speedup against the state-of-the-art using comparable hardware. In addition, although some optimization designs are specific for the problem, techniques proposed in this work including the indexed compact buffer scheme and approximate sorting would also be beneficial and applicable to other network applications leveraging GPU acceleration. Jianhua Sun 0002, Hao Chen 0002, Ligang He, Huailiang Tan |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | VMCD: A Virtual Multi-Channel Disk I/O Scheduling Method for Virtual MachinesabstractIn the era of cloud computing and big data, virtualization is gaining great popularity in storage systems. Since multiple guest virtual machines (DomUs) are running on a single physical device, disk I/O fairness among DomUs and aggregated throughput remain the challenges in virtualized environments. Although several methods have been developed for disk I/O performance virtualization among multiple DomUs, most of them suffer from one or more of the following drawbacks. (1) A fair scheduling mechanism is missing when requests converge together from multiple queues. (2) Existing methods rely on better performance of the underlying storage system such as solid state drive (SSD). (3) Throughput and latency are not considered simultaneously. To address these disadvantages, this paper presents a virtual multi-channel of disk I/O (VMCD) method that can be built on top of an ordinary storage utility, which mitigates the interference among multiple DomUs by using separated virtual channel (V-Channel) and an I/O request queue for each DomU. In our VMCD, several mechanisms are employed to enhance the I/O performance, including a credit allocation mechanism, a global monitoring strategy, and a virtual multi-channel fair scheduling algorithm. The proposed techniques are implemented on the Xen virtual disk and evaluated on Linux guest operating systems. Experiments results show that VMCD increases fairness by 70 percent approximately compared with CFQ and Anticipatory schedulers, by 30 percent approximately compared with Deadline scheduler; and enhances bandwidth utilization by 28 percent approximately compared with CFQ and Anticipatory schedulers, by 37 percent compared with Deadline in the case of three or more virtual DomUs running on the same physical host. Huailiang Tan, Zaihong He, Keqin Li 0001, Kai Hwang 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2014 | DMVL: An I/O bandwidth dynamic allocation method for virtual networks
Huailiang Tan, Lianjun Huang, Zaihong He, Youyou Lu, Xubin He |
J. Netw. Comput. Appl. | 1 |
| 2014 | BAG: Managing GPU as Buffer Cache in Operating SystemsabstractThis paper presents the design, implementation and evaluation of BAG, a system that manages GPU as the buffer cache in operating systems. Unlike previous uses of GPUs, which have focused on the computational capabilities of GPUs, BAG is designed to explore a new dimension in managing GPUs in heterogeneous systems where the GPU memory is an exploitable but always ignored resource. With the carefully designed data structures and algorithms, such as concurrent hashtable, log-structured data store for the management of GPU memory, and highly-parallel GPU kernels for garbage collection, BAG achieves good performance under various workloads. In addition, leveraging the existing abstraction of the operating system not only makes the implementation of BAG non-intrusive, but also facilitates the system deployment. Hao Chen 0002, Jianhua Sun 0002, Ligang He, Kenli Li 0001, Huailiang Tan |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2012 | SDM: A Stripe-Based Data Migration Scheme to Improve the Scalability of RAID-6abstractIn large scale data storage systems, RAID-6 has received more attention due to its capability to tolerate concurrent failures of any two disks, providing a higher level of reliability. However, a challenging issue is its scalability, or how to efficiently expand the disks. The main reason causing this problem is the typical fault tolerant scheme of most RAID-6 systems known as Maximum Distance Separable (MDS) codes, which offer data protection against disk failures with optimal storage efficiency but they are difficult to scale. To address this issue, we propose a novel Stripe-based Data Migration (SDM) scheme for large scale storage systems based on RAID-6 to achieve higher scalability. SDM is a stripe-level scheme, and the basic idea of SDM is optimizing data movements according to the future parity layout, which minimizes the overhead of data migration and parity modification. SDM scheme also provides uniform data distribution, fast data addressing and migration. We have conducted extensive mathematical analysis of applying SDM to various popular RAID-6 coding methods such as RDP, P-Code, H-Code, HDP, X-Code, and EVENODD. The results show that, compared to existing scaling approaches, SDM decreases more than 72.7% migration I/O operations and saves the migration time by up to 96.9%, which speeds up the scaling process by a factor of up to 32. Chentao Wu, Xubin He, Jizhong Han, Huailiang Tan, Changsheng Xie 0001 |
CLUSTER | 4 |
| 2008 | A Dynamic Rebuild Strategy of Multi-host System Volume on Transparence Computing ModeabstractBy statistical analysis of a few isomorphic hosts booting remotely and deployed separately from a shared system volume, a similar model is proposed on the transparence computing mode. A kind of information left on the rewritten SCE (similar collection elements) will instruct the hosts to quickly sense and rapidly position the rewritten SCE when they rebuild their system volumes. Based on the two dimension local characteristic (TDLC) of rewritten blocks, a system rebuild strategy (SRS) is proposed. The time complexity of the SRS in positioning rewritten address is O(1). An optimistic forecast algorithm and a prefetching strategy to accelerate the I/O processing are proposed. The experiment result shows that the multi-host system rebuild strategy (SRS) takes on a favorable performance and stability. Huailiang Tan, Jianhua Sun 0002 |
HPCC | 1 |