Jianhang Huang

dblp:152/9329 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 An improved you only look once model for the multi-scale steel surface defect detection with multi-level alignment and cross-layer redistribution features
Jianhang Huang, Lijie Jia, Yitian Zhou
Eng. Appl. Artif. Intell.1
2023 Improving 3-D NAND SSD Read Performance by Parallelizing Read-Retry
abstract
With the bit density improvement and the 3-D flash techniques, modern NAND flash-memory-based solid-state disks (SSDs) dramatically increase the flash storage capacity. However, in high-density SSDs, the long read latency overheads due to massive read-retry steps have become a serious performance concern to develop flash memory in storage devices. In this article, we proposed a parallel read-retry scheme (PaRR) to utilize the read-retry characteristics among flash memory cells. Specifically, when reading the multiple data pages simultaneously, PaRR reduces the read-response time by parallelizing read-retry operations. PaRR reduces the number of read-retry operations by parallelizing read-retry steps with the different aggressive read reference voltages if simultaneously reading the data pages that exhibit virtually equivalent reliability characteristic. Our evaluation shows that PaRR improves the I/O performance by 33.77% on average compared with the state-of-the-art scheme.
Jinhua Cui 0001, Zhimin Zeng, Jianhang Huang, Weiqi Yuan, Laurence T. Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Policy Optimization with Stochastic Mirror Descent
abstract
Improving sample efficiency has been a longstanding goal in reinforcement learning. This paper proposes VRMPO algorithm: a sample efficient policy gradient method with stochastic mirror descent. In VRMPO, a novel variance-reduced policy gradient estimator is presented to improve sample efficiency. We prove that the proposed VRMPO needs only O(ε−3) sample trajectories to achieve an ε-approximate first-order stationary point, which matches the best sample complexity for policy optimization. Extensive empirical results demonstrate that VRMP outperforms the state-of-the-art policy gradient methods in various settings.
Long Yang 0004, Yu Zhang 0009, Gang Zheng 0005, Pengfei Li 0005, Jianhang Huang, Gang Pan 0001
AAAI6
2022 ADS: Leveraging Approximate Data for Efficient Data Sanitization in SSDs
abstract
NAND flash memory has been widely adopted in emerging storage systems. To ensure data security, the support of data sanitization in NAND flash memory-based storage systems is widely employed. Although some existing studies made efforts in employing encryption-based, erasure-based, and scrubbing-based secure deletion approaches to achieve the security requirement, they suffer from the risk of being deciphered, the severe performance, and the scrubbing disturbance problems. Meanwhile, 3-D NAND flash technology, which stacks flash cells in vertical direction, is gaining traction in the modern systems. This made the problems more severe because of the increased number of scrubbing disturbance directions in 3-D NAND flash memory. To address the above issue, this work proposes an approximate-data-aware data sanitization scheme (ADS) with the assistance of the error-resilient data of modern applications, which guarantees the highest degree of security for security-sensitive data sanitization (i.e., storage systems do not keep any old version of secure data once secure data are updated). ADS classifies request data into approx-secure (AS), precise-secure (PS), approx-unsecure (AU), and precise-unsecure (PU) data by considering two factors, including data error resilience and data privacy. Then, a novel data allocation strategy is proposed to selectively interleave secure data and approximate data within the flash blocks, which creates the scrubbing friendly data patterns to minimize the overhead of secure deletion. Our experimental results show that ADS reduces the average secure deletion latency by 58.93% over the state of the art.
Jinhua Cui 0001, Jianhang Huang, Laurence T. Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Exploiting Uncorrectable Data Reuse for Performance Improvement of Flash Memory
abstract
With the increased density and the technology scaling, flash memory is more vulnerable to noise effects, overwhelmingly lowering the data retention time. The periodic data refresh technique is commonly used to retain the long-term data integrity. However, the frequent refresh requests introduce the increased access conflict and severe write amplification, leading to suboptimal performance and lifetime improvement of flash memory-based storage systems. In this article, we propose ApproxRefresh, which enables the uncorrectable data reuse with the assistance of approximate-computing applications to reduce the refresh costs. Specifically, we design a lightweight remapping-free refresh technique, called RFR, to periodically correct, compute an enhanced ECC, and only remap new ECC parity bits, which dramatically reduces the refresh costs. Then, with approximate-read and precise-read hotness awareness, ApproxRefresh selectively adopts RFR or the traditional remapping-based refresh technique to reduce the refresh costs and the read disturbance. Besides, ApproxRefresh further improves flash access performance based on data hotness and process variations. In addition, a corresponding ApproxRefresh-aware garbage collection algorithm is proposed to complete the design. Evaluations show that ApproxRefresh reduces the refresh latency by 36.19% over the state of the art.
Jinhua Cui 0001, Jianhang Huang, Laurence T. Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 SmartHeating: On the Performance and Lifetime Improvement of Self-Healing SSDs
abstract
In NAND flash memory-based solid-state drives (SSDs), during the idle time between the consecutive program/erase cycles (dwell time), the dielectric damage of flash cell can be partially repaired, also known as the self-recovery effect. As the effectiveness of the self-recovery effect can be improved under high temperature, self-healing SSDs are proven feasible to extend the flash endurance significantly. However, current self-healing SSDs perform the heating operations on all the worn-out blocks without considering the data retention requirement, and measures the lifetime of flash memory based on the worst-case self-recovery effect, leading to some unnecessary heating operations and the degraded performance. We propose SmartHeating, a smart heating scheme that exploits the dwell time variation and the write hotness variation to improve the I/O performance and the lifetime of self-healing SSDs. SmartHeating tracks the dwell time of all worn-out flash blocks, predicts their self-recovery effect and reliability, and avoids performing heating operations on the worn-out flash blocks that still have strong flash reliability. In addition, by exploiting the data hotness variation, SmartHeating only heats the worn-out flash blocks that store write-cold data, while allocating write-hot data to a small portion of worn-out flash blocks with negligible refresh overhead. The experimental results show that SmartHeating reduces the number of heating operations by 12.5% on average, boosts I/O performance of flash storage systems by 21.0%, and improves the lifetime of flash memory by $1.20\times $ compared with conventional heating scheme.
Jinhua Cui 0001, Jianhang Huang, Laurence T. Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Exploiting Disturbance-Aware Read Redirection for Performance Improvement in 3D Flash Memory
abstract
Three-dimensional (3D) NAND flash memory has been widely adopted in emerging storage systems, ranging from mobile devices to cloud systems, due to its performance benefits and scalability. While 3D flash improves capacity through stacking more cells in the vertical direction, the read disturbance has become a serious reliability concern to develop 3D flash memory in storage devices, especially for read-intensive applications. To address the above challenge, we propose a novel disturbance-aware read redirection scheme (DRR). By exploiting parity-based redundancy and data duplication to regenerate data content, DRR redirects reads from strong disturbed flash blocks to weak ones without any data replication. DRR makes the redirection decision based on the accumulative read disturbance, data redundancy, data duplication, and the current I/O traffic. Results show that DRR provides as much as 24.6% performance improvement in the average read/write response time, and extends the endurance of 3D flash devices by 4.7%.
Jinhua Cui 0001, Jianhang Huang, Laurence T. Yang
ACM Great Lakes Symposium on VLSI3
2020 ApproxRefresh: Enabling Uncorrectable Data Reuse on Flash Memory with Approximate Read Awareness
abstract
With the increased density and the technology scaling, flash memory is more vulnerable to noise effects, overwhelmingly lowering the data retention time. The periodic data refresh technique is commonly used to retain the long-term data integrity. However, the frequent refresh requests also introduce the increased access conflict and severe write amplification, leading to sub-optimal performance and lifetime improvement.
Jinhua Cui 0001, Jianhang Huang, Laurence T. Yang
LCTES3
2018 ShadowGC: Cooperative garbage collection with multi-level buffer for performance improvement in NAND flash-based SSDs
abstract
Garbage collection, an essential background activity in NAND flash based SSDs, often introduces large runtime overhead. Recent studies showed that it is beneficial to separate the flash pages that have dirty copies in the write buffers from those that do not. However, the existing schemes exploring this observation have limitations, which prevent them from maximizing the performance improvement. In this paper, we address the above challenge through ShadowGC, a novel GC design that exploits the pages in both host-side and device-side write buffers and adopts different read and write strategies to minimize the GC overhead. When garbage collecting flash pages that have dirty copies in the device-side write buffer, ShadowGC reads data from the write buffer. When garbage collecting flash pages that have dirty copies in the host-side write buffer, ShadowGC moves them to dedicated blocks and speeds up the movement with fast-write operations. Our experimental results show that, on average, ShadowGC reduces the write amplification by 16.2% and the GC latency by 20.5% over the state-of-the-art.
Jinhua Cui 0001, Youtao Zhang, Jianhang Huang, Weiguo Wu, Jun Yang 0002
DATE3
2018 DLV: Exploiting Device Level Latency Variations for Performance Improvement on Flash Memory Storage Systems
abstract
NAND flash has been widely adopted in storage systems due to its better read and write performance and lower power consumption over traditional mechanical hard drives. To meet the increasing performance demand of modern applications, recent studies speed up flash accesses by exploiting access latency variations at the device level. Unfortunately, existing flash access schedulers are still oblivious to such variations, leading to suboptimal I/O performance improvements. In this paper, we propose DLV, a novel flash access scheduler for exploring scheduling opportunities due to device level access latency variations. DLV improves flash access speeds based on process variations and data retention time difference across flash blocks. More importantly, DLV integrates access speed optimization with access scheduling such that the average access response time can be effectively reduced on flash memory storage systems. Our experimental results show that DLV achieves an average of 41.5% performance improvement over the state-of-the-art.
Jinhua Cui 0001, Youtao Zhang, Weiguo Wu, Jun Yang 0002, Yinfeng Wang, Jianhang Huang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2016 Exploiting latency variation for access conflict reduction of NAND flash memory
abstract
NAND flash memory has been widely used in storage systems by offering greater read/write performance and lower power consumption than mechanical hard drives. Recently, the tradeoff between endurance, write speed, and read speed has been exploited from many ways for I/O performance improvement, which also induce the read/write latency variation. In this paper, the latency variation is exploited in I/O scheduling for access characteristic guided read and write latency minimization. First, with the understanding of the relationship among read latency, write latency and raw bit error rates (RBER), different ways to exploit the relationship for read and write latency reduction is discussed. Then, an I/O scheduling scheme is proposed by using hotness and retention age of accessed data to determine the speed of writes or reads, giving scheduling priority to fast writes and fast reads for conflict reduction. Experiments with various traces reveal that the proposed technique achieves significant read and write performance improvements.
Jinhua Cui 0001, Weiguo Wu, Xingjun Zhang, Jianhang Huang, Yinfeng Wang
MSST4
2016 VIOS: A Variation-Aware I/O Scheduler for Flash-Based Storage Systems
Jinhua Cui 0001, Weiguo Wu, Shiqiang Nie, Jianhang Huang, Zhuang Hu, Nianjun Zou, Yinfeng Wang
NPC4
2016 Exploiting Cross-Layer Hotness Identification to Improve Flash Memory System Performance
Jinhua Cui 0001, Weiguo Wu, Shiqiang Nie, Jianhang Huang, Zhuang Hu, Nianjun Zou, Yinfeng Wang
NPC4
2016 A Statistics Based Prediction Method for Rendering Application
Qian Li 0013, Weiguo Wu, Jianhang Huang, Mingxia Feng
NPC4
2015 Economy-Oriented Deadline Scheduling Policy for Render System Using IaaS Cloud
Qian Li 0013, Weiguo Wu, Zeyu Sun 0002, Jianhang Huang
ICA3PP (3)5
2014 A Utility-Maximizing Tasks Assignment Method for Rendering Cluster System
abstract
A utility-maximizing tasks assignment method for rendering cluster system based on feedback is proposed to solve the problem that traditional task-centered assignment method and naïve load balancing strategy cannot make full use of resources. The method uses feedback on resource usage to choose an appropriate number of threads for the renderer and then divides computing nodes of the rendering cluster system into fine grain computing units. After that, the method takes advantage of frame-to-frame coherence to assign tasks to computing units with a new static load balancing strategy. Experiments on two scene models with different complexity and comparisons with the naïve method and the fixed-threads method show that the utility-maximizing method not only reduces the rendering time of a rendering job, but also balances the load between computing units. The scalability of the proposed method is also verified on different number of computing nodes.
Jianhang Huang, Weiguo Wu, Qian Li 0013
ISPA1