EDBT 2026 Demo / reviewers in the wild / expert
Shuo-Han Chen
dblp:155/4842
· DBLP profile ↗
60ranked-venue papers
20as first author
26since 2021 · last 2026
0000-0002-1619-4335ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 44 · 16 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 9 · 6 since 2021Computer networks · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sk-Art: Optimizing Adaptive Radix Trees for the Shift-Based Skyrmion Racetrack Memory
Zhen-Yang Guo, Tien-Hsin Hsieh, Chieh-Jen Wang, Shuo-Han Chen |
COMPSAC | 5 |
| 2026 | Exploiting Port-Level Parallelism and Inter-Track Skyrmion Reuse for Energy-Efficient Updates in Skyrmion Racetrack Memory
Cheng-Zhi Hou, Tzu-Ying Yu, Wei-Kuan Shih, Shuo-Han Chen |
ISLPED | 4 |
| 2025 | Facilitating the Merging Process of Pull Requests in Automated-Testing Robot FrameworkabstractRobot Framework is a keyword-driven test automation framework that is widely used for acceptance test-driven development to verify software systems’ functionality and quality. While automated testing scripts are developed to verify the functionalities of software systems (i.e., web services), ensuring the maintainability and correctness of automated testing scripts is vital in large-scale software projects, especially when multiple teams collaboratively develop test scripts in parallel. In this paper, we focus on how to facilitate the merging of pull requests (PRs) for projects that use the Robot Framework. We present findings from a two-year industry-academia collaboration involving three concurrently working teams, each consisting of six to eight members developing testing scripts for web services. Through detailed analysis of PR merging workflows, keyword usage patterns, and team coordination practices, we identify common bottlenecks, such as conflicting keyword dependency and unaligned script styles, that hinder efficient merging. Although the teams employ a Kanban approach, which helps limit tasks in progress, visualize workflows, and aim for continuous improvement, significant delays arose at the “Waiting for Merge” stage. Some pull requests remained unmerged for over a month, while more complex test cases could take up to two months. Our study proposes a set of best practices and an automated tool to streamline PR reviews, reduce merge conflicts, and enhance script maintainability. We discuss the observed outcomes, including improved collaboration metrics and reduced integration overhead, offering actionable insights for practitioners and researchers working with Robot Framework in multi-team environments. Zhen-Yang Guo, Shuo-Han Chen, Andrew Garland, Wei-Hao Chen, Yu-Pei Liang |
COMPSAC | 2 |
| 2025 | High-Performance Address Translation for Solid-State Drives via Asymmetry DecompositionabstractIn the post-AI era, the growing complexity of computer systems has increased the demand for energy-efficient, high-performance storage solutions. Solid-state drives (SSDs), widely used in various applications, face performance bottlenecks in address translation under intensive workloads due to their serial processing architecture. Additionally, under intensive random or parallel I/O workloads, the address translation process in SSDs has emerged as one of the primary bottlenecks due to its serial processing architecture. As NAND flash access latencies decrease with enhanced parallelism, the impact of address translation overhead becomes more pronounced. Although it is possible to accelerate address translation in SSDs by utilizing more processing cores or increasing processing frequency, this performance improvement typically comes at the cost of energy efficiency. Moreover, since the sub-tasks involved in address translation are inherently asymmetric, decomposing the address translation procedure with a symmetric assumption can lead to inefficiencies or contention due to dependencies. These observations motivate the proposal of an asymmetric multiprocessing address translation (AMP-AT) strategy that decomposes address translation into asymmetric sub-tasks and pipelines them to improve performance and energy efficiency. Our approach leverages the inherent asymmetry of sub-tasks to minimize contention and dependencies, achieving significant gains in both metrics, as validated by experimental results. Yu-Ta Liu, Shuo-Han Chen |
COMPSAC | 2 |
| 2025 | Enabling Data-Deduplication-Assisted Data Relocation for Interlaced Magnetic RecordingabstractInterlaced Magnetic Recording (IMR) drives have been regarded as promising hard disk drives (HDDs) to meet the ever-growing storage demand in the post-AI era. Among various technologies, IMR drives achieve increased storage capacity by overlapping tracks in an interlaced fashion; however, updating data on overlapped tracks necessitates rewriting up to two tracks, which can degrade performance. Previous strategies have attempted to reduce track rewrites by either delaying the timing of overlapping based on storage space usage or relocating frequently updated data to non-overlapped tracks through update-frequency-based hot/cold data separation. In contrast, this paper proposes a novel approach that leverages data deduplication as a natural solution to the rewrite challenge in IMR drives. Unlike conventional hot/cold separation, this paper utilizes deduplication metadata, specifically reference counts, as a direct indicator for data relocation. The rationale behind the proposed approach is that deduplicated data, which remains static unless reference counts drop to one, can be stored or migrated to overlapped tracks without inducing future track rewrites. The proposed approach diverges from traditional frequency tracking methods by leveraging data deduplication information to guide data placement and reduce rewriting. Evaluation results show that the proposed scheme effectively reduces the accumulated read/write latency by $37.57 \%$ on average compared with state-of-the-art data management on IMR drives with data deduplication enabled. Chen-Jui Tu, Shuo-Han Chen |
DAC | 2 |
| 2025 | Exploiting LDPC Syndrome for Multidimensional Hard-Decoding Read Retry on NAND FlashabstractNAND-flash-based solid-state drives (SSDs) are under constant pressure to deliver higher storage density while minimizing power and performance overhead. As the number of bits stored per NAND flash cell has scaled from single-level cells (SLC) to triple-level cells (TLC) and soon to penta-level cells (PLC), the reduced voltage margins between cell states challenge data reliability, requiring stronger decoding techniques. To maintain reliability and correct error data bits, low-density parity-check (LDPC) codes are widely deployed on these high-density devices and can operate in two modes: hard decoding, which uses threshold-based bit decisions and is relatively power-efficient, and soft decoding, which leverages additional reliability information but imposes higher computational and energy costs. In practice, NAND flash controllers initiate soft decoding when hard decoding fails, thereby preserving data integrity at the expense of latency and power overhead. Current approaches employ read-retry tables to adjust reference voltages and maximize hard decoding success rates; however, such tables cannot fully address diverse bit-error patterns, often unnecessarily invoking soft decoding and incurring significant performance overhead. To overcome this limitation, we propose a novel LDPC-syndrome-based loss function that adaptively adjusts multidimensional reference voltages, significantly reducing unnecessary soft decoding triggers without relying on predetermined read-retry tables or iterative voltage adjustments. Experimental results demonstrate that our proposed loss function effectively reduces the soft decoding trigger rate and the number of page reads, substantially minimizing the performance and power costs associated with soft decoding. Szu-Wei Chen, Shuo-Han Chen |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2024 | Facilitating the Process Rate of Difference Determination within HTML after Web Page EventsabstractThe goal of developing automated web testing scripts is to assess the functionality and responsiveness of developed web pages via simulating user behaviors and interactions. Automated web testing uses HTML content that changes according to input events to determine the success of an interaction. When writing automation scripts, developers usually find it challenging to locate stable element constraints due to the rapid change of elements on the web pages or due to the wrong functionality assumption with a limited understanding of the test requirements. In this paper, to avoid labor and facilitate the process of checking the attributes or labels changed after input events, we propose an extension for the browser developer tool to identify HTML changes before and after automated input events. Within the developed tool, developers can access all HTML differences in the panel of the browser developer tool to check if any detail of HTML elements is changed because of automated input events. In addition, timing and filtering functions are also included in the developed tool to avoid unnecessary element comparisons and improve the stability of developed automated scripts. After real-world testing with both seasoned and junior developers, the developed tool extension can effectively reduce the development time by an average of 20 % to 30%. Zhen-Yang Guo, Yu-Siang Liao, Shuo-Han Chen |
COMPSAC | 3 |
| 2024 | An Automation Tool Converting Test Steps into Keywords for Automated Test DevelopmentabstractRobot Framework is an extensible, keyword-driven, and Python-based test automation framework that is widely used for end-to-end acceptance testing or acceptance-driven test development (ATDD). The simplicity of its syntax also contributes to its excellent readability, making the naming of each keyword in test scripts particularly important. The name of a keyword should accurately describe its function, as it affects the time that code reviewers spend to understand its purpose or even the entire test case. Selecting a suitable name for keywords often requires a considerable amount of time. Therefore, a fixed naming format is a method to shorten this process. With a set of fixed naming rules, the structure of previously named similar actions' keywords can be directly copied to the current keyword being named. Additionally, a fixed naming format can significantly reduce the time spent searching for reused keywords. In addition to the aforementioned advantages, this study points out that the fixed naming format can also be used to directly convert test steps into corresponding action keywords during the test case design phase and proposes an effective automation tool to convert test steps into keywords for enhancing the value of test steps and saving test case development time. In situations where the test steps are well-structured, the conversion tool can accelerate test case development speed by extracting the action keywords of each test step and converting them into corresponding keywords. Zhen-Yang Guo, Zi-Long You, Shuo-Han Chen, Shiang-Jiun Chen |
COMPSAC | 3 |
| 2024 | OC-DLRM: Minimizing the I/O Traffic of DLRM Between Main Memory and OCSSDabstractDue to the exponential growth of data in computing, DRAM-based main memory is now insufficient for data-intensive applications like machine learning and recommendation systems. This has led to a performance issue involving data transfer between main memory and storage devices. Conventional NAND-based SSDs are unable to efficiently handle this problem as they can't distinguish between data types from the host system. In contrast, open-channel SSDs (OCSSD) offer a solution by optimizing data placement from the host-side system. This research focuses on developing a new data access model for deep learning recommendation systems (DLRM) using OCSSD storage drives, called OC-DLRM. OC-DLRM reduces I/O traffic to flash memory by aggregating frequently-accessed data using the I/O unit of a flash memory drive. Our experiments show that OC-DLRM has significant performance improvement compared with traditional swapping space management techniques. Shang-Hung Ti, Tseng-Yi Chen, Tsung Tai Yeh, Shuo-Han Chen, Yu-Pei Liang |
DATE | 4 |
| 2024 | Securing Deep Neural Networks on Edge from Membership Inference Attacks Using Trusted Execution EnvironmentsabstractPrivacy concerns arise from malicious attacks on Deep Neural Network (DNN) applications during sensitive data inference on edge devices. Membership Inference Attack (MIA) is developed by adversaries to determine whether sensitive data is used to train the DNN applications. Prior work uses Trusted Execution Environments (TEEs) to hide DNN model inference from adversaries on edge devices. Unfortunately, existing methods have two major problems. First, due to the restricted memory of TEEs, prior work cannot secure large-size DNNs from gradient-based MIAs. Second, prior work is ineffective on output-based MIAs. To mitigate the problems, we present a depth-wise layer partitioning method to run large sensitive layers inside TEEs. We further propose a model quantization strategy to improve the defense capability of DNNs against output-based MIAs and accelerate the computation. We also automate the process of securing PyTorch-based DNN models inside TEEs. Experiments on Raspberry Pi 3B+ show that our method can reduce the accuracy of gradient-based MIAs on AlexNet, VGG-16, and ResNet-20 evaluated on the CIFAR-100 dataset by 28.8%, 11%, and 35.3%. The accuracy of output-based MIAs on the three models is also reduced by 18.5%, 13.4%, and 29.6%, respectively. Cheng-Yun Yang, Gowri Ramshankar, Nicholas Eliopoulos, Purvish Jajal, Sudarshan Nambiar, Evan Miller, Jing (Dave) Tian, Shuo-Han Chen, Chiy-Ferng Perng, Yung-Hsiang Lu |
ISLPED | 9 |
| 2023 | Skyrmion Vault: Maximizing Skyrmion Lifespan for Enabling Low-Power Skyrmion Racetrack MemoryabstractSkyrmion racetrack memory (SK-RM) has demonstrated great potential as a high-density and low-cost nonvolatile memory. Nevertheless, even though random data accesses are supported on SK-RM, data accesses can not be carried out on individual data bit directly. Instead, special skyrmion manipulations, such as injecting and shifting, are required to support random information update and deletion. With such special manipulations, the latency and energy consumption of skyrmion manipulations could quickly accumulate and induce additional overhead on the data read/write path of SK-RM. Meanwhile, injection operation consumes more energy and has higher latency than any other manipulations. Although prior arts have tried to alleviate the overhead of skyrmion manipulations, the possibility of minimizing injections through buffering skyrmions for future reuse and energy conservation receives much less attention. Such observation motivates us to propose the concept of skyrmion vault to effectively utilize the skyrmion buffer track structure for energy conservation through maximizing the lifespan of injected skyrmions and minimizing the number of skyrmion injections. Experimental results have shown promising improvements in both energy consumption and skyrmions' lifespan. Syue-Wei Lu, Shuo-Han Chen, Yu-Pei Liang, Yuan-Hao Chang 0001, Wang Kang 0001, Tseng-Yi Chen, Wei-Kuan Shih |
ASP-DAC | 2 |
| 2023 | Enabling Highly-Efficient DNA Sequence Mapping via ReRAM-based TCAMabstractIn the post-pandemic era, third-generation DNA sequencing (TGS) has received increasing attention from both academics and industries. As TGS technologies have become a requisite for extracting DNA sequences, the DNA sequence mapping, which is the most basic bioinformatics application and the core of polymerase chain reaction (PCR) tests, receives great challenges, due to the large size and noisy nature of TGS technologies. In addition, the ever-increasing data volume of DNA sequences also induces the issue of memory wall while large datasets are moved between the memory and the computing units. However, much less effort has been devoted to DNA sequence mapping acceleration while considering both the memory wall issue and the challenges of TGS technologies. To enable highly-efficient DNA sequence mapping, this study proposes a novel resistive random-access memory (ReRAM)-based ternary content-addressable memory (TCAM) and exploits the intrinsic parallelity of ReRAM crossbar for efficient mapping acceleration. Promising results have been demonstrated through a series of experiments with different scales of datasets. Yu-Shao Lai, Shuo-Han Chen, Yuan-Hao Chang 0001 |
ISLPED | 2 |
| 2023 | Sky-NN: Enabling Efficient Neural Network Data Processing with Skyrmion Racetrack MemoryabstractThe thriving of artificial intelligence has brought numerous efforts to build strengthened and sophisticated neural network models to resolve almost all kinds of problems in different academic fields. Owing to the growing complexity and size of neural networks, nonvolatile random access memory (NVRAM) has been utilized to avoid excessive data movements between volatile memory and persistent storage. Among various NVRAM alternatives, skyrmion racetrack memory (SK-RM) is regarded as a promising candidate owing to its high memory density and efficient reads and writes. Nevertheless, due to the distinct shift operation of SK-RM, directly applying existing data process methods of neural networks on SK-RM hinders the benefits and performance of both SK-RM and neural networks. To resolve this issue, this paper proposes Sky-NN to enable efficient NN data processing methods on SK-RM by utilizing the distinct shift and re-assemblability capability of skyrmions. A series of experiments were conducted to demonstrate the capability of Sky-NN. Yong-Cheng Liaw, Shuo-Han Chen, Yuan-Hao Chang 0001, Yu-Pei Liang |
ISLPED | 2 |
| 2022 | On Minimizing the Read Latency of Flash Memory to Preserve Inter-Tree Locality in Random ForestabstractMany prior research works have been widely discussed how to bring machine learning algorithms to embedded systems. Because of resource constraints, embedded platforms for machine learning applications play the role of a predictor. That is, an inference model will be constructed on a personal computer or a server platform, and then integrated into embedded systems for just-in-time inference. With the consideration of the limited main memory space in embedded systems, an important problem for embedded machine learning systems is how to efficiently move inference model between the main memory and a secondary storage (e.g., flash memory). For tackling this problem, we need to consider how to preserve the locality inside the inference model during model construction. Therefore, we have proposed a solution, namely locality-aware random forest (LaRF), to preserve the inter-locality of all decision trees within a random forest model during the model construction process. Owing to the locality preservation, LaRF can improve the read latency by 81.5% at least, compared to the original random forest library. Yu-Pei Liang, Tseng-Yi Chen, Yuan-Hao Chang 0001, Shuo-Han Chen, Wei-Kuan Shih |
ICCAD | 5 |
| 2022 | Evolving Skyrmion Racetrack Memory as Energy-Efficient Last-Level Cache DevicesabstractSkyrmion racetrack memory (SK-RM) has been regarded as a promising alternative to replace static random-access memory (SRAM) as a large-size on-chip cache device with high memory density. Different from other nonvolatile random-access memories (NVRAMs), data bits of SK-RM can only be altered or detected at access ports, and shift operations are required to move data bits across access ports along the racetrack. Owing to these special characteristics, word-based mapping and bit-interleaved mapping architectures have been proposed to facilitate reading and writing on SK-RM with different data layouts. Nevertheless, when SK-RM is used as an on-chip cache device, existing mapping architectures lead to the concerns of unpredictable access performance or excessive energy consumption during both data reads and writes. To resolve such concerns, this paper proposes extracting the merits of existing mapping architectures for allowing SK-RM to seamlessly switch its data update policy by considering the write latency requirement of cache accesses. Promising results have been demonstrated through a series of benchmark-driven experiments. Ya-Hui Yang, Shuo-Han Chen, Yuan-Hao Chang 0001 |
ISLPED | 2 |
| 2022 | KVSTL: An Application Support to LSM-Tree Based Key-Value Store via Shingled Translation Layer Data ManagementabstractLSM-tree based Key-value (KV) stores greatly fit the needs of write-intensive applications with its efficient data store and retrieval operations to datasets. To accommodate ever-growing datasets, shingled magnetic recording (SMR) drives have become a popular option to provide large storage capacity for KV stores at low cost. SMR drives achieve high storage density via overlapping tracks on the disk surface. However, the overlapped track layout induces the sequential-write constraint and prevents KV stores from storing and rearranging KV pairs efficiently. In this paper, we present KVSTL, a KV store aware Shingled Translation Layer (STL), to preserve the merits of existing KV stores, while exploiting the high storage density of SMR drives. KVSTL is proposed as an application support to hide the management complexity of SMR drives and facilitates the management of SMR drives via passing only the “level” and “invalidation” information of LSM-tree based KV stores onto SMR drives. The proposed KVSTL achieves its performance enhancement via managing key-value pairs with level awareness and enabling efficient storage space management with the invalidation information. The results show that KVSTL can reduce the written data amount for 69.45 percent on average and the latency for up to 62.72 percent when compared with SMR-based LevelDB. Shuo-Han Chen, Yuhong Liang, Ming-Chang Yang |
IEEE Trans. Computers | 1 |
| 2022 | MAGIC: Making IMR-Based HDD Perform Like CMR-Based HDDabstractThe past decades have witnessed the tremendous success of Conventional Magnetic Recording (CMR)-based Hard Disk Drives (HDDs) in data storage. To eliminate the bottleneck of CMR-based HDDs in providing higher areal density, an emerging Interlaced Magnetic Recording (IMR) is capable of achieving higher areal density with limited changes to disk makeup. Nevertheless, existing approaches for IMR-based HDDs may suffer serious read and write performance degradation as compared with CMR-based HDDs. Thus, this article presents a device-level solution, namelyMAGICtranslation layer, which aims atMAkinGIMR-based HDDs perform likeCMR-based HDDs in terms of comparable access performance. Specifically, not merely trying to improve the performance of raw IMR-based HDDs, this work, for the first time, moves one step forward to minimize the performance gap between IMR and CMR-based HDDs. Technically, by 1) fully utilizing two special CMR-like potentials of IMR and 2) gracefully trading the sequential access performance as space usage increases, MAGIC minimizes track rewriting overheads to achieve CMR-like performance. Our results reveal that MAGIC not only improves the write performance compared with existing designs, but also has potential to approach read and write performance of CMR-based HDD. Yuhong Liang, Ming-Chang Yang, Shuo-Han Chen |
IEEE Trans. Computers | 3 |
| 2022 | A File-Oriented Fast Secure Deletion Strategy for Shingled Magnetic Recording DrivesabstractNowadays, securely erasing deleted files has become one of the necessary tasks for users who want to protect their deleted data from malicious attackers. Nevertheless, existing secure deletion approaches are considered inefficient for erasing deleted files permanently because the file systems and storage devices do not share their file information or data layout with each other. On the emerging shingled magnetic recording (SMR) drives, the inefficiency of existing secure deletion approaches is exaggerated by the inherent sequential-write constraint of the high storage density SMR technology. On SMR drives, tracks are overlapped via utilizing the size difference between disk read/write heads to increase the storage density. Due to the overlapped track layout, secure deletion requests may induce a significant amount of write amplification and serious performance degradation if the data layout is not properly configured. Such observation motivates this article to come up with a file-oriented fast secure deletion (FFSD) strategy to deal with the sequential-write constraint of SMR drives and improve the efficiency of secure deletion operations on SMR drives. The experimental results show that the proposed strategy can effectively reduce the secure deletion latency by$286.15\times $on average when compared with the conventional approach. Shuo-Han Chen, Chun-Feng Wu, Ming-Chang Yang, Yuan-Hao Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | How to Enable Index Scheme for Reducing the Writing Cost of DNA Storage on Insertion and DeletionabstractRecently, the requirement of storing digital data has been growing rapidly; however, the conventional storage medium cannot satisfy these huge demands. Fortunately, thanks to biological technology development, storing digital data into deoxyribonucleic acid (DNA) has become possible in recent years. Furthermore, because of the attractive features (e.g., high storing density, long-term durability, and stability), DNA storage has been regarded as a potential alternative storage medium to store massive digital data in the future. Nevertheless, reading and writing digital data over DNA requires a series of extremely time-consuming processes (i.e., DNA sequencing and DNA synthesis). More specifically, among the two costs, the writing cost is the predominant cost of a DNA data storage system. Therefore, to enable efficient DNA storage, this article proposes an index management scheme for reducing the number of accesses to DNA storage. Additionally, this article introduces a new DNA data encoding format with VERA (Version Editing Recovery Approach) to reduce the total writing bits while inserting and deleting the data. To the best of our knowledge, this work is the first work to provide a total data management solution for DNA storage. According to the experimental results, the proposed design with VERA can reduce the cost by 77% and improve the performance by 71% compared to the append-only methods. Yi-Syuan Lin, Yu-Pei Liang, Tseng-Yi Chen, Yuan-Hao Chang 0001, Shuo-Han Chen, Hsin-Wen Wei, Wei-Kuan Shih |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2021 | Facilitating the Efficiency of Secure File Data and Metadata Deletion on SMR-based Ext4 File SystemabstractThe efficiency of secure deletion is highly dependent on the data layout of underlying storage devices. In particular, owing to the sequential-write constraint of the emerging Shingled Magnetic Recording (SMR) technology, an improper data layout could lead to serious write amplification and hinder the performance of secure deletion. The performance degradation of secure deletion on SMR drives is further aggravated with the need to securely erase the file system metadata of deleted files due to the small-size nature of file system metadata. Such an observation motivates us to propose a secure-deletion and SMR-aware space allocation (SSSA) strategy to facilitate the process of securely erasing both the deleted files and their metadata simultaneously. The proposed strategy is integrated within the widely-used extended file system 4 (ext4) and is evaluated through a series of experiments to demonstrate the effectiveness of the proposed strategy. The evaluation results show that the proposed strategy can reduce the secure deletion latency by 91.3% on average when compared with naive SMR-based ext4 file system. Ping-Xiang Chen, Shuo-Han Chen, Yuan-Hao Chang 0001, Yu-Pei Liang, Wei-Kuan Shih |
ASP-DAC | 2 |
| 2021 | Eco-feller: Minimizing the Energy Consumption of Random Forest Algorithm by an Eco-pruning Strategy over MLC NVRAMabstractRandom forest has been widely used to classifying objects recently because of its efficiency and accuracy. On the other hand, nonvolatile memory has been regarded as a promising candidate to be a part of a hybrid memory architecture. For achieving the higher accuracy, random forest tends to construct lots of decision trees, and then conducts some post-pruning methods to fell low contribution trees for increasing the model accuracy and space utilization. However, the cost of writing operations is always very high on non-volatile memory. Therefore, writing the to-be-pruned trees into non-volatile memory will significantly waste both energy and time. This work proposed a framework to ease such hurt of training a random forest model. The main spirit of this work is to evaluate the importance of trees before constructing it, and then adopts different writing modes to write the trees to the non-volatile memory space. The experimental results show the proposed framework can significantly mitigate the waste of energy with high accuracy. Yu-Pei Liang, Yung-Han Hsu, Tseng-Yi Chen, Shuo-Han Chen, Hsin-Wen Wei, Tsan-sheng Hsu, Wei-Kuan Shih |
DAC | 4 |
| 2021 | Brief Industry Paper: An Energy-Reduction On-Chip Memory Management for Intermittent SystemsabstractIntermittent systems enable continuous and accumulative process execution under constraint or unstable power supply. To enable intermittent computing, process status and data are typically checkpointed from volatile memory (VM) to nonvolatile memory (NVM) before running out of power. After power resumes, these logged data can be loaded back from NVM to VM for continuous execution. Nevertheless, existing approaches rarely considered the energy consumed during moving data and may waste precious power resource over data movement, instead of computation. Such observation motivates us to propose an energy-reduction on-chip memory management (ERCM2) scheme to utilize the high cell density and non-volatility of SpinTransfer Torque RAM (STT-RAM) for enabling a hybrid on chip memory architecture. The experimental results show that the proposed scheme can achieve the access performance close to conventional SRAM-based on-chip memory architecture with lower energy consumption. Yu-Pei Liang, Yu-Ting Fang, Shuo-Han Chen, Yen-Ting Chen, Tseng-Yi Chen, Wei-Lin Wang, Wei-Kuan Shih, Yuan-Hao Chang 0001 |
RTAS | 3 |
| 2021 | Facilitating external sorting on SMR-based large-scale storage systems
Chih-Hsuan Chen, Shuo-Han Chen, Yu-Pei Liang, Tseng-Yi Chen, Tsan-sheng Hsu, Hsin-Wen Wei, Wei-Kuan Shih |
Future Gener. Comput. Syst. | 2 |
| 2021 | On Minimizing Internal Data Migrations of Flash Devices via Lifetime-Retention HarmonizationabstractWith the emerge of high-density triple-level-cell (TLC) and 3D NAND flash, the access performance and endurance of flash devices are degraded due to the downscaling of flash cells. In addition, we observe that the mismatch between data lifetime requirement and flash block retention capability could further worsen the access performance and endurance. This is because the “lifetime-retention mismatch” could result in massive internal data migrations during garbage collection and data refreshing, and further aggravate the already-worsened access performance and endurance of high-density NAND flash devices. Such an observation motivates us to resolve the lifetime-retention mismatch problem by proposing a “time harmonization strategy”, which coordinates the flash block retention capability with the data lifetime requirement to enhance the performance of flash devices with very limited endurance degradation. Specifically, this study aims to lower the amount of internal data migrations caused by garbage collection and data refreshing via storing data of different lifetime requirement in flash blocks with suitable retention capability. The trace-driven evaluation results reveal that the proposed design can effectively reduce the average response time by about 99 percent on average without sacrificing the overall endurance, as compared with the state-of-the-art designs. Ming-Chang Yang, Chun-Feng Wu, Shuo-Han Chen, Yuan-Hao Chang 0001 |
IEEE Trans. Computers | 3 |
| 2021 | Optimizing Lifetime Capacity and Read Performance of Bit-Alterable 3-D NAND FlashabstractWith the technology advance of bit-alterable 3-D NAND flash, bit-level program and erase operations have been realized and provide the possibility of “bit-level rewrite.” Bit-level rewrite is predicted to be highly beneficial to the performance of the densely packed, bit-error-prone 3-D NAND flash because bit-level rewrites can remove error bits at bit-level granularity, shorten the error correction latency, and boost the read performance. Distinctly, bit-level rewrite can curtail the lifetime expense of refresh operations via correcting the error bit stored in the individual flash cell directly without a full-page rewrite, which is employed by previous refresh techniques. However, because bit-level rewrite is predicted to have similar latency and wearing as conventional full-page rewrites, the throughput of bit-level rewrites needs to be examined to avoid low rewrite efficiency. This observation inspires us to investigate and propose the bit-level error removal (BER) scheme to utilize the bit-level rewrites for optimizing both the read performance and lifetime capacity in a most-efficient way. The experimental results are encouraging and showed that the read performance can be improved by an average of 25.22% with 40.39% reduction of lifetime expense. Shuo-Han Chen, Ming-Chang Yang, Yuan-Hao Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Modular Neural Networks for Low-Power Image Classification on Embedded DevicesabstractEmbedded devices are generally small, battery-powered computers with limited hardware resources. It is difficult to run deep neural networks (DNNs) on these devices, because DNNs perform millions of operations and consume significant amounts of energy. Prior research has shown that a considerable number of a DNN’s memory accesses and computation are redundant when performing tasks like image classification. To reduce this redundancy and thereby reduce the energy consumption of DNNs, we introduce the Modular Neural Network Tree architecture. Instead of using one large DNN for the classifier, this architecture uses multiple smaller DNNs (called modules ) to progressively classify images into groups of categories based on a novel visual similarity metric. Once a group of categories is selected by a module, another module then continues to distinguish among the similar categories within the selected group. This process is repeated over multiple modules until we are left with a single category. The computation needed to distinguish dissimilar groups is avoided, thus reducing redundant operations, memory accesses, and energy. Experimental results using several image datasets reveal the effectiveness of our proposed solution to reduce memory requirements by 50% to 99%, inference time by 55% to 95%, energy consumption by 52% to 94%, and the number of operations by 15% to 99% when compared with existing DNN architectures, running on two different embedded systems: Raspberry Pi 3 and Raspberry Pi Zero. Abhinav Goel, Sarah Aghajanzadeh, Caleb Tung, Shuo-Han Chen, George K. Thiruvathukal, Yung-Hsiang Lu |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2020 | Boosting the Profitability of NVRAM-based Storage Devices via the Concept of Dual-Chunking Data DeduplicationabstractWith the latest advance in the non-volatile random-access memory (NVRAM), NVRAM is widely considered as the mainstream for the next-generation storage mediums. NVRAM has numerous attractive features, which include byte addressability, limited idle energy consumption, and great read/write access speed. However, owing to the high manufacturing cost of NVRAM, the incentive of deploying NVRAM in consumer electronics is lowered due to the consideration of profitability. To resolve the profitability issue and bring the benefits of NVRAM into the design of consumer electronics, avoiding storing duplicate data on NVRAM becomes a crucial task for lowering the demand and deployment cost of NVRAM. Such observation motivates us to propose a data deduplication extended file system design (DeEXT) to boost the profitability of NVRAM via the concept of dual-chunking data deduplication while considering the characteristics of NVRAM and duplicate data content. The proposed DeEXT was then evaluated by real-world data deduplication traces with encouraging results. Shuo-Han Chen, Yu-Pei Liang, Yuan-Hao Chang 0001, Hsin-Wen Wei, Wei-Kuan Shih |
ASP-DAC | 1 |
| 2020 | A Real-Time Feature Indexing System on Live Video StreamsabstractMost of the existing video storage systems rely on offline processing to support the feature-based indexing on video streams. The feature-based indexing technique provides an effective way for users to search video content through visual features, such as object categories (e.g., cars and persons). However, due to the reliance on offline processing, video streams along with their captured features cannot be searchable immediately after video streams are recorded. According to our investigation, buffering and storing live video steams are more time-consuming than the YOLO v3 object detector. Such observation motivates us to propose a real-time feature indexing (RTFI) system to enable instantaneous feature-based indexing on live video streams after video streams are captured and processed through object detectors. RTFI achieves its real-time goal via incorporating the novel design of metadata structure and data placement, the capability of modern object detector (i.e., YOLO v3), and the deduplication techniques to avoid storing repetitive video content. Notably, RTFI is the first system design for realizing real-time feature-based indexing on live video streams. RTFI is implemented on a Linux server and can improve the system throughput by upto 10.60x, compared with the base system without the proposed design. In addition, RTFI is able to make the video content searchable within 20 milliseconds for 10 live video streams after the video content is received by the proposed system, excluding the network transfer latency. Aditya Chakraborty, Akshay Pawar, Hojoung Jang, Shunqiao Huang, Sripath Mishra, Shuo-Han Chen, Yuan-Hao Chang 0001, George K. Thiruvathukal, Yung-Hsiang Lu |
COMPSAC | 6 |
| 2020 | Camera Placement Meeting Restrictions of Computer VisionabstractIn the blooming era of smart edge devices, surveillance cameras have been deployed in many locations. Surveillance cameras are most useful when they are spaced out to maximize coverage of an area. However, deciding where to place cameras is an NP-hard problem and researchers have proposed heuristic solutions. Existing work does not consider a significant restriction of computer vision: in order to track a moving object, the object must occupy enough pixels. The number of pixels depends on many factors (How far away is the object? What is the camera resolution? What is the focal length?). In this study, we propose a camera placement method that identifies effective camera placement in arbitrary spaces and can account for different camera types as well. Our strategy represents spaces as polygons, then uses a greedy algorithm to partition the polygons and determine the cameras' locations to provide the desired coverage. Our solution also makes it possible to perform object tracking via overlapping camera placement. Our method is evaluated against complex shapes and real-world museum floor plans, achieving up to 85% coverage and 25% overlap. Sarah Aghajanzadeh, Roopasree Naidu, Shuo-Han Chen, Caleb Tung, Abhinav Goel, Yung-Hsiang Lu, George K. Thiruvathukal |
ICIP | 3 |
| 2020 | Crowdsourcing Detection of Sampling Biases in Image DatasetsabstractDespite many exciting innovations in computer vision, recent studies reveal a number of risks in existing computer vision systems, suggesting results of such systems may be unfair and untrustworthy. Many of these risks can be partly attributed to the use of a training image dataset that exhibits sampling biases and thus does not accurately reflect the real visual world. Being able to detect potential sampling biases in the visual dataset prior to model development is thus essential for mitigating the fairness and trustworthy concerns in computer vision. In this paper, we propose a three-step crowdsourcing workflow to get humans into the loop for facilitating bias discovery in image datasets. Through two sets of evaluation studies, we find that the proposed workflow can effectively organize the crowd to detect sampling biases in both datasets that are artificially created with designed biases and real-world image datasets that are widely used in computer vision research and system development. Xiao Hu 0004, Anirudh Vegesana, Somesh Dube, Kaiwen Yu, Gore Kao, Shuo-Han Chen, Yung-Hsiang Lu, George K. Thiruvathukal, Ming Yin 0001 |
WWW | 7 |
| 2020 | A Partial Page Cache Strategy for NVRAM-Based Storage DevicesabstractNonvolatile random access memory (NVRAM) is becoming a popular alternative as the memory and storage medium in battery-powered embedded systems because of its fast read/write performance, byte-addressability, and nonvolatility. A well-known example is phase-change memory (PCM) that has much longer life expectancy and faster access performance than NAND flash. When NVRAM is considered as both main memory and storage in battery-powered embedded systems, existing page cache mechanisms have too many unnecessary data movements between main memory and storage. To tackle this issue, we propose the concept of “union page cache,” to jointly manage data of the page cache in both main memory and storage. To realize this concept, we design a partial page cache strategy that considers both main memory and storage as its management space. This strategy can eliminate unnecessary data movements between main memory and storage without sacrificing the data integrity of file systems. A series of experiments was conducted on an embedded platform. The results show that the proposed strategy can improve the file accessing performance up to 85.62% when PCM used as a case study. Shuo-Han Chen, Tseng-Yi Chen, Yuan-Hao Chang 0001, Hsin-Wen Wei, Wei-Kuan Shih |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | Beyond Address Mapping: A User-Oriented Multiregional Space Management Design for 3-D NAND Flash MemoryabstractDue to the ever-growing demands of larger capacity of flash storage devices, various new manufacturing techniques have been proposed to provide high-density and large-capacity NAND flash devices. Among these new techniques, 3-D NAND flash is regarded as one of the most promising candidates for the next-generation flash storage devices. 3-D NAND flash brings high bit density and significant cost saving via stacking memory cells vertically. However, the read/write and erase units of 3-D NAND flash also grow larger than those of traditional planner flash devices. This growing trend of read/write and erase units for 3-D NAND flash imposes significant management difficulties, such as the grown size of mapping information, decreased garbage collection efficiency, and worsened write amplification issue. To alleviate these negative impacts of the growing read/write and erase units, this paper proposes a multiregional space management design to achieve subpage-level management while adaptively adjusting mapping granularity by considering the user behaviors. The proposed design was evaluated by a series of experiments, and results show that the access performance can be improved by 64%. Shuo-Han Chen, Che-Wei Tsao, Yuan-Hao Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | B*-Sort: Enabling Write-Once Sorting for Nonvolatile MemoryabstractNonvolatile random access memory (NVRAM) has been regarding a promising technology to replace DRAM as the main memory in embedded systems owing to its nonvolatility and low idle power consumption. However, due to the asymmetric read/write costs and limited lifetime of NVRAM, most of the existing fundamental algorithms are not NVRAM-friendly with their write pattern and write intensiveness. Thus, existing fundamental algorithms for NVRAM embedded devices has been revealed. For instance, as the sorting algorithm is one of the most fundamental algorithms, most of the existing sorting algorithms are not NVRAM-friendly because they impose heavy write traffic [i.e., O(n lgn)] on main memory, where n is the number of unsorted elements. To resolve this issue, this article proposes a write-once sorting algorithm, namely B*-sort, to reduce the amount of write traffic on NVRAM-based main memory. B*sort adopts a brand-new concept, i.e., tree-based sort, inspired by the binary-search-tree structure to achieve the write-once property which can guarantee the optimal endurance during the sorting process. According to the experimental results, B*-sort can achieve significant performance improvement for sorting on NVRAM-based systems. Yu-Pei Liang, Tseng-Yi Chen, Yuan-Hao Chang 0001, Shuo-Han Chen, Hsin-Wen Wei, Wei-Kuan Shih |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | DSTL: A Demand-Based Shingled Translation Layer for Enabling Adaptive Address Mapping on SMR DrivesabstractShingled magnetic recording (SMR) is regarded as a promising technology for resolving the areal density limitation of conventional magnetic recording hard disk drives. Among different types of SMR drives, drive-managed SMR (DM-SMR) requires no changes on the host software and is widely used in today’s consumer market. DM-SMR employs a shingled translation layer (STL) to hide its inherent sequential-write constraint from the host software and emulate the SMR drive as a block device via maintaining logical to physical block address mapping entries. However, because most existing STL designs do not simultaneously consider the access pattern and the data update frequency of incoming workloads, those mapping entries maintained within the STL cannot be effectively managed, thus inducing unnecessary performance overhead. To resolve the inefficiency of existing STL designs, this article proposes a demand-based STL (DSTL) to simultaneously consider the access pattern and update frequency of incoming data streams to enhance the access performance of DM-SMR. The proposed design was evaluated by a series of experiments, and the results show that the proposed DSTL can outperform other SMR management approach by up to 86.69% in terms of read/write performance. Yi-Jing Chuang, Shuo-Han Chen, Yuan-Hao Chang 0001, Yu-Pei Liang, Hsin-Wen Wei, Wei-Kuan Shih |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2019 | The Best of Both Worlds: On Exploiting Bit-Alterable NAND Flash for Lifetime and Read Performance OptimizationabstractWith the emergence of bit-alterable 3D NAND flash, programming and erasing a flash cell at bit-level granularity have become a reality. Bit-level operations can benefit the high density, high bit-error-rate 3D NAND flash via realizing the "bit-level rewrite operation," which can refresh error bits at bit-level granularity for reducing the error correction latency and improving the read performance with minimal lifetime expense. Different from existing refresh techniques, bit-level operations can lower the lifetime expense via removing error bits directly without page-based rewrites. However, since bit-level rewrites may induce a similar amount of latency as conventional page-based rewrites and thus lead to low rewrite throughput, the efficiency of bit-level rewrites should be carefully considered. Such observation motivates us to propose a bit-level error removal (BER) scheme to derive the most-efficient way of utilizing the bit-level operations for both lifetime and read performance optimization. A series of experiments was conducted to demonstrate the capability of the BER scheme with encouraging results. Shuo-Han Chen, Ming-Chang Yang, Yuan-Hao Chang 0001 |
DAC | 1 |
| 2019 | Enabling File-Oriented Fast Secure Deletion on Shingled Magnetic Recording DrivesabstractExisting secure deletion approaches are inefficient in erasing data permanently because file systems have no knowledge of the data layout on the storage device, nor is the storage device aware of file information within the file systems. This inefficiency is exaggerated on the emerging shingled magnetic recording (SMR) drive due to its inherent sequential-write constraint. On SMR drives, secure deletion requests may lead to serious write amplification and performance degradation if the data layout is not properly configured. Such observation motivates us to propose a file-oriented fast secure deletion (FFSD) strategy to alleviate the negative impacts of SMR drives' sequential-write constraint and improve the efficiency of secure deletion operations on SMR drives. A series of experiments was conducted to demonstrate the capability of the proposed strategy on improving the efficiency of secure deletion on SMR drives. Shuo-Han Chen, Ming-Chang Yang, Yuan-Hao Chang 0001, Chun-Feng Wu |
DAC | 1 |
| 2019 | Rethinking Last-level-cache Write-back Strategy for MLC STT-RAM Main Memory with Asymmetric Write EnergyabstractTo meet the requirement of low-power consumption, multi-level-cell STT-RAM (MLC STT-RAM) has been widely regarded as a potential candidate for replacing DRAM-based main memory in the next generation computer architectures because of its high memory cell density, fast read/write performance and zero refresh power consumption. However, MLC STT-RAM has higher power consumption than DRAM while a write operation is performed because MLC STT-RAM sometimes needs to perform a two-step transition to change the originally stored bits to another specifically written bit patterns. As a result, MLC STT-RAM has different power consumption while different bit patterns are written to a memory cell. To the best of our knowledge, a few or none of the previous studies rethink a cache replacement policy to overcome the asymmetric write energy issue of MLC STT-RAM-based main memory. Thus, this study proposes an energy-aware cache replacement policy, namely E-cache, which considers asymmetric write-back power consumption on MLC STT-RAM-based main memory to evict a proper cached data from the last-level cache, so as to minimize system power consumption. The experimental results show that the proposed solution reduces the energy consumption by 36% on average, compared with the LRU. Yu-Pei Liang, Tseng-Yi Chen, Yuan-Hao Chang 0001, Shuo-Han Chen, Wei-Kuan Shih |
ISLPED | 4 |
| 2019 | 1+1>2: variation-aware lifetime enhancement for embedded 3D NAND flash systemsabstractThree-dimensional (3D) NAND flash has been developed to boost the storage capacity by stacking memory cells vertically. One critical characteristic of 3D NAND flash is its large endurance variation. With this characteristic, the lifetime will be determined by the unit with the worst endurance. However, few works can exploit the variations with acceptable overhead for lifetime improvement. In this paper, a variation-aware lifetime improvement framework is proposed. The basic idea is motivated by an observation that there is an elegant matching between unit endurance and wearing variations when wear leveling and implicit compression are applied together. To achieve the matching goal, the framework is designed from three-type-unit levels, including cell, line, and block, respectively. Series of evaluations are conducted, and the evaluation results show that the lifetime improvement is encouraging, better than that of the combination with the state-of-the-art schemes. Yejia Di, Liang Shi 0001, Shuo-Han Chen, Chun Jason Xue, Edwin H.-M. Sha |
LCTES | 3 |
| 2019 | Mitigating write amplification issue of SMR drives via the design of sequential-write-constrained cache
Yu-Pei Liang, Shuo-Han Chen, Yuan-Hao Chang 0001, Yong-Chin Lin, Hsin-Wen Wei, Wei-Kuan Shih |
J. Syst. Archit. | 2 |
| 2019 | Enabling Sequential-write-constrained B+-tree Index Scheme to Upgrade Shingled Magnetic Recording Storage PerformanceabstractWhen a shingle magnetic recording (SMR) drive has been widely applied to modern computer systems (e.g., archive file systems, big data computing systems, and large-scale database systems), storage system developers should thoroughly review whether current designs (e.g., index schemes and data placements) are appropriate for an SMR drive because of its sequential write constraint. Through many prior works excellently manage data in an SMR drive by integrating their proposed solutions into the driver layer, an index scheme over an SMR drive has never been optimized by any previous works because managing index over the SMR drive needs to jointly consider the properties of B + -tree and SMR natures (e.g., sequential write constraint and zone partitions) in a host storage system. Moreover, poor index management will result in terrible storage performance because an index manager is extensively used in file systems and database applications. For optimizing the B + -tree index structure over an SMR storage, this work identifies performance overheads caused by the B + -tree index structure in an SMR drive. By such observation, this study proposes a sequential-write-constrained B + -tree index scheme, namely SW-B + tree, which consists of an address redirection data structure, an SMR-aware node allocation mechanism, and a frequency-aware garbage collection strategy. According to our experiments, the SW-B + tree can improve the SMR storage performance 55% on average. Yu-Pei Liang, Tseng-Yi Chen, Yuan-Hao Chang 0001, Shuo-Han Chen, Kam-yiu Lam, Wei-Hsin Li, Wei-Kuan Shih |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2019 | mwJFS: A Multiwrite-Mode Journaling File System for MLC NVRAM StoragesabstractAt present, nonvolatile random access memory (NVRAM) is widely considered as a promising candidate for the next-generation storage medium due to its appealing characteristics, including short read/write latency, byte addressability, and low idle energy consumption. In addition, to provide a higher bit density, multilevel-cell (MLC) NVRAM has also been proposed. Nevertheless, when compared with conventional single-level-cell (SLC) NVRAM, MLC NVRAM has longer write latency and higher energy consumption. Hence, the performance of MLC NVRAM-based storage systems could be degraded due to the lengthened write latency. The performance degradation is further magnified by existing journaling file systems (JFS) on MLC NVRAM-based storage devices due to the JFS's fail-safe policy of writing the same data twice. Such observations motivate us to propose multiwrite-mode JFSs (mwJFSs) to alleviate the drawbacks of MLC NVRAM and boost the performance of MLC NVRAM-based JFS. The proposed mwJFS differentiates the data retention requirement of journaled data and applies different write modes to enhance the access performance with lower energy consumption. A series of experiments was conducted to demonstrate the capability of mwJFS on MLC NVRAM-based storage systems. Shuo-Han Chen, Yuan-Hao Chang 0001, Yu-Ming Chang, Wei-Kuan Shih |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Enabling union page cache to boost file access performance of NVRAM-based storage deviceabstractDue to the fast access performance, byte-addressability, and non-volatility of non-volatile random access memory (NVRAM), NVRAM has emerged as a popular candidate for the design of memory/storage systems on mobile computing systems. For example, the latest 3D xPoint memory could be a kind of NVRAM with much longer life expectancy than NAND flash and could ease the possible endurance issue. When NVRAM is considered as both main memory and storage in mobile computing systems, existing page cache mechanisms introduce too many unnecessary data movements between main memory and storage. To resolve this issue, we propose the concept of "union page cache," which jointly manages data of the page cache in both main memory and storage. To realize this concept, a partial page cache strategy is designed to consider both main memory and storage as its management space and to eliminate unnecessary data movements between main memory and storage without sacrificing the data consistency of file systems. Experimental results show that the proposed strategy can boost the file accessing performance upto 85.62% when using PCM as a case study. Shuo-Han Chen, Tseng-Yi Chen, Yuan-Hao Chang 0001, Hsin-Wen Wei, Wei-Kuan Shih |
DAC | 1 |
| 2018 | Enhancing the Energy Efficiency of Journaling File System via Exploiting Multi-Write Modes on MLC NVRAMabstractNon-volatile random-access memory (NVRAM) is regarded as a great alternative storage medium owing to its attractive features, including low idle energy consumption, byte addressability, and short read/write latency. In addition, multi-level-cell (MLC) NVRAM has also been proposed to provide higher bit density. However, MLC NVRAM has lower energy efficiency and longer write latency when compared with single-level-cell (SLC) NVRAM. These drawbacks could lead to higher energy consumption of MLC NVRAM-based storage systems. The energy consumption is magnified by existing journaling file systems (JFS) on MLC NVRAM-based storage devices due to the JFS's fail-safe policy of writing the same data twice. Such observations motivate us to propose a multi-write-mode journaling file systems (mwJFS) to alleviate the drawbacks of MLC NVRAM and lower the energy consumption of MLC NVRAM-based JFS. The proposed mwJFS differentiates the data retention requirement of journaled data and applies different write modes to enhance the energy efficiency with better access performance. A series of experiments was conducted to demonstrate the capability of mwJFS on a MLC NVRAM-based storage system. Shuo-Han Chen, Yuan-Hao Chang 0001, Tseng-Yi Chen, Yu-Ming Chang, Pei-Wen Hsiao, Hsin-Wen Wei, Wei-Kuan Shih |
ISLPED | 1 |
| 2018 | wrJFS: A Write-Reduction Journaling File System for Byte-addressable NVRAMabstractNon-volatile random-access memory (NVRAM) becomes a mainstream storage device in embedded systems due to its favorable features, such as small size, low power consumption, and short read/write latency. Unlike dynamic random access memory (DRAM), NVRAM has asymmetric performance and energy consumption on read/write operations. Generally, on NVRAM, a write operation consumes more energy and time than a read operation. Unfortunately, current mobile/embedded file systems, such as EXT2/3 and EXT4, are very unfriendly for NVRAM devices. The reason is that current mobile/embedded file systems employ a journaling mechanism for increasing its data reliability. Although a journaling mechanism raises the safety of data in a file system, it also repeatedly writes data to a data storage while data is committed and checkpointed. Though several related works have been proposed to reduce the amount of write traffic to NVRAM, they still cannot effectively minimize the write amplification of a journaling mechanism. Such observations motivate us to design a two-phase write reduction journaling file system called wrJFS. In the first phase, wrJFS classified data into two categories: Metadata and user data. As the size of metadata is usually very small (few bytes), byte-enabled journaling strategy will handle metadata during commit and checkpoint stages. In contrast, the size of user data is very large relative to metadata; thus, user data will be processed in the second phase. In the second phase, user data will be compressed by hardware encoder to reduce the write size and managed compressed-enabled journaling strategy to avoid the write amplification on NVRAM. Moreover, we analyze the overhead of wrJFS and show that the overhead is negligible. According to the experimental results, the proposed wrJFS outperforms other journaling file systems even though the experiments include the overhead of data compression. Tseng-Yi Chen, Yuan-Hao Chang 0001, Shuo-Han Chen, Chih-Ching Kuo, Ming-Chang Yang, Hsin-Wen Wei, Wei-Kuan Shih |
IEEE Trans. Computers | 3 |
| 2018 | An Erase Efficiency Boosting Strategy for 3D Charge Trap NAND FlashabstractOwing to the fast-growing demands of larger and faster NAND flash devices, new manufacturing techniques have accelerated the down-scaling process of NAND flash memory. Among these new techniques, 3D charge trap flash is considered to be one of the most promising candidates for the next-generation NAND flash devices. However, the long erase latency of 3D charge trap flash becomes a critical issue. This issue is exacerbated because the distinct transient voltage shift phenomenon is worsened when the number of program/erase cycle increases. In contrast to existing works that aim to tackle the erase latency issue by reducing the number of block erases, we tackle this issue by utilizing the “multi-block erase” feature. In this work, an erase efficiency boosting strategy is proposed to boost the garbage collection efficiency of 3D charge trap flash via enabling multi-block erase inside flash chips. A series of experiments was conducted to demonstrate the capability of the proposed strategy on improving the erase efficiency and access performance of 3D charge trap flash. The results show that the erase latency of 3D charge trap flash memory is improved by 75.76 percent on average even when the P/E cycle reaches$10^{4}$. Shuo-Han Chen, Yuan-Hao Chang 0001, Yu-Pei Liang, Hsin-Wen Wei, Wei-Kuan Shih |
IEEE Trans. Computers | 1 |
| 2018 | UnistorFS: A Union Storage File System Design for Resource Sharing between Memory and Storage on Persistent RAM-Based SystemsabstractWith the advanced technology in persistent random access memory (PRAM), PRAM such as three-dimen-sional XPoint memory and Phase Change Memory (PCM) is emerging as a promising candidate for the next-generation medium for both (main) memory and storage. Previous works mainly focus on how to overcome the possible endurance issues of PRAM while both main memory and storage own a partition on the same PRAM device. However, a holistic software-level system design should be proposed to fully exploit the benefit of PRAM. This article proposes a union storage file system (UnistorFS), which aims to jointly manage the PRAM resource for main memory and storage. The proposed UnistorFS realizes the concept of using the PRAM resource as memory and storage interchangeably to achieve resource sharing while main memory and storage coexist on the same PRAM device with no partition or logical boundary. This approach not only enables PRAM resource sharing but also eliminates unnecessary data movements between main memory and storage since they are already in the same address space and can be accessed directly. At the same time, the proposed UnistorFS ensures the persistence of file data and sanity of the file system after power recycling. A series of experiments was conducted on a modified Linux kernel. The results show that the proposed UnistorFS can eliminate unnecessary memory accesses and outperform other PRAM-based file systems for 0.2--8.7 times in terms of read/write performance. Shuo-Han Chen, Tseng-Yi Chen, Yuan-Hao Chang 0001, Hsin-Wen Wei, Wei-Kuan Shih |
ACM Trans. Storage | 1 |
| 2018 | A Progressive Performance Boosting Strategy for 3-D Charge-Trap NAND Flash
Shuo-Han Chen, Yen-Ting Chen, Yuan-Hao Chang 0001, Hsin-Wen Wei, Wei-Kuan Shih |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | Mitigating the Write Amplification Problem of Write-Optimized File Systems on Flash StorageabstractAs the volume of data stored by Big data and Cloud services continues to grow, both academia and industry are seeking for high-performance storage systems. Recently, with the recent advances in write-optimized indexes (WOI), WOI-based file systems can now outperform conventional file systems with orders of magnitude on random writes, metadata updates, and small file creation. Based on the B-tree structure, WOI-based file systems can not only process data faster than the conventional B-tree but also improve the range query performance. However, the write amplification of these WOI-based file systems becomes a serious performance overhead when adopting flash storage as underlying storage devices due to the recursive entry update behavior. To mitigate the write amplification problem of WOIbased file systems, we propose a flash-friendly WOI design to reduce the number of write requests on flash storage. To evaluate the performance of the proposed design, we adapt B+-tree as a case study and the experimental results are promising. Shuo-Han Chen, Jun-Long Lin, Tseng-Yi Chen, Tsan-sheng Hsu, Hsin-Wen Wei, Wei-Kuan Shih |
CLUSTER | 1 |
| 2017 | xB+-Tree: Access-Pattern-Aware Cache-Line-Based Tree for Non-volatile Main Memory ArchitectureabstractNon-volatile memory (NVM) has widely participated in the evolution of the next-generation memory architecture by way of being the substitution of the main memory. To cope with the problem of asymmetric read/write speeds of NVM, several excellent researches have been proposed to reduce the number of writes to the NVM-based main memory. Nevertheless, most of these existing approaches do not take the cache-line-based access behavior between the processor and the main memory into consideration. Thus, in order to essentially improve the access performance of the NVM-based memory architecture, this work aims to optimize the cache-line-based access performance over the NVM-based memory architecture based on the special access patterns in many popular internet of things (IoT) and in-memory database applications. Our experiments based on the well-known Gem5 full system simulator reveal that, compared to other existing representative approaches, the proposed design can effectively reduce the total execution time of insertion by 20.92~55.20% and improve the execution time of query by 2.06~23.36%. Li-Zheng Liang, Ming-Chang Yang, Yuan-Hao Chang 0001, Tseng-Yi Chen, Shuo-Han Chen, Hsin-Wen Wei, Wei-Kuan Shih |
COMPSAC (1) | 5 |
| 2017 | Enabling Write-Reduction Strategy for Journaling File Systems over Byte-addressable NVRAMabstractNon-volatile random-access memory (NVRAM) becomes a mainstream storage device in embedded systems due to its favorable features, such as small size, low power consumption, and short read/write latency. On NVRAM, a write operation consumes more energy and time than a read operation. However, current mobile/embedded file systems (e.g., EXT2/3 and EXT4) are very unfriendly for NVRAM devices. The reason is that a journaling mechanism writes the same data twice during data commitment and checkpoint. Such observations motivate this paper to design a two-phase write reduction journaling file system called wrJFS. In the first phase, wrJFS classified data into two categories: Metadata and user data. Metadata will be handled by partial byte-enabled journaling strategy, and user data will be processed in the second phase. In the second phase, user data will be compressed by hardware encoder so as to reduce the write size, and managed compressed-enabled journaling strategy to avoid the write amplification. The experimental results show that the proposed wrJFS can reduce the size of the write request by 89.7% on average, compared with the original EXT3. Tseng-Yi Chen, Yuan-Hao Chang 0001, Shuo-Han Chen, Chih-Ching Kuo, Ming-Chang Yang, Hsin-Wen Wei, Wei-Kuan Shih |
DAC | 3 |
| 2017 | Boosting the Performance of 3D Charge Trap NAND Flash with Asymmetric Feature Process Size CharacteristicabstractThe growing demands of large capacity fash-based storages have facilitated the down-scaling process of NAND fash memory. Among NAND fash technologies, 3D charge trap fash is regarded as one of the most promising candidates. Owing to the cylindrical geometry of vertical channels, the access performance of each page in one block is distinctive, and this situation is exaggerated in the 3D charge trap fash with the fast-growing number of layers. In this study, a progressive performance boosting strategy is proposed to boost the performance of 3D charge trap fash by utilizing its asymmetric page access speed feature. A series of experiments was conducted to demonstrate the capability of the proposed strategy on improving access performance of 3D charge trap flash. Shuo-Han Chen, Yen-Ting Chen, Hsin-Wen Wei, Wei-Kuan Shih |
DAC | 1 |
| 2017 | Enhancing Usability for the Wireless Charging Vehicle SimulatorabstractIn our previous work, a simulation framework was proposed to focus on imitating the behavior of wireless charging vehicles (WCVs) and wireless sensor networks (WSNs) because current mainstream simulators have very limited support on simulation of WCVs. The WCV is an integration of a mobile vehicle and a wireless power transfer broadcaster, which is used to recharge sensors wirelessly to prolong the lifetime of sensor networks. In the study of WCVs, simulators are extensively used to study the routing algorithms and behaviors of WCVs because it is very costly to build a WSN testbed and many specifications are not standardized. However, mainstream WSN simulators require researchers to develop and integrate their WCV modules. Therefore, the previously proposed framework aims to provide a simple framework for simulating of wireless power transfer and mobile vehicles. In this study, to further strength the usability of the proposed simulation framework, a graphic user interface is introduced to allow users to assign sensors' location and specify simulation parameters. Shuo-Han Chen, I-Ju Wang, Tseng-Yi Chen, Hsin-Wen Wei, Tsan-sheng Hsu, Wei-Kuan Shih |
ICCCN | 1 |
| 2017 | An update-overhead-aware caching policy for write-optimized file systems on SMR disksabstractTo accommodate the sheer volume of data in the era of Big Data and Cloud Computing, both new storage medium technologies and high-performance file systems are proposed. For storage medium, Shingled Magnetic Recording (SMR) increases the areal density by overlapping adjacent tracks so as to provide larger storage capacity. On the other hand, write-optimized indexes (WOI) file systems are also studied and can now outperform conventional file systems with orders of magnitude. However, the main drawback of SMR is the random-write restriction because random-write operations will cause the extra overhead of rewriting data stored in overlapped tracks. The rewriting overhead is amplified by the recursive entry update behavior of WOI-based file systems. Therefore, the rewriting issue becomes a serious performance overhead when adopting SMR drives as underlying storage devices for WOI-based file systems. To mitigate the write amplification problem when deploying WOI-based file systems on SMR disks, this paper proposes the update-overhead-aware caching policy to reduce the update overhead with the help of flash-based storage devices. To evaluate the performance of the proposed design, the B+-tree as a case study. The experimental results are promising. Shuo-Han Chen, Wei-Shin Li, Min-Hong Shen, Yi-Han Lien, Tseng-Yi Chen, Tsan-sheng Hsu, Hsin-Wen Wei, Wei-Kuan Shih |
IPCCC | 1 |
| 2017 | A wireless sensor network simulator focuses on imitating wireless charging vehicle: demo abstractabstractIn this live demonstration, we would like to present a simulation framework for simulating the behaviors of Wireless Sensor Network (WSN) and Wireless Charging Vehicle (WCV). Different to general purpose WSN simulators, the proposed simulation framework focuses on the simulation of mobile vehicles and wireless power transfer techniques. Besides, the proposed framework provides an easy-to-use user interface and a well structured system architecture. The proposed framework aims to eliminate the need for researchers to build their own simulator or integrate wireless charging modules, allowing them directly to evaluate their algorithm and compare the simulation results. Shuo-Han Chen, Yu-Pei Liang, Chi-Heng Lee, I-Ju Wang, Wei-Kuan Shih |
IPSN | 1 |
| 2017 | On Space Utilization Enhancement of File Systems for Embedded Storage SystemsabstractSince the mid-2000s, mobile/embedded computing systems conventionally have limited computing power, Random Access Memory (RAM) space, and storage capacity due to the consideration of their cost, energy consumption, and physical size. Recently, some of these systems, such as mobile phone and embedded consumer electronics, have more powerful computing capability, so they manage their data in small flash storage devices (e.g., Embedded Multi Media Card (eMMC) and Secure Digital (SD) cards) with a simple file system. However, the existing file systems usually have low space utilization for managing small files and the tail data of large files. In this work, we thus propose a dynamic tail packing scheme to enhance the space utilization of file systems over flash storage devices in embedded computing systems by dynamically aggregating/packing the tail data of (small) files together. To evaluate the benefits and overheads of the proposed scheme, we theoretically formulate analysis equations for obtaining the best settings in the dynamic tail packing scheme. Additionally, the proposed scheme was implemented in the file system of Linux operating systems to evaluate its capability. The results demonstrate that the proposed scheme could significantly improve the space utilization of existing file systems. Tseng-Yi Chen, Yuan-Hao Chang 0001, Shuo-Han Chen, Nien-I Hsu, Hsin-Wen Wei, Wei-Kuan Shih |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2016 | Enabling sub-blocks erase management to boost the performance of 3D NAND flash memoryabstract3D NAND has been proposed to provide a large capacity storage with low-cost consideration due to its high density memory architecture. However, 3D NAND needs to consume enormous time for garbage collection because of live-page copying overhead and long block erase time. To alleviate the impact of live-page copying on the performance of 3D NAND, a sub-block erase design has been designed. With sub-block erase design, this paper proposes a performance booster strategy to extremely boost the performance of garbage collection. As experimental results shows, the proposed strategy has a significant improvement on the average response time. Tseng-Yi Chen, Yuan-Hao Chang 0001, Chien-Chung Ho, Shuo-Han Chen |
DAC | 4 |
| 2015 | A QoS-Aware Data Reconstruction Strategy for a Data Fault-Tolerant Storage SystemabstractRecently, many applications and users rely on cloud storage services, such as Google drive, Dropbox, iCloud and Sky drive, to store private files and system data, and cloud storage services must thus be reliable and secure. To increase reliability, previous studies have proposed a variety of erasure coding algorithms for data fault tolerance for use in storage systems. Although these data fault tolerance mechanisms increase data reliability, implementation also increases storage system costs and energy consumption due to data redundancy. However, to date energy-efficient schemes have only been developed based on a RAID architecture, and none have been implemented using an erasure coding algorithm. To address this issue, this study proposes an energy-aware I/O framework with a quality-of-service (QoS) aware data reconstruction scheduler for erasure coding algorithms, called the EEC-scheme. This approach reduces storage system energy consumption and decreases response times for user requests when the system restores failed disks. A series of experiments show that the proposed scheme can significantly reduce power consumption in storage systems. Hsin-Wen Wei, Tseng-Yi Chen, Shuo-Han Chen, Nai-Yuan Jhang, Li-Zheng Liang, Chih-Ching Kuo, Tsan-sheng Hsu, Wei-Kuan Shih |
CloudCom | 3 |
| 2015 | Design a Hash-Based Control Mechanism in vSwitch for Software-Defined Networking EnvironmentabstractUnlike a traditional network architecture, a software-defined networking architecture is divided into the control plane and the data plane. Network administrators use the centralized control plane to manage network authority and determine where network traffic is to be sent in the data plane. However, a centralized control structure causes a bottleneck with an overloading flow or under a DDoS attack. Under such conditions, the probability of network misconfiguration may increase rapidly and network performance may decline rapidly. This work paper proposes a hash-based mechanism that operates in the control plane to increase the reliability and scalability of the network. The hash function is utilized to assign incoming packets to queues in the control plane. The controller schedules the queues using a round-robin method to reduce the probability of failure in response to malicious attacks and to reduce transmission delay when network congestion occurs. The experimental results reveal that the proposed mechanism effectively distributes the workload and increases the reliability of the network under high-density data transmission. Shih-Wen Hsu, Tseng-Yi Chen, Yung-Chun Chang, Shuo-Han Chen, Han-Chieh Chao, Tsen-Yeh Lin, Wei-Kuan Shih |
CLUSTER | 4 |
| 2015 | Prolong Lifetime of Dynamic Sensor Network by an Intelligent Wireless Charging VehicleabstractThe lifetime of wireless sensor networks are constrained by its limited battery capacity. Therefore, the lifetime is widely regard as a bottleneck of technique of wireless sensor network. Recently, the emerging breakthrough in wireless power transfer technique is expected to eliminate the power constraint bottleneck. In this paper, we propose an intelligent wireless charging vehicle (IWCV) strategy to resolve above problem in a dynamic and scalable approach. The IWCV strategy includes an intelligent routing strategy to traverse the sensor network topology and charging their battery to prolong their lifetime. What makes IWCV different to previous studies is that IWCV can still work even if the topology changes by re-computing the traversing route and stop time for each node in a relative short amount of time, compared with the time needed to find the shortest Hamiltoaian cycle. With the scalable and dynamic feature of IWCV, one can change their sensor network topology without down time to reconfigure the wireless charging vehicle while still maintain low energy consumption during traveling and charging in dynamic network topologies. Shuo-Han Chen, Yung-Chun Chang, Tseng-Yi Chen, Yu-Chun Cheng, Hsin-Wen Wei, Tsan-sheng Hsu, Wei-Kuan Shih |
VTC Fall | 1 |
| 2014 | An enhanced user interface design with Auto-Adjusting Icon Placement on foldable devicesabstractFlexible electronics appear in the consumer, medical, and military sectors. Thanks to the development of flexible electronics, flexible touchscreens have been widely carried on in various devices, such as mobile phones, wearable devices and hand-held tablets. Flexible touchscreens not only bring the technique of displays to next generation, but also significantly alter the interactive behaviors of users and devices. On the flexible touchscreens, when the displays are folded, some touch area around the folded line is not touchable in users' operation, and this is a critical research problem for the flexible touch screens. However, to our knowledge, little or no user interface research has solved this problem. To resolve this critical problem, in this study, we design a novel user interface, called the Auto-Adjusting Placement, which can dynamically adjust objects, such as icons, texts and pictures, on the flexible touch screens to avoid the area around the folded line and to keep the high availability/readability of the objects. Our demonstrations show that the Auto-Adjusting Placement is well performed on flexible touchscreens and therefore users have a consistent interface to use. We also filed a patent for this Auto-Adjusting Placement. Tseng-Yi Chen, Shuo-Han Chen, Heng-Yin Chen, Wei-Kuan Shih |
SMC | 2 |