VLDB 2026 Research / reviewers in the wild / expert
Che-Wei Tsao
dblp:130/1387
· DBLP profile ↗
19ranked-venue papers
6as first author
6since 2021 · last 2026
0009-0002-0434-7290ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Index for Square Pattern MatchingabstractA string S is called a square if it can be written as the concatenation of two identical strings. Two strings P and Q of the same length are said to square match if, for every substring of P, it is a square if and only if the corresponding substring of Q is also a square. The square pattern matching problem asks for locating all substrings of a given text T of length n that square match a query pattern P of length m. This notion captures similarity in repetition structures and is motivated by applications in areas such as bioinformatics and music structure analysis. In this paper, we introduce a novel technique, called the longest prefix square (LPS) encoding, which represents the square structure of a string as an integer array of the same length. We show that two strings square match if and only if they have identical LPS encodings. Based on this result, we construct an index solving the square pattern matching problem in time O(m lg m + occ) using O(nlg²n) bits of space, where occ denotes the number of occurrences of substrings in T that square match P. If the LPS encoding of P is precomputed, the query time improves to O(m + occ). Po-Chun Chen, Che-Wei Tsao, Wing-Kai Hon, Dominik Köppl |
CPM | 2 |
| 2026 | Enumeration of Unbordered Words in Compressed Representation
Che-Wei Tsao, Yi-Hua Lin, Wing-Kai Hon, Dominik Köppl |
DCC | 1 |
| 2026 | τ λ -Index: A framework for locating rare patterns in repetitive corpora
Che-Wei Tsao, Jin Jie Deng, Long-Qi Chen, Wing-Kai Hon, Dominik Köppl, Kunihiko Sadakane |
Inf. Syst. | 1 |
| 2024 | τλ-Index: Locating Rare Patterns in Similar StringsabstractWe propose a space-efficient index, called τλ-index, for locating rare patterns among similar strings, which has potential usage in drug discovery. Experiments are conducted on real data to compare our index with the state-of-the-art r-index. Long-Qi Chen, Che-Wei Tsao, Jin Jie Deng, Wing-Kai Hon, Kunihiko Sadakane |
DCC | 2 |
| 2023 | Retention-Aware Read Acceleration Strategy for LDPC-Based NAND Flash MemoryabstractWith the strong demand for stable and great quality of service in many network and multimedia services, flash-memory storage systems have been widely adopted in the storage I/O stack in servers and data centers to provide greater access performance. In these services, a huge-size storage system is essential. However, the huge-size flash storage system is very expensive. Flash storage vendors gradually adopt the high-density, low-reliability, and cost-efficient multiple-level cell (MLC) NAND flash memory chip as the major storage medium. Unfortunately, MLC NAND flash memory also brings about the critical issue of the high raw bit error rate. To resolve this issue, vendors adopt the more complex error correction code [such as low-density parity-check (LDPC)]. However, LDPC also results in significant read performance degradation due to its multiple read-retry sensing and decoding steps. To resolve this issue, we proposed a retention-aware read acceleration design (referred to as RRA) for the LDPC-based flash storage system to maintain stable and great read performance without significantly affecting the lifetime. Without significantly modifying the existing flash translation layer (FTL) design, we proposed a retention-aware management module to the existing FTL design. This module can efficiently identify and predict the data access characteristics and precisely allocate the suitable blocks for different data. The proposed design was evaluated with a series of experiments. The experiment results demonstrate that it could effectively reduce average read response time without significantly increasing the number of total live-page copying compared to the typical wear-leveling strategy. Tse-Yuan Wang, Che-Wei Tsao, Yuan-Hao Chang 0001, Tei-Wei Kuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Rethinking the Interactivity of OS and Device Layers in Memory ManagementabstractIn the big data era, a huge number of services has placed a fast-growing demand on the capacity of DRAM-based main memory. However, due to the high hardware cost and serious leakage power/energy consumption, the growth rate of DRAM capacity cannot meet the increased rate of the required main memory space when the energy or hardware cost is a critical concern. To tackle this issue, hybrid main-memory devices/modules have been proposed to replace the pure DRAM main memory with a hybrid main memory module that provides a large main memory space by integrating a small-sized DRAM and a large-sized non-volatile memory (NVM) into the same memory module. Although NVMs have high-density and low-cost features, they suffer from the low read/write performance and low endurance issue, compared to DRAM. Thus, inside the hybrid main-memory module, it also includes a memory management design to use DRAM as the cache of NVMs to enhance its performance and lifetime. However, it also introduces new design challenges in both the OS and the memory module. In this work, we rethink the interactivity of OS and hybrid main-memory module, and propose a cross-layer cache design that (1) utilizes the information from the operating system to optimize the hit ratio of the DRAM cache inside the memory module, and (2) takes advantage of the bulk-size (or block-based) read/write feature of NVM to minimize the time overhead on the data movement between DRAM and NVM. At the same time, this cross-layer cache design is very lightweight and only introduces limited runtime management overheads. A series of experiments was conducted to evaluate the effectiveness of the proposed cross-layer cache design. The results show that the proposed design could improve access performance for up to 88%, compared to the investigated well-known page replacement algorithms. Tse-Yuan Wang, Chun-Feng Wu, Che-Wei Tsao, Yuan-Hao Chang 0001, Tei-Wei Kuo, Xue (Steve) Liu |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | Determinizing Crash Behavior with a Verified Snapshot-Consistent Flash Translation Layer
Yun-Sheng Chang, Yao Hsiao, Tzu-Chi Lin, Che-Wei Tsao, Chun-Feng Wu, Yuan-Hao Chang 0001, Hsiang-Shang Ko, Yu-Fang Chen 0001 |
OSDI | 4 |
| 2020 | Beyond Address Mapping: A User-Oriented Multiregional Space Management Design for 3-D NAND Flash MemoryabstractDue to the ever-growing demands of larger capacity of flash storage devices, various new manufacturing techniques have been proposed to provide high-density and large-capacity NAND flash devices. Among these new techniques, 3-D NAND flash is regarded as one of the most promising candidates for the next-generation flash storage devices. 3-D NAND flash brings high bit density and significant cost saving via stacking memory cells vertically. However, the read/write and erase units of 3-D NAND flash also grow larger than those of traditional planner flash devices. This growing trend of read/write and erase units for 3-D NAND flash imposes significant management difficulties, such as the grown size of mapping information, decreased garbage collection efficiency, and worsened write amplification issue. To alleviate these negative impacts of the growing read/write and erase units, this paper proposes a multiregional space management design to achieve subpage-level management while adaptively adjusting mapping granularity by considering the user behaviors. The proposed design was evaluated by a series of experiments, and results show that the access performance can be improved by 64%. Shuo-Han Chen, Che-Wei Tsao, Yuan-Hao Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | Proactive channel adjustment to improve polar code capability for flash storage devicesabstractWith the low encoding/decoding complexity and the high error correction capability, polar code with the support of list-decoding and cyclic redundancy check can outperform LDPC code in the area of data communication. Thus, it also draws a lot of attentions on how to adopt and enable polar codes in storage applications. However, the code construction and encoding length limitation issues obstruct the adoption of polar codes in flash storage devices. To enable polar codes in flash storage devices, we propose a proactive channel adjustment design to extend the effective time of a code construction to improve the error correction capability of polar codes. This design pro-actively tunes the quality of the critical flash cells to maintain the correctness of the code construction and relax the constraint of the encoding length limitation, so that polar codes can be enabled in flash storage devices. A series of experiments demonstrates that the proposed design can effectively improve the error correction capability of polar codes in flash storage devices. Kun-Cheng Hsu, Che-Wei Tsao, Yuan-Hao Chang 0001, Tei-Wei Kuo |
DAC | 2 |
| 2018 | Boosting NVDIMM Performance With a Lightweight Caching Algorithm
Che-Wei Tsao, Yuan-Hao Chang 0001, Tei-Wei Kuo |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | Distillation: A light-weight data separation design to boost performance of NVDIMM main memoryabstractIn the big data era, data-intensive applications have growing demand for the capacity of DRAM main memory, but the frequent DRAM refresh, high leakage power, and high unit cost bring serious design issues on scaling up DRAM capacity. To address this issue, NVDIMM, which is a hybrid memory module, becomes a possible alternative to replace DRAM as main memory in some data-intensive applications. NVDIMM that consists of a small-sized high-speed DRAM and a large-sized low-cost non-volatile memory (i.e., flash memory) has the serious performance issue on accessing data stored in the flash memory because of the huge performance gap between DRAM and flash memory. However, there is no or limited room to adopt a complex caching algorithm for using DRAM as the cache of flash memory in NVDIMM main memory because a complex caching algorithm itself would already cause too much performance degradation on handling each request to NVDIMM main memory. In this paper, we present a light-weight data separation design to boost NVDIMM performance with limited data separation overhead for reducing the data accesses to flash memory. A series of experiments was conducted based on popular benchmarks, and the results demonstrate that the proposed design can effectively improve the performance of NVDIMM main memory. Che-Wei Tsao, Yuan-Hao Chang 0001, Tei-Wei Kuo, Shau-Yin Tseng |
RTCSA | 1 |
| 2016 | Multi-version checkpointing for flash file systemsabstractReliability has become a critical design issue in flash storage systems, because of the adoption of the low-cost, high-error-rate flash chips to fulfill the needs of the fast-growing storage capacity. In this paper, a multi-version checkpointing strategy is proposed to resolve the reliability issue of flash storage systems from the perspective of flash file systems. The proposed strategy can efficiently and effectively utilize checkpoints of file systems to guarantee the integrity and consistency of flash file systems after files or flash pages are corrupted. By utilizing the coexistence fact of multiple versions of the same data in flash memory, a control/recovery mechanism is presented to maintain checkpoints and to recover file systems with minimized management and recovery time overheads. A series of experiments was conducted based on realistic traces that were collected from benchmarks running over flash file systems in Linux operating systems. The results illustrate that the proposed strategy can significantly improve the reliability of flash file systems, as compared with other existing designs. Shih-Chun Chou, Yuan-Hao Chang 0001, Yuan-Hung Kuan, Po-Chun Huang, Che-Wei Tsao |
ASP-DAC | 5 |
| 2016 | Graceful Space Degradation: An Uneven Space Management for Flash Storage DevicesabstractThe high cell density, multilevel-cell programming, and manufacturing process variance force the new coming flash memory to have large bit-error-rate variance among blocks and pages, where a flash chip consists of multiple blocks and each block consists of a fixed number of pages. In order to avoid storing the crucial user data in more fragile pages, conventional flash management software tends to aggressively discard the high bit-error-rate area in the unit of a block. However, together with the aggressive discarding strategies and the enlarging sizes of pages/blocks of next generation flash memory, the available space of flash devices might encounter a very sharp degradation and therefore result in rapidly-shortened device lifespan. Thus, we advocate the concept of “graceful space degradation” to mitigate this problem by discarding the high bit-error-rate (or worn-out) area in the unit of pages (instead of blocks). To furthermore realize this concept, we are the pioneer to put forward an “uneven space management” to manage flash blocks containing different number of bad pages. Our design especially focuses on placing data with different access behaviors to make the best uses of blocks with different available space so as to ultimately prolong the device lifespan with good access performance. The experiments were conducted based on representative realistic workloads, and the results reveal that the proposed design can extend the device lifetime by at least 2.38 times of that of existent approaches, with very limited performance overheads. Ming-Chang Yang, Yuan-Hao Chang 0001, Yuan-Hung Kuan, Che-Wei Tsao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Byte-Addressable Update Scheme to Minimize the Energy Consumption of PCM-Based Storage SystemsabstractIn recent years, phase-change memory (PCM) has generated a great deal of interest because of its byte addressability and nonvolatility properties. It is regarded as a good alternative storage medium that can reduce the performance gap between the main memory and the secondary storage in computing systems. However, its high energy consumption on writes is a challenging issue in the design of battery-powered mobile computing systems. To reduce the energy consumption, we exploit the byte addressability and the asymmetric read-write energy/latency of PCM in an energy-efficient update scheme for journaling file systems. We also introduce a concept called the 50% rule to determine/recommend the best update strategy for block updates. The proposed scheme only writes modified data, instead of the whole updated block, to PCM-based storage devices without extra hardware support. Moreover, it guarantees the sanity/integrity of file systems even if the computing system crashes or there is a power failure during the data update process. We implemented the proposed scheme on the Linux system and conducted a series of experiments to evaluate the scheme. The results are very encouraging. Ming-Chang Yang, Yuan-Hao Chang 0001, Che-Wei Tsao |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | Efficient Victim Block Selection for Flash Storage DevicesabstractMotivated by the needs to enhance the performance of garbage collection in low-cost flash storage devices, we propose a victim block selection design to efficiently identify the blocks for erases and reclaim the space of invalid data without extensively scanning flash memory for the data status stored in the storage, so as to improve the garbage collection performance on reclaiming the space of invalid data. Moreover, this design could easily identify and reclaim the space released by file systems. Experiments based on benchmark traces show significant performance improvement of garbage collection with limited system overheads. Che-Wei Tsao, Yuan-Hao Chang 0001, Ming-Chang Yang, Po-Chun Huang |
IEEE Trans. Computers | 1 |
| 2014 | Garbage collection and wear leveling for flash memory: Past and futureabstractRecently, storage systems have observed a great leap in performance, reliability, endurance, and cost, due to the advance in non-volatile memory technologies, such as NAND flash memory. However, although delivering better performance, shock resistance, and energy efficiency than mechanical hard disks, NAND flash memory comes with unique characteristics and operational constraints, and cannot be directly used as an ideal block device. In particular, to address the notorious write-once property, garbage collection is necessary to clean the outdated data on flash memory. However, garbage collection is very time-consuming and often becomes the performance bottleneck of flash memory. Moreover, because flash memory cells endure very limited writes (as compared to mechanical hard disks) before they are worn out, the wear-leveling design is also indispensable to equalize the use of flash memory space and to prolong the flash memory lifetime. In response, this paper surveys state-of-the-art garbage collection and wear-leveling designs, so as to assist the design of flash memory management in various application scenarios. The future development trends of flash memory, such as the widespread adoption of higher-level flash memory and the emerging of three-dimensional (3D) flash memory architectures, are also discussed. Ming-Chang Yang, Yu-Ming Chang, Che-Wei Tsao, Po-Chun Huang, Yuan-Hao Chang 0001, Tei-Wei Kuo |
SMARTCOMP | 3 |
| 2013 | Performance enhancement of garbage collection for flash storage devices: an efficient victim block selection designabstractMotivated by the needs to enhance the performance of garbage collection in low-cost flash storage devices, we propose a victim block selection design to efficiently identify the blocks for erases and reclaim the space of invalid data without extensively scanning flash memory for the status of data stored in the storage, so as to achieve improved performance of garbage collection on reclaiming space of invalid data. At the same time, this design could also easily identify and reclaim the space released by file systems. A series of experiments based on benchmark traces demonstrates the significantly improved performance of garbage collection with limited system overheads. Che-Wei Tsao, Yuan-Hao Chang 0001, Ming-Chang Yang |
DAC | 1 |
| 2013 | New ERA: new efficient reliability-aware wear leveling for endurance enhancement of flash storage devicesabstractAs the program/erase (P/E) cycles of flash memory keep decreasing, improving the lifetime/endurance of flash memory has become a fundamental issue in the design of flash devices. This work is motivated by the observation that flash blocks endured the same P/E cycles usually have different bit error rates. In contrast to the existing wear-leveling techniques that try to distribute erases to flash blocks as evenly as possible, we propose an efficient reliability-aware wear-leveling scheme to distribute block erases based on the bit error rates of blocks so as to even out the error rate among flash blocks, to maximize the number of good blocks, and thus to ultimately prolong the lifetime of flash storage devices. The experiments were conducted based on representative realistic workloads to evaluate the efficacy of the proposed scheme, for which the results are very encouraging. Ming-Chang Yang, Yuan-Hao Chang 0001, Che-Wei Tsao, Po-Chun Huang |
DAC | 3 |
| 2013 | A fifty-percent rule to minimize the energy consumption of PCM-based storage systemsabstractIn recent years, phase-change memory (PCM) has drawn a lot of attention because of its byte-addressability and non-volatility. It has become a good alternative storage medium to reduce the performance gap between main memory and secondary storage, but its high energy consumption on writes is a challenging issue in the design of battery-powered mobile computing systems. By utilizing the byte-addressability and the asymmetric read-write energy/latency of PCM, we propose an energy-efficient update scheme with a fifty-percent rule for journaling file systems to reduce the energy consumption. This scheme only writes the modified data, instead of the whole updated block, to PCM-based storage devices with the guarantee of the sanity/integrity of file systems even if the system crashes or power failure occurs during the process of data updates. A series of experiments based on the implementation on the Linux system was conducted to evaluate the capability of the proposed scheme, and the results are very encouraging. Ming-Chang Yang, Martin Kuo, Che-Wei Tsao, Yuan-Hao Chang 0001 |
RTCSA | 3 |