EDBT 2026 Demo / reviewers in the wild / expert
Yun-Chih Chen
dblp:25/9202
· DBLP profile ↗
16ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-2665-9490ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Zone-aware metadata placement in B-tree filesystem
Ming-Feng Wei, Yun-Chih Chen, Yuan-Hao Chang 0001, Tei-Wei Kuo |
ASP-DAC | 2 |
| 2026 | Anytime ROS 2: Timely Task Completion in Non-Preemptive Robotic Systems
Harun Teper, Daniel Kuhse, Yun-Chih Chen, Georg von der Brüggen, Zhishan Guo, Jian-Jia Chen |
RTAS | 3 |
| 2026 | Bloom Merge Strategies for Sustainable SSD Endurance in Write-Intensive LSM-TreesabstractLSM-Tree is a critical data structure designed for write-optimized, user-facing key-value databases. However, LSM-Trees must frequently perform data merge operations to maintain read efficiency and discard obsolete data. These operations generate a considerable amount of write activity on the storage device (e.g., an SSD), which can drastically reduce the device’s lifespan. Often, these merges involve rewriting data that has not changed, a process that could be avoided. Recognizing this, we introduce “Bloom Merge,” an innovative merge strategy for LSM-Trees specifically developed for SSD. Based on the key distribution in LSM-Tree’s SSTables, this method selectively and efficiently perform merges, only when necessary. It also mitigates the potential negative impact on read performance through the strategic use of in-memory Bloom Filters. We present several key insights into determining the optimal conditions for merging and outline strategies that achieve a balanced improvement in both read and write performance. Our evaluation demonstrates that Bloom Merge significantly enhances write efficiency while reducing unnecessary operations. Yi-Hua Chen, Wei-Chun Cheng, Yun-Chih Chen, Wei-Kuan Shih, Yuan-Hao Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Beyond Bandwidth Doubling: Embrace Bit-Flips and Unlock Processing-in-NANDabstractNVMe SSDs offer unprecedented capacity and bandwidth and upcoming PCIe standards promise even more. However, the underlying technology, NAND memory, already struggles with significant heat and power consumption challenges. Just like microprocessors before, NAND also experiences Dark Silicon, preventing performance from improving at the same pace as capacity. Much of the power (and thus heat) within a NAND chip results from transferring data at a high rate, another symptom of a compute-centric style of processing. Therefore, we argue for data-centric Processing-in-NAND (PiN). However, PiN comes with significant challenges, such as limited capabilities and the need to cope with bit-flip errors. Even beyond Processing-in-Memory (PiM), databases may soon have to accept that memory is not error-free, an assumption that comes at a significant cost in power, capacity and performance. Our discussion indicates that no PiN design will serve as a singular, universally applicable solution to the limit of bandwidth scaling. Instead, successful integration into database architecture requires carefully identifying PiN-compatible functionality and abstractions, and cooperation with other innovations, such as Computational Storage and CXL. Lastly, we analyze the fundamental error tolerance of Bloom filters and binary sketches as PiM-compatible data structures, which we believe may be of independent interest. Maximilian Berens, Yun-Chih Chen, Jian-Jia Chen, Jens Teubner |
ICDE | 2 |
| 2024 | Search-in-Memory (SiM): Conducting Data-Bound Computations on Flash Chip for Enhanced EfficiencyabstractLarge-scale data systems utilize indexes like hash tables and trees for efficient data retrieval. These indexes are stored on disk and loaded into DRAM on demand, where they are post-processed and analyzed by the CPU. This method incurs substantial data 110, especially when optimizations like prefetching is used. This issue is inherent in the von Neumann architecture, where storage systems are dedicated solely to data storage, while CPUs handle all computations. However, data indexing primarily involves filtering tasks, which require only simple equality tests and not the complex arithmetic capabilities of a CPU. This inefficiency in the von Neumann architecture has led to a growing interest in in-memory computing, initially centered on DRAM. Recently, NAND flash-based in-storage computing has gained attention due to its ability to compute over larger working sets without requiring initial memory loading. In response, we propose the Search-in-Memory (SiM) chip, which minimally modifies an existing flash memory chip to allow it to conduct equality tests internally and send only relevant search results, not the entire data page. Specifically, we implement data filtering by using the existing logic gates in a flash memory chip's peripheral circuits for bit-serial equality tests, which processes all bits on a page simultaneously. Additionally, we introduce a versatile SIMD interface with two primary commands: search and gather, making SiM adaptable to different application scenarios. We use “Optimistic Error Correction” to efficiently ensure data accuracy. Our evaluations show that this new architecture could significantly improve throughput over traditional CPU -centric architectures. Yun-Chih Chen, Yuan-Hao Chang 0001, Tei-Wei Kuo |
DATE | 1 |
| 2024 | LUTIN: Efficient Neural Network Inference with Table LookupabstractDNN models are becoming increasingly large and complex, but they are also being deployed on commodity devices that require low power and latency but lack specialized accelerators. We introduce LUTIN (LUT-based INference), which reduces the amount of matrix multiplication in DNN inference by converting it into table lookups. LUTIN's innovation is its use of hyperparameter optimization to refine the quantization process and vector partitioning, allowing it to run efficiently on a variety of hardware. By reducing off-chip memory lookups and designing a cache-efficient data layout, LUTIN reduces energy consumption while increasing the use of available CPU cache, even on devices with limited processing power. Our approach goes beyond the traditional limitations of 8-bit quantization, investigating lower bit-widths to further reduce LUT size while meeting accuracy requirements. Experimental results show that LUTIN achieves up to a 2.34x speedup in latency and a 2.04x improvement in energy efficiency over full-precision models. Shi-Zhe Lin, Yun-Chih Chen, Yuan-Hao Chang 0001, Tei-Wei Kuo, Hsiang-Pang Li |
ISLPED | 2 |
| 2024 | Search-in-Memory: Reliable, Versatile, and Efficient Data Matching in SSD's NAND Flash Memory Chip for Data Indexing AccelerationabstractTo index the increasing volume of data, modern data indexes are typically stored on solid-state drives and cached in DRAM. However, searching such an index has resulted in significant I/O traffic due to limited access locality and inefficient cache utilization. At the heart of index searching is the operation of filtering through vast data spans to isolate a small, relevant subset, which involves basic equality tests rather than the complex arithmetic provided by modern CPUs. This article demonstrates the feasibility of performing data filtering directly within a NAND flash memory chip, transmitting only relevant search results rather than complete pages. Instead of adding complex circuits, we propose repurposing existing circuitry for efficient and accurate bitwise parallel matching. We demonstrate how different data structures can use our flexible SIMD command interface to offload index searches. This strategy not only frees up the CPU for more computationally demanding tasks, but it also optimizes DRAM usage for write buffering, significantly lowering energy consumption associated with I/O transmission between the CPU and DRAM. Extensive testing across a wide range of workloads reveals up to a$9\times $speedup in write-heavy workloads and up to 45% energy savings due to reduced read and write I/O. Furthermore, we achieve significant reductions in median and tail read latencies of up to 89% and 85%, respectively. Yun-Chih Chen, Yuan-Hao Chang 0001, Tei-Wei Kuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | AttentionRC: A Novel Approach to Improve Locality Sensitive Hashing Attention on Dual-Addressing MemoryabstractAttention is a crucial component of the Transformer architecture and a key factor in its success. However, it suffers from quadratic growth in time and space complexity as input sequence length increases. One popular approach to address this issue is the Reformer model, which uses locality-sensitive hashing (LSH) attention to reduce computational complexity. LSH attention hashes similar tokens in the input sequence to the same bucket and attends tokens only within the same bucket. Meanwhile, a new emerging nonvolatile memory (NVM) architecture, row column NVM (RC-NVM), has been proposed to support row- and column-oriented addressing (i.e., dual addressing). In this work, we present AttentionRC, which takes advantage of RC-NVM to further improve the efficiency of LSH attention. We first propose an LSH-friendly data mapping strategy that improves memory write and read cycles by 60.9% and 4.9%, respectively. Then, we propose a sort-free RC-aware bucket access and a swap strategy that utilizes dual-addressing to reduce 38% of the data access cycles in attention. Finally, by taking advantage of dual-addressing, we propose transpose-free attention to eliminate the transpose operations that were previously required by the attention, resulting in a 51% reduction in the matrix multiplication time. Chun-Lin Chu, Yun-Chih Chen, Wei Cheng 0006, Ing-Chao Lin, Yuan-Hao Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | GEAR: Graph-Evolving Aware Data Arranger to Enhance the Performance of Traversing Evolving Graphs on SCMabstractIn the era of big data, social network services continuously modify social connections, leading to dynamic and evolving graph data structures. These evolving graphs, vital for representing social relationships, pose significant memory challenges as they grow over time. To address this, storage-class-memory (SCM) emerges as a cost-effective solution alongside DRAM. However, contemporary graph evolution processes often scatter neighboring vertices across multiple pages, causing weak graph spatial locality and high-TLB misses during traversals. This article introduces SCM-Based graph-evolving aware data arranger (GEAR), a joint management middleware optimizing data arrangement on SCMs to enhance graph traversal efficiency. SCM-based GEAR comprises multilevel page allocation, locality-aware data placement, and dual-granularity wear leveling techniques. Multilevel page allocation prevents scattering of neighbor vertices relying on managing each page in a finer-granularity, while locality-aware data placement reserves space for future updates, maintaining strong graph spatial locality. The dual-granularity wear leveler evenly distributes updates across SCM pages with considering graph traversing characteristics. Evaluation results demonstrate SCM-based GEAR’s superiority, achieving 23% to 70% reduction in traversal time compared to state-of-the-art frameworks. Wen-Yi Wang, Chun-Feng Wu, Yun-Chih Chen, Tei-Wei Kuo, Yuan-Hao Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | APP: Enabling Soft Real-Time Execution on Densely-Populated Hybrid Memory SystemabstractMemory swapping was considered slow and evil, but swapping to Ultra Low-Latency storage like Optane has become a promising solution to save power and cost, helping densely-populated edge server to overcome its DRAM capacity bottleneck. However, the lack of integration between CPU scheduling and memory paging causes soft real-time tasks running on edge servers to miss deadlines under heavy memory multiplexing. We propose APP (Adaptive Page Pinning), lightweight protection of working set memory to ensure meeting soft real-time task deadlines without starving other non-real-time tasks. Experiments show that APP alleviates thrashing in memory-intensive tasks and upholds soft real-time task deadlines. Zheng-Wei Wu, Yun-Chih Chen, Yuan-Hao Chang 0001, Tei-Wei Kuo |
DAC | 2 |
| 2023 | HAPIC: A Scalable, Lightweight and Reactive Cache for Persistent-Memory-Based IndexabstractIn-memory index delivers low-latency responses for data services. It has been ported to high-capacity persistent memory (PM) to accommodate more data. However, read-heavy, extremely-skewed, and highly-dynamic workloads can suffer from degraded performance on PM-based indexes. We present HAPIC, a scalable cache over PM-based indexes to capture the constantly-changing query hotspots in skewed workloads. HAPIC embodies the data access frequency gradient in a hierarchy of hash tables to efficiently identify hotspots and reacts quickly to workload changes with epoch-based promotion. Compared with the state-of-the-art strategy, HAPIC reacts to hotspot shifts significantly faster, with up to 14% higher stable read throughput, 26% lower median latency, and 13% lower P99 latency. Chih-Ting Lo, Yun-Chih Chen, Yuan-Hao Chang 0001, Tei-Wei Kuo |
ICCAD | 2 |
| 2023 | REFROM: Responsive, Energy-Efficient Frame Rendering for Mobile DevicesabstractThe increasing demand for high-quality graphics on mobile devices necessitates a high frame rate for display refresh. However, current process scheduling and memory management policies fail to consider the computation demands of frame rendering because they are optimized for saving energy and resource utilization. This leads to unresponsive displays for mobile users due to rendering delays. Accurately estimating computation demands is challenging for the mobile operating system, particularly under memory pressure, without display-specific semantics from user space. Moreover, the complexity of frame rendering makes it infeasible to schedule them with real-time policies. To address these issues, we propose a new framework called REFROM that utilizes a history-based frame time estimator to analyze frame time samples from UI threads and predict the computation requirements of upcoming frames. Experimental results demonstrate that REFROM reduces the number of delayed frames by up to 40% and improves up to 4% energy efficiency compared to the existing approaches. Tsung-Yen Hsu, Yi-Shen Chen, Yun-Chih Chen, Yuan-Hao Chang 0001, Tei-Wei Kuo |
ISLPED | 3 |
| 2023 | ZoneLife: How to Utilize Data Lifetime Semantics to Make SSDs SmarterabstractFrom cloud databases to large-scale data analytics, modern applications exploit solid state drives (SSD)’s low latency to write an enormous amount of short-lived data. These data do not require the strong data protection typical SSDs use to reliably store data for a guaranteed period. In recent years, SSD’s density has been growing rapidly at the cost of degraded reliability, forcing SSD vendors to trade endurance and performance for stronger error protection. An intuitive question to ask is, “What if the SSD can identify these short-lived data to save the tax of over-protection?” In this article, we answer affirmatively with a novel co-design called, ZoneLife, which exposes the data lifetime semantics from applications to the SSD. ZoneLife enables the SSD to select the optimal error-correction code (ECC) out of multiple codes of different strengths. As a result, the SSD can store short-lived data with significantly less resources. ZoneLife efficiently translates the data addresses of different lifetimes with a multigranularity flash-translation-layer (FTL). Existing systems can easily adopt ZoneLife with localized modifications because ZoneLife’s host driver API generalizes Linux’s write hint interface, and its device firmware utilizes the popular Zone Namespace interface. ZoneLife is evaluated with several representative database and cloud workloads, and the results show noticeable improvements in SSD’s endurance and write throughput. Yun-Chih Chen, Chun-Feng Wu, Yuan-Hao Chang 0001, Tei-Wei Kuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Reptail: Cutting Storage Tail Latency with Inherent RedundancyabstractMission-critical edge applications require both low latency and strict data safety. Although emerging ultra-dense solid-state drives (SSDs) can extend the amount of data edge servers can process, the reduced parallelism can worsen read tail latency and even violate the deadline of mission-critical edge applications. To cut ultra-dense SSDs’ read tail latency, we propose Reptail, a co-design of host OS and SSD, that exploits the inherent redundancy in transactional systems. We use journaling file system to show how exposing SSD’s internals to host OS’s redundancy semantics can improve its read scheduling, thus reducing read tail latency. We evaluate Reptail with diverse workloads and find more than 20% latency improvements in the 95th and 99th percentile. Yun-Chih Chen, Chun-Feng Wu, Yuan-Hao Chang 0001, Tei-Wei Kuo |
DAC | 1 |
| 2021 | RVO: Unleashing SSD's Parallelism by Harnessing the Unused PowerabstractAnalytic video surveillance system is one of the fastest-growing cyber-physical applications worldwide. A video surveillance system must have a scalable storage backend to simultaneously ingest video frames and serve read requests for video analytics. As a result, 3D NAND flash-based storage devices, i.e., Solid-State Drives (SSD), are gradually regarded as promising candidates thanks to their rapidly growing density and parallelism. However, when deployed in power-constrained environments like the network edges, SSDs’ parallelism often cannot be fully unleashed. In edge video analytics systems, the lost parallelism can cause video frame drop and untimely analytics. To tackle this limitation, we first reveal that a flash program operation’s power usage is over-estimated in the conventional SSD design, leading to a limited degree of I/O parallelism. Based on the observation, we propose a novel command set, RVO (Read-Verify Overlap), which reclaims the unused power from the overestimation to amend the lost parallelism. To realize feasible fine-grained power management, we further accompany RVO with a generic power-aware scheduler. Through experiments, we show how video analytics systems equipped with RVO can achieve zero frame drop while ensuring compliance with industrial read latency requirements, even in write-intensive workloads. Hasan Alhasan, Yun-Chih Chen, Chien-Chung Ho |
ISLPED | 2 |
| 2019 | Chronic Kidney Disease Stage Classification Using Renal Artery Doppler-Derived ParametersabstractIn renal medicine, Estimated Glomerular Filtration Rate (eGFR) based method is a standard for the diagnosis of chronic kidney disease. However, this method is invasive, uncomfortable, costly, and could be dangerous because it requires to draw blood from the artery vessels. Researchers have developed several non-invasive Doppler-derived measures based chronic kidney disease (CKD) stage diagnosing or prognosing approaches; however, there is no adequate automatic renal artery Doppler-derived CKD stage classification method in the literature. Thus, we propose a non-invasive, safer, faster, and low cost, SVM-based CKD stage classification method from a sonogram of the renal artery blood flow. The proposed method extracts kurtosis and curvature parameters of the probability distribution that generated from renal artery blood flow waveform. Kurtosis and curvatures are employed to measure the tailedness and curvedness of the probability distribution. We collected a total of 528 sonograms from 110 (49 males) CKD patients during 2010-2013. The experimental results revealed a statistically significant correlation between the parameters and CKD progress stages. Post-voting results revealed the best f1score of 0.956 for Positive (stages 1-5) CKD stages. Munkhjargal Gochoo, Jun-Wei Hsieh, Chien-Hung Lee, Yun-Chih Chen, Yu-Chi Shih |
SMC | 4 |