EDBT 2026 Demo / reviewers in the wild / expert
Chun-Feng Wu
dblp:90/597
· DBLP profile ↗
48ranked-venue papers
11as first author
31since 2021 · last 2026
0000-0002-6367-0517ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 7 first-author · 28 since 2021Software engineering, systems software and programming languages · 5 · 3 since 2021Computer networks · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Security and privacy · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GECKO: Graph-Evolving aware ChecKpOinter for Intermittent SystemsabstractGraph workloads increasingly rely on large, continuously evolving datasets, where SSD data placement and migration strongly influence query and update efficiency. Prior SSD based graph management schemes, including log based designs that follow a read modify write pattern and GraphSSD, target servers and PCs with stable power. When deployed on intermittently powered systems, these power unaware designs often scatter frequently updated hub node edges across many flash pages, which increases over read and triggers excessive flash I/O. The resulting energy overhead is further amplified after each power recovery because the system must reload graph data from NAND into DRAM, repeatedly paying for unnecessary reads and reducing the number of graph queries completed per charge cycle. We propose Graph Evolving aware ChecKpOinter (GECKO), which leverages graph evolution awareness to improve hub edge locality and applies power aware I/O coordination to reduce redundant flash accesses under intermittent power. Across evolving graph updates and queries, GECKO significantly lowers flash I/O, enabling better energy efficiency on energy constrained intermittent systems. Pin-Hong Li, Yan-Han Chang, Chun-Feng Wu, Pi-Cheng Hsiu |
ISLPED | 3 |
| 2025 | Design and Optimization for AI/ML Acceleration on Resource-constrained and Edge SystemsabstractThe rapid advancement of AI (from foundational machine learning to Large Language Models) and edge computing has placed unprecedented demands on computation, memory, and storage on resource-constrained edge devices. As AI models scale, the ability to efficiently manage computing resources, utilize memory and storage, and reduce energy consumption has become critical. This paper introduces contributions on 4 topics related to deploying AI on resource-constrained edge devices: 1) unlocking training of foundational machine learning algorithms on the edge, 2) exploring hardware-aware DNN architecture and mapping co-optimization for inference on heterogeneous systems, 3) scaling RAG by leveraging advanced memory, storage, and energy-efficient designs, and 4) investigating cost-effective and high-performance large-scale graph processing. Jalil Boukhobza, Alessio Burrello, Yuan-Hao Chang 0001, Yawei Li 0001, Daniele Jahier Pagliari, Chun-Feng Wu, Ming-Chang Yang, Tsun-Yu Yang |
CASES | 6 |
| 2025 | P-DAC: Power-Efficient Photonic Accelerators for LLM InferenceabstractAs traditional electronic hardware encounters the limitations of Moore’s Law, optical computing is emerging as a promising alternative, delivering high data transmission rates, especially beneficial for big data and AI applications. Photonic accelerators, such as the LighteningTransformer, utilize optical analog signals to accelerate Transformerbased models, achieving exceptional speed and low energy consumption. However, controlling modern optical intensity modulators (e.g., MachZehnder Modulators) requires using electrical analog signals (e.g., voltage values) to adjust the optical signal intensity for realizing optical-based vector inner product calculations. Managing this modulation consumes significant power, as it involves selecting optimal electrical values through an electrical controller and converting digital signals to analog using digital-to-analog converters (DACs). In this work, we introduce P-DAC, a solution designed to reduce DAC power consumption, significantly enhancing the energy efficiency of optical accelerators for Transformer models. Wen-Tse Chang, Chun-Feng Wu, Yun-Chen Lo |
DAC | 2 |
| 2025 | UPVSS: Jointly Managing Vector Similarity Search with Near-Memory Processing SystemsabstractVector similarity search plays a pivotal role in modern applications, including recommendation systems, image search, large language models (LLMs), and high-dimensional data retrieval. As data size scales, our research reveals that the search phase imposes substantial demands on DRAM bandwidth, leading to performance limitations in conventional von Neumann architecture with shared memory buses. This data movement bottleneck restricts the efficiency and scalability of vector similarity search due to insufficient memory bandwidth. To mitigate this issue, we leverage UPMEM, an off-the-shelf near-memory processing (NMP) system, to minimize the data movement between memory and compute units. However, UPMEM’s computing engine has certain limitations and requires thorough application integration to unleash its high-parallelism capabilities. In this work, we introduce UPMEM-aware Vector Similarity Search (UPVSS), an architecture-aware system that jointly manages vector similarity search and UPMEM’s NMP technology. UPVSS prioritizes offloading operations based on their strengths and capabilities, effectively alleviating the data movement bottleneck and improving overall system performance. Chun-Chien Liu, Chun-Feng Wu, Yunho Jin |
DAC | 2 |
| 2025 | GAIA: Glass-Aware I/O MiddlewareabstractAs cloud-scale services and data-centric applications continue to generate massive volumes of data, the need for ultra-durable, energy-efficient, and cost-effective archival storage becomes increasingly urgent. Quartz glass has recently emerged as a promising archival medium, offering multi-century durability, radiation and thermal resistance, and support for three-dimensional data encoding using femtosecond laser writing. However, the hybrid mechanical-optical architecture of glass storage—requiring mechanical movement along the X and Y axes and optical focal tuning along the Z axis—introduces unique performance bottlenecks during data access, which conventional I/O scheduling strategies are not equipped to handle.In this work, we present GAIA, a Glass-Aware I/O middlewAre designed to optimize data access in quartz glass storage systems. GAIA features three coordinated strategies: (1) Zigzag Data Placement, which aligns data with the mechanical stage’s natural motion to minimize direction-switching latency; (2) Z-Axis First Placement, which prioritizes low-latency optical traversal along the depth dimension; and (3) Shortest Moving Time First (SMTF) scheduling, which selects I/O operations based on predicted movement time rather than geometric distance. Through trace-driven simulations using enterprise-scale workloads and various glass sizes, GAIA reduces data read latency by up to 82% compared to traditional baseline schedulers. These results demonstrate the critical importance of middleware-level co-design in unlocking the performance potential of next-generation glass-based storage systems. Hung-Yuna Chen, Chun-Feng Wu, Yuan-Hao Chang 0001, David Hung-Chang Du |
ICCAD | 2 |
| 2025 | Demeter: A Scalable and Elastic Tiered Memory Solution for Virtualized Cloud via Guest DelegationabstractMemory scalability has emerged as a critical bottleneck in virtualized cloud environments. Tiered memory architectures that combine limited fast memory with abundant slower memory offer a promising solution, but existing hypervisor-based approaches suffer from significant performance penalties. We present Demeter, introducing a paradigm shift through guest-delegated tiered memory management based on two key insights: (1) delegation to guests eliminates both expensive access tracking at the hypervisor level and frequent TLB flushes that severely degrade memory virtualization performance under two-dimensional address translation, and (2) Processor Event-Based Sampling, which cannot be effectively utilized by hypervisor-based solutions, remains fully functional and highly efficient when properly leveraged within the guest. Building on these insights, Demeter designs an efficient range-based tiered memory management scheme in guest virtual address space to preserve locality information and employs a double balloon-based provisioning mechanism that maintains cloud elasticity while enabling vendor-specific QoS control. Our evaluation with seven real-world workloads across DRAM+PMEM and DRAM+CXL.mem configurations demonstrates that Demeter improves performance by up to 2× compared to existing hypervisor-based approaches and by 28% on average compared to the next best guest-based alternative. Our implementation is fully open source and publicly available at Zenodo. Junliang Hu, Zhisheng Hu, Chun-Feng Wu, Ming-Chang Yang |
SOSP | 3 |
| 2024 | How to Steal CPU Idle Time When Synchronous I/O Mode Becomes PromisingabstractThe advent of Ultra-Low-Latency storage devices has narrowed the performance gap between storage and CPU in computing platforms, facilitating synchronous I/O adoption. Yet, this approach introduces substantial busy waiting time and underutilizes computing units. To address this, we propose a light-weighted Idle-Time-Stealing (ITS) design. This involves a self-improving thread conducting prefetching for high-priority processes during synchronous I/O, and an I/O-waiting process continuing subsequent instruction executions when justifiable. Another thread, the self-sacrificing thread, proactively switches low-priority process I/O requests from synchronous to asynchronous mode, prioritizing high-priority executions. Experimental results demonstrate the effectiveness of our ITS design in reducing CPU idle time. Chun-Feng Wu, Yuan-Hao Chang 0001, Ming-Chang Yang, Tei-Wei Kuo |
DAC | 1 |
| 2024 | ALISA: An Adaptive Learned Index Structure for Spatial Data on Solid-State DrivesabstractSpatial learned index is becoming popular as a solution to relieve the intense storage demands and high I/O costs of spatial databases. LISA, the original and most prominent spatial learned index structure, is tailored for HDD-resident spatial data and comes with strict data arrangement requirements. Given that direct SSD migration may drastically impair SSD durability especially when page utilization is low, this work aims to adapt this innovative index structure to SSDs to leverage the faster performance and expand application possibilities. We propose an Adaptive Learned Index structure for Spatial dAta on SSDs (ALISA), with mechanisms to persistently monitor updated data distribution and adaptively restructure to align with SSD access characteristics. The evaluation results show that ALISA can significantly extend SSD lifespan and improve low page utilization, thereby enhancing query performance. Che-Wei Lin, Chun-Feng Wu |
ICCAD | 2 |
| 2024 | GEAR: Graph-Evolving Aware Data Arranger to Enhance the Performance of Traversing Evolving Graphs on SCMabstractIn the era of big data, social network services continuously modify social connections, leading to dynamic and evolving graph data structures. These evolving graphs, vital for representing social relationships, pose significant memory challenges as they grow over time. To address this, storage-class-memory (SCM) emerges as a cost-effective solution alongside DRAM. However, contemporary graph evolution processes often scatter neighboring vertices across multiple pages, causing weak graph spatial locality and high-TLB misses during traversals. This article introduces SCM-Based graph-evolving aware data arranger (GEAR), a joint management middleware optimizing data arrangement on SCMs to enhance graph traversal efficiency. SCM-based GEAR comprises multilevel page allocation, locality-aware data placement, and dual-granularity wear leveling techniques. Multilevel page allocation prevents scattering of neighbor vertices relying on managing each page in a finer-granularity, while locality-aware data placement reserves space for future updates, maintaining strong graph spatial locality. The dual-granularity wear leveler evenly distributes updates across SCM pages with considering graph traversing characteristics. Evaluation results demonstrate SCM-based GEAR’s superiority, achieving 23% to 70% reduction in traversal time compared to state-of-the-art frameworks. Wen-Yi Wang, Chun-Feng Wu, Yun-Chih Chen, Tei-Wei Kuo, Yuan-Hao Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | A digital 3D TCAM accelerator for the inference phase of Random ForestabstractRandom forest is a popular ensemble machine-learning algorithm for classification and regression tasks. However, the irregular tree shapes and non-deterministic memory access patterns make it hard for the current von Neumann architecture to handle random forest efficiently. This paper proposes a digital 3D TCAM-based accelerator for the random forest, adopting the idea of processing-in-memory (PIM) to reduce data movement. By utilizing this accelerator, we propose a TCAM-based approach to provide real-time inference with low energy consumption, making it suitable for edge or embedded environments. In the experiments, the proposed approach achieves an average of 3.13 times higher throughput with 22 times more energy saving than the GPU approach. Chieh-Lin Tsai, Chun-Feng Wu, Yuan-Hao Chang 0001, Han-Wen Hu, Yung-Chun Lee, Hsiang-Pang Li, Tei-Wei Kuo |
DAC | 2 |
| 2023 | Data Freshness Optimization on Networked Intermittent SystemsabstractA networked intermittent system (NIS) is often deployed in the field for environmental monitoring, where sink nodes are responsible for relaying the data captured by sensors to a central system. To evaluate the quality of the captured monitoring data, Age of Information (AoI) is adopted to quantify the freshness of the data received by the central server. As the sink nodes are powered by ambient energy sources (e.g., solar and wind), the energy-efficient design of the sink nodes is crucial in order to improve the system-wide AoI. This work proposes the energy-efficient sink node design to save energy and extend system uptime. We devise an AoI-aware data forwarding algorithm based on the branch-and-bound (B&B) paradigm for deriving the optimal solution offline. In addition, an AoI-aware data forwarding algorithm is developed to approximate the optimal solution during runtime. The experimental results show that our solution can greatly improve the average data freshness for 148% against existing well-known strategies and achieves 91 % performance of the optimal solution. Compared with the state-of-the-art algorithm, our energy-efficient design can deliver better$A^{3}oI$results by up to 9.6%. Hao-Jan Huang, Wen Sheng Lim, Chia-Heng Tu, Chun-Feng Wu, Yuan-Hao Chang 0001 |
DATE | 4 |
| 2023 | S3: Increasing GPU Utilization during Generative Inference for Higher ThroughputabstractGenerating texts with a large language model (LLM) consumes massive amounts of memory. Apart from the already-large model parameters, the key/value (KV) cache that holds information about previous tokens in a sequence can grow to be even larger than the model itself. This problem is exacerbated in one of the current LLM serving frameworks which reserves the maximum sequence length of memory for the KV cache to guarantee generating a complete sequence as they do not know the output sequence length. This restricts us to use a smaller batch size leading to lower GPU utilization and above all, lower throughput. We argue that designing a system with a priori knowledge of the output sequence can mitigate this problem. To this end, we propose $S^3$, which predicts the output sequence length, schedules generation queries based on the prediction to increase device resource utilization and throughput, and handle mispredictions. Our proposed method achieves 6.49× throughput over those systems that assume the worst case for the output sequence length. Yunho Jin, Chun-Feng Wu, David Brooks 0001, Gu-Yeon Wei |
NeurIPS | 2 |
| 2023 | DeepWare: Imaging Performance Counters With Deep Learning to Detect RansomwareabstractIn the year passed, rarely a month passes without a ransomware incident being published in a newspaper or social media. In addition to the rise in the frequency of ransomware attacks, emerging attacks are very effective as they utilize sophisticated techniques to bypass existing organizational security perimeter. To tackle this issue, this paper presents “DeepWare,” which is a ransomware detection model inspired by deep learning and hardware performance counter (HPC). Different from previous works aiming to check all HPC results returned from a single timing for every running process, DeepWare carries out a simple yet effective concept of “imaging hardware performance counters with deep learning to detect ransomware,” so as to identify ransomware efficiently and effectively. To be more specific, DeepWare monitors the system-wide change in the distribution of HPC data. By imaging the HPC values and restructuring the conventional CNN model, DeepWare can address HPC’s nondeterminism issue by extracting the event-specific and event-wise behavioral features, which allows it to distinguish the ransomware activity from the benign one effectively. The experiment results across ransomware families show that the proposed DeepWare is effective at detecting different classes of ransomware with the 98.6% recall score, which is 84.41%, 60.93%, and 21% improvement overRATAFIA,OC-SVM, andEGBmodels respectively. DeepWare achieves an average MCC score of 96.8% and nearly zero false-positive rates by using just a 100 ms snapshot of HPC data. This timeliness of DeepWare is critical on the ground that organizations and individuals have the opportunity to take countermeasures in the first stage of the attack. Besides, the experiment conducted on unseen ransomware families such as CoronaVirus, Ryuk, and Dharma demonstrates that DeepWare has excellent potential to be a useful tool for zero-day attack detection. Gaddisa Olani Ganfure, Chun-Feng Wu, Yuan-Hao Chang 0001, Wei-Kuan Shih |
IEEE Trans. Computers | 2 |
| 2023 | Accelerating Random Forest on Memory-Constrained Devices Through Data Storage OptimizationabstractRandom forests is a widely used classification algorithm. It consists of a set of decision trees each of which is a classifier built on the basis of a random subset of the training data-set. In an environment where the memory work-space is low in comparison to the data-set size, when training a decision tree, a large proportion of the execution time is related to I/O operations. These are caused by data blocks transfers between the storage device and the memory work-space (in both directions). Our analysis of random forests training algorithms showed that there are two major issues :(1)Block Under-utilization: data blocks are poorly used when loaded into memory and have to be reloaded multiple times, meaning that the algorithm exhibits a poor spatial locality;(2)Data Over-read: the data-set is supposed to be fully loaded in memory whereas a large proportion of data are not effectively useful when building a decision tree. Our proposed solution is structured to address these two issues. First, we propose to reorganize the data-set in such a way to enhance spatial locality and second, to remove the assumption that the data-set is entirely loaded into memory and access data only when effectively needed. Our experiments show that this method made it possible to reduce random forest building time by 51 to 95% in comparison to a state-of-the-art method. Camélia Slimani, Chun-Feng Wu, Stéphane Rubini, Yuan-Hao Chang 0001, Jalil Boukhobza |
IEEE Trans. Computers | 2 |
| 2023 | ZoneLife: How to Utilize Data Lifetime Semantics to Make SSDs SmarterabstractFrom cloud databases to large-scale data analytics, modern applications exploit solid state drives (SSD)’s low latency to write an enormous amount of short-lived data. These data do not require the strong data protection typical SSDs use to reliably store data for a guaranteed period. In recent years, SSD’s density has been growing rapidly at the cost of degraded reliability, forcing SSD vendors to trade endurance and performance for stronger error protection. An intuitive question to ask is, “What if the SSD can identify these short-lived data to save the tax of over-protection?” In this article, we answer affirmatively with a novel co-design called, ZoneLife, which exposes the data lifetime semantics from applications to the SSD. ZoneLife enables the SSD to select the optimal error-correction code (ECC) out of multiple codes of different strengths. As a result, the SSD can store short-lived data with significantly less resources. ZoneLife efficiently translates the data addresses of different lifetimes with a multigranularity flash-translation-layer (FTL). Existing systems can easily adopt ZoneLife with localized modifications because ZoneLife’s host driver API generalizes Linux’s write hint interface, and its device firmware utilizes the popular Zone Namespace interface. ZoneLife is evaluated with several representative database and cloud workloads, and the results show noticeable improvements in SSD’s endurance and write throughput. Yun-Chih Chen, Chun-Feng Wu, Yuan-Hao Chang 0001, Tei-Wei Kuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Energy Efficiency Enhancement of SCM-Based Systems: Write-Friendly CodingabstractWith the advent of the Internet of Things (IoT), more and more wearable devices have been developed and integrated into our daily lives. Energy efficiency is critical for these devices because they are typically run on energy-constrained resources like batteries or energy harvesters. The storage-class memory (SCM) technologies and data compression techniques could improve their energy efficiency via reducing data movements and squeezing the data size, respectively, where data compression is especially important for IoT and embedded systems to reduce the volume of data for the space-constrained memory/storage devices. Nevertheless, both of them cannot be aware of their inherent characteristics for further minimization of energy consumption. To this end, a write-friendly coding scheme is proposed in this work that jointly manages both techniques to yield energy-efficient SCM-based systems. Moreover, a novel design of ignorable bits is presented to partially skip write operations after completing data compression without sacrificing data accuracy. We evaluate the proposed scheme via a series of intensive experiments, the experimental results of which indicate that the proposed coding scheme reduces energy consumption by up to 45% under the investigated benchmarks. Yi-Shen Chen, Chun-Feng Wu, Yuan-Hao Chang 0001, Tei-Wei Kuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | RTrap: Trapping and Containing Ransomware With Machine LearningabstractWith advances in social engineering tricks and other technical shortcomings, ransomware attacks have become a severe cybercrime affecting organizations of all shapes and sizes. Although the security teams are making plenty of ransomware detection tools, the ransomware incident report shows they are ineffective in detecting emerging ransomware attacks. This work presents “RTrap,” a systematic framework to detect and contain ransomware efficiently and effectively via machine learning-generated deceptive files. Using a data-driven decoy file selection and generation strategy, RTrap plants deceptive decoy files across the directory to lure the ransomware to access it. RTrap also introduced a lightweight decoy watcher to monitor generated decoy files in real time. As the timing of the ransomware attack is not known to the victim in advance, and the ransomware encryption process is speedy, the proposed decoy-watcher executes an automatic/automated response after the detection promptly. The experiment shows that RTrap can detect ransomware with an average 18 file loss per 10311 legitimate user files. Gaddisa Olani Ganfure, Chun-Feng Wu, Yuan-Hao Chang 0001, Wei-Kuan Shih |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | A joint management middleware to improve training performance of deep recommendation systems with SSDsabstractAs the sizes and variety of training data scale over time, data preprocessing is becoming an important performance bottleneck for training deep recommendation systems. This challenge becomes more serious when training data is stored in Solid-State Drives (SSDs). Due to the access behavior gap between recommendation systems and SSDs, unused training data may be read and filtered out during preprocessing. This work advocates a joint management middleware to avoid reading unused data by bridging the access behavior gap. The evaluation results show that our middleware can effectively improve the performance of the data preprocessing phase so as to boost training performance. Chun-Feng Wu, Carole-Jean Wu, Gu-Yeon Wei, David Brooks 0001 |
DAC | 1 |
| 2022 | GraphRC: Accelerating Graph Processing on Dual-Addressing Memory with Vertex MergingabstractArchitectural innovation in graph accelerators attracts research attention due to foreseeable inflation in data sizes and the irregular memory access pattern of graph algorithms. Conventional graph accelerators ignore the potential of Non-Volatile Memory (NVM) crossbar as a dual-addressing memory and treat it as a traditional single-addressing memory with higher density and better energy efficiency. In this work, we present GraphRC, a graph accelerator that leverages the power of dual-addressing memory by mapping in-edge/out-edge requests to column/row-oriented memory accesses. Although the capability of dual-addressing memory greatly improves the performance of graph processing, some memory accesses still suffer from low-utilization issues. Therefore, we propose a vertex merging (VM) method that improves cache block utilization rate by merging memory requests from consecutive vertices. VM reduces the execution time of all 6 graph algorithms on all 4 datasets by 24.24% on average. We then identify the data dependency inherent in a graph limits the usage of VM, and its effectiveness is bounded by the percentage of mergeable vertices. To overcome this limitation, we propose an aggressive vertex merging (AVM) method that outperforms VM by ignoring the data dependency inherent in a graph. AVM significantly reduces the execution time of ranking-based algorithms on all 4 datasets while preserving the correct ranking of the top 20 vertices. Wei Cheng 0006, Chun-Feng Wu, Yuan-Hao Chang 0001, Ing-Chao Lin |
ICCAD | 2 |
| 2022 | Leveraging Write Heterogeneity of Phase Change Memory on Supporting Self-Balancing Binary TreeabstractWith the increasing demand of massive/big data applications, nonvolatile memory (NVM), such as phase-change memory (PCM), has become a promising candidate to replace DRAM because of its low leakage power, nonvolatility, and high density. However, most of the existing memory read/write intensive algorithms and data structures are not aware of the PCM write heterogeneity in terms of both energy consumption and latency. In particular, self-balancing binary search trees, which are widely used to manage massive data in the big-data era, were designed without the consideration of PCM characteristics. Thus, the multiple rotations of the tree balancing process would degrade the memory performance. This work explores the relations among nodes and analyzes tree operations, and the node indexing and address mapping are redesigned to reduce the tree management overhead on single-level cell (SLC) PCM by decreasing the number of bit flips of tree rotations. When multilevel cell (MLC) PCM is included, our address mapping algorithm is developed to reduce the total energy consumption and latency with considerations of the heterogeneous write operations of different cell states. Experimental results show that our solution significantly outperforms the original implementation of a self-balancing binary search tree when the amount of data is large. Chun-Feng Wu, Yuan-Hao Chang 0001, Ming-Chang Yang, Chieh-Fu Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Exploring Synchronous Page Fault HandlingabstractThe advance of nonvolatile memory in storage technology has presented challenges in redefining the ways in handling the main memory and the storage. This work is motivated by the strong demands in effective handling of page faults over ultralow-latency storage devices. In particular, we propose synchronous and asynchronous prefetching strategies to satisfy process executions with different memory demands in supporting of synchronous page fault handling. An adaptive CPU scheduling strategy is also proposed to cope with the needs of processes in maintaining their working sets in the main memory. Six representative benchmarks and applications were evaluated. It was shown that our strategy can effectively save 12.33% of the total execution time and reduce 13.33% of page faults, compared to the conventional demand paging strategy with nearly no sacrificing of process fairness. Yin-Chiuan Chen, Chun-Feng Wu, Yuan-Hao Chang 0001, Tei-Wei Kuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | A File-Oriented Fast Secure Deletion Strategy for Shingled Magnetic Recording DrivesabstractNowadays, securely erasing deleted files has become one of the necessary tasks for users who want to protect their deleted data from malicious attackers. Nevertheless, existing secure deletion approaches are considered inefficient for erasing deleted files permanently because the file systems and storage devices do not share their file information or data layout with each other. On the emerging shingled magnetic recording (SMR) drives, the inefficiency of existing secure deletion approaches is exaggerated by the inherent sequential-write constraint of the high storage density SMR technology. On SMR drives, tracks are overlapped via utilizing the size difference between disk read/write heads to increase the storage density. Due to the overlapped track layout, secure deletion requests may induce a significant amount of write amplification and serious performance degradation if the data layout is not properly configured. Such observation motivates this article to come up with a file-oriented fast secure deletion (FFSD) strategy to deal with the sequential-write constraint of SMR drives and improve the efficiency of secure deletion operations on SMR drives. The experimental results show that the proposed strategy can effectively reduce the secure deletion latency by$286.15\times $on average when compared with the conventional approach. Shuo-Han Chen, Chun-Feng Wu, Ming-Chang Yang, Yuan-Hao Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Performance Enhancement of SMR-Based Deduplication SystemsabstractDue to the fast-growing amount of data and cost consideration, shingled-magnetic-recording (SMR) drives are developed to provide low-cost and high-capacity data storage by enhancing the areal-density of hard disk drives, and (data) deduplication techniques are getting popular in data-centric applications to reduce the amount of data that need to be stored in storage devices by eliminating the duplicate data chunks. However, directly applying deduplication techniques on SMR drives could significantly decrease the runtime performance of the deduplication system because of the time-consuming SMR space reclamation caused by the sequential write constraint of SMR drives. In this article, an SMR-aware deduplication scheme is proposed to improve the runtime performance of SMR-based deduplication systems with the consideration of the sequential write constraint of SMR drives. Moreover, to bridge the information gap between the deduplication system and the SMR drive, the lifetime information of data chunks is extracted to separate data chunks of different lifetimes in different places of SMR drives, so as to further reduce the SMR space reclamation overhead. A series of experiments was conducted with a set of realistic deduplication workloads. The results show that the proposed scheme can significantly improve the runtime performance of the SMR-based deduplication system with limited system overheads. Chun-Feng Wu, Martin Kuo, Ming-Chang Yang, Yuan-Hao Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Rethinking the Interactivity of OS and Device Layers in Memory ManagementabstractIn the big data era, a huge number of services has placed a fast-growing demand on the capacity of DRAM-based main memory. However, due to the high hardware cost and serious leakage power/energy consumption, the growth rate of DRAM capacity cannot meet the increased rate of the required main memory space when the energy or hardware cost is a critical concern. To tackle this issue, hybrid main-memory devices/modules have been proposed to replace the pure DRAM main memory with a hybrid main memory module that provides a large main memory space by integrating a small-sized DRAM and a large-sized non-volatile memory (NVM) into the same memory module. Although NVMs have high-density and low-cost features, they suffer from the low read/write performance and low endurance issue, compared to DRAM. Thus, inside the hybrid main-memory module, it also includes a memory management design to use DRAM as the cache of NVMs to enhance its performance and lifetime. However, it also introduces new design challenges in both the OS and the memory module. In this work, we rethink the interactivity of OS and hybrid main-memory module, and propose a cross-layer cache design that (1) utilizes the information from the operating system to optimize the hit ratio of the DRAM cache inside the memory module, and (2) takes advantage of the bulk-size (or block-based) read/write feature of NVM to minimize the time overhead on the data movement between DRAM and NVM. At the same time, this cross-layer cache design is very lightweight and only introduces limited runtime management overheads. A series of experiments was conducted to evaluate the effectiveness of the proposed cross-layer cache design. The results show that the proposed design could improve access performance for up to 88%, compared to the investigated well-known page replacement algorithms. Tse-Yuan Wang, Chun-Feng Wu, Che-Wei Tsao, Yuan-Hao Chang 0001, Tei-Wei Kuo, Xue (Steve) Liu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2021 | A Write-friendly Arithmetic Coding Scheme for Achieving Energy-Efficient Non-Volatile Memory SystemsabstractIn the era of the Internet of Things (IoT), wearable IoT devices become popular and closely related to our life. Most of these devices are based on the embedded systems that have to operate on limited energy resources, such as batteries or energy harvesters. Therefore, energy efficiency is one of the critical issues for these devices. To relieve the energy consumption by reducing the total accesses on memory and storage layers, the technologies of storage-class memory (SCM) and data compression techniques are applied to eliminate the data movements and squeeze the data size, respectively. However, the information gap between them hinders the cooperation among the two techniques for achieving further optimizations on minimizing energy consumption. This work proposes a write-friendly arithmetic coding with joint managing both techniques to achieve energy-efficient non-volatile memory (NVM) systems. In particular, the concept of "ignorable bits" is introduced to further skip the write operations while storing the compressed data into SCM devices. The proposed design was evaluated by a series of intensive experiments, and the results are encouraging. Yi-Shen Chen, Chun-Feng Wu, Yuan-Hao Chang 0001, Tei-Wei Kuo |
ASP-DAC | 2 |
| 2021 | Reptail: Cutting Storage Tail Latency with Inherent RedundancyabstractMission-critical edge applications require both low latency and strict data safety. Although emerging ultra-dense solid-state drives (SSDs) can extend the amount of data edge servers can process, the reduced parallelism can worsen read tail latency and even violate the deadline of mission-critical edge applications. To cut ultra-dense SSDs’ read tail latency, we propose Reptail, a co-design of host OS and SSD, that exploits the inherent redundancy in transactional systems. We use journaling file system to show how exposing SSD’s internals to host OS’s redundancy semantics can improve its read scheduling, thus reducing read tail latency. We evaluate Reptail with diverse workloads and find more than 20% latency improvements in the 95th and 99th percentile. Yun-Chih Chen, Chun-Feng Wu, Yuan-Hao Chang 0001, Tei-Wei Kuo |
DAC | 2 |
| 2021 | Future Computing Platform Design: A Cross-Layer Design ApproachabstractFuture computing platforms are facing a paradigm shift with the emerging resistive memory technologies. First, they offer fast memory accesses and data persistence in a single large-capacity device deployed on the memory bus, blurring the boundary between memory and storage. Second, they enable computing-in-memory for neuromorphic computing to mitigate costly data movements. Due to the non-ideality of these resistive memory devices at the moment, we envision that cross-layer design is essential to bring such a system into practice. In this paper, we showcase a few examples to demonstrate how cross-layer design can be developed to fully exploit the potential of resistive memories and accelerate its adoption for future computing platforms. Hsiang-Yun Cheng, Chun-Feng Wu, Christian Hakert, Kuan-Hsun Chen, Yuan-Hao Chang 0001, Jian-Jia Chen, Chia-Lin Yang, Tei-Wei Kuo |
DATE | 2 |
| 2021 | Space-efficient Graph Data Placement to Save Energy of ReRAM CrossbarabstractAiming to extract the information behind messy data, graph computation is one of the popular big data analysis applications. During running graph computation, large numbers of vertices and edges will be moved between memory and computing units, and these intensive data movements lead to a performance bottleneck. To break the bottleneck, Resistive Random-Access Memory (ReRAM) based crossbar accelerators, which can act as both computing and memory units simultaneously on one chip, are a promising solution to eliminate these data movements. However, running graph computation on crossbar accelerators incurs high power consumption because real-world graphs are too sparse and discrete to unleash the computation capability provided by crossbar accelerators. In contrast to previous works which require extra general-purpose computing units to work with crossbar accelerators, this work proposes a software strategy, called graph-aware crossbar placement strategy, to improve the utilization of crossbar accelerators by clustering graph nodes with strong graph spatial locality. The evaluation results show that the proposed graph-aware crossbar placement strategy can efficiently save the energy consumption of crossbar accelerators. Ting-Shan Lo, Chun-Feng Wu, Yuan-Hao Chang 0001, Tei-Wei Kuo, Wei-Chen Wang 0002 |
ISLPED | 2 |
| 2021 | On Minimizing Internal Data Migrations of Flash Devices via Lifetime-Retention HarmonizationabstractWith the emerge of high-density triple-level-cell (TLC) and 3D NAND flash, the access performance and endurance of flash devices are degraded due to the downscaling of flash cells. In addition, we observe that the mismatch between data lifetime requirement and flash block retention capability could further worsen the access performance and endurance. This is because the “lifetime-retention mismatch” could result in massive internal data migrations during garbage collection and data refreshing, and further aggravate the already-worsened access performance and endurance of high-density NAND flash devices. Such an observation motivates us to resolve the lifetime-retention mismatch problem by proposing a “time harmonization strategy”, which coordinates the flash block retention capability with the data lifetime requirement to enhance the performance of flash devices with very limited endurance degradation. Specifically, this study aims to lower the amount of internal data migrations caused by garbage collection and data refreshing via storing data of different lifetime requirement in flash blocks with suitable retention capability. The trace-driven evaluation results reveal that the proposed design can effectively reduce the average response time by about 99 percent on average without sacrificing the overall endurance, as compared with the state-of-the-art designs. Ming-Chang Yang, Chun-Feng Wu, Shuo-Han Chen, Yuan-Hao Chang 0001 |
IEEE Trans. Computers | 2 |
| 2021 | Beyond Write-Reduction Consideration: A Wear-Leveling-Enabled B⁺-Tree Indexing Scheme Over an NVRAM-Based ArchitectureabstractRecently, nonvolatile random-access memory (NVRAM) has been regarded as the most up-and-coming main memory technology in embedded and Internet-of-Things (IoT) systems due to its attractive features: zero-static power consumption and high memory cell density. However, the endurance issue as a “nightmare” always haunts NVRAM system developers. Worse still, NVRAM’s lifespan will wear out soon in embedded applications because their data management systems usually utilize an indexing scheme to maintain small data. Plus, a node structure within the indexing scheme will be frequently updated because of data creation and deletion. Therefore, many previous works rethink B+-tree indexing scheme on an NVRAM-based system. The most previous studies focused on reducing the amount of write traffic to memory. Unfortunately, they are failed to extend the NVRAM lifespan because their solution cannot evenly distribute the amount of write traffic to each memory cell. Additionally, prior solutions have not considered that all nodes within B+-tree indexing structure have different update frequencies. Based on such the observation, this work proposes a wear-leveling-aware B+-tree design, namely, waB+-tree, to consider the update frequency of each node within the B+-tree structure, so as to evenly scatter the amount of write traffic to the NVRAM cells. According to our experiments, the proposed waB+-tree shows the encouraging results of endurance improvement. Dharamjeet, Tseng-Yi Chen, Yuan-Hao Chang 0001, Chun-Feng Wu, Chi-Heng Lee, Wei-Kuan Shih |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | iCheck: Progressive Checkpointing for Intermittent SystemsabstractEnergy harvesting devices powered by ambient energies, instead of batteries, have been drawn lots of attention due to their advantages of energy saving, easy deployment without relying on stable power sources, and smaller sizes, facilitating promising applications, such as environmental and health monitoring. These devices perform the computations intermittently, where the code executions are halted and resumed depending on the availability of the harvested energy. On such devices, the capacitors are present and served as the energy buffers for preserving the program states when sudden power outages occur. Nevertheless, the capacitors have relatively shorter lifetimes, compared with the rest of hardware components on the devices, and larger capacitors, which are desired by the systems requiring complex computations, hamper the achievement of device miniaturization, e.g., for medical implants or smart dust. In this article, we propose a new intermittent checkpointing strategy,iCheck, to tackle the issues raised for the program-state retaining when the capacitors are not functioning correctly (or when the capacitor-less devices are adopted). The proposediCheckis designed to perform the checkpointing-based program-state preserving progressively with being aware of the power-failure characteristics of the harvested energy source to maximize the progress forwarding and to ensure data consistency while encountering incomplete checkpoints caused by sudden power losses. The proposed design is evaluated with a series of experiments with encouraging results. Wen Sheng Lim, Chia-Heng Tu, Chun-Feng Wu, Yuan-Hao Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | DeepGuard: Deep Generative User-behavior Analytics for Ransomware DetectionabstractIn the last couple of years, the move to cyberspace provides a fertile environment for ransomware criminals like ever before. Notably, since the introduction of WannaCry, numerous ransomware detection solution has been proposed. However, the ransomware incidence report shows that most organizations impacted by ransomware are running state of the art ransomware detection tools. Hence, an alternative solution is an urgent requirement as the existing detection models are not sufficient to spot emerging ransomware treat. With this motivation, our work proposes "DeepGuard," a novel concept of modeling user behavior for ransomware detection. The main idea is to log the file-interaction pattern of typical user activity and pass it through deep generative autoencoder architecture to recreate the input. With sufficient training data, the model can learn how to reconstruct typical user activity (or input) with minimal reconstruction error. Hence, by applying the three-sigma limit rule on the model's output, DeepGuard can distinguish the ransomware activity from the user activity. The experiment result shows that DeepGuard effectively detects a variant class of ransomware with minimal false-positive rates. Overall, modeling the attack detection with user-behavior permits the proposed strategy to have deep visibility of various ransomware families. Gaddisa Olani Ganfure, Chun-Feng Wu, Yuan-Hao Chang 0001, Wei-Kuan Shih |
ISI | 2 |
| 2020 | Determinizing Crash Behavior with a Verified Snapshot-Consistent Flash Translation Layer
Yun-Sheng Chang, Yao Hsiao, Tzu-Chi Lin, Che-Wei Tsao, Chun-Feng Wu, Yuan-Hao Chang 0001, Hsiang-Shang Ko, Yu-Fang Chen 0001 |
OSDI | 5 |
| 2020 | Joint Management of CPU and NVDIMM for Breaking Down the Great Memory WallabstractTo provide larger memory space with lower costs, NVDIMM is a production-ready device. However, directly placing NVDIMM as the main memory would seriously degrade the system performance because of the “great memory wall” caused by the fact that in NVDIMM, the slow memory (e.g., flash memory) is several orders of magnitude slower than the fast memory (e.g., DRAM). In this article, we present a joint management framework of host/CPU and NVDIMM to break down the great memory wall by bridging the process information gap between host/CPU and NVDIMM. In this framework, a page semantic-aware strategy is proposed to precisely predict, mark, and relocate data or memory pages to the fast memory in advance by exploiting the process access patterns, so that the frequency of the slow memory accesses can be further reduced. The proposed framework with the proposed strategy was evaluated with several well-known benchmarks and the results are encouraging. Chun-Feng Wu, Yuan-Hao Chang 0001, Ming-Chang Yang, Tei-Wei Kuo |
IEEE Trans. Computers | 1 |
| 2020 | Request Flow Coordination for Growing-Scale Solid-State DrivesabstractPerformance-intensive applications have led both interface and architecture changes of high-end, growing-scale solid-state drives (SSDs). However, we observe that most of the time, the actual drive performance could not be easily scaled or boosted up with the increasing of internal resources of growing-scale SSDs due to the potential congestion of I/O requests. Such observation inspires this article to look for a request flow coordination design to appropriately control and throttle the I/O request over the increasingly-complicated SSD internal organization with manageable coordination overhead. The main objective is to avoid overloading or congesting any sub-module of growing-scale SSDs by making good use of the abundant internal resources, so as to effectively improve the drive performance in terms of the request-response time. The capability of the proposed design was evaluated with realistic and intensive I/O workloads, and the results are very encouraging. Ming-Chang Yang, Yuan-Hao Chang 0001, Tei-Wei Kuo, Chun-Feng Wu |
IEEE Trans. Computers | 4 |
| 2020 | DeepPrefetcher: A Deep Learning Framework for Data Prefetching in Flash Storage DevicesabstractIn today's information-driven world, data access latency accounts for the expensive part of processing user requests. One potential solution to access latency is prefetching, a technique to speculate and move future requests closer to the processing unit. However, the block access requests received by the storage device show poor spatial locality because most file-related locality is absorbed in the higher layers of the memory hierarchy, including the CPU cache and main memory. Besides, the utilization of multithreading results in an interleaved access request making prefetching at the storage level more picky using existing prefetching techniques. Toward this, we propose and assess DeepPrefetcher, a novel deep neural network inspired context-aware prefetching method that adapts to arbitrary memory access patterns. DeepPrefetcher learns the block access pattern contexts using distributed representation and leverage long short-term memory learning model for context-aware data prefetching. Instead of using the logical block address (LBA) value directly, we model the difference between successive access requests, which contains more patterns than LBA value for modeling. By targeting access pattern sequence in this manner, the DeepPrefetcher can learn the vital context from a long input LBA sequence and learn to predict both the previously seen and unseen access patterns. The experimental result reveals that DeepPrefetcher can increase an average prefetch accuracy, coverage, and speedup by 21.5%, 19.5%, and 17.2%, respectively, contrasted with the baseline prefetching strategies. Overall, the proposed prefetching approach surpasses other schemes in all benchmarks, and the outcomes are promising. Gaddisa Olani Ganfure, Chun-Feng Wu, Yuan-Hao Chang 0001, Wei-Kuan Shih |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | On Minimizing Analog Variation Errors to Resolve the Scalability Issue of ReRAM-Based Crossbar AcceleratorsabstractCrossbar accelerators with a resistive random-access memory (ReRAM) are a promising solution for accelerating neural network applications. The advantages of achieving high computation throughput per watt make ReRAM-based crossbar accelerators become a potential solution for accelerating inference operations in the Internet of Things and edge devices. Due to the analog variation errors, the launched ReRAM-based crossbar accelerators can only perform well when each ReRAM cell is used to represent a limited number of data bits. To make such ReRAM-based crossbar accelerators applicable in wide application scenarios, several proposed researches target at binary neural networks and focus on the chip designs in relieving the implementation challenges on computation accuracy for realizing single-bit ReRAM-based crossbar accelerators. Even though several small-sized ReRAM-based crossbar accelerators are announced, the scalability issue hinders ReRAM-based crossbar accelerators from being scaled up. That is, when there are more and more wordline in an ReRAM-based crossbar accelerator, the analog variation error is amplified and thus seriously degrades the computation accuracy. In this work, we propose an adaptive data manipulation strategy to substantially reduce analog variation errors so as to fill up the gap on scaling up the ReRAM-based crossbar accelerators. In particular, a weightrounding design is proposed to manipulate data to minimize overlapping variation so that the number of wordlines can be scaled up. In addition, an input subcycling design is proposed to further trade tolerable errors with neural networks' execution time. Moreover, a bitline redundant design is proposed to trade acceptable space overhead for eliminating the analog variation errors. The emulation experiments show that the proposed adaptive data manipulation strategy can improve the accuracy in running MNIST and CIFAR-10 by 1.3× and 2.6× with nearly no management penalty and hardware cost. The experimental results also show the close-to-ideal-case accuracy by substantially reducing analog variation errors. Yao-Wen Kang, Chun-Feng Wu, Yuan-Hao Chang 0001, Tei-Wei Kuo, Shu-Yin Ho |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | When Storage Response Time Catches Up With Overall Context Switch Overhead, What Is Next?abstractThe virtual memory technique provides a large and cheap memory space by extending the memory space with storage devices. It applies context switch to asynchronously swapping pages between memory and storage devices for hiding the long response time of storage devices when a page fault occurs. However, the overall context switch overhead is high because the context switch itself is a complex function and would further incur TLB shootdown/flush and compulsory CPU cache misses after context switches. On the contrary, as the rapid responsiveness improvement of high-end storage devices, we observe that the response time of high-end storage devices catches up and gradually becomes smaller than the overall context switch overhead. At this turning point, to further enhance the system responsiveness, we advocate adopting synchronous swapping rather than context switch in response to page faults. Meanwhile, we propose a strategy, called shadow huge page management, to further improve the overall system performance by minimizing the overall time overheads caused by page faults and page swappings. Evaluation results show that the proposed system can efficiently reduce the total CPU wasting time. Chun-Feng Wu, Yuan-Hao Chang 0001, Ming-Chang Yang, Tei-Wei Kuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Enabling File-Oriented Fast Secure Deletion on Shingled Magnetic Recording DrivesabstractExisting secure deletion approaches are inefficient in erasing data permanently because file systems have no knowledge of the data layout on the storage device, nor is the storage device aware of file information within the file systems. This inefficiency is exaggerated on the emerging shingled magnetic recording (SMR) drive due to its inherent sequential-write constraint. On SMR drives, secure deletion requests may lead to serious write amplification and performance degradation if the data layout is not properly configured. Such observation motivates us to propose a file-oriented fast secure deletion (FFSD) strategy to alleviate the negative impacts of SMR drives' sequential-write constraint and improve the efficiency of secure deletion operations on SMR drives. A series of experiments was conducted to demonstrate the capability of the proposed strategy on improving the efficiency of secure deletion on SMR drives. Shuo-Han Chen, Ming-Chang Yang, Yuan-Hao Chang 0001, Chun-Feng Wu |
DAC | 4 |
| 2018 | Improving runtime performance of deduplication system with host-managed SMR storage drivesabstractDue to the cost consideration for data storage, high-areal-density shingled-magnetic-recording (SMR) drives and data deduplication techniques are getting popular in many data storage services for the improvement of profit per storage unit. However, naively applying deduplication techniques upon SMR drives may dramatically downgrade the runtime performance of data storage services, because of the time-consuming SMR space reclamation processes. This work advocates a vertical integration solution by jointly managing the host-managed SMR drives with deduplication system, in order to essentially relieve the time-consuming SMR space reclamation issue. The proposed design was evaluated by a series of realistic deduplication workloads with encouraging results. Chun-Feng Wu, Ming-Chang Yang, Yuan-Hao Chang 0001 |
DAC | 1 |
| 2018 | Hot-Spot Suppression for Resource-Constrained Image Recognition Devices With Nonvolatile MemoryabstractResource-constrained devices with convolutional neural networks (CNNs) for image recognition are becoming popular in various Internet of Things and surveillance applications. They usually have a low-power CPU and limited CPU cache space. In such circumstances, nonvolatile memory (NVM) has great potential to replace DRAM as main memory to improve overall energy efficiency and provide larger main-memory space. However, due to the iterative access pattern, performing CNN-based image recognition may introduce some write hot-spots on the NVM main memory. These write hot-spots may lead to reliability issues due to limited write endurance of NVM. In order to improve the endurance of NVM main memory, this paper leverages the CPU cache pinning technique and exploits the iterative access pattern of CNN to resolve the write hot-spot effect. In particular, we present a CNN-aware self-bouncing pinning strategy to minimize the maximal write cycles in NVM cells by proactively fastening CPU cache lines, so as to effectively suppress the write hot-spots to NVM main memory with limited performance degradation. The proposed strategy was evaluated by a series of intensive experiments and the results are encouraging. Chun-Feng Wu, Ming-Chang Yang, Yuan-Hao Chang 0001, Tei-Wei Kuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2015 | Distributed Metaserver Mechanism and Recovery Mechanism Support in Quantcast File SystemabstractWith the need of data storage increases tremendously nowadays, distributed file system becomes the most important data storage system in cloud computing. In distributed file system development, there are many researchers work hard to refine the architecture to provide scalability and reliability. In our work, we propose a distributed metaserver system including metaserver scale-out, metadata replication, metaserver recovery, and metaserver management recovery mechanisms. In our experiments, the proposed system can increase the capacity of metadata and increase the reliability by fault tolerance mechanism. The overhead of read/write data is very little in the proposed system as well. Su-Shien Ho, Chun-Feng Wu, Jiazheng Zhou, Ching-Hsien Hsu, Hung-Chang Hsiao, Yeh-Ching Chung |
COMPSAC | 2 |
| 2014 | Sensing phone use of motorcycle driversabstractDue to safety reasons, using mobile phones while driving is prohibited in many countries. Research has also shown that motorcycle riders are 20 times more likely to be killed in a crash than vehicle occupants. Therefore, it is more critical to restrict the use of mobile phones of motorcycle drivers than car drivers. There are some studies that focus on how to distinguish phone use between a driver and other passengers in a car. The techniques used for cars, however, are not always applicable to motorcycles. In this paper, we propose a way to detect phone use of motorcycle drivers. By using two low-cost Bluetooth emitters, mobile phones of the driver and the passenger can measure the signal strengths and decide their locations. We have conducted extensive experiments with various smartphones. The results show that on average we can achieve 96% accuracy. Jyh-Cheng Chen, Chun-Feng Wu, Wei-Ho Chung, Ping-Fan Ho |
GLOBECOM | 2 |
| 2014 | Cooperative distributed erasure code scheduling for smart grid communicationsabstractThe performance of smart grid applications such as advanced metering infrastructure (AMI) and demand response management (DRM) can be improved by exploiting wireless communication technologies. The communication data loading and data losses in smart grid impact the system. In this work, we propose the routing protocol which adopts the cooperative transmission architecture in smart grid communications. Our proposed scheme enables the cooperative nodes to encode and forward the packets through the distributed packet-level erasure coding. The experimental results indicate that our proposed erasure code embedded routing protocol can obtain the better throughput than the conventional routing protocols. Chun-Feng Wu, Wei-Ho Chung |
PIMRC | 1 |
| 2014 | Iterative Symbol Decoding of Distributed Channel Encoded Fixed-Length Codes for Cooperative CommunicationsabstractIn this paper, we study the iterative decoding of distributed channel encoded fixed-length codes for power-efficient cooperative transmissions. By exploiting the technique of iterative source channel decoding (ISCD), the iterative decoding for cooperative transmissions exploits two types of the redundant information: the residual redundancy inherently remaining in the source codes, and the artificial redundant information provided by the distributed channel coder. To avoid the information loss on the bit-to-symbol conversion, we adopt symbol-level ISCD scheme consisting of the symbol-level channel decoder and the symbol-level source decoder. Our proposed iterative decoding can improve parameter signal-to-noise ratio (PSNR) performances by iteratively refining the extrinsic information for the source and the distributed channel decoders. Furthermore, the proposed source decoding algorithm based on the symbol-level BCJR algorithm enables the iterative decoding to efficiently exploit the extrinsic information conveyed by relay and source nodes. Simulation results show that the proposed iterative decoding scheme with the proposed source decoder and the symbol-level BCJR channel decoder delivers the robust power-efficient performance, and outperforms conventional decoding schemes. Chun-Feng Wu, Wei-Ho Chung |
IEEE Trans. Wirel. Commun. | 1 |
| 2012 | Symbol-based iterative decoding of convolutionally encoded multiple descriptionsabstractTransmission of convolutionally encoded multiple descriptions over noisy channels can benefit from the use of iterative source-channel decoding. The authors first modified the BCJR algorithm in a way that symbol a posteriori probabilities can be derived and used as extrinsic information to improve the iterative decoding between the source and channel decoders. The authors also formulate a recursive implementation for the source decoder that processes reliability information received on different channels and combines them with inter-description correlation to estimate the transmitted quantiser index. Simulation results are presented for two-channel scalar quantisation of Gauss–Markov sources which demonstrate the error-resilience capabilities of symbol-based iterative decoding. Chun-Feng Wu, Wen-Whei Chang |
IET Commun. | 1 |
| 2009 | Iterative Decoding of Convolutionally Encoded Multiple DescriptionsabstractIn this work, we attempt to capitalize more fully on the source residual redundancy and then develop an MD-ISCD scheme which permits to exchange between its two constituent decoders the whole symbol extrinsic information. The first step toward realization is to derive the modified BCJR algorithm based on sectionalized code trellises that provides reliability information on each transmitted symbol rather than on index-bits. To further reduce the computation, we also apply the concept of Jacobian logarithm to formulate the algorithm in the logarithmic domain. Kuang-Yi Yen, Chun-Feng Wu, Wen-Whei Chang |
DCC | 2 |
| 2007 | Perceptual-based playout mechanisms for multi-stream voice over IP networks
Chun-Feng Wu, Cheng-Lung Lee, Wen-Whei Chang |
INTERSPEECH | 1 |