Chih-Chieh Chou

dblp:22/6676 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
1since 2021 · last 2022
0000-0002-3094-6951ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 80% Hardware accelerators and domain-specific architectures · 20%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%
Computer networks
1 paper
Optical networks · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory management › virtual memory
address translation
0.612022
Reducing Minor Page Fault Overheads through Enhanced Page Walker · ACM Trans. Archit. Code Optim. 2022
Memory systems › memory management › virtual memory
page fault handling
0.612022
Reducing Minor Page Fault Overheads through Enhanced Page Walker · ACM Trans. Archit. Code Optim. 2022
Memory systems › memory management › virtual memory › address translation
TLB
0.612022
Reducing Minor Page Fault Overheads through Enhanced Page Walker · ACM Trans. Archit. Code Optim. 2022
Memory systems › memory management
virtual memory
0.612022
Reducing Minor Page Fault Overheads through Enhanced Page Walker · ACM Trans. Archit. Code Optim. 2022
Operating systems › resource management
memory management
0.212022
Reducing Minor Page Fault Overheads through Enhanced Page Walker · ACM Trans. Archit. Code Optim. 2022
Operating systems › resource management › memory management
page allocation
0.212022
Reducing Minor Page Fault Overheads through Enhanced Page Walker · ACM Trans. Archit. Code Optim. 2022
Optical networks › optical buffer
fiber delay lines
0.112006
A Necessary and Sufficient Condition for the Construction of 2-to-1 Optical FIFO Multiplexers by a Single Crossbar Switch and Fiber Delay Lines · IEEE Trans. Inf. Theory 2006
Optical networks
optical buffer
0.112006
A Necessary and Sufficient Condition for the Construction of 2-to-1 Optical FIFO Multiplexers by a Single Crossbar Switch and Fiber Delay Lines · IEEE Trans. Inf. Theory 2006
Optical networks › optical communication components
optical multiplexer
0.112006
A Necessary and Sufficient Condition for the Construction of 2-to-1 Optical FIFO Multiplexers by a Single Crossbar Switch and Fiber Delay Lines · IEEE Trans. Inf. Theory 2006

Methods — techniques the papers use, named apart from their topics

hardware-software co-design · 1.1gem5 simulation · 1.1packet routing · 0.1crossbar switch · 0.1c-transform · 0.1
YearPublicationVenuePosition
2022 Reducing Minor Page Fault Overheads through Enhanced Page Walker
abstract
Application virtual memory footprints are growing rapidly in all systems from servers down to smartphones. To address this growing demand, system integrators are incorporating ever larger amounts of main memory, warranting rethinking of memory management. In current systems, applications produce page fault exceptions whenever they access virtual memory regions that are not backed by a physical page. As application memory footprints grow, they induce more and more minor page faults. Handling of each minor page fault can take a few thousands of CPU cycles and blocks the application till the OS kernel finds a free physical frame. These page faults can be detrimental to the performance when their frequency of occurrence is high and spread across application runtime. Specifically, lazy allocation-induced minor page faults are increasingly impacting application performance. Our evaluation of several workloads indicates an overhead due to minor page faults as high as 29% of execution time. In this article, we propose to mitigate this problem through a hardware, software co-design approach. Specifically, we first propose to parallelize portions of the kernel page allocation to run ahead of fault time in a separate thread. Then we propose the Minor Fault Offload Engine (MFOE), a per-core hardware accelerator for minor fault handling. MFOE is equipped with a pre-allocated page frame table that it uses to service a page fault. On a page fault, MFOE quickly picks a pre-allocated page frame from this table, makes an entry for it in the TLB, and updates the page table entry to satisfy the page fault. The pre-allocation frame tables are periodically refreshed by a background kernel thread, which also updates the data structures in the kernel to account for the handled page faults. We evaluate this system in the gem5 architectural simulator with a modified Linux kernel running on top of simulated hardware containing the MFOE accelerator. Our results show that MFOE improves the average critical path fault handling latency by 33× and tail critical path latency by 51×. Among the evaluated applications, we observed an improvement of runtime by an average of 6.6%.
Chandrahas Tirumalasetty, Chih-Chieh Chou, A. L. Narasimha Reddy, Paul Gratz, Ayman Abouelwafa
ACM Trans. Archit. Code Optim.2
2020 Virtualize and share non-volatile memories in user space
Chih-Chieh Chou, Jaemin Jung, A. L. Narasimha Reddy, Paul Gratz, Doug Voigt
CCF Trans. High Perform. Comput.1
2019 vNVML: An Efficient User Space Library for Virtualizing and Sharing Non-Volatile Memories
abstract
The emerging non-volatile memory (NVM) has attractive characteristics such as DRAM-like, low-latency together with the non-volatility of storage devices. Recently, byte-addressable, memory bus-attached NVM has become available. This paper addresses the problem of combining a smaller, faster byte-addressable NVM with a larger, slower storage device, like SSD, to create the impression of a larger and faster byte-addressable NVM which can be shared across many applications. In this paper, we propose vNVML, a user space library for virtualizing and sharing NVM. vNVML provides for applications transaction like memory semantics that ensures write ordering and persistency guarantees across system failures. vNVML exploits DRAM for read caching, to enable improvements in performance and potentially to reduce the number of writes to NVM, extending the NVM lifetime. vNVML is implemented and evaluated with realistic workloads to show that our library allows applications to share NVM, both in a single O/S and when docker like containers are employed. The results from the evaluation show that vNVML incurs less than 10% overhead while providing the benefits of an expanded virtualized NVM space to the applications, allowing applications to safely share the virtual NVM.
Chih-Chieh Chou, Jaemin Jung, A. L. Narasimha Reddy, Paul Gratz, Doug Voigt
MSST1
2006 A Necessary and Sufficient Condition for the Construction of 2-to-1 Optical FIFO Multiplexers by a Single Crossbar Switch and Fiber Delay Lines
abstract
In this paper, we prove a necessary and sufficient condition for the construction of 2-to-1 optical buffered first-in–first-out (FIFO) multiplexers by a single crossbar switch and fiber delay lines. We consider a feedback system consisting of an$(M+2)times (M+2)$crossbar switch and$M$fiber delay lines with delays$d_1, d_2,ldots, d_M$. These$M$fiber delay lines are connected from$M$outputs of the crossbar switch back to$M$inputs of the switch, leaving two inputs (respectively, two outputs) of the switch for the two inputs (respectively, two outputs) of the 2-to-1 multiplexer. The main contribution of this paper is the formal proof that$d_1=1$and$d_i le d_i+1 le 2d_i$,$i=1,2, ldots, M-1$, is a necessary and sufficient condition on the delays$d_1, d_2,ldots,d_M$for such a feedback system to be operated as a 2-to-1 FIFO multiplexer with buffer$sum _i=1^M d_i$under a simple packet routing policy. Specifically, the routing of a packet is according to a specific decomposition of the packet delay, called the$cal C$-transform in this paper. Our result shows that under such a feedback architecture a 2-to-1 FIFO multiplexer can be constructed with$M=O(log B)$, where$B$is the buffer size. Therefore, our construction improves on a more complicated construction recently proposed by Sarwate and Anantharam that requires$M=O(sqrt B)$under the same feedback architecture (we note that their design is more general and works for priority queues).
Chih-Chieh Chou, Cheng-Shang Chang, Duan-Shin Lee, Jay Cheng
IEEE Trans. Inf. Theory1
2005 Communication-driven task binding for multiprocessor with latency insensitive network-on-chip
abstract
Network-on-Chip is a new design paradigm for designing core based System-on-Chip. It features high degree of reusability and scalability. In this paper, we propose a switch which employs the latency insensitive concepts and applies the round-robin scheduling techniques to achieve high communication resource utilization. Based on the assumptions of the 2D-mesh network topology constructed by the switch, this work not only models the communication and the contention effect of the network, but develops a communication-driven task binding algorithm that employs the divide and conquer strategy to map applications onto the multiprocessor system-on-chip. The algorithm attempts to derive a binding of tasks such that the overall system throughput is maximized. To compare with the task binding without consideration of communication and contention effect, the experimental results demonstrate that the overall improvement of the system throughput is 20% for 844 test cases.
Liang-Yu Lin, Cheng-Yeh Wang, Pao-Jui Huang, Chih-Chieh Chou, Jing-Yang Jou
ASP-DAC4
2003 Bounding the Execution Times of DMA I/O Tasks on Hard-Real-Time Embedded Systems
Chih-Chieh Chou, Po-Yuan Chen
RTCSA2