Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Kuan-Chung Chen

dblp:14/6257 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
0since 2021 · last 2019
0000-0002-9699-2993ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-authorComputer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Processor architecture and microarchitecture · 48% GPUs and heterogeneous computing · 38% Parallel and multicore computing · 14%
Computer networks
1 paper
Internet of things and sensor networks · 67% Routing and switching · 33%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Internet of things and sensor networks › wireless sensor network › mobile sink
mobile sink data collection
0.412019
Efficient Path Planning for a Mobile Sink to Reliably Gather Data from Sensors with Diverse Sensing Rates and Limited Buffers · IEEE Trans. Mob. Comput. 2019
Routing and switching
path planning
0.412019
Efficient Path Planning for a Mobile Sink to Reliably Gather Data from Sensors with Diverse Sensing Rates and Limited Buffers · IEEE Trans. Mob. Comput. 2019
Internet of things and sensor networks
wireless sensor network
0.412019
Efficient Path Planning for a Mobile Sink to Reliably Gather Data from Sensors with Diverse Sensing Rates and Limited Buffers · IEEE Trans. Mob. Comput. 2019
Processor architecture and microarchitecture › multithreading
fine-grain multithreading
0.312018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018
Processor architecture and microarchitecture
multithreading
0.312018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018
GPUs and heterogeneous computing › GPU microarchitecture
SIMT execution
0.312018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018
Graph algorithms and graph theory
spanning tree
0.112019
Efficient Path Planning for a Mobile Sink to Reliably Gather Data from Sensors with Diverse Sensing Rates and Limited Buffers · IEEE Trans. Mob. Comput. 2019
GPUs and heterogeneous computing
control flow divergence
0.112018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018
Parallel and multicore computing
data parallelism
0.112018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018
GPUs and heterogeneous computing › control flow divergence
reconvergence
0.112018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018
Parallel and multicore computing › data parallelism
SIMD vectorization
0.112018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018

Methods — techniques the papers use, named apart from their topics

heuristic algorithm · 0.8NP-completeness reduction · 0.8explicit vectorization · 0.3OpenCL · 0.3
YearPublicationVenuePosition
2019 Efficient Path Planning for a Mobile Sink to Reliably Gather Data from Sensors with Diverse Sensing Rates and Limited Buffers
abstract
Wireless sensor networks are vulnerable to energy holes, where sensors close to a static sink are fast drained of their energy. Using a mobile sink (MS) can conquer this predicament and extend sensor lifetime. How to schedule a traveling path for the MS to efficiently gather data from sensors is critical in performance. Some studies select a subset of sensors as rendezvous points (RPs). Non-RP sensors send data to the nearest RPs and the MS visits RPs to retrieve data. However, these studies assume that sensors produce data with the same speed and have no limitation on buffer size. When the two assumptions are invalid, they may encounter serious packet loss due to buffer overflow at RPs. In the paper, we show that the path planning problem is NP-complete and propose an efficient path planning for reliable data gathering (EARTH) algorithm by relaxing these impractical assumptions. It forms a spanning tree to connect all sensors and then selects each RP based on hop count and distance in the tree and the amount of forwarding data from other sensors. An enhanced EARTH (eEARTH) algorithm is also developed to further reduce path length. Both EARTH and eEARTH incur less computational overhead and can flexibly recompute new paths when sensors change sensing rates. Simulation results verify that they can find short traveling paths for the MS to collect sensing data without packet loss, as compared with existing methods.
You-Chiun Wang, Kuan-Chung Chen
IEEE Trans. Mob. Comput.2
2018 Enabling SIMT Execution Model on Homogeneous Multi-Core System
abstract
Single-instruction multiple-thread (SIMT) machine emerges as a primary computing device in high-perfor-mance computing, since the SIMT execution paradigm can exploit data-level parallelism effectively. This article explores the SIMT execution potential on homogeneous multi-core processors, which generally run in multiple-instruction multiple-data (MIMD) mode when utilizing the multi-core resources. We address three architecture issues in enabling SIMT execution model on multi-core processor, including multithreading execution model, kernel thread context placement, and thread divergence. For the SIMT execution model, we propose a fine-grained multithreading mechanism on an ARM-based multi-core system. Each of the processor cores stores the kernel thread contexts in its L1 data cache for per-cycle thread-switching requirement. For divergence-intensive kernels, an Inner Conditional Statement First (ICS-First) mechanism helps early re-convergence to occur and significantly improves the performance. The experiment results show that effectiveness in data-parallel processing reduces on average 36% dynamic instructions, and boosts the SIMT executions to achieve on average 1.52× and up to 5× speedups over the MIMD counterpart for OpenCL benchmarks for single issue in-order processor cores. By using the explicit vectorization optimization on the kernels, the SIMT model gains further benefits from the SIMD extension and achieves 1.71× speedup over the MIMD approach. The SIMT model using in-order superscalar processor cores outperforms the MIMD model that uses superscalar out-of-order processor cores by 40%. The results show that, to exploit data-level parallelism, enabling the SIMT model on homogeneous multi-core processors is important.
Kuan-Chung Chen, Chung-Ho Chen
ACM Trans. Archit. Code Optim.1
2015 A memory-efficient NoC system for OpenCL many-core platform
abstract
We present a memory-efficient NoC system for OpenCL many-core platforms. By offloading the OpenCL kernel programs into the many-core processor, we reveal that the memory contention overheads have dramatically increased with the scaling of the system, resulting to the poor performance scalability of the many-core system. We explore a memory-efficient NoC system design which includes a mesh network, a hybrid network interface for packet composition and decomposition, and a memory controller with access scheduling capability. Our experimental results show that a simple memory access scheduling and caching approach can easily boost the performance of the NoC and memory system up to 20 percent by eliminating the memory controller contention problem.
Chien-Hsuan Yen, Chung-Ho Chen, Kuan-Chung Chen
ISCAS3
2014 An OpenCL runtime system for a heterogeneous many-core virtual platform
abstract
We present a many-core full system simulation platform and its OpenCL runtime system. The OpenCL runtime system includes an on-the-fly compiler and resource manager for the ARM-based many-core platform. Using this platform, we evaluate approaches of work-item scheduling and memory management in OpenCL memory hierarchy. Our experimental results show that scheduling work-items on a many-core system using general purpose RISC CPU should avoid per work-item context switching. Data deployment and work-item coalescing are the two keys for significant speedup.
Kuan-Chung Chen, Chung-Ho Chen
ISCAS1
2014 Virtualization Technology for TCP/IP Offload Engine
abstract
Network I/O virtualization plays an important role in cloud computing. This paper addresses the system-wide virtualization issues of TCP/IP Offload Engine (TOE) and presents the architectural designs. We identify three critical factors that affect the performance of a TOE: I/O virtualization architectures, quality of service (QoS), and virtual machine monitor (VMM) scheduler. In our device emulation based TOE, the VMM manages the socket connections in the TOE directly and thus can eliminate packet copy and demultiplexing overheads as appeared in the virtualization of a layer 2 network card. To further reduce hypervisor intervention, the direct I/O access architecture provides the per VM-based physical control interface that helps removing most of the VMM interventions. The direct I/O access architecture out-performs the device emulation architecture as large as 30 percent, or achieves 80 percent of the native 10 Gbit/s TOE system. To continue serving the TOE commands for a VM, no matter the VM is idle or switched out by the VMM, we decouple the TOE I/O command dispatcher from the VMM scheduler. We found that a VMM scheduler with preemptive I/O scheduling and a programmable I/O command dispatcher with deficit weighted round robin (DWRR) policy are able to ensure service fairness and at the same time maximize the TOE utilization.
En-Hao Chang, Chen-Chieh Wang, Chien-Te Liu, Kuan-Chung Chen, Chung-Ho Chen
IEEE Trans. Cloud Comput.4
2013 CASL hypervisor and its virtualization platform
abstract
In this paper, we present an ARM-based hardwareassisted hypervisor, named CASL-Hypervisor, and a full system virtualization platform developed in SystemC which enables software/hardware co-simulation of virtual machine systems. CASL-Hypervisor takes advantage of an additional processor mode, extended memory management unit, configurable hardware traps and specialized hardware devices to virtualize unmodified Linux-based guest operating systems. By utilizing hardware extensions, development effort of CASL-Hypervisor can be greatly reduced and the hypervisor has achieved relatively low virtualization overhead. Evaluation is demonstrated on an approximately-timed manner so it is able to do fast software/hardware co-simulation and evaluations. We use the ARM-v7A instruction set simulator as the host processor. The hypervisor overhead can be quantified through instruction count ratio of guest operating system to the hypervisor. The results show that CASL-Hypervisor successfully virtualizes four guest operating systems with about 9.78% overhead.
Chien-Te Liu, Kuan-Chung Chen, Chung-Ho Chen
ISCAS2