Chao-Chin Wu

dblp:79/6694 · DBLP profile ↗
← Back
26ranked-venue papers
11as first author
4since 2021 · last 2025
0000-0002-0469-9707ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 8 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Efficient configuration of heterogeneous resources and task scheduling strategies in deep learning auto-tuning systems
Pao-Yi Ken, Chao-Chin Wu
J. Supercomput.2
2022 Optimizing 2-opt-based heuristics on GPU for solving the single-row facility layout problem
Ping Chou, Chorng-Shiuh Koong, Chao-Chin Wu, Liang-Rui Chen
Future Gener. Comput. Syst.4
2021 Design and Implementation of an iOS APP: Multimedia Interactive System and Items for Woodworking Teaching
Chorng-Shiuh Koong, Hung-Chang Lin, Chao-Chin Wu
ICCE3
2021 A lightweight BLASTP and its implementation on CUDA GPUs
Liang-Tsung Huang, Kai-Cheng Wei, Chao-Chin Wu, Jian-An Wang
J. Supercomput.3
2020 A Variational Autoencoder-Based Secure Transceiver Design Using Deep Learning
abstract
To achieve new applications for 5G communications, physical layer security has recently drawn significant attention. In a wiretap channel system, our goal is to minimize information leakage to an eavesdropper while maximizing the performance of transmission to the desired or legitimate receiver. Complicated systems or channel models make it difficult to design secrecy systems based on the information theory. In this paper, we propose a deep learning-based transceiver design for secrecy systems as an alternative. Specifically, we modify the loss function design of a variational autoencoder, which is a special type of neural network, making it possible to provide both robust data transmission and security in an unsupervised fashion. We further investigate the impact of an imperfect channel state information and use simulation results to prove that our approach can outperform the existing learning-based methods.
Chao-Chin Wu, Kuan-Fu Chen, Ta-Sung Lee
GLOBECOM2
2017 Reconstructing permutation table to improve the Tabu Search for the PFSP on GPU
Kai-Cheng Wei, Hsun Chu, Chao-Chin Wu
J. Supercomput.4
2014 Solving the Permutation Problem Efficiently for Tabu Search on CUDA GPUs
Liang-Tsung Huang, Syun-Sheng Jhan, Yun-Ju Li, Chao-Chin Wu
ICCCI4
2013 Using CUDA GPU to Accelerate the Ant Colony Optimization Algorithm
abstract
Graph Processing Units (GPUs) have recently evolved into a super multi-core and a fully programmable architecture. In the CUDA programming model, the programmers can simply implement parallelism ideas of a task on GPUs. The purpose of this paper is to accelerate Ant Colony Optimization (ACO) for Traveling Salesman Problems (TSP) with GPUs. In this paper, we propose a new parallel method, which is called the Transition Condition Method. Experimental results are extensively compared and evaluated on the performance side and the solution quality side. The TSP problems are used as a standard benchmark for our experiments. In terms of experimental results, our new parallel method achieves the maximal speed-up factor of 4.74 than the previous parallel method. On the other hand, the quality of solutions is similar to the original sequential ACO algorithm. It proves that the quality of solutions does not be sacrificed in the cause of speed-up.
Kai-Cheng Wei, Chao-Chin Wu, Chien-Ju Wu
PDCAT2
2012 Optimizing Dynamic Programming on Graphics Processing Units Via Data Reuse and Data Prefetch with Inter-Block Barrier Synchronization
abstract
Our previous study focused on accelerating an important category of DP problems, called nonserial polyadic dynamic programming (NPDP), on a graphics processing unit (GPU). In NPDP applications, the degree of parallelism varies significantly in different stages of computation, making it difficult to fully utilize the compute power of hundreds of pro-cessing cores in a GPU. To address this challenge, we proposed a methodology that can adaptively adjust the thread-level parallelism in mapping a NPDP problem onto the GPU, thus providing sufficient and steady degrees of parallelism across different compute stages. This work aims at further improving the performance of NPDP problems. Sub problems and data are tiled to make it possible to fit small data regions into shared memory and reuse the buffered data for each tile of sub problems, thus reducing the amount of global memory access. However, we found invoking the same kernel many times, due to data consistency enforcement across different stages, makes it impossible to reuse the tiled data in shared memory after the kernel is invoked again. Fortunately, the inter-block synchronization technique allows us to invoke the kernel exactly one time with the restriction that the maximum number of blocks is equal to the total number of streaming multiprocessors. In addition to data reuse, invoking the kernel only one time also enables us to prefetch data to shared memory across inter-block synchronization point, which improves the performance more than data reuse. We realize our approach in a real-world NPDP application â" the optimal matrix parenthesization problem. Experimental results demonstrate invoking a kernel only one time cannot guarantee performance improvement unless we also reuse and prefetch data across barrier synchronization points.
Chao-Chin Wu, Kai-Cheng Wei, Ting-Hong Lin
ICPADS1
2012 Extending FuzzyCLIPS for parallelizing data-dependent fuzzy expert systems
Chao-Chin Wu, Lien Fu Lai, Yu-Shuo Chang
J. Supercomput.1
2012 Performance evaluation of enhancement of the layered self-scheduling approach for heterogeneous multicore cluster systems
Chao-Chin Wu, Lien Fu Lai, Liang-Tsung Huang, Ming-Lung Chen
J. Supercomput.1
2012 Using hybrid MPI and OpenMP programming to optimize communications in parallel loop self-scheduling schemes for multicore PC clusters
Chao-Chin Wu, Lien Fu Lai, Chao-Tung Yang, Po-Hsun Chiu
J. Supercomput.1
2012 Designing parallel loop self-scheduling schemes using the hybrid MPI and OpenMP programming model for multi-core grid systems
Chao-Chin Wu, Chao-Tung Yang, Kuan-Chou Lai, Po-Hsun Chiu
J. Supercomput.1
2011 Developing a fuzzy search engine based on fuzzy ontology and semantic search
abstract
Most of existing search engines retrieve web pages by means of finding exact keywords. Traditional keyword-based search engines suffer several problems. First, synonyms and terms similar to keywords are not taken into consideration to search web pages. Users may need to input several similar keywords individually to complete a search. Second, traditional search engines treat all keywords as the same importance and cannot differentiate the importance of one keyword from that of another. Third, traditional search engines lack an applicable classification mechanism to reduce the search space and improve the search results. In this paper, we develop a fuzzy search engine, called Fuzzy-Go. First, a fuzzy ontology is constructed by using fuzzy logic to capture the similarities of terms in the ontology, which offering appropriate semantic distances between terms to accomplish the semantic search of keywords. The Fuzzy Go search engine can thus automatically retrieve web pages that contain synonyms or terms similar to keywords. Second, users can input multiple keywords with different degrees of importance based on their needs. The totally satisfactory degree of keywords can be aggregated based on their degrees of importance and degrees of satisfaction. Third, the domain classification of web pages offers users to select the appropriate domain for searching web pages, which excludes web pages in the inappropriate domains to reduce the search space and to improve the search results.
Lien Fu Lai, Chao-Chin Wu, Pei-Ying Lin, Liang-Tsung Huang
FUZZ-IEEE2
2011 First Report of Knowledge Discovery in Predicting Protein Folding Rate Change upon Single Mutation
Lien Fu Lai, Chao-Chin Wu, Liang-Tsung Huang
ICIC (3)2
2011 Optimizing Dynamic Programming on Graphics Processing Units via Adaptive Thread-Level Parallelism
abstract
Dynamic programming (DP) is an important computational method for solving a wide variety of discrete optimization problems such as scheduling, string editing, packaging, and inventory management. In general, DP is classified into four categories based on the characteristics of the optimization equation. Because applications that are classified in the same category of DP have similar program behavior, the research community has sought to propose general solutions for parallelizing each category of DP. However, most existing studies focus on running DP on CPU-based parallel systems rather than on accelerating DP algorithms on the graphics processing unit (GPU). This paper presents the GPU acceleration of an important category of DP problems called nonserial polyadic dynamic programming (NPDP). In NPDP applications, the degree of parallelism varies significantly in different stages of computation, making it difficult to fully utilize the compute power of hundreds of processing cores in a GPU. To address this challenge, we propose a methodology that can adaptively adjust the thread-level parallelism in mapping a NPDP problem onto the GPU, thus providing sufficient and steady degrees of parallelism across different compute stages. We realize our approach in a real-world NPDP application -- the optimal matrix parenthesization problem. Experimental results demonstrate our method can achieve a speedup of 13.40 over the previously published GPU algorithm.
Chao-Chin Wu, Jenn-Yang Ke, Heshan Lin, Wu-chun Feng
ICPADS1
2011 Performance-based parallel loop self-scheduling using hybrid OpenMP and MPI programming on multicore SMP clusters
abstract
Abstract Parallel loop self‐scheduling on parallel and distributed systems has been a critical problem and it is becoming more difficult to deal with in the emerging heterogeneous cluster computing environments. In the past, some self‐scheduling schemes have been proposed as applicable to heterogeneous cluster computing environments. In recent years, multicore computers have been widely included in cluster systems. However, previous researches into parallel loop self‐scheduling did not consider certain aspects of multicore computers; for example, it is more appropriate for shared‐memory multiprocessors to adopt Open Multi‐Processing (OpenMP) for parallel programming. In this paper, we propose a performance‐based approach using hybrid OpenMP and MPI parallel programming, which partition loop iterations according to the performance weighting of multicore nodes in a cluster. Because iterations assigned to one MPI process are processed in parallel by OpenMP threads run by the processor cores in the same computational node, the number of loop iterations allocated to one computational node at each scheduling step depends on the number of processor cores in that node. Experimental results show that the proposed approach performs better than previous schemes. Copyright © 2010 John Wiley & Sons, Ltd.
Chao-Tung Yang, Chao-Chin Wu, Jen-Hsiang Chang
Concurr. Comput. Pract. Exp.2
2011 Parallelizing a CLIPS-based course timetabling expert system
Chao-Chin Wu
Expert Syst. Appl.1
2010 Predicting Protein Stability Change upon Double Mutation from Partial Sequence Information Using Data Mining Approach
Lien Fu Lai, Chao-Chin Wu, Liang-Tsung Huang
ICIC (1)2
2010 An integrated security-aware job scheduling strategy for large-scale computational grids
Chao-Chin Wu, Ren-Yi Sun
Future Gener. Comput. Syst.1
2010 Development of knowledge-based system for predicting the stability of proteins upon point mutations
Liang-Tsung Huang, Lien Fu Lai, Chao-Chin Wu, M. Michael Gromiha
Neurocomputing3
2009 Developing the KMKE Knowledge Management System Based on Design Patterns and Parallel Processing
Lien Fu Lai, Chao-Chin Wu, Liang-Tsung Huang, Ya-Chin Chang
ICIC (1)2
2008 Parallel Loop Self-Scheduling for Heterogeneous Cluster Systems with Multi-core Computers
abstract
Multicore computers have been widely included in cluster systems. They are shared memory architecture. However, previous research on parallel loop self-scheduling did not consider the feature of multicore computers. It is more suitable for shared-memory multiprocessors to adopt OpenMP for parallel programming. Therefore, in this paper, we propose to adopt hybrid programming model MPI+OpenMP to design loop self-scheduling schemes for cluster systems with multicore computers. Initially, each computer runs only one MPI process no matter how many cores it has. A MPI process will fork OpenMP threads depending on the number of cores in the computer. Each idle slave MPI-process will request tasks from the master process. The tasks dispatched to a process will be executed in parallel by OpenMP threads. According to the experimental results, our method outperforms the previous work by 18.66% or 29.76% depending on the problem size. Moreover, the performance improvement is very stable no matter our method is based on which traditional scheme.
Chao-Chin Wu, Lien Fu Lai, Po-Hsun Chiu
APSCC1
2008 GA-Based Job Scheduling Strategies for Fault Tolerant Grid Systems
abstract
This work mainly aims at the designs of the genetic algorithm based scheduling strategies by considering four different fault tolerance techniques in the Grid environment, including Retry, Migration, Checkpoint, Replication. We also take into account the risk relationship between jobs and nodes to improve the system reliability in the scheduling algorithm. According to the simulation results, we can find out that the performance of fault tolerant algorithms is better than risky algorithm whether in makespan, average turnaround time, or the job failure rate. Checkpoint algorithm has the best performance in all algorithms. On the other hand, retry algorithm is recommended for the system where the job sizes are usually smaller because of its simplicity. Finally, replicated algorithm is not suitable for the Grid since it imposes too much overhead.
Chao-Chin Wu, Kuan-Chou Lai, Ren-Yi Sun
APSCC1
2001 Design of a scalable multiprocessor architecture and its simulation
Der-Lin Pean, Chao-Chin Wu, Huey-Ting Chua
J. Syst. Softw.2
1998 Look-Ahead Memory Consistency Model
abstract
We propose a hardware-centric look-ahead memory consistency model that makes the data consistent according to the special ordering requirement of memory accesses for critical sections. The novel model imposes fewer restrictions on event ordering than previously proposed models thus offering the potential of higher performance. The architecture has the following features: blocking and waking up processes by hardware; allowing instructions to be executed out-of-order; until having acquired the lock can the processor allow the requests for accessing the protected data to be evicted to the memory subsystem. The advantages of the look-ahead model include: more program segments are allowed parallel execution; locks can be released earlier, resulting in reduced waiting times for acquiring locks; and less network traffic because more write requests are merged by using two write caches.
Chao-Chin Wu, Der-Lin Pean
ICPADS1