EDBT 2026 Demo / reviewers in the wild / expert
Chao-Chin Wu
dblp:79/6694
· DBLP profile ↗
26ranked-venue papers
11as first author
4since 2021 · last 2025
0000-0002-0469-9707ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 8 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient configuration of heterogeneous resources and task scheduling strategies in deep learning auto-tuning systems
Pao-Yi Ken, Chao-Chin Wu |
J. Supercomput. | 2 |
| 2022 | Optimizing 2-opt-based heuristics on GPU for solving the single-row facility layout problem
Ping Chou, Chorng-Shiuh Koong, Chao-Chin Wu, Liang-Rui Chen |
Future Gener. Comput. Syst. | 4 |
| 2021 | Design and Implementation of an iOS APP: Multimedia Interactive System and Items for Woodworking Teaching
Chorng-Shiuh Koong, Hung-Chang Lin, Chao-Chin Wu |
ICCE | 3 |
| 2021 | A lightweight BLASTP and its implementation on CUDA GPUs
Liang-Tsung Huang, Kai-Cheng Wei, Chao-Chin Wu, Jian-An Wang |
J. Supercomput. | 3 |
| 2020 | A Variational Autoencoder-Based Secure Transceiver Design Using Deep LearningabstractTo achieve new applications for 5G communications, physical layer security has recently drawn significant attention. In a wiretap channel system, our goal is to minimize information leakage to an eavesdropper while maximizing the performance of transmission to the desired or legitimate receiver. Complicated systems or channel models make it difficult to design secrecy systems based on the information theory. In this paper, we propose a deep learning-based transceiver design for secrecy systems as an alternative. Specifically, we modify the loss function design of a variational autoencoder, which is a special type of neural network, making it possible to provide both robust data transmission and security in an unsupervised fashion. We further investigate the impact of an imperfect channel state information and use simulation results to prove that our approach can outperform the existing learning-based methods. Chao-Chin Wu, Kuan-Fu Chen, Ta-Sung Lee |
GLOBECOM | 2 |
| 2017 | Reconstructing permutation table to improve the Tabu Search for the PFSP on GPU
Kai-Cheng Wei, Hsun Chu, Chao-Chin Wu |
J. Supercomput. | 4 |
| 2014 | Solving the Permutation Problem Efficiently for Tabu Search on CUDA GPUs
Liang-Tsung Huang, Syun-Sheng Jhan, Yun-Ju Li, Chao-Chin Wu |
ICCCI | 4 |
| 2013 | Using CUDA GPU to Accelerate the Ant Colony Optimization AlgorithmabstractGraph Processing Units (GPUs) have recently evolved into a super multi-core and a fully programmable architecture. In the CUDA programming model, the programmers can simply implement parallelism ideas of a task on GPUs. The purpose of this paper is to accelerate Ant Colony Optimization (ACO) for Traveling Salesman Problems (TSP) with GPUs. In this paper, we propose a new parallel method, which is called the Transition Condition Method. Experimental results are extensively compared and evaluated on the performance side and the solution quality side. The TSP problems are used as a standard benchmark for our experiments. In terms of experimental results, our new parallel method achieves the maximal speed-up factor of 4.74 than the previous parallel method. On the other hand, the quality of solutions is similar to the original sequential ACO algorithm. It proves that the quality of solutions does not be sacrificed in the cause of speed-up. Kai-Cheng Wei, Chao-Chin Wu, Chien-Ju Wu |
PDCAT | 2 |
| 2012 | Optimizing Dynamic Programming on Graphics Processing Units Via Data Reuse and Data Prefetch with Inter-Block Barrier SynchronizationabstractOur previous study focused on accelerating an important category of DP problems, called nonserial polyadic dynamic programming (NPDP), on a graphics processing unit (GPU). In NPDP applications, the degree of parallelism varies significantly in different stages of computation, making it difficult to fully utilize the compute power of hundreds of pro-cessing cores in a GPU. To address this challenge, we proposed a methodology that can adaptively adjust the thread-level parallelism in mapping a NPDP problem onto the GPU, thus providing sufficient and steady degrees of parallelism across different compute stages. This work aims at further improving the performance of NPDP problems. Sub problems and data are tiled to make it possible to fit small data regions into shared memory and reuse the buffered data for each tile of sub problems, thus reducing the amount of global memory access. However, we found invoking the same kernel many times, due to data consistency enforcement across different stages, makes it impossible to reuse the tiled data in shared memory after the kernel is invoked again. Fortunately, the inter-block synchronization technique allows us to invoke the kernel exactly one time with the restriction that the maximum number of blocks is equal to the total number of streaming multiprocessors. In addition to data reuse, invoking the kernel only one time also enables us to prefetch data to shared memory across inter-block synchronization point, which improves the performance more than data reuse. We realize our approach in a real-world NPDP application â" the optimal matrix parenthesization problem. Experimental results demonstrate invoking a kernel only one time cannot guarantee performance improvement unless we also reuse and prefetch data across barrier synchronization points. Chao-Chin Wu, Kai-Cheng Wei, Ting-Hong Lin |
ICPADS | 1 |
| 2012 | Extending FuzzyCLIPS for parallelizing data-dependent fuzzy expert systems
Chao-Chin Wu, Lien Fu Lai, Yu-Shuo Chang |
J. Supercomput. | 1 |
| 2012 | Performance evaluation of enhancement of the layered self-scheduling approach for heterogeneous multicore cluster systems
Chao-Chin Wu, Lien Fu Lai, Liang-Tsung Huang, Ming-Lung Chen |
J. Supercomput. | 1 |
| 2012 | Using hybrid MPI and OpenMP programming to optimize communications in parallel loop self-scheduling schemes for multicore PC clusters
Chao-Chin Wu, Lien Fu Lai, Chao-Tung Yang, Po-Hsun Chiu |
J. Supercomput. | 1 |
| 2012 | Designing parallel loop self-scheduling schemes using the hybrid MPI and OpenMP programming model for multi-core grid systems
Chao-Chin Wu, Chao-Tung Yang, Kuan-Chou Lai, Po-Hsun Chiu |
J. Supercomput. | 1 |
| 2011 | Developing a fuzzy search engine based on fuzzy ontology and semantic searchabstractMost of existing search engines retrieve web pages by means of finding exact keywords. Traditional keyword-based search engines suffer several problems. First, synonyms and terms similar to keywords are not taken into consideration to search web pages. Users may need to input several similar keywords individually to complete a search. Second, traditional search engines treat all keywords as the same importance and cannot differentiate the importance of one keyword from that of another. Third, traditional search engines lack an applicable classification mechanism to reduce the search space and improve the search results. In this paper, we develop a fuzzy search engine, called Fuzzy-Go. First, a fuzzy ontology is constructed by using fuzzy logic to capture the similarities of terms in the ontology, which offering appropriate semantic distances between terms to accomplish the semantic search of keywords. The Fuzzy Go search engine can thus automatically retrieve web pages that contain synonyms or terms similar to keywords. Second, users can input multiple keywords with different degrees of importance based on their needs. The totally satisfactory degree of keywords can be aggregated based on their degrees of importance and degrees of satisfaction. Third, the domain classification of web pages offers users to select the appropriate domain for searching web pages, which excludes web pages in the inappropriate domains to reduce the search space and to improve the search results. Lien Fu Lai, Chao-Chin Wu, Pei-Ying Lin, Liang-Tsung Huang |
FUZZ-IEEE | 2 |
| 2011 | First Report of Knowledge Discovery in Predicting Protein Folding Rate Change upon Single Mutation
Lien Fu Lai, Chao-Chin Wu, Liang-Tsung Huang |
ICIC (3) | 2 |
| 2011 | Optimizing Dynamic Programming on Graphics Processing Units via Adaptive Thread-Level ParallelismabstractDynamic programming (DP) is an important computational method for solving a wide variety of discrete optimization problems such as scheduling, string editing, packaging, and inventory management. In general, DP is classified into four categories based on the characteristics of the optimization equation. Because applications that are classified in the same category of DP have similar program behavior, the research community has sought to propose general solutions for parallelizing each category of DP. However, most existing studies focus on running DP on CPU-based parallel systems rather than on accelerating DP algorithms on the graphics processing unit (GPU). This paper presents the GPU acceleration of an important category of DP problems called nonserial polyadic dynamic programming (NPDP). In NPDP applications, the degree of parallelism varies significantly in different stages of computation, making it difficult to fully utilize the compute power of hundreds of processing cores in a GPU. To address this challenge, we propose a methodology that can adaptively adjust the thread-level parallelism in mapping a NPDP problem onto the GPU, thus providing sufficient and steady degrees of parallelism across different compute stages. We realize our approach in a real-world NPDP application -- the optimal matrix parenthesization problem. Experimental results demonstrate our method can achieve a speedup of 13.40 over the previously published GPU algorithm. Chao-Chin Wu, Jenn-Yang Ke, Heshan Lin, Wu-chun Feng |
ICPADS | 1 |
| 2011 | Performance-based parallel loop self-scheduling using hybrid OpenMP and MPI programming on multicore SMP clustersabstractAbstract Parallel loop self‐scheduling on parallel and distributed systems has been a critical problem and it is becoming more difficult to deal with in the emerging heterogeneous cluster computing environments. In the past, some self‐scheduling schemes have been proposed as applicable to heterogeneous cluster computing environments. In recent years, multicore computers have been widely included in cluster systems. However, previous researches into parallel loop self‐scheduling did not consider certain aspects of multicore computers; for example, it is more appropriate for shared‐memory multiprocessors to adopt Open Multi‐Processing (OpenMP) for parallel programming. In this paper, we propose a performance‐based approach using hybrid OpenMP and MPI parallel programming, which partition loop iterations according to the performance weighting of multicore nodes in a cluster. Because iterations assigned to one MPI process are processed in parallel by OpenMP threads run by the processor cores in the same computational node, the number of loop iterations allocated to one computational node at each scheduling step depends on the number of processor cores in that node. Experimental results show that the proposed approach performs better than previous schemes. Copyright © 2010 John Wiley & Sons, Ltd. Chao-Tung Yang, Chao-Chin Wu, Jen-Hsiang Chang |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | Parallelizing a CLIPS-based course timetabling expert system
Chao-Chin Wu |
Expert Syst. Appl. | 1 |
| 2010 | Predicting Protein Stability Change upon Double Mutation from Partial Sequence Information Using Data Mining Approach
Lien Fu Lai, Chao-Chin Wu, Liang-Tsung Huang |
ICIC (1) | 2 |
| 2010 | An integrated security-aware job scheduling strategy for large-scale computational grids
Chao-Chin Wu, Ren-Yi Sun |
Future Gener. Comput. Syst. | 1 |
| 2010 | Development of knowledge-based system for predicting the stability of proteins upon point mutations
Liang-Tsung Huang, Lien Fu Lai, Chao-Chin Wu, M. Michael Gromiha |
Neurocomputing | 3 |
| 2009 | Developing the KMKE Knowledge Management System Based on Design Patterns and Parallel Processing
Lien Fu Lai, Chao-Chin Wu, Liang-Tsung Huang, Ya-Chin Chang |
ICIC (1) | 2 |
| 2008 | Parallel Loop Self-Scheduling for Heterogeneous Cluster Systems with Multi-core ComputersabstractMulticore computers have been widely included in cluster systems. They are shared memory architecture. However, previous research on parallel loop self-scheduling did not consider the feature of multicore computers. It is more suitable for shared-memory multiprocessors to adopt OpenMP for parallel programming. Therefore, in this paper, we propose to adopt hybrid programming model MPI+OpenMP to design loop self-scheduling schemes for cluster systems with multicore computers. Initially, each computer runs only one MPI process no matter how many cores it has. A MPI process will fork OpenMP threads depending on the number of cores in the computer. Each idle slave MPI-process will request tasks from the master process. The tasks dispatched to a process will be executed in parallel by OpenMP threads. According to the experimental results, our method outperforms the previous work by 18.66% or 29.76% depending on the problem size. Moreover, the performance improvement is very stable no matter our method is based on which traditional scheme. Chao-Chin Wu, Lien Fu Lai, Po-Hsun Chiu |
APSCC | 1 |
| 2008 | GA-Based Job Scheduling Strategies for Fault Tolerant Grid SystemsabstractThis work mainly aims at the designs of the genetic algorithm based scheduling strategies by considering four different fault tolerance techniques in the Grid environment, including Retry, Migration, Checkpoint, Replication. We also take into account the risk relationship between jobs and nodes to improve the system reliability in the scheduling algorithm. According to the simulation results, we can find out that the performance of fault tolerant algorithms is better than risky algorithm whether in makespan, average turnaround time, or the job failure rate. Checkpoint algorithm has the best performance in all algorithms. On the other hand, retry algorithm is recommended for the system where the job sizes are usually smaller because of its simplicity. Finally, replicated algorithm is not suitable for the Grid since it imposes too much overhead. Chao-Chin Wu, Kuan-Chou Lai, Ren-Yi Sun |
APSCC | 1 |
| 2001 | Design of a scalable multiprocessor architecture and its simulation
Der-Lin Pean, Chao-Chin Wu, Huey-Ting Chua |
J. Syst. Softw. | 2 |
| 1998 | Look-Ahead Memory Consistency ModelabstractWe propose a hardware-centric look-ahead memory consistency model that makes the data consistent according to the special ordering requirement of memory accesses for critical sections. The novel model imposes fewer restrictions on event ordering than previously proposed models thus offering the potential of higher performance. The architecture has the following features: blocking and waking up processes by hardware; allowing instructions to be executed out-of-order; until having acquired the lock can the processor allow the requests for accessing the protected data to be evicted to the memory subsystem. The advantages of the look-ahead model include: more program segments are allowed parallel execution; locks can be released earlier, resulting in reduced waiting times for acquiring locks; and less network traffic because more write requests are merged by using two write caches. Chao-Chin Wu, Der-Lin Pean |
ICPADS | 1 |