Hung-Hsin Chen

dblp:68/8805 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
3since 2021 · last 2023
0000-0003-4171-7991ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
GPUs and heterogeneous computing · 28% Cloud and datacenter computing · 25% Electronic design automation · 19%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
genomics
0.712023
IMMerge: merging imputation data at scale · Bioinform. 2023
Bioinformatics and computational biology › genomics › genotyping
genotype imputation
0.712023
IMMerge: merging imputation data at scale · Bioinform. 2023
Performance modeling and evaluation › parallel system performance
strong and weak scaling
0.512021
Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From National Tsing Hua University · IEEE Trans. Parallel Distributed Syst. 2021
Cloud and datacenter computing
cluster resource management and scheduling
0.412020
KubeShare: A Framework to Manage GPUs as First-Class and Shared Resources in Container Cloud · HPDC 2020
Cloud and datacenter computing › cloud platform
container cloud
0.412020
KubeShare: A Framework to Manage GPUs as First-Class and Shared Resources in Container Cloud · HPDC 2020
GPUs and heterogeneous computing
GPU resource management
0.412020
KubeShare: A Framework to Manage GPUs as First-Class and Shared Resources in Container Cloud · HPDC 2020
GPUs and heterogeneous computing
GPU sharing
0.412020
KubeShare: A Framework to Manage GPUs as First-Class and Shared Resources in Container Cloud · HPDC 2020
Electronic design automation › hardware verification and test
fault modeling
0.212013
Fault Models and Test Methods for Subthreshold SRAMs · IEEE Trans. Computers 2013
Electronic design automation
hardware verification and test
0.212013
Fault Models and Test Methods for Subthreshold SRAMs · IEEE Trans. Computers 2013
Electronic design automation › hardware verification and test
memory testing
0.212013
Fault Models and Test Methods for Subthreshold SRAMs · IEEE Trans. Computers 2013
Electronic design automation › hardware verification and test › memory testing
SRAM testing
0.212013
Fault Models and Test Methods for Subthreshold SRAMs · IEEE Trans. Computers 2013
High-performance computing › numerical linear algebra
eigensolver
0.112021
Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From National Tsing Hua University · IEEE Trans. Parallel Distributed Syst. 2021
High-performance computing › numerical linear algebra › eigensolver
polynomial filtering eigensolver
0.112021
Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From National Tsing Hua University · IEEE Trans. Parallel Distributed Syst. 2021
Machine learning › Efficient and distributed learning
distributed training
0.112020
KubeShare: A Framework to Manage GPUs as First-Class and Shared Resources in Container Cloud · HPDC 2020
GPUs and heterogeneous computing
GPU computing
0.112020
KubeShare: A Framework to Manage GPUs as First-Class and Shared Resources in Container Cloud · HPDC 2020
Memory systems › random-access memory
SRAM
0.012013
Fault Models and Test Methods for Subthreshold SRAMs · IEEE Trans. Computers 2013

Methods — techniques the papers use, named apart from their topics

fine-grained allocation · 0.9GPU virtualization · 0.9multiprocessing · 0.7fisher's z transformation · 0.7polynomial filtering · 0.5parallel eigensolver · 0.5sense amplifier analysis · 0.2open defect analysis · 0.2address decoder fault analysis · 0.2
YearPublicationVenuePosition
2023 IMMerge: merging imputation data at scale
abstract
SUMMARY: Genomic data are often processed in batches and analyzed together to save time. However, it is challenging to combine multiple large VCFs and properly handle imputation quality and missing variants due to the limitations of available tools. To address these concerns, we developed IMMerge, a Python-based tool that takes advantage of multiprocessing to reduce running time. For the first time in a publicly available tool, imputation quality scores are correctly combined with Fisher's z transformation. AVAILABILITY AND IMPLEMENTATION: IMMerge is an open-source project under MIT license. Source code and user manual are available at https://github.com/belowlab/IMMerge.
Wanying Zhu, Hung-Hsin Chen, Alexander S. Petty, Lauren E. Petty, Hannah G. Polikowsky, Eric R. Gamazon, Jennifer E. Below, Heather M. Highland
Bioinform.2
2023 Gemini: Enabling Multi-Tenant GPU Sharing Based on Kernel Burst Estimation
abstract
Recent years have seen rapid adoption of GPUs in various types of platforms because of the tremendous throughput powered by massive parallelism. However, as the computing power of GPU continues to grow at a rapid pace, it also becomes harder to utilize these additional resources effectively with the support of GPU sharing. In this work, we designed and implementedGemini, a user-space runtime scheduling framework to enable fine-grained GPU allocation control with support for multi-tenancy and elastic allocation, which are critical for cloud and resource providers. Our key idea is to introduce the concept ofkernel burst, which refers to a group of consecutive kernels launched together without being interrupted by synchronous events. Based on the characteristics of kernel burst, we proposed a low overheadevent-driven monitorand adynamic time-sharing schedulerto achieve our goals. Our experiment evaluations using five types of GPU applications show that Gemini enabled multi-tenant and elastic GPU allocation with less than 5% performance overhead. Furthermore, compared to static scheduling, Gemini achieved 20%$\sim$30% performance improvement without requiring prior knowledge of applications.
Hung-Hsin Chen, En-Te Lin, Yu-Min Chou, Jerry Chou 0001
IEEE Trans. Cloud Comput.1
2021 Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From National Tsing Hua University
abstract
As a special activity of the Student Cluster Competition at SC19 conference, we made an attempt to reproduce the scalability evaluations of a highly paralleled polynomial filtering eigensolver for computing planetary interior normal modes. Our experiments were conducted on a Mars dataset using a small scale 4-node cluster with Intel Skylake CPU architecture, while the original article's were conducted on a Moon dataset using a large scale 256-node supercomputer with Intel CPU Skylake and KNL architectures. This article shares our experiences and observations from our reproducibility activity and discusses our findings on three main sections: the weak scalability, the strong scalability, and the relationships between variables. The results of weak scalability and strong scalability were successfully reproduced. But due to the differences on the problem scale, input dataset, and system architecture, different behaviors regarding the polynomial degree were observed.
Wei-Fang Sun, Hung-Hsin Chen, ShaoFu Lin, YuanChing Lin, Jing-Wei Wu, En-Te Lin, Jerry Chou 0001
IEEE Trans. Parallel Distributed Syst.2
2020 KubeShare: A Framework to Manage GPUs as First-Class and Shared Resources in Container Cloud
abstract
Container has emerged as a new technology in clouds to replace virtual machines~(VM) for distributed applications deployment and operation. With the increasing number of new cloud-focused applications, such as deep learning and high performance applications, started to reply on the high computing throughput of GPUs, efficiently supporting GPU in container cloud becomes essential. While GPU virtualization has been extensively studied for VM, limited work has been done for containers. One of the key challenges is the lack of support for GPU sharing between multiple concurrent containers. This limitation leads to low resource utilization when a GPU device cannot be fully utilized by a single application due to the burstiness of GPU workload and the limited memory bandwidth. To overcome this issue, we designed and implemented KubeShare, which extends Kubernetes to enable GPU sharing with fine-grained allocation. KubeShare is the first solution for Kubernetes to make GPU device as a first class resources for scheduling and allocations. Using real deep learning workloads, we demonstrated KubeShare can significantly increase GPU utilization and overall system throughput around 2x with less than 10% performance overhead during container initialization and execution.
Ting-An Yeh, Hung-Hsin Chen, Jerry Chou 0001
HPDC2
2019 Student Cluster Competition 2018, team NTHU: Reproducing performance of multi-physics simulations of the tsunamigenic 2004 sumatra megathrust earthquake on the Intel Skylake architecture
ShaoFu Lin, ChiChen Yang, Scott Cheng, KengJui Hsu, Hung-Hsin Chen, YuanChing Lin, Jerry Chou 0001
Parallel Comput.5
2018 Trading Decision of Taiwan Stocks with the Help of United States Stock Market
abstract
The paper studies how to improve the trading decision of Taiwan stocks with the information of US stock market. Our method first aligns the trading days between Taiwan and US stock markets. Next, the similarity between the portfolio index (PI, constructed from 100 Taiwan stocks) and one of the US stock indices, the Dow Jones Industrial Average (DJIA), NASDAQ composite index (NASDAQ), or Standard & Poor’s 500 (S&P 500), is computed, respectively. The trading signals of PI or each US stock index are generated by the method of Lee et al. Finally, the consensus signals of PI are determined by the majority vote scheme with the weighted functions, calculated from the similarity. The testing period of PI starts from 2000/1/4 to 2017/12/29, totally 4480 days. As the experimental results show, the index combination (PI, DJIA, NASDAQ) with the weighted function W(4) is considered to be the best combination for trading PI. Its average annualized return (cumulative return) achieves 15.03% (1170.42%), which is better than the method of Lee et al. 13.88% (947.65%), and the buy-and-hold strategy 9.85% (442.90%).
Shih-Chan Huang, Chang-Biau Yang, Hung-Hsin Chen
KES3
2014 Taiwan Stock Investment with Gene Expression Programming
abstract
Abstract In this paper, we first find out some good trading strategies from the historical series and apply them in the future. The profitable strategies are trained out by the gene expression programming (GEP), which involves some well-known stock technical indicators as features. Our data set collects the 100 stocks with the top capital from the listed companies in the Taiwan stock market. Accordingly, we build a new series called portfolio index as the investment target. For each trading day, we search for some similar template intervals from the historical data and pick out the pertained trading strategies from the strategy pool. These strategies are validated by the return during a few days before the trading day to check whether each of them is suitable or not. Then these suitable strategies decide the buying or selling consensus signal with the majority vote on the trading day. The training period is from 1996/1/6 to 2012/12/28, and the testing period is from 2000/1/4 to 2012/12/28. Two simulation experiments are performed. In experiment 1, the best average accumulated return is 548.97% (average annualized return is 15.47%). In experiment 2, we increase the diversity of trading strategies with more training. The best average accumulated return is increased to 685.31% (average annualized return is 17.18%). These two results are much better than that of the buy-and-hold strategy, whose return is 287.00%.
Cheng-Han Lee, Chang-Biau Yang, Hung-Hsin Chen
KES3
2013 Fault Models and Test Methods for Subthreshold SRAMs
abstract
Due to the increasing demand of an extra-low-power system, a great amount of research effort has been spent in the past to develop an effective and economic subthreshold SRAM design. However, the test methods regarding those newly developed subthreshold SRAM designs have not yet been fully discussed. In this paper, we first categorize the subthreshold SRAM designs into three types, study the faulty behavior of open defects and address decoders faults on each type of designs, and then identify the faults which may not be covered by a traditional SRAM test method. We will also discuss the impact of open defects and threshold-voltage mismatch on sense amplifiers under subthreshold operations. A discussion about the temperature at test is also provided.
Chen-Wei Lin, Hung-Hsin Chen, Hao-Yu Yang, Chin-Yuan Huang, Mango Chia-Tso Chao, Rei-Fu Huang
IEEE Trans. Computers2
2012 Testing strategies for a 9T sub-threshold SRAM
abstract
Due to the increasing demands of lower-power devices, a lot of research effort has been devoted to develop new SRAM cell designs that can be effectively and economically operated at the subthreshold region. However, each new SRAM cell design has its own cell structure and design techniques, which may result in different faulty behaviors than the conventional 6T SRAMs and require specialized test methods to detect those uncovered fault models. In this paper, we focus on developing the test methods for testing a new 9T subthreshold SRAM design, which utilizes single bit-line read/write, two write word-lines for writing different values, and a separate read path. A mixed march algorithm with different background and address-traverse directions is proposed to detect various uncovered fault models and validated through real test chips. A new specialized technique of floating bit-line attacking is also presented to detect the stability faults, which cannot be effectively detected by applying the conventional test methods, for the new 9T SRAM design.
Hao-Yu Yang, Chen-Wei Lin, Hung-Hsin Chen, Mango Chia-Tso Chao, Ming-Hsien Tu, Shyh-Jye Jou, Ching-Te Chuang
ITC3
2011 Detecting stability faults in sub-threshold SRAMs
abstract
Detecting stability faults has been a crucial task and a hot research topic for the testing of conventional super-threshold 6T SRAM in the past. When lowering the supply voltage of SRAM to the subthreshold region, the impact of stability faults may significantly change, and hence the test methods developed in the past for detecting stability faults may no longer be effective. In this paper, we first categorize the subthreshold-SRAM designs into different types according to their bit-cell structures. Based on each type, we then analyze the difference of its stability faults compared to the conventional super-threshold 6T SRAM, and discuss how the stability-fault test methods should be modified accordingly. A series of experiments are conducted to validate the effectiveness of each stability-fault test method for different types of subthreshold-SRAM designs.
Chen-Wei Lin, Hao-Yu Yang, Chin-Yuan Huang, Hung-Hsin Chen, Mango Chia-Tso Chao
ICCAD4
2010 Fault models and test methods for subthreshold SRAMs
abstract
Due to the increasing demand of an extra-low-power system, a great amount of research effort has been spent in the past to develop an effective and economic subthreshold-SRAM design. However, the test methods regarding those newly developed subthreshold-SRAM designs have not yet been fully discussed. In this paper, we first categorize the subthreshold-SRAM designs into three types, study the faulty behavior of different open defects for each type of designs, and then identify the faults which may or may not be covered by a traditional SRAM test method. For those hard-to-detect faults, we will further discuss the corresponding test method according to different each type of subthreshold-SRAM designs. At last, a discussion about the temperature at test will also be provided.
Chen-Wei Lin, Hung-Hsin Chen, Hao-Yu Yang, Mango Chia-Tso Chao, Rei-Fu Huang
ITC2