Ming-Hung Chen

dblp:66/4912 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 5 since 2021Computer networks · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Vela: A Virtualized LLM Training System with GPU Direct RoCE
abstract
Vela is a cloud-native system designed for LLM training workloads built using off-the-shelf hardware, Linux KVM-based virtualization, and a virtualized RDMA over Converged Ethernet (RoCE) network. Vela virtual machines (VMs) support peer-to-peer DMA between the GPUs and SRIOV-based network interface. In this paper, we share Vela's key architectural aspects with details from an NVIDIA A100 GPU-based deployment in one of the IBM Cloud data centers. Throughout the paper, we share insights and experiences from designing, building, and operating the system over a ~2.5 year timeframe to highlight the capabilities of readily available software and hardware technologies and the improvement opportunities for future AI systems, thereby making AI infrastructure more accessible to a broader community. As we evaluated the system for performance at ~1500 GPU scale, we achieved ~80% of the ideal throughput while training a 50 billion parameter decoder model using model parallelism, and ~70% per GPU FLOPS compared to a single VM with the High-Performance Linpack benchmark.
Apoorve Mohan, Robert Walkup, Bengi Karaçali, Ming-Hung Chen, Abdullah Kayi, Liran Schour, Shweta Salaria, Sophia Wen, I-Hsin Chung, Abdul Alim, Constantinos Evangelinos, Lixiang Luo, Marc Dombrowa, Laurent Schares, Ali Sydney, Pavlos Maniotis, Sandhya Koteshwara, Brent Tang, Joel Belog, Rei Odaira, Vasily Tarasov, Eran Gampel, Drew Thorstensen, Talia Gershon, Seetharami Seelam
ASPLOS (2)4
2025 Fast Malicious Packets Inspection Framework Using Converged Accelerator
Chuan-Ming Ou, Yong-Xuan Huang, Ming-Hung Chen, I-Hsin Chung, Jerry Chou 0001
HPC Asia3
2025 PCIe Bandwidth-Aware Scheduling for Multi-Instance GPUs
Yan-Mei Tang, Wei-Fang Sun, Hsu-Tzu Ting, Ming-Hung Chen, I-Hsin Chung, Jerry Chou 0001
HPC Asia4
2023 Chic-sched: a HPC Placement-Group Scheduler on Hierarchical Topologies with Constraints
abstract
Efficient placement of advanced HPC and AI workloads with application constraints is raising challenges for resource schedulers on shared infrastructures, such as the Cloud. In this work, we propose a novel Constraints- and Heuristics-based scheduler on HIerarchical Topologies for High-Performance Computing workloads in the Cloud (chic-sched, for short). Our heuristics-based algorithm enables placement across multiple levels in a network hierarchy with loosely specified constraints, and it works without retries by providing suboptimal placements to minimize placement failures. This allows for fast scheduling at scale, and the O(N log N) complexity enables placement decisions within tens of milliseconds for groups of hundreds of virtual machines (VM). We introduce a new and simple metric to quantify the goodness of group placements. With this metric, in terms of deviation from ideal placements, we show that chic-sched is 20-50% better than the common bestFit or worstFit algorithms in all scenarios of two-level placements with spreading and packing constraints. We evaluate chic-sched with publicly available VM-request traces from a production Cloud, and, comparing against bestFit, we show that it achieves 8% lower placement failure rates and more than 40% better placement locality. Finally, to quantify the goodness of constraints-based placements, we conduct experiments with a realistic MPI workload on synthetically allocated VM clusters in a public cloud. We measure a 9% performance improvement over an adverse placement in a scenario where our heuristics-based scheduler would return a good, but not perfect, placement.
Laurent Schares, Asser N. Tantawi, Pavlos Maniotis, Ming-Hung Chen, Claudia Misale, Seetharami Seelam, Hao Yu 0008
IPDPS4
2022 Enabling Software-Defined RF Convergence with a Novel Coarse-Scale Heterogeneous Processor
abstract
RF system development is traditionally constrained by a restrictive trade-off between power efficiency and programmatic flexibility. We outline a path towards achieving both, thereby enabling a range of new system concepts that better utilize limited resources. As an example, for many future applications, we consider RF convergence – reusing the same spectrum and waveforms to achieve multiple distributed system functions and goals, simultaneously. To enable this next step in processing, we develop a novel framework that includes both software and the system-on-chip (SoC) design.
Daniel W. Bliss, Tutu Ajayi, Ali Akoglu, Ilkin Aliyev, Toygun Basaklar, Leul Belayneh, David T. Blaauw, John S. Brunhaver, Chaitali Chakrabarti, Liangliang Chang, Kuan-Yu Chen 0001, Ming-Hung Chen, Xing Chen 0004, Alex R. Chiriyath, Alhad Daftardar, Ronald G. Dreslinski, Arindam Dutta, Allen-Jasmin Farcas, Yukang Fu, A. Alper Goksoy, Xin He 0011, Md Sahil Hassan, Andrew Herschfelt, Jacob Holtom, Hun-Seok Kim, Anish Krishnakumar, Owen Ma, Joshua Mack, Saurav Mallik, Sumit K. Mandal, Radu Marculescu, Brittany M. McCall, Trevor N. Mudge, Ümit Y. Ogras, Vishrut Pandey, Saquib Ahmad Siddiqui, Yu-Hsiu Sun, Adarsh A. Venkataramani, Xiangdong Wei, Benjamin R. Willis, Hanguang Yu, Yufan Yue
ISCAS12
2022 NVMe Virtualization for Cloud Virtual Machines
abstract
Public clouds are rapidly moving to support Non-Volatile Memory Express (NVMe) based storage to meet the ever-increasing I/O throughput and latency demands of modern workloads. They provide NVMe storage through virtual machines (VMs) where multiple VMs running on a host may share a physical NVMe device. The virtualization method used to share the NVMe capability has important performance, usability and security implications. In this paper, we propose three NVMe storage virtualization methods: PCI device passthrough, virtual block device method, and Storage Performance Development Kit (SPDK) virtual host target method. We evaluate these virtualization methods in terms of performance, scalability, CPU overhead, technology maturity, security, and availability to use one or more of these methods in IBM public cloud.
Lixiang Luo, I-Hsin Chung, Seetharami R. Seelam, Ming-Hung Chen, Yun Joon Soh
ICPE4
2021 Performance-Driven Simultaneous Partitioning and Routing for Multi-FPGA Systems
abstract
A multi-FPGA system consists of multiple FPGAs connected by physical wires, and a circuit is partitioned to fit each FPGA and routed on the system by such physical wires. Due to the limited numbers of input/output (I/O) pins in an FPGA, however, not all signals can be transmitted between FPGAs directly. Moreover, the routing resource may not be sufficient to accommodate many cross-FPGA signals from circuit partitioning. As a result, input/output time-division multiplexing (TDM) is introduced to send a group of cross-FPGA signals in a routing channel with a timing penalty. To optimize the performance of such a system, we shall develop a simultaneous partitioning and routing algorithm considering the timing penalty caused by I/O TDM. Considering the TDM delay penalty, we propose a simultaneous partitioning and routing algorithm to remedy the insufficiency of the two-stage flow of partitioning followed by routing. Our algorithm consists of two major steps: (1) a novel routing-aware partitioning framework to obtain an initial solution considering irregular, asymmetric connections, and (2) a partition-aware routing scheme to optimize routing in each partitioning pass. Experimental results show that our proposed algorithm can achieve better timing than the classical flow.
Ming-Hung Chen, Yao-Wen Chang
DAC1
2021 Time-Division Multiplexing Based System-Level FPGA Routing
abstract
Multi-FPGA system prototyping has become popular for modern VLSI logic verification, but such a system realization is often limited by its number of inter-FPGA connections. As a result, time-division multiplexing (TDM) is employed to accommodate more inter-FPGA signals than the connections in a multi-FPGA system. However, the inter-FPGA signal delay induced by TDM becomes significant due to time-multiplexing. Researchers have shown that TDM ratios (signal time-multiplexing ratios) significantly affect the performance of a multi-FPGA system and inter-FPGA routing highly influences the quality of this system. This paper presents a framework to minimize the system clock period for a system-level FPGA while considering the inter-FPGA routing topology and the timing criticality of nets. Our framework consists of two stages: (1) a distributed profiling scheme to generate the desired net-ordering and then alleviate the routing congestion, and (2) a net-/edge-based refinement to assign TDM ratios efficiently with a strict decrease in the ratios. Based on the 2019 CAD contest at ICCAD benchmarks and the contest evaluation metric with both quality and efficiency, experimental results show that our framework achieves the best overall score among all the participating teams and published works.
Wei-Kai Liu, Ming-Hung Chen, Chen-Chia Chang, Yao-Wen Chang
ICCAD2
2019 Banded Pair-HMM Algorithm for DNA Variant Calling and Its Hardware Accelerator Design
abstract
In this paper, we propose a new pair hidden Markov model (Pair-HMM) algorithm, namely Banded Pair-HMM, which is a heuristic approach for variant calling applications. When compared to the conventional Pair-HMM, our Banded Pair-HMM can reduce the execution time at a minor cost in accuracy. In addition, a hardware accelerator is implemented using TSMC 40nm technology based on the proposed algorithm. As demonstrated later in the paper, the proposed hardware accelerator runs 4× faster than the conventional Pair-HMM hardware, and over 17,000× faster than the original Pair-HMM software.
Ming-Hung Chen, Mao-Jan Lin, Yi-Chang Lu
BIBE1
2019 Towards VR/AR Multimedia Content Multicast over Wireless LAN
abstract
Virtual reality (VR) and augmented reality (AR) are expected to change people's daily life in the future, but it still faces many challenges. New advanced coding techniques like depth image based rendering (DIBR) mitigate the need to complete content by synthesizing the scene with nearby views. However, due to the limitation of wireless spectrum, delivering VR/AR content to many people doing some activities (e.g., watching sport in a stadium or seeing an exhibition in a museum) in mobile hotspot areas is still very challenging in terms of available bandwidth. In this work, we analyze how to multicast multi-view VR/AR 3D multimedia streams over multiple wireless access points (APs) to multiple users, and then propose a multi-view allocation (MVA) algorithm to improve the delivered content quality. The simulation results show that our MVA algorithm can achieve about 90% performance of the optimal solution and outperform non-DIBR and SSA model, by 48% and 25%, respectively.
Ming-Hung Chen, Kai-Wen Hu, I-Hsin Chung, Cheng-Fu Chou
CCNC1
2019 Temporal-based Load Adaptive SDN Controller Failover Mechanism
abstract
Recently distributed multiple SDN controller architecture is proposed to improve the scalability of SDN architecture. Unfortunately, the state-of-art SDN controller failover mechanism, i.e., switch leader competition, did not consider the characteristics of data center network and the time cost, resource usage, and load balancing after failure. In this paper, we propose a temporal-based load adaptive SDN controller failover mechanism (TLACF) based on the temporal pattern of traffic. The results demonstrate the proposed TLACF can save up to 46% time cost of failover and 66% memory usage for state backup.
Ming-Hung Chen, Zhi-Qiang Zhong, I-Hsin Chung, Cheng-Fu Chou
CCNC1
2018 FlexProtect: A SDN-based DDoS Attack Protection Architecture for Multi-tenant Data Centers
abstract
With the recent advances in software-defined networking (SDN), the multi-tenant data centers provide more efficient and flexible cloud platform to their subscribers. However, as the number, scale, and diversity of distributed denial-of-service (DDoS) attack is dramatically escalated in recent years, the availability of those platforms is still under risk. We note that the state-of-art DDoS protection architectures did not fully utilize the potential of SDN and network function virtualization (NFV) to mitigate the impact of attack traffic on data center network. Therefore, in this paper, we exploit the flexibility of SDN and NFV to propose FlexProtect, a flexible distributed DDoS protection architecture for multi-tenant data centers. In FlexProtect, the detection virtual network functions (VNFs) are placed near the service provider and the defense VNFs are placed near the edge routers for effectively detection and avoid internal bandwidth consumption, respectively. Based on the architecture, we then propose FP-SYN, an anti-spoofing SYN flood protection mechanism. The emulation and simulation results with real-world data demonstrates that, compared with the traditional approach, the proposed architecture can significantly reduce 46% of the additional routing path and save 60% internal bandwidth consumption. Moreover, the proposed detection mechanism for anti-spoofing can achieve 98% accuracy.
Ming-Hung Chen, Jyun-Yan Ciou, I-Hsin Chung, Cheng-Fu Chou
HPC Asia1
2018 Towards a Single-Host Many-GPU System
abstract
As computation-intensive tasks such as deep learning and big data analysis take advantage of GPU based accelerators, the interconnection links may become a bottleneck. In this paper, we investigate the upcoming performance bottleneck of multi-accelerator systems, as the number of accelerators equipped with single host grows. We instrumented the host PCIe fabric to measure the data transfer and compared it with the measurements from the software tool. It shows how the data transfer (P2P) helps to avoid the bottleneck on the interconnection links, but multi-GPU performance does not scale up as expected due to the control messages. We quantify the impact of host control messages with suggestions to remedy scalability bottlenecks. We also implement the proposed strategy on Lulesh to validate the concept. The result shows our strategy can save 59.86% time cost of the kernel and 13.32% PCIe H2D payload.
Ming-Hung Chen, I-Hsin Chung, Bülent Abali, Paul Crumley
SBAC-PAD1
2017 ETMP-BGP: Effective tunnel-based multi-path BGP routing using software-defined networking
abstract
Border Gateway Protocol (BGP) has been the de facto inter-domain routing protocol since it was introduced. Its destination-based routing nature, which is not able to choose a specific end-to-end AS-level route, might overload some popular peering (or inter-AS) links, while making some others underused. This could result in congestion and cause service disruption. To cope with the issue of destination-based routing of BGP, we resort to Software-Defined Networking (SDN) techniques to propose Effective Tunnel-based Multi-path BGP (ETMP-BGP), which is able to obtain the whole control of tunnel-based multi-path BGP routing in terms of AS-level routing. In ETMP-BGP, destinations provide the feedback of path quality to the source. This feedback is used to detect congestions on AS-level routes and adjust such routes accordingly in a per-source basis. That is, ETMP-BGP with SDN is able to get the AS-level global view of Internet and choose a better AS-level end-to-end BGP routing. Through simulations, our results show that ETMP-BGP could mitigate the AS-level congestions and outperforms existing BGP schemes.
Jose Luis Garcia Gomez, Ming-Hung Chen, Cheng-Fu Chou
SMC3
2011 Striking the Balance between Content Diversity and Content Importance in Swarm-Based P2P Streaming System
abstract
During recent years, the success of live swarm-based P2P streaming system has been witnessed. Nevertheless, how to design an effective mechanism in mitigating video quality degradation in a lossy network environment is still not thoroughly resolved yet. Unlike conventional client-server paradigm, there is data availability problem in swarm-based P2P streaming system. If we directly conduct the importance-first scheduling strategy (i.e. let the chunks with most distortion-rate efficiency get scheduled first) in swarm-based P2P streaming system, the serious content bottleneck for low priority chunks occurs, particularly when the population size is large. In this paper, we first identify the unique problem and understand the rationale behind. After that, a dynamic strategy-switching approach that combines the advantages of random and importance-first scheduling strategy is proposed. Simulation results indicate that compared with existing approaches our approach not only provides better scheduling efficiency, but also is scalable even if population size is large.
Chun-Yuan Chang, Cheng-Fu Chou, Ming-Hung Chen
HPCC3
2011 Striking the balance between content diversity and content importance in swarm-based P2P streaming
abstract
During recent years, the success of live swarm-based P2P streaming system has been witnessed. Nevertheless, how to design an effective mechanism for mitigating video quality degradation in a lossy network environment is still not thoroughly resolved yet. Unlike conventional client-server paradigm, there is data availability problem in swarm-based P2P streaming system. If we directly conduct the importance-first scheduling strategy in swarm-based P2P streaming system, the serious content bottleneck for low priority chunks occurs, particularly when the population size is large. In this work, we propose a dynamic strategy-switching approach that combines the advantages of random scheduling and importance first scheduling to deal with the problem. Simulation results indicate that our approach not only provides better scheduling efficiency, but also is scalable even if population size is large.
Chun-Yun Chang, Cheng-Fu Chou, Ming-Hung Chen
ICNP3
2011 Towards energy-efficient streaming system for mobile hotspots
abstract
Modern mobile devices have become an important part of our daily life but the performance of multimedia applications still suffers from the constrained energy supply and communication bandwidth of the mobile devices. In this work, we develop an energy-efficient streaming system for mobile hotspots to achieve better Quality-of-Experience. Our main idea is (a) to avoid redundant 3G transmissions as well as reduce the usage of 3G links for those low residual-energy users, and (b) to enable nearby mobile users cooperatively to share the downloaded data via short-range interfaces. The experiment results shows our scheme can improve the system lifetime by 27%, and provide better throughput as well as lower loss rate than conversional 3G systems do.
Ming-Hung Chen, Chun-Yuan Chang, Ming-Yuan Hsu, Ke-Han Lee, Cheng-Fu Chou
SIGCOMM1
2011 Towards quality-oriented scheduling for live swarm-based P2P streaming
abstract
During recent years, the success of live swarm-based P2P streaming system has been witnessed. Nevertheless, how to design an effective mechanism for mitigating video quality degradation in a lossy network environment is still not thoroughly resolved yet. Unlike conventional client-server paradigm, there is data availability problem in swarm-based P2P streaming system. If we directly conduct the importance-first scheduling strategy (i.e. let the chunks with most distortion-rate efficiency get scheduled first) in swarm-based P2P streaming system, the serious content bottleneck for low priority chunks occurs, particularly when the population size is large. To cope with the above issue, in this work, we propose a dynamic strategy-switching approach that combines the advantages of random scheduling and importance-first scheduling. Our simulation results indicate that compared with existing approaches our approach not only provides better scheduling efficiency, but also is scalable even if population size is large.
Chun-Yuan Chang, Cheng-Fu Chou, Ming-Hung Chen
VCIP3
2010 On the Design of the Semantic P2P System for Music Recommendation
abstract
The pervasive use of Peer-to-Peer (P2P) systems and the growing demand on personalization for consumers has made future business focus on the niche market instead of the mass market. The recommender system, which is able to timely select data of the interest to each individual user, has become the key to any successful business. However, currently most recommendation systems are based on a centralized architecture; nonetheless, they are not suitable for P2P environments. In this article, we propose a distributed semantic P2P overlay, which can provide music search and recommendation services by considering both of user preference and diversity of interests. We then propose a 3C hybrid recommendation procedure that adapts three traditional filtering techniques to fitting the requirements in a distributed semantic overlay. First, we choose a set of proper meta-data to represent a music object and use them to construct the characteristic-vector-based content filter. Second, a dominant attribute, which is one of the attributes in the characteristic vector of a music object, is used to build the profile of a peer. With the idea of the social network, a P2P profile-based collaborative filter is proposed. Finally, we explore the item-to-item relationship to construct a history-based cooperative filter. We use simulations and a real database called Audio Scrobbler, which tracks users' listening habits, to evaluate the performance of the recommendation system. The results demonstrate the effectiveness of the proposed approach compared with existing recommendation systems.
Ming-Hung Chen, Kate Ching-Ju Lin, Chui-Chiu Kung, Cheng-Fu Chou, Chang-Jen Tu
ISPA1
2009 Improved min-cost flow scheduler for mesh-based P2P streaming system
abstract
Perceptual visual quality guarantee is of vital concern for the P2P streaming system. One of challenging issues for assuring overall video quality is to design an effective media chunk scheduler. To the best of our knowledge, there is still a lack of studies on integrating video characteristics into P2P media chunk scheduler. In this work, we attempt to integrate video characteristics, i.e. rate-distortion impact, into P2P media chunk scheduler. No more "rarest-first" strategy, "rate-distortion first" strategy (RD-first) is deployed so that the chunks with larger rate-distortion score can be spread more effectively. Predictably, such a strategy can be better against video quality degradation by losing some less important chunks under a bandwidth constrained network. Moreover, to adapt for dynamic networks, we propose a hybrid available bandwidth prober. The simulation results show that received video quality can be substantially improved up to 1.7 dB compared with existing P2P schedulers.
Chun-Yuan Chang, Ti-Chia Chiu, You-Ming Chen, Ming-Hung Chen, Cheng-Fu Chou
ICME4
2008 A q-Domain Characteristic-Based Bit-Rate Model for Video Transmission
abstract
For low-delay video transmission, we introduce aq-domain characteristic-based bit-rate model. Specifically, three characteristics are efficiently extracted from the quantized DCT spectra to construct the bit-rate model. Extensive experimental results show that our rate model can provide more accuracy with lower complexity than existing models.
Chun-Yuan Chang, Cheng-Fu Chou, Din-Yuen Chan, Tsungnan Lin, Ming-Hung Chen
IEEE Trans. Circuits Syst. Video Technol.5
2007 A Two-Layer Characteristic-based Rate Control Framework for Low Delay Video Transmission
abstract
In this paper, we present a two-layer rate control framework for low delay video transmission based on a characteristic-based rate-quantization (R-Q) model. Specifically, with the frame-layer characteristic-based R-Q model, we are able to figure out the most appropriate quantization parameter (QP) to maintain targeted end-to-end delay by jointly considering the current buffer status, network bandwidth and the variation of bit rate of the input video source. Next, in macroblock-layer (MB), a greed-based rate-distortion (R-D) refinement of MBs approach is proposed to efficiently utilize the remaining available bit rate in order to enhance visual quality as much as possible. The experiments show that our two-layer characteristic-based rate control framework can quickly smooth out the impact of varying bit rate of video source on the buffer, meet targeted buffer delay time well, and substantially enhance the quality of the perceived video with lower computational complexity.
Chun-Yuan Chang, Ming-Hung Chen, Cheng-Fu Chou, Din-Yuen Chan
ICC2