EDBT 2026 Demo / reviewers in the wild / expert
Sangyoon Oh 0001
dblp:65/912-1 · also Oh Sangyoon 0001
· DBLP profile ↗
48ranked-venue papers
5as first author
26since 2021 · last 2026
0000-0001-5854-149XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 3 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 1 first-authorComputer networks · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reducing Backfill Failures From Workload Drift with Lightweight Uncertainty Buffers in HPC Job Scheduling
Jiheon Choi, Sangyoon Oh 0001 |
CCGrid | 2 |
| 2026 | GASched: Goal-Adaptive Hierarchical Reinforcement Learning for Multi-Objective HPC Job Scheduling
Minsol Choo, Sangyoon Oh 0001 |
CCGrid | 2 |
| 2026 | S-CQR: Stratified Calibration for Runtime Prediction in HPC Backfill Scheduling
Jiheon Choi, Sangyoon Oh 0001 |
Euro-Par (2) | 2 |
| 2026 | UARP: uncertainty-aware runtime prediction for preventing scheduler termination under Wallclock constraints in HPCabstractEffective resource allocation has become a critical issue in high-performance computing (HPC) systems. To effectively allocate resources (e.g., CPU/GPU cores), recent studies focus on predicting each workload’s runtime using machine learning and deep learning models. These methods in HPC often suffer from underestimation, as 33–64% of jobs terminate due to wallclock time limits, whereas user-provided estimates achieve 78–99% success. This failure stems from minimizing mean squared error, which biases predictions toward average-case performance and underestimates jobs in high-skewed (i.e., long-tail) runtime distributions. Specifically, HPC workloads exhibit long-tail runtime distributions, with most jobs completing quickly while a small fraction runs for extremely long durations. To overcome these challenges, we introduce an uncertainty-aware runtime prediction (UARP) method based on multi-quantile regression. Our method directly addresses the underestimation problem by quantifying uncertainty by modeling the conditional distribution without distributional assumptions. Our approach uses the highest-quantile (99th) model and the residual model from the median quantile. The 99th model primarily provides conservative bounds and protection against job underestimates, while the residual uncertainty model protects against unpredictable workloads by estimating prediction variance. In particular, the expected predicted error (i.e., uncertainty) from the residual model plays a critical role in our adaptive safety margin calculation. UARP adds a conservative prediction (99th quantile) and an additional safety margin from our formula, enabling our adaptive margin approach specifically tailored to each job’s characteristics. Evaluation on four production HPC systems (SDSC DataStar, KIT FH2, ANL Interpid, KISTI NURION), UARP achieves 92–99% job success rates while maintaining resource utilization within 1–2% of EASY backfilling. Our method deploys with identical parameters across all systems. This parameter-free deployment eliminates the per-system tuning that fixed-margin approaches require. In addition, our approach integrates with existing schedulers through minimal modification, utilizing uncertainty-aware predictions to prevent timeout-based job termination and preserve system efficiency. Jiheon Choi, Sangyoon Oh 0001 |
J. Supercomput. | 2 |
| 2025 | When HPC Scheduling Meets Active Learning: Maximizing The Performance with Minimal Data
Jiheon Choi, Minsol Choo, Oh-Kyoung Kwon, Sangyoon Oh 0001 |
HPC Asia | 6 |
| 2025 | Lightweight multi-layered de-identification architecture: Secure client selection in federated learning
Jiheon Choi, Sangyoon Oh 0001 |
J. Syst. Archit. | 2 |
| 2024 | Preserving Near-Optimal Gradient Sparsification Cost for Scalable Distributed Deep LearningabstractCommunication overhead is a major obstacle to scaling distributed training systems. Gradient sparsification is a potential optimization approach to reduce the communication volume without significant loss of model fidelity. However, existing gradient sparsification methods have low scalability owing to inefficient design of their algorithms, which raises the communication overhead significantly. In particular, gradient build-up and inadequate sparsity control methods degrade the sparsification performance considerably. Moreover, communication traffic increases drastically owing to workload imbalance of gradient selection between workers.To address these challenges, we propose a novel gradient sparsification scheme called ExDyna. In ExDyna, the gradient tensor of the model comprises fined-grained blocks, and contiguous blocks are grouped into non-overlapping partitions. Each worker selects gradients in its exclusively allocated partition so that gradient build-up never occurs. To balance the workload of gradient selection between workers, ExDyna adjusts the topology of partitions by comparing the workloads of adjacent partitions. In addition, ExDyna supports online threshold scaling, which estimates the accurate threshold of gradient selection on-the-fly. Accordingly, ExDyna can satisfy the user-required sparsity level during a training period regardless of models and datasets. Therefore, ExDyna can enhance the scalability of distributed training systems by preserving near-optimal gradient sparsification cost. In experiments, ExDyna outperformed state-of-the-art sparsifiers in terms of training speed and sparsification performance while achieving high accuracy. Daegun Yoon, Sangyoon Oh 0001 |
CCGrid | 2 |
| 2024 | Staleness aware semi-asynchronous federated learning
Miri Yu, Jiheon Choi, Sangyoon Oh 0001 |
J. Parallel Distributed Comput. | 4 |
| 2023 | MiCRO: Near-Zero Cost Gradient Sparsification for Scaling and Accelerating Distributed DNN TrainingabstractGradient sparsification is a communication optimisation technique for scaling and accelerating distributed deep neural network (DNN) training. It reduces the increasing communication traffic for gradient aggregation. However, existing sparsifiers have poor scalability because of the high computational cost of gradient selection and/or increase in communication traffic. In particular, an increase in communication traffic is caused by gradient build-up and inappropriate threshold for gradient selection. To address these challenges, we propose a novel gradient sparsification method called MiCRO. In MiCRO, the gradient vector is partitioned, and each partition is assigned to the corresponding worker. Each worker then selects gradients from its partition, and the aggregated gradients are free from gradient build-up. Moreover, MiCRO estimates the accurate threshold to maintain the communication traffic as per user requirement by minimising the compression ratio error. MiCRO enables near-zero cost gradient sparsification by solving existing problems that hinder the scalability and acceleration of distributed DNN training. In our extensive experiments, MiCRO outperformed state-of-the-art sparsifiers with an outstanding convergence rate. Daegun Yoon, Sangyoon Oh 0001 |
HiPC | 2 |
| 2023 | Addressing Client Heterogeneity in Synchronous Federated Learning: The CHAFL ApproachabstractFederated learning (FL) is proposed to address the security vulnerabilities of conventional distributed deep learning. Since the capabilities of participating FL clients are highly variable in terms of both statistical and system aspects, FL training will face diminished convergence accuracy and speed. Hence, we propose CHAFL (client heterogeneity aware federated learning) to address client heterogeneity (i.e., statistical and system heterogeneity) in synchronous FL. CHAFL selects clients based on global loss with contribution, defined as local loss, enabling higher round-to-accuracy for the global model than previous studies. Additionally, to handle system heterogeneity, it proposes a lightweight algorithm that eliminates the profiling process previously employed to calculate adaptive local epochs in existing studies, thereby improving time-to-accuracy. In order to verify the effectiveness of CHAFL, we exploit three benchmark datasets on non-IID and system heterogeneous setting for empirical evaluation. Compared to the baseline, CHAFL achieves an accuracy improvement of 3.1 to 5.7%, along with 1.04 to 1.73x higher round-to-accuracy and 1.03 to 1.98x higher time-to-accuracy. Miri Yu, Oh-Kyoung Kwon, Sangyoon Oh 0001 |
ICPADS | 3 |
| 2023 | DEFT: Exploiting Gradient Norm Difference between Model Layers for Scalable Gradient SparsificationabstractGradient sparsification is a widely adopted solution for reducing the excessive communication traffic in distributed deep learning. However, most existing gradient sparsifiers have relatively poor scalability because of considerable computational cost of gradient selection and/or increased communication traffic owing to gradient build-up. To address these challenges, we propose a novel gradient sparsification scheme, DEFT, that partitions the gradient selection task into sub tasks and distributes them to workers. DEFT differs from existing sparsifiers, wherein every worker selects gradients among all gradients. Consequently, the computational cost can be reduced as the number of workers increases. Moreover, gradient build-up can be eliminated because DEFT allows workers to select gradients in partitions that are non-intersecting (between workers). Therefore, even if the number of workers increases, the communication traffic can be maintained as per user requirement. Daegun Yoon, Sangyoon Oh 0001 |
ICPP | 2 |
| 2023 | Crossover-SGD: A gossip-based communication in distributed deep learning for alleviating large mini-batch problem and enhancing scalabilityabstractSummary Distributed deep learning is an effective way to reduce the training time for large datasets as well as complex models. However, the limited scalability caused by network‐overheads makes it difficult to synchronize the parameters of all workers and gossip‐based methods that demonstrate stable scalability regardless of the number of workers have been proposed. However, to use gossip‐based methods in general cases, the validation accuracy for a large mini‐batch needs to be verified. For this, we first empirically study the characteristics of gossip methods in a large mini‐batch problem and observe that gossip methods preserve higher validation accuracy than AllReduce‐SGD (stochastic gradient descent) when the number of batch sizes is increased, and the number of workers is fixed. However, the delayed parameter propagation of the gossip‐based models decreases validation accuracy in large node scales. To cope with this problem, we propose Crossover‐SGD that alleviates the delay propagation of weight parameters via segment‐wise communication and random network topology with fair peer selection. We also adapt hierarchical communication to limit the number of workers in gossip‐based communication methods. To validate the effectiveness of our method, we conduct empirical experiments and observe that our Crossover‐SGD shows higher node scalability than stochastic gradient push. Sangho Yeo, Minho Bae, Minjoong Jeong, Oh-Kyoung Kwon, Sangyoon Oh 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | WAVE: designing a heuristics-based three-way breadth-first search on GPUs
Daegun Yoon, Minjoong Jeong, Sangyoon Oh 0001 |
J. Supercomput. | 3 |
| 2023 | SAGE: toward on-the-fly gradient compression ratio scaling
Daegun Yoon, Minjoong Jeong, Sangyoon Oh 0001 |
J. Supercomput. | 3 |
| 2023 | User Preference-Based Hierarchical Offloading for Collaborative Cloud-Edge ComputingabstractCloud computing and mobile edge computing techniques supply efficient ways to solve the contradiction between the increasing computing and storage demands of portable terminals and the limited capacity. In this paper, we conduct a three-tier hierarchical service system with multiple UEs, multiple MECs, and a single cloud center. It's worth noting that multiple UEs with personalized options generate a large number of different tasks in real time. To deal with this offloading problem, a response ratio offloading strategy (RROS) centered on user preference and real-time nature is designed to make MECs or CC serve as many UEs as possible. Therefore, a MEC-choosing preference list of each UE is created based on its past experiences at first. Then, each MEC iteratively sorts UEs with its ranking in the UEs' preference list. In order to avoid that the first task arriving at MEC occupies too many resources of MEC and cannot achieve global optimization, we also adopt loop iterative sequencing for multiple tasks arriving within a stipulated time. Lastly, by comparing the optimal response ratio on different MECs and CC, multiple MECs and the CC collaborative offload computing tasks of multiple UEs. Experimental results show that the algorithm significantly outperforms conventional techniques. Shujuan Tian, Chi Chang, Saiqin Long, Sangyoon Oh 0001, Zhetao Li |
IEEE Trans. Serv. Comput. | 4 |
| 2022 | Is Ant Colony System better than FFD for VM placement in a heterogeneous cluster?abstractFirst fit decreasing (FFD) is the most popular heuristic for virtual machine (VM) placement problems. However, FFD does not perform as much in a heterogeneous cluster environment. Moreover, FFD and other heuristics, such as best fit decreasing (BFD), are limited to handle the VM placement problem effectively when multiple resources are considered together. In this study, we analyze the reason why the ant colony system performs better than FFD for VM placement in a heterogeneous cluster. We verified our logical observations through experimental comparisons with other heuristics. Minjoong Jeong, Sangyoon Oh 0001 |
IC2E | 3 |
| 2022 | AR-CNN: an attention ranking network for learning urban perception
Zhetao Li, Wei-Shi Zheng 0001, Sangyoon Oh 0001, Kien Nguyen 0002 |
Sci. China Inf. Sci. | 4 |
| 2022 | AMBLE: Adjusting mini-batch and local epoch for federated learning with heterogeneous devices
Juwon Park, Daegun Yoon, Sangho Yeo, Sangyoon Oh 0001 |
J. Parallel Distributed Comput. | 4 |
| 2021 | A dynamic task offloading algorithm based on greedy matching in vehicle network
Shujuan Tian, Xianghong Deng, Tingrui Pei, Sangyoon Oh 0001, Weiping Xue |
Ad Hoc Networks | 5 |
| 2021 | Novel data-placement scheme for improving the data locality of Hadoop in heterogeneous environmentsabstractSummary To address the challenging needs of high‐performance big data processing, parallel‐distributed frameworks such as Hadoop are being utilized extensively. However, in heterogeneous environments, the performance of Hadoop clusters is below par. This is primarily because the blocks of the clusters are allocated equally to all nodes without regard to differences in the capability of individual nodes. This results in reduced data locality. Thus, a new data‐placement scheme that enhances data locality is required for Hadoop in heterogeneous environments. This article proposes a new data placement scheme that preserves the same degree of data locality in heterogeneous environments as that of the standard Hadoop, with only a small amount of replicated data. In the proposed scheme, only those blocks with the highest probability of being accessed remotely are selected and replicated. The results of experiments conducted indicate that the proposed scheme incurs only a 20% disk space overhead and has virtually the same data locality ratio as the standard Hadoop, which has a replication factor of three and 200% disk space overhead. Minho Bae, Sangho Yeo, Gyudong Park, Sangyoon Oh 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | Parallel Programming Models in High-Performance Cloud (ParaMo 2019)
Sangyoon Oh 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Exploring a system architecture of content-based publish/subscribe system for efficient on-the-fly data disseminationabstractSummary In a cloud‐scale publish/subscribe messaging system, it is difficult to partition subscription data among several servers. Without a sophisticated scheme and a system architecture, the messaging system would either waste resources or fail to deliver messages on time. In this study, we propose DRDA, a dynamic replication degree adjustment technology, for efficient message delivery. The technology calculates and maintains the number of subscription replications at a reasonable level by monitoring the statuses of servers, based on the number of subscription replications and the frequency of event dissemination. To verify the effectiveness of our proposed scheme and system architecture, we build a prototype of a content‐based publish/subscribe system that dynamically adjusts the number of replications among brokers. Furthermore, we compare the load balance, resource overhead, and performance of a publish/subscribe system with DRDA with a publish/subscribe system without DRDA. The experimental results show that DRDA outperforms other approaches under various parameter configurations. We have added the prototype code to a GitHub repository to make it publicly available. Daegun Yoon, Gyudong Park, Sangyoon Oh 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Balanced content space partitioning for pub/sub: a study on impact of varying partitioning granularity
Daegun Yoon, Zhetao Li, Sangyoon Oh 0001 |
J. Supercomput. | 3 |
| 2021 | A low redundancy and high time efficiency large-scale task assignment strategy for heterogeneous service-oriented cloud computing systems
Lizan Wang, Guoqi Xie, Tingrui Pei, Sangyoon Oh 0001, Zhetao Li |
J. Supercomput. | 5 |
| 2021 | Lightweight Single Image Super-resolution with Dense Connection Distillation NetworkabstractSingle image super-resolution attempts to reconstruct a high-resolution (HR) image from its corresponding low-resolution (LR) image, which has been a research hotspot in computer vision and image processing for decades. To improve the accuracy of super-resolution images, many works adopt very deep networks to model the translation from LR to HR, resulting in memory and computation consumption. In this article, we design a lightweight dense connection distillation network by combining the feature fusion units and dense connection distillation blocks (DCDB) that include selective cascading and dense distillation components. The dense connections are used between and within the distillation block, which can provide rich information for image reconstruction by fusing shallow and deep features. In each DCDB, the dense distillation module concatenates the remaining feature maps of all previous layers to extract useful information, the selected features are then assessed by the proposed layer contrast-aware channel attention mechanism, and finally the cascade module aggregates the features. The distillation mechanism helps to reduce training parameters and improve training efficiency, and the layer contrast-aware channel attention further improves the performance of model. The quality and quantity experimental results on several benchmark datasets show the proposed method performs better tradeoff in term of accuracy and efficiency. Yanchun Li, Jianglian Cao, Zhetao Li, Sangyoon Oh 0001, Nobuyoshi Komuro |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | Accelerated deep reinforcement learning with efficient demonstration utilization techniques
Sangho Yeo, Sangyoon Oh 0001, Minsu Lee 0001 |
World Wide Web | 2 |
| 2020 | An Effective Design to Improve the Efficiency of DPUs on FPGAabstractConvolutional neural networks (CNNs) have been widely used in various complicated problems, such as image classification, objection detection, semantic segmentation. To meet diversified CNN structures, the deep learning processing unit (DPU) is designed as a general accelerator on field programmable gate array (FPGA) to support various CNN layers, such as convolution, pooling, activation, etc. However, low DPU utilization and schedule efficiency appear when DPU used to multitask application completed by CNN models. In this paper, an effective design including multi-core with different size (MCDS) and DPU Plus is proposed to improve the efficiency of DPUs usage from the two dimensions of time and space. Through increasing the number of DPU cores on an FPGA and the utilization of single DPU core, the design of MCDS can effectively improve the overall throughput with restricted on-chip resources. Furthermore, the design of DPU Plus is proposed to improve the schedule efficiency of DPUs through simultaneously implementing DPU with other significant auxiliary modules of the application system on the same FPGA. Finally, a color space conversion module is implemented cooperate to the DPU cores to testify its performance, and the experimen shows that compared with running on the the CPU completely, it achieves16.2x acceleration, and increases the throughput of the entire system by 3.0x. Qingyong Deng, Saiqin Long, Shaohui Liu, Sangyoon Oh 0001 |
ICPADS | 5 |
| 2020 | A similarity clustering-based deduplication strategy in cloud storage systemsabstractDeduplication is a data redundancy elimination technique, designed to save system storage resources by reducing redundant data in cloud storage systems. With the development of cloud computing technology, deduplication has been increasingly applied to cloud data centers. However, traditional technologies face great challenges in big data deduplication to properly weigh the two conflicting goals of deduplication throughput and high duplicate elimination ratio. This paper proposes a similarity clustering-based deduplication strategy (named SCDS), which aims to delete more duplicate data without significantly increasing system overhead. The main idea of SCDS is to narrow the query range of fingerprint index by data partitioning and similarity clustering algorithms. In the data preprocessing stage, SCDS uses data partitioning algorithm to classify similar data together. In the data deletion stage, the similarity clustering algorithm is used to divide the similar data fingerprint superblock into the same cluster. Repetitive fingerprints are detected in the same cluster to speed up the retrieval of duplicate fingerprints. Experiments show that the deduplication ratio of SCDS is better than some existing similarity deduplication algorithms, but the overhead is only slightly higher than some high throughput but low deduplication ratio methods. Saiqin Long, Zhetao Li, Qingyong Deng, Sangyoon Oh 0001, Nobuyoshi Komuro |
ICPADS | 5 |
| 2020 | Modeling Analysis and Cost-Performance Ratio Optimization of Virtual Machine Scheduling in Cloud ComputingabstractAs an essential feature of cloud computing, dynamic scalability enables the cloud system to dynamically expand or shrink resources according to user needs at runtime. Effectively predicting and optimizing the cost and performance of cloud computing platforms have become one of the key research challenges in the field of cloud computing. In this article, to quantitatively predict the cost and performance of cloud computing platforms, we propose a cloud computing resource analysis model considering both hot/cold startup and hot/cold shutdown of virtual machines (VMs), and use the M/M/N/oo queuing model to analyze cloud computing platform and acquire accurate performance indicators, such as elasticity indicators, cost indicators, performance indicators, cost-performance ratios, etc. In addition, we establish a multi-objective optimization model to optimize both performance and cost of cloud computing platform. Then the optimal stopping and cost-performance optimization algorithm are applied to obtain the optimal configurations, including the number of hot startup VMs, the system service rate, the hot/cold startup rate of VMs, and the hot/cold shutdown rate. By comparing with existing optimization methods, we demonstrate the superiority of our cost-performance ratio optimization method. Jiale Dang, Zhetao Li, Hongfang Gong, Feng Zhang 0007, Sangyoon Oh 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2018 | Decentralized Message Broker Federation Architecture with Multiple DHT Rings for High Survivability
Minsub Kim, Minho Bae, Sangho Yeo, Gyudong Park, Sangyoon Oh 0001 |
ICCSA (5) | 5 |
| 2017 | High Performance Query Processing for Web Scale RDF Data using BSP Style Communication and Balanced DistributionabstractTo overcome scalability and performance issues for process queries over a web-scale RDF data, various studies have proposed RDF SPARQL query processing algorithm using parallel processing manners. However, it is hard to resolve the scalability and performance issues together because the problem of communication overhead between nodes is closely related to the data distribution for parallel processing. For efficient RDF query parallel processing, it is essential to distribute and process data evenly while reducing communication overhead. In this paper, we propose RDF query parallel processing algorithms with RDF data partitioning technique to guarantee evenly distributed data over the cluster. We also propose our in-memory RDF query processing system as a form of Bulk Synchronization Parallel system to reduce network overhead. Our empirical evaluation results show that the proposed system outperforms a popular RDF-3X on LUBM benchmark and UniProt queries from 2.20 to 43.08 times. Especially, the effectiveness of the system improves significantly when the SPARQL queries are complex with high input and select. Minho Bae, Junho Eum, Sangyoon Oh 0001 |
ICPP | 4 |
| 2016 | MGEScan: a Galaxy-based system for identifying retrotransposons in genomesabstractUNLABELLED: : MGEScan-long terminal repeat (LTR) and MGEScan-non-LTR are successfully used programs for identifying LTRs and non-LTR retrotransposons in eukaryotic genome sequences. However, these programs are not supported by easy-to-use interfaces nor well suited for data visualization in general data formats. Here, we present MGEScan, a user-friendly system that combines these two programs with a Galaxy workflow system accelerated with MPI and Python threading on compute clusters. MGEScan and Galaxy empower researchers to identify transposable elements in a graphical user interface with ready-to-use workflows. MGEScan also visualizes the custom annotation tracks for mobile genetic elements in public genome browsers. A maximum speed-up of 3.26× is attained for execution time using concurrent processing and MPI on four virtual cores. MGEScan provides four operational modes: as a command line tool, as a Galaxy Toolshed, on a Galaxy-based web server, and on a virtual cluster on the Amazon cloud. AVAILABILITY AND IMPLEMENTATION: MGEScan tutorials and source code are available at http://mgescan.readthedocs.org/ CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hyungro Lee, Minsu Lee 0001, Wazim Mohammed Ismail, Mina Rho, Geoffrey C. Fox, Sangyoon Oh 0001, Haixu Tang |
Bioinform. | 6 |
| 2016 | Energy-efficient multisite offloading policy using Markov decision process for mobile cloud computing
Mati B. Terefe, Heezin Lee, Nojung Heo, Geoffrey C. Fox, Sangyoon Oh 0001 |
Pervasive Mob. Comput. | 5 |
| 2015 | Evaluating ARM HPC clusters for scientific workloadsabstractSummary The power consumption of modern high‐performance computing (HPC) systems that are built using power hungry commodity servers is one of the major hurdles for achieving Exascale computation. Several efforts have been made by the HPC community to encourage the use of low‐powered system‐on‐chip (SoC) embedded processors in large‐scale HPC systems. These initiatives have successfully demonstrated the use of ARM SoCs in HPC systems, but there is still a need to analyze the viability of these systems for HPC platforms before a case can be made for Exascale computation. The major shortcomings of current ARM‐HPC evaluations include a lack of detailed insights about performance levels on distributed multicore systems and performance levels for benchmarking in large‐scale applications running on HPC. In this paper, we present a comprehensive evaluation of results that covers major aspects of server and HPC benchmarking for ARM‐based SoCs. For the experiments, we built an unconventional cluster of ARM Cortex‐A9s that is referred to as Weiser and ran single‐node benchmarks (STREAM, Sysbench, and PARSEC) and multi‐node scientific benchmarks (High‐performance Linpack (HPL), NASA Advanced Supercomputing (NAS) Parallel Benchmark, and Gadget‐2) in order to provide a baseline for performance limitations of the system. Based on the experimental results, we claim that the performance of ARM SoCs depends heavily on the memory bandwidth, network latency, application class, workload type, and support for compiler optimizations. During server‐based benchmarking, we observed that when performing memory intensive benchmarks for database transactions, x86 performed 12% better for multithreaded query processing. However, ARM performed four times better for performance to power ratios for a single core and 2.6 times better on four cores. We noticed that emulated double precision floating point in Java resulted in three to four times slower performance as compared with the performance in C for CPU‐bound benchmarks. Even though Intel x86 performed slightly better in computation‐oriented applications, ARM showed better scalability in I/O bound applications for shared memory benchmarks. We incorporated the support for ARM in the MPJ‐Express runtime and performed comparative analysis of two widely used message passing libraries. We obtained similar results for network bandwidth, large‐scale application scaling, floating‐point performance, and energy‐efficiency for clusters in message passing evaluations (NBP and Gadget 2 with MPJ‐Express and MPICH). Our findings can be used to evaluate the energy efficiency of ARM‐based clusters for server workloads and scientific workloads and to provide a guideline for building energy‐efficient HPC clusters. Copyright © 2015 John Wiley & Sons, Ltd. Maqbool Jahanzeb, Sangyoon Oh 0001, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | Ontology-based quantitative similarity metric for event matching in publish/subscribe system
Hongjae Kim, Sanggil Kang, Sangyoon Oh 0001 |
Neurocomputing | 3 |
| 2015 | An approach to mitigate DoS attack based on routing misbehavior in wireless ad hoc networks
Wonil Kim, Kangseok Kim, Sangyoon Oh 0001, Dong-Kyoo Kim |
Peer-to-Peer Netw. Appl. | 4 |
| 2014 | Semantic similarity method for keyword query system on RDF
Minho Bae, Sanggil Kang, Sangyoon Oh 0001 |
Neurocomputing | 3 |
| 2013 | Effective Hotspot Removal System Using Neural Network Predictor
Sangyoon Oh 0001, Mun-Young Kang, Sanggil Kang |
ACIIDS (2) | 1 |
| 2013 | K-depth RDF Keyword Search Algorithm Based on Structure IndexingabstractInformation retrieval from large scale RDF datasets is a challenging task. Because it takes much time to process query and it is hard to store a large collection of RDFs, it requires an efficient method to index and query. In this paper, we propose a novel RDF management system architecture along with indexing and querying algorithms. Our empirical experiments show our system performs substantially better than the conventional system. Also we verify the effectiveness of k-depth concept in our design with additional experiments. Minho Bae, Sanggil Kang, Sangyoon Oh 0001 |
KES-AMSTA | 4 |
| 2012 | An Intelligent RDF Management System with Hybrid Querying Approach
Jangsu Kihm, Minho Bae, Sanggil Kang, Sangyoon Oh 0001 |
ICCCI (1) | 4 |
| 2011 | Ensemble Learning with Active Example Selection for Imbalanced Biomedical Data ClassificationabstractIn biomedical data, the imbalanced data problem occurs frequently and causes poor prediction performance for minority classes. It is because the trained classifiers are mostly derived from the majority class. In this paper, we describe an ensemble learning method combined with active example selection to resolve the imbalanced data problem. Our method consists of three key components: 1) an active example selection algorithm to choose informative examples for training the classifier, 2) an ensemble learning method to combine variations of classifiers derived by active example selection, and 3) an incremental learning scheme to speed up the iterative training procedure for active example selection. We evaluate the method on six real-world imbalanced data sets in biomedical domains, showing that the proposed method outperforms both the random under sampling and the ensemble with under sampling methods. Compared to other approaches to solving the imbalanced data problem, our method excels by 0.03-0.15 points in AUC measure. Sangyoon Oh 0001, Minsu Lee 0001, Byoung-Tak Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2010 | Real-time performance analysis for publish/subscribe systems
Sangyoon Oh 0001, Jai-Hoon Kim, Geoffrey C. Fox |
Future Gener. Comput. Syst. | 1 |
| 2009 | Ensemble Learning Based on Active Example Selection for Solving Imbalanced Data Problem in Biomedical DataabstractThe imbalanced data problem is popular in biomedical classification tasks. Since trained classifiers using imbalanced data are mostly derived from the majority class, their prediction performance is poor for the minority class. In this paper, we propose a novel ensemble learning method based on an active example selection algorithm to resolve the imbalanced data problem. To compensate a possible sub-optimal classifier, our proposed ensemble learning methods aggregates classifiers built by the active example selection algorithm. We implement this ensemble learning method based on the active example selection algorithm using incremental naive Bayes classifiers. Our empirical results show that we greatly improve the performance of classification models trained by five real world imbalanced biomedical data. The proposed ensemble learning methods outperforms other approaches by 0.03~0.15 in terms of AUC which solve imbalanced data problem. Minsu Lee 0001, Sangyoon Oh 0001, Byoung-Tak Zhang |
BIBM | 2 |
| 2008 | XML Metadata ServicesabstractAbstract As service‐oriented architecture principles have gained importance, an emerging need has appeared for methodologies to locate desired services that provide access to their capability descriptions. These services must typically be assembled into short‐term service collections that, together with code execution services, are combined into a meta‐application to perform a particular task. To address the metadata requirements of these problems, we introduce a hybrid Information Service to manage both stateless and stateful (transient) metadata. We leverage the two widely used Web Service standards: Universal Description, Discovery and Integration (UDDI) and Web Services Context (WS‐Context), in our design. We describe our approach and experiences when designing ‘semantics’. We report the results from a prototype of the system that is applied to a mobile environment for optimizing Web Service communications. Copyright © 2007 John Wiley & Sons, Ltd. Mehmet S. Aktas, Geoffrey C. Fox, Marlon E. Pierce, Sangyoon Oh 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2007 | Optimizing Web Service messaging performance in mobile computing
Sangyoon Oh 0001, Geoffrey C. Fox |
Future Gener. Comput. Syst. | 1 |
| 2006 | Multi-module Image Classification System
Wonil Kim, Sangyoon Oh 0001, Sanggil Kang, Dongkyun Kim |
FQAS | 2 |
| 2003 | A Web Service Approach to Universal Accessibility in Collaboration Services
Sangmi Lee, Sung Hoon Ko, Geoffrey C. Fox, Kangseok Kim, Sangyoon Oh 0001 |
ICWS | 5 |
| 2002 | Grid services for earthquake scienceabstractAbstract We describe an information system architecture for the ACES (Asia–Pacific Cooperation for Earthquake Simulation) community. It addresses several key features of the field—simulations at multiple scales that need to be coupled together; real‐time and archival observational data, which needs to be analyzed for patterns and linked to the simulations; a variety of important algorithms including partial differential equation solvers, particle dynamics, signal processing and data analysis; a natural three‐dimensional space (plus time) setting for both visualization and observations; the linkage of field to real‐time events both as an aid to crisis management and to scientific discovery. We also address the need to support education and research for a field whose computational sophistication is rapidly increasing and spans a broad range. The information system assumes that all significant data is defined by an XML layer which could be virtual, but whose existence ensures that all data is object‐based and can be accessed and searched in this form. The various capabilities needed by ACES are defined as grid services, which are conformant with emerging standards and implemented with different levels of fidelity and performance appropriate to the application. Grid Services can be composed in a hierarchical fashion to address complex problems. The real‐time needs of the field are addressed by high‐performance implementation of data transfer and simulation services. Further, the environment is linked to real‐time collaboration to support interactions between scientists in geographically distant locations. Copyright © 2002 John Wiley & Sons, Ltd. Geoffrey C. Fox, Sung Hoon Ko, Marlon E. Pierce, Ozgur Balsoy, Jake Kim, Sangmi Lee, Kangseok Kim, Sangyoon Oh 0001, Xi Rao, Mustafa Varank, Hasan Bulut, Gurhan Gunduz, Xiaohong Qiu, Shrideep Pallickara, Ahmet Uyar, Choon-Han Youn |
Concurr. Comput. Pract. Exp. | 8 |