Sangyoon Oh 0001

dblp:65/912-1 · also Oh Sangyoon 0001 · DBLP profile ↗
← Back
48ranked-venue papers
5as first author
26since 2021 · last 2026
0000-0001-5854-149XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 3 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 1 first-authorComputer networks · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Reducing Backfill Failures From Workload Drift with Lightweight Uncertainty Buffers in HPC Job Scheduling
Jiheon Choi, Sangyoon Oh 0001
CCGrid2
2026 GASched: Goal-Adaptive Hierarchical Reinforcement Learning for Multi-Objective HPC Job Scheduling
Minsol Choo, Sangyoon Oh 0001
CCGrid2
2026 S-CQR: Stratified Calibration for Runtime Prediction in HPC Backfill Scheduling
Jiheon Choi, Sangyoon Oh 0001
Euro-Par (2)2
2026 UARP: uncertainty-aware runtime prediction for preventing scheduler termination under Wallclock constraints in HPC
abstract
Effective resource allocation has become a critical issue in high-performance computing (HPC) systems. To effectively allocate resources (e.g., CPU/GPU cores), recent studies focus on predicting each workload’s runtime using machine learning and deep learning models. These methods in HPC often suffer from underestimation, as 33–64% of jobs terminate due to wallclock time limits, whereas user-provided estimates achieve 78–99% success. This failure stems from minimizing mean squared error, which biases predictions toward average-case performance and underestimates jobs in high-skewed (i.e., long-tail) runtime distributions. Specifically, HPC workloads exhibit long-tail runtime distributions, with most jobs completing quickly while a small fraction runs for extremely long durations. To overcome these challenges, we introduce an uncertainty-aware runtime prediction (UARP) method based on multi-quantile regression. Our method directly addresses the underestimation problem by quantifying uncertainty by modeling the conditional distribution without distributional assumptions. Our approach uses the highest-quantile (99th) model and the residual model from the median quantile. The 99th model primarily provides conservative bounds and protection against job underestimates, while the residual uncertainty model protects against unpredictable workloads by estimating prediction variance. In particular, the expected predicted error (i.e., uncertainty) from the residual model plays a critical role in our adaptive safety margin calculation. UARP adds a conservative prediction (99th quantile) and an additional safety margin from our formula, enabling our adaptive margin approach specifically tailored to each job’s characteristics. Evaluation on four production HPC systems (SDSC DataStar, KIT FH2, ANL Interpid, KISTI NURION), UARP achieves 92–99% job success rates while maintaining resource utilization within 1–2% of EASY backfilling. Our method deploys with identical parameters across all systems. This parameter-free deployment eliminates the per-system tuning that fixed-margin approaches require. In addition, our approach integrates with existing schedulers through minimal modification, utilizing uncertainty-aware predictions to prevent timeout-based job termination and preserve system efficiency.
Jiheon Choi, Sangyoon Oh 0001
J. Supercomput.2
2025 When HPC Scheduling Meets Active Learning: Maximizing The Performance with Minimal Data
Jiheon Choi, Minsol Choo, Oh-Kyoung Kwon, Sangyoon Oh 0001
HPC Asia6
2025 Lightweight multi-layered de-identification architecture: Secure client selection in federated learning
Jiheon Choi, Sangyoon Oh 0001
J. Syst. Archit.2
2024 Preserving Near-Optimal Gradient Sparsification Cost for Scalable Distributed Deep Learning
abstract
Communication overhead is a major obstacle to scaling distributed training systems. Gradient sparsification is a potential optimization approach to reduce the communication volume without significant loss of model fidelity. However, existing gradient sparsification methods have low scalability owing to inefficient design of their algorithms, which raises the communication overhead significantly. In particular, gradient build-up and inadequate sparsity control methods degrade the sparsification performance considerably. Moreover, communication traffic increases drastically owing to workload imbalance of gradient selection between workers.To address these challenges, we propose a novel gradient sparsification scheme called ExDyna. In ExDyna, the gradient tensor of the model comprises fined-grained blocks, and contiguous blocks are grouped into non-overlapping partitions. Each worker selects gradients in its exclusively allocated partition so that gradient build-up never occurs. To balance the workload of gradient selection between workers, ExDyna adjusts the topology of partitions by comparing the workloads of adjacent partitions. In addition, ExDyna supports online threshold scaling, which estimates the accurate threshold of gradient selection on-the-fly. Accordingly, ExDyna can satisfy the user-required sparsity level during a training period regardless of models and datasets. Therefore, ExDyna can enhance the scalability of distributed training systems by preserving near-optimal gradient sparsification cost. In experiments, ExDyna outperformed state-of-the-art sparsifiers in terms of training speed and sparsification performance while achieving high accuracy.
Daegun Yoon, Sangyoon Oh 0001
CCGrid2
2024 Staleness aware semi-asynchronous federated learning
Miri Yu, Jiheon Choi, Sangyoon Oh 0001
J. Parallel Distributed Comput.4
2023 MiCRO: Near-Zero Cost Gradient Sparsification for Scaling and Accelerating Distributed DNN Training
abstract
Gradient sparsification is a communication optimisation technique for scaling and accelerating distributed deep neural network (DNN) training. It reduces the increasing communication traffic for gradient aggregation. However, existing sparsifiers have poor scalability because of the high computational cost of gradient selection and/or increase in communication traffic. In particular, an increase in communication traffic is caused by gradient build-up and inappropriate threshold for gradient selection. To address these challenges, we propose a novel gradient sparsification method called MiCRO. In MiCRO, the gradient vector is partitioned, and each partition is assigned to the corresponding worker. Each worker then selects gradients from its partition, and the aggregated gradients are free from gradient build-up. Moreover, MiCRO estimates the accurate threshold to maintain the communication traffic as per user requirement by minimising the compression ratio error. MiCRO enables near-zero cost gradient sparsification by solving existing problems that hinder the scalability and acceleration of distributed DNN training. In our extensive experiments, MiCRO outperformed state-of-the-art sparsifiers with an outstanding convergence rate.
Daegun Yoon, Sangyoon Oh 0001
HiPC2
2023 Addressing Client Heterogeneity in Synchronous Federated Learning: The CHAFL Approach
abstract
Federated learning (FL) is proposed to address the security vulnerabilities of conventional distributed deep learning. Since the capabilities of participating FL clients are highly variable in terms of both statistical and system aspects, FL training will face diminished convergence accuracy and speed. Hence, we propose CHAFL (client heterogeneity aware federated learning) to address client heterogeneity (i.e., statistical and system heterogeneity) in synchronous FL. CHAFL selects clients based on global loss with contribution, defined as local loss, enabling higher round-to-accuracy for the global model than previous studies. Additionally, to handle system heterogeneity, it proposes a lightweight algorithm that eliminates the profiling process previously employed to calculate adaptive local epochs in existing studies, thereby improving time-to-accuracy. In order to verify the effectiveness of CHAFL, we exploit three benchmark datasets on non-IID and system heterogeneous setting for empirical evaluation. Compared to the baseline, CHAFL achieves an accuracy improvement of 3.1 to 5.7%, along with 1.04 to 1.73x higher round-to-accuracy and 1.03 to 1.98x higher time-to-accuracy.
Miri Yu, Oh-Kyoung Kwon, Sangyoon Oh 0001
ICPADS3
2023 DEFT: Exploiting Gradient Norm Difference between Model Layers for Scalable Gradient Sparsification
abstract
Gradient sparsification is a widely adopted solution for reducing the excessive communication traffic in distributed deep learning. However, most existing gradient sparsifiers have relatively poor scalability because of considerable computational cost of gradient selection and/or increased communication traffic owing to gradient build-up. To address these challenges, we propose a novel gradient sparsification scheme, DEFT, that partitions the gradient selection task into sub tasks and distributes them to workers. DEFT differs from existing sparsifiers, wherein every worker selects gradients among all gradients. Consequently, the computational cost can be reduced as the number of workers increases. Moreover, gradient build-up can be eliminated because DEFT allows workers to select gradients in partitions that are non-intersecting (between workers). Therefore, even if the number of workers increases, the communication traffic can be maintained as per user requirement.
Daegun Yoon, Sangyoon Oh 0001
ICPP2
2023 Crossover-SGD: A gossip-based communication in distributed deep learning for alleviating large mini-batch problem and enhancing scalability
abstract
Summary Distributed deep learning is an effective way to reduce the training time for large datasets as well as complex models. However, the limited scalability caused by network‐overheads makes it difficult to synchronize the parameters of all workers and gossip‐based methods that demonstrate stable scalability regardless of the number of workers have been proposed. However, to use gossip‐based methods in general cases, the validation accuracy for a large mini‐batch needs to be verified. For this, we first empirically study the characteristics of gossip methods in a large mini‐batch problem and observe that gossip methods preserve higher validation accuracy than AllReduce‐SGD (stochastic gradient descent) when the number of batch sizes is increased, and the number of workers is fixed. However, the delayed parameter propagation of the gossip‐based models decreases validation accuracy in large node scales. To cope with this problem, we propose Crossover‐SGD that alleviates the delay propagation of weight parameters via segment‐wise communication and random network topology with fair peer selection. We also adapt hierarchical communication to limit the number of workers in gossip‐based communication methods. To validate the effectiveness of our method, we conduct empirical experiments and observe that our Crossover‐SGD shows higher node scalability than stochastic gradient push.
Sangho Yeo, Minho Bae, Minjoong Jeong, Oh-Kyoung Kwon, Sangyoon Oh 0001
Concurr. Comput. Pract. Exp.5
2023 WAVE: designing a heuristics-based three-way breadth-first search on GPUs
Daegun Yoon, Minjoong Jeong, Sangyoon Oh 0001
J. Supercomput.3
2023 SAGE: toward on-the-fly gradient compression ratio scaling
Daegun Yoon, Minjoong Jeong, Sangyoon Oh 0001
J. Supercomput.3
2023 User Preference-Based Hierarchical Offloading for Collaborative Cloud-Edge Computing
abstract
Cloud computing and mobile edge computing techniques supply efficient ways to solve the contradiction between the increasing computing and storage demands of portable terminals and the limited capacity. In this paper, we conduct a three-tier hierarchical service system with multiple UEs, multiple MECs, and a single cloud center. It's worth noting that multiple UEs with personalized options generate a large number of different tasks in real time. To deal with this offloading problem, a response ratio offloading strategy (RROS) centered on user preference and real-time nature is designed to make MECs or CC serve as many UEs as possible. Therefore, a MEC-choosing preference list of each UE is created based on its past experiences at first. Then, each MEC iteratively sorts UEs with its ranking in the UEs' preference list. In order to avoid that the first task arriving at MEC occupies too many resources of MEC and cannot achieve global optimization, we also adopt loop iterative sequencing for multiple tasks arriving within a stipulated time. Lastly, by comparing the optimal response ratio on different MECs and CC, multiple MECs and the CC collaborative offload computing tasks of multiple UEs. Experimental results show that the algorithm significantly outperforms conventional techniques.
Shujuan Tian, Chi Chang, Saiqin Long, Sangyoon Oh 0001, Zhetao Li
IEEE Trans. Serv. Comput.4
2022 Is Ant Colony System better than FFD for VM placement in a heterogeneous cluster?
abstract
First fit decreasing (FFD) is the most popular heuristic for virtual machine (VM) placement problems. However, FFD does not perform as much in a heterogeneous cluster environment. Moreover, FFD and other heuristics, such as best fit decreasing (BFD), are limited to handle the VM placement problem effectively when multiple resources are considered together. In this study, we analyze the reason why the ant colony system performs better than FFD for VM placement in a heterogeneous cluster. We verified our logical observations through experimental comparisons with other heuristics.
Minjoong Jeong, Sangyoon Oh 0001
IC2E3
2022 AR-CNN: an attention ranking network for learning urban perception
Zhetao Li, Wei-Shi Zheng 0001, Sangyoon Oh 0001, Kien Nguyen 0002
Sci. China Inf. Sci.4
2022 AMBLE: Adjusting mini-batch and local epoch for federated learning with heterogeneous devices
Juwon Park, Daegun Yoon, Sangho Yeo, Sangyoon Oh 0001
J. Parallel Distributed Comput.4
2021 A dynamic task offloading algorithm based on greedy matching in vehicle network
Shujuan Tian, Xianghong Deng, Tingrui Pei, Sangyoon Oh 0001, Weiping Xue
Ad Hoc Networks5
2021 Novel data-placement scheme for improving the data locality of Hadoop in heterogeneous environments
abstract
Summary To address the challenging needs of high‐performance big data processing, parallel‐distributed frameworks such as Hadoop are being utilized extensively. However, in heterogeneous environments, the performance of Hadoop clusters is below par. This is primarily because the blocks of the clusters are allocated equally to all nodes without regard to differences in the capability of individual nodes. This results in reduced data locality. Thus, a new data‐placement scheme that enhances data locality is required for Hadoop in heterogeneous environments. This article proposes a new data placement scheme that preserves the same degree of data locality in heterogeneous environments as that of the standard Hadoop, with only a small amount of replicated data. In the proposed scheme, only those blocks with the highest probability of being accessed remotely are selected and replicated. The results of experiments conducted indicate that the proposed scheme incurs only a 20% disk space overhead and has virtually the same data locality ratio as the standard Hadoop, which has a replication factor of three and 200% disk space overhead.
Minho Bae, Sangho Yeo, Gyudong Park, Sangyoon Oh 0001
Concurr. Comput. Pract. Exp.4
2021 Parallel Programming Models in High-Performance Cloud (ParaMo 2019)
Sangyoon Oh 0001
Concurr. Comput. Pract. Exp.1
2021 Exploring a system architecture of content-based publish/subscribe system for efficient on-the-fly data dissemination
abstract
Summary In a cloud‐scale publish/subscribe messaging system, it is difficult to partition subscription data among several servers. Without a sophisticated scheme and a system architecture, the messaging system would either waste resources or fail to deliver messages on time. In this study, we propose DRDA, a dynamic replication degree adjustment technology, for efficient message delivery. The technology calculates and maintains the number of subscription replications at a reasonable level by monitoring the statuses of servers, based on the number of subscription replications and the frequency of event dissemination. To verify the effectiveness of our proposed scheme and system architecture, we build a prototype of a content‐based publish/subscribe system that dynamically adjusts the number of replications among brokers. Furthermore, we compare the load balance, resource overhead, and performance of a publish/subscribe system with DRDA with a publish/subscribe system without DRDA. The experimental results show that DRDA outperforms other approaches under various parameter configurations. We have added the prototype code to a GitHub repository to make it publicly available.
Daegun Yoon, Gyudong Park, Sangyoon Oh 0001
Concurr. Comput. Pract. Exp.3
2021 Balanced content space partitioning for pub/sub: a study on impact of varying partitioning granularity
Daegun Yoon, Zhetao Li, Sangyoon Oh 0001
J. Supercomput.3
2021 A low redundancy and high time efficiency large-scale task assignment strategy for heterogeneous service-oriented cloud computing systems
Lizan Wang, Guoqi Xie, Tingrui Pei, Sangyoon Oh 0001, Zhetao Li
J. Supercomput.5
2021 Lightweight Single Image Super-resolution with Dense Connection Distillation Network
abstract
Single image super-resolution attempts to reconstruct a high-resolution (HR) image from its corresponding low-resolution (LR) image, which has been a research hotspot in computer vision and image processing for decades. To improve the accuracy of super-resolution images, many works adopt very deep networks to model the translation from LR to HR, resulting in memory and computation consumption. In this article, we design a lightweight dense connection distillation network by combining the feature fusion units and dense connection distillation blocks (DCDB) that include selective cascading and dense distillation components. The dense connections are used between and within the distillation block, which can provide rich information for image reconstruction by fusing shallow and deep features. In each DCDB, the dense distillation module concatenates the remaining feature maps of all previous layers to extract useful information, the selected features are then assessed by the proposed layer contrast-aware channel attention mechanism, and finally the cascade module aggregates the features. The distillation mechanism helps to reduce training parameters and improve training efficiency, and the layer contrast-aware channel attention further improves the performance of model. The quality and quantity experimental results on several benchmark datasets show the proposed method performs better tradeoff in term of accuracy and efficiency.
Yanchun Li, Jianglian Cao, Zhetao Li, Sangyoon Oh 0001, Nobuyoshi Komuro
ACM Trans. Multim. Comput. Commun. Appl.4
2021 Accelerated deep reinforcement learning with efficient demonstration utilization techniques
Sangho Yeo, Sangyoon Oh 0001, Minsu Lee 0001
World Wide Web2
2020 An Effective Design to Improve the Efficiency of DPUs on FPGA
abstract
Convolutional neural networks (CNNs) have been widely used in various complicated problems, such as image classification, objection detection, semantic segmentation. To meet diversified CNN structures, the deep learning processing unit (DPU) is designed as a general accelerator on field programmable gate array (FPGA) to support various CNN layers, such as convolution, pooling, activation, etc. However, low DPU utilization and schedule efficiency appear when DPU used to multitask application completed by CNN models. In this paper, an effective design including multi-core with different size (MCDS) and DPU Plus is proposed to improve the efficiency of DPUs usage from the two dimensions of time and space. Through increasing the number of DPU cores on an FPGA and the utilization of single DPU core, the design of MCDS can effectively improve the overall throughput with restricted on-chip resources. Furthermore, the design of DPU Plus is proposed to improve the schedule efficiency of DPUs through simultaneously implementing DPU with other significant auxiliary modules of the application system on the same FPGA. Finally, a color space conversion module is implemented cooperate to the DPU cores to testify its performance, and the experimen shows that compared with running on the the CPU completely, it achieves16.2x acceleration, and increases the throughput of the entire system by 3.0x.
Qingyong Deng, Saiqin Long, Shaohui Liu, Sangyoon Oh 0001
ICPADS5
2020 A similarity clustering-based deduplication strategy in cloud storage systems
abstract
Deduplication is a data redundancy elimination technique, designed to save system storage resources by reducing redundant data in cloud storage systems. With the development of cloud computing technology, deduplication has been increasingly applied to cloud data centers. However, traditional technologies face great challenges in big data deduplication to properly weigh the two conflicting goals of deduplication throughput and high duplicate elimination ratio. This paper proposes a similarity clustering-based deduplication strategy (named SCDS), which aims to delete more duplicate data without significantly increasing system overhead. The main idea of SCDS is to narrow the query range of fingerprint index by data partitioning and similarity clustering algorithms. In the data preprocessing stage, SCDS uses data partitioning algorithm to classify similar data together. In the data deletion stage, the similarity clustering algorithm is used to divide the similar data fingerprint superblock into the same cluster. Repetitive fingerprints are detected in the same cluster to speed up the retrieval of duplicate fingerprints. Experiments show that the deduplication ratio of SCDS is better than some existing similarity deduplication algorithms, but the overhead is only slightly higher than some high throughput but low deduplication ratio methods.
Saiqin Long, Zhetao Li, Qingyong Deng, Sangyoon Oh 0001, Nobuyoshi Komuro
ICPADS5
2020 Modeling Analysis and Cost-Performance Ratio Optimization of Virtual Machine Scheduling in Cloud Computing
abstract
As an essential feature of cloud computing, dynamic scalability enables the cloud system to dynamically expand or shrink resources according to user needs at runtime. Effectively predicting and optimizing the cost and performance of cloud computing platforms have become one of the key research challenges in the field of cloud computing. In this article, to quantitatively predict the cost and performance of cloud computing platforms, we propose a cloud computing resource analysis model considering both hot/cold startup and hot/cold shutdown of virtual machines (VMs), and use the M/M/N/oo queuing model to analyze cloud computing platform and acquire accurate performance indicators, such as elasticity indicators, cost indicators, performance indicators, cost-performance ratios, etc. In addition, we establish a multi-objective optimization model to optimize both performance and cost of cloud computing platform. Then the optimal stopping and cost-performance optimization algorithm are applied to obtain the optimal configurations, including the number of hot startup VMs, the system service rate, the hot/cold startup rate of VMs, and the hot/cold shutdown rate. By comparing with existing optimization methods, we demonstrate the superiority of our cost-performance ratio optimization method.
Jiale Dang, Zhetao Li, Hongfang Gong, Feng Zhang 0007, Sangyoon Oh 0001
IEEE Trans. Parallel Distributed Syst.6
2018 Decentralized Message Broker Federation Architecture with Multiple DHT Rings for High Survivability
Minsub Kim, Minho Bae, Sangho Yeo, Gyudong Park, Sangyoon Oh 0001
ICCSA (5)5
2017 High Performance Query Processing for Web Scale RDF Data using BSP Style Communication and Balanced Distribution
abstract
To overcome scalability and performance issues for process queries over a web-scale RDF data, various studies have proposed RDF SPARQL query processing algorithm using parallel processing manners. However, it is hard to resolve the scalability and performance issues together because the problem of communication overhead between nodes is closely related to the data distribution for parallel processing. For efficient RDF query parallel processing, it is essential to distribute and process data evenly while reducing communication overhead. In this paper, we propose RDF query parallel processing algorithms with RDF data partitioning technique to guarantee evenly distributed data over the cluster. We also propose our in-memory RDF query processing system as a form of Bulk Synchronization Parallel system to reduce network overhead. Our empirical evaluation results show that the proposed system outperforms a popular RDF-3X on LUBM benchmark and UniProt queries from 2.20 to 43.08 times. Especially, the effectiveness of the system improves significantly when the SPARQL queries are complex with high input and select.
Minho Bae, Junho Eum, Sangyoon Oh 0001
ICPP4
2016 MGEScan: a Galaxy-based system for identifying retrotransposons in genomes
abstract
UNLABELLED: : MGEScan-long terminal repeat (LTR) and MGEScan-non-LTR are successfully used programs for identifying LTRs and non-LTR retrotransposons in eukaryotic genome sequences. However, these programs are not supported by easy-to-use interfaces nor well suited for data visualization in general data formats. Here, we present MGEScan, a user-friendly system that combines these two programs with a Galaxy workflow system accelerated with MPI and Python threading on compute clusters. MGEScan and Galaxy empower researchers to identify transposable elements in a graphical user interface with ready-to-use workflows. MGEScan also visualizes the custom annotation tracks for mobile genetic elements in public genome browsers. A maximum speed-up of 3.26× is attained for execution time using concurrent processing and MPI on four virtual cores. MGEScan provides four operational modes: as a command line tool, as a Galaxy Toolshed, on a Galaxy-based web server, and on a virtual cluster on the Amazon cloud. AVAILABILITY AND IMPLEMENTATION: MGEScan tutorials and source code are available at http://mgescan.readthedocs.org/ CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hyungro Lee, Minsu Lee 0001, Wazim Mohammed Ismail, Mina Rho, Geoffrey C. Fox, Sangyoon Oh 0001, Haixu Tang
Bioinform.6
2016 Energy-efficient multisite offloading policy using Markov decision process for mobile cloud computing
Mati B. Terefe, Heezin Lee, Nojung Heo, Geoffrey C. Fox, Sangyoon Oh 0001
Pervasive Mob. Comput.5
2015 Evaluating ARM HPC clusters for scientific workloads
abstract
Summary The power consumption of modern high‐performance computing (HPC) systems that are built using power hungry commodity servers is one of the major hurdles for achieving Exascale computation. Several efforts have been made by the HPC community to encourage the use of low‐powered system‐on‐chip (SoC) embedded processors in large‐scale HPC systems. These initiatives have successfully demonstrated the use of ARM SoCs in HPC systems, but there is still a need to analyze the viability of these systems for HPC platforms before a case can be made for Exascale computation. The major shortcomings of current ARM‐HPC evaluations include a lack of detailed insights about performance levels on distributed multicore systems and performance levels for benchmarking in large‐scale applications running on HPC. In this paper, we present a comprehensive evaluation of results that covers major aspects of server and HPC benchmarking for ARM‐based SoCs. For the experiments, we built an unconventional cluster of ARM Cortex‐A9s that is referred to as Weiser and ran single‐node benchmarks (STREAM, Sysbench, and PARSEC) and multi‐node scientific benchmarks (High‐performance Linpack (HPL), NASA Advanced Supercomputing (NAS) Parallel Benchmark, and Gadget‐2) in order to provide a baseline for performance limitations of the system. Based on the experimental results, we claim that the performance of ARM SoCs depends heavily on the memory bandwidth, network latency, application class, workload type, and support for compiler optimizations. During server‐based benchmarking, we observed that when performing memory intensive benchmarks for database transactions, x86 performed 12% better for multithreaded query processing. However, ARM performed four times better for performance to power ratios for a single core and 2.6 times better on four cores. We noticed that emulated double precision floating point in Java resulted in three to four times slower performance as compared with the performance in C for CPU‐bound benchmarks. Even though Intel x86 performed slightly better in computation‐oriented applications, ARM showed better scalability in I/O bound applications for shared memory benchmarks. We incorporated the support for ARM in the MPJ‐Express runtime and performed comparative analysis of two widely used message passing libraries. We obtained similar results for network bandwidth, large‐scale application scaling, floating‐point performance, and energy‐efficiency for clusters in message passing evaluations (NBP and Gadget 2 with MPJ‐Express and MPICH). Our findings can be used to evaluate the energy efficiency of ARM‐based clusters for server workloads and scientific workloads and to provide a guideline for building energy‐efficient HPC clusters. Copyright © 2015 John Wiley & Sons, Ltd.
Maqbool Jahanzeb, Sangyoon Oh 0001, Geoffrey C. Fox
Concurr. Comput. Pract. Exp.2
2015 Ontology-based quantitative similarity metric for event matching in publish/subscribe system
Hongjae Kim, Sanggil Kang, Sangyoon Oh 0001
Neurocomputing3
2015 An approach to mitigate DoS attack based on routing misbehavior in wireless ad hoc networks
Wonil Kim, Kangseok Kim, Sangyoon Oh 0001, Dong-Kyoo Kim
Peer-to-Peer Netw. Appl.4
2014 Semantic similarity method for keyword query system on RDF
Minho Bae, Sanggil Kang, Sangyoon Oh 0001
Neurocomputing3
2013 Effective Hotspot Removal System Using Neural Network Predictor
Sangyoon Oh 0001, Mun-Young Kang, Sanggil Kang
ACIIDS (2)1
2013 K-depth RDF Keyword Search Algorithm Based on Structure Indexing
abstract
Information retrieval from large scale RDF datasets is a challenging task. Because it takes much time to process query and it is hard to store a large collection of RDFs, it requires an efficient method to index and query. In this paper, we propose a novel RDF management system architecture along with indexing and querying algorithms. Our empirical experiments show our system performs substantially better than the conventional system. Also we verify the effectiveness of k-depth concept in our design with additional experiments.
Minho Bae, Sanggil Kang, Sangyoon Oh 0001
KES-AMSTA4
2012 An Intelligent RDF Management System with Hybrid Querying Approach
Jangsu Kihm, Minho Bae, Sanggil Kang, Sangyoon Oh 0001
ICCCI (1)4
2011 Ensemble Learning with Active Example Selection for Imbalanced Biomedical Data Classification
abstract
In biomedical data, the imbalanced data problem occurs frequently and causes poor prediction performance for minority classes. It is because the trained classifiers are mostly derived from the majority class. In this paper, we describe an ensemble learning method combined with active example selection to resolve the imbalanced data problem. Our method consists of three key components: 1) an active example selection algorithm to choose informative examples for training the classifier, 2) an ensemble learning method to combine variations of classifiers derived by active example selection, and 3) an incremental learning scheme to speed up the iterative training procedure for active example selection. We evaluate the method on six real-world imbalanced data sets in biomedical domains, showing that the proposed method outperforms both the random under sampling and the ensemble with under sampling methods. Compared to other approaches to solving the imbalanced data problem, our method excels by 0.03-0.15 points in AUC measure.
Sangyoon Oh 0001, Minsu Lee 0001, Byoung-Tak Zhang
IEEE ACM Trans. Comput. Biol. Bioinform.1
2010 Real-time performance analysis for publish/subscribe systems
Sangyoon Oh 0001, Jai-Hoon Kim, Geoffrey C. Fox
Future Gener. Comput. Syst.1
2009 Ensemble Learning Based on Active Example Selection for Solving Imbalanced Data Problem in Biomedical Data
abstract
The imbalanced data problem is popular in biomedical classification tasks. Since trained classifiers using imbalanced data are mostly derived from the majority class, their prediction performance is poor for the minority class. In this paper, we propose a novel ensemble learning method based on an active example selection algorithm to resolve the imbalanced data problem. To compensate a possible sub-optimal classifier, our proposed ensemble learning methods aggregates classifiers built by the active example selection algorithm. We implement this ensemble learning method based on the active example selection algorithm using incremental naive Bayes classifiers. Our empirical results show that we greatly improve the performance of classification models trained by five real world imbalanced biomedical data. The proposed ensemble learning methods outperforms other approaches by 0.03~0.15 in terms of AUC which solve imbalanced data problem.
Minsu Lee 0001, Sangyoon Oh 0001, Byoung-Tak Zhang
BIBM2
2008 XML Metadata Services
abstract
Abstract As service‐oriented architecture principles have gained importance, an emerging need has appeared for methodologies to locate desired services that provide access to their capability descriptions. These services must typically be assembled into short‐term service collections that, together with code execution services, are combined into a meta‐application to perform a particular task. To address the metadata requirements of these problems, we introduce a hybrid Information Service to manage both stateless and stateful (transient) metadata. We leverage the two widely used Web Service standards: Universal Description, Discovery and Integration (UDDI) and Web Services Context (WS‐Context), in our design. We describe our approach and experiences when designing ‘semantics’. We report the results from a prototype of the system that is applied to a mobile environment for optimizing Web Service communications. Copyright © 2007 John Wiley & Sons, Ltd.
Mehmet S. Aktas, Geoffrey C. Fox, Marlon E. Pierce, Sangyoon Oh 0001
Concurr. Comput. Pract. Exp.4
2007 Optimizing Web Service messaging performance in mobile computing
Sangyoon Oh 0001, Geoffrey C. Fox
Future Gener. Comput. Syst.1
2006 Multi-module Image Classification System
Wonil Kim, Sangyoon Oh 0001, Sanggil Kang, Dongkyun Kim
FQAS2
2003 A Web Service Approach to Universal Accessibility in Collaboration Services
Sangmi Lee, Sung Hoon Ko, Geoffrey C. Fox, Kangseok Kim, Sangyoon Oh 0001
ICWS5
2002 Grid services for earthquake science
abstract
Abstract We describe an information system architecture for the ACES (Asia–Pacific Cooperation for Earthquake Simulation) community. It addresses several key features of the field—simulations at multiple scales that need to be coupled together; real‐time and archival observational data, which needs to be analyzed for patterns and linked to the simulations; a variety of important algorithms including partial differential equation solvers, particle dynamics, signal processing and data analysis; a natural three‐dimensional space (plus time) setting for both visualization and observations; the linkage of field to real‐time events both as an aid to crisis management and to scientific discovery. We also address the need to support education and research for a field whose computational sophistication is rapidly increasing and spans a broad range. The information system assumes that all significant data is defined by an XML layer which could be virtual, but whose existence ensures that all data is object‐based and can be accessed and searched in this form. The various capabilities needed by ACES are defined as grid services, which are conformant with emerging standards and implemented with different levels of fidelity and performance appropriate to the application. Grid Services can be composed in a hierarchical fashion to address complex problems. The real‐time needs of the field are addressed by high‐performance implementation of data transfer and simulation services. Further, the environment is linked to real‐time collaboration to support interactions between scientists in geographically distant locations. Copyright © 2002 John Wiley & Sons, Ltd.
Geoffrey C. Fox, Sung Hoon Ko, Marlon E. Pierce, Ozgur Balsoy, Jake Kim, Sangmi Lee, Kangseok Kim, Sangyoon Oh 0001, Xi Rao, Mustafa Varank, Hasan Bulut, Gurhan Gunduz, Xiaohong Qiu, Shrideep Pallickara, Ahmet Uyar, Choon-Han Youn
Concurr. Comput. Pract. Exp.8