EDBT 2026 Demo / reviewers in the wild / expert
Seung-Jong Park
dblp:42/2979
· DBLP profile ↗
37ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0001-7821-7793ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 15 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 11 · 1 since 2021Systems, architecture and hardware · 9 · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Artificial intelligence and machine learning · 4Graphics, computer vision, multimedia, augmented reality and games · 2Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | RainbowCake: Mitigating Cold-starts in Serverless with Layer-wise Container Caching and SharingabstractServerless computing has grown rapidly as a new cloud computing paradigm that promises ease-of-management, cost-efficiency, and auto-scaling by shipping functions via self-contained virtualized containers. Unfortunately, serverless computing suffers from severe cold-start problems---starting containers incurs non-trivial latency. Full container caching is widely applied to mitigate cold-starts, yet has recently been outperformed by two lines of research: partial container caching and container sharing. However, either partial container caching or container sharing techniques exhibit their drawbacks. Partial container caching effectively deals with burstiness while leaving cold-start mitigation halfway; container sharing reduces cold-starts by enabling containers to serve multiple functions while suffering from excessive memory waste due to over-packed containers. Hanfei Yu, Rohan Basu Roy, Christian Fontenot, Devesh Tiwari, Jian Li 0008, Hong Zhang 0025, Hao Wang 0022, Seung-Jong Park |
ASPLOS (1) | 8 |
| 2024 | Stellaris: Staleness-Aware Distributed Reinforcement Learning with Serverless ComputingabstractDeep reinforcement learning (DRL) has achieved remarkable success in diverse areas, including gaming AI, scientific simulations, and large-scale (HPC) system scheduling. DRL training, which involves a trial-and-error process, demands considerable time and computational resources. To overcome this challenge, distributed DRL algorithms and frameworks have been developed to expedite training by leveraging large-scale resources. However, existing distributed DRL solutions rely on synchronous learning with serverful infrastructures, suffering from low training efficiency and overwhelming training costs. This paper proposes Stellaris, the first to introduce a generic asynchronous learning paradigm for distributed DRL training with serverless computing. We devise an importance sampling truncation technique to stabilize DRL training and develop a staleness-aware gradient aggregation method tailored to the dynamic staleness in asynchronous serverless DRL training. Experiments on AWS EC2 regular testbeds and HPC clusters show that Stellaris outperforms existing state-of-the-art DRL baselines by achieving $2.2 \times$ higher rewards (i.e., training quality) and reducing 41% training costs. Hanfei Yu, Hao Wang 0022, Devesh Tiwari, Jian Li 0008, Seung-Jong Park |
SC | 5 |
| 2024 | Nitro: Boosting Distributed Reinforcement Learning with Serverless ComputingabstractDeep reinforcement learning (DRL) has demonstrated significant potential in various applications, including gaming AI, robotics, and system scheduling. DRL algorithms produce, sample, and learn from training data online through a trial-and-error process, demanding considerable time and computational resources. To address this, distributed DRL algorithms and paradigms have been developed to expedite training using extensive resources. Through carefully designed experiments, we are the first to observe that strategically increasing the actor-environment interactions by spawning more concurrent actors at certain training rounds within ephemeral time frames can significantly enhance training efficiency. Yet, current distributed DRL solutions, which are predominantly server-based (or serverful), fail to capitalize on these opportunities due to their long startup times, limited adaptability, and cumbersome scalability. This paper proposes Nitro , a generic training engine for distributed DRL algorithms that enforces timely and effective boosting with concurrent actors instantaneously spawned by serverless computing. With serverless functions, Nitro adjusts data sampling strategies dynamically according to the DRL training demands. Nitro seizes the opportunity of real-time boosting by accurately and swiftly detecting an empirical metric. To achieve cost efficiency, we design a heuristic actor scaling algorithm to guide Nitro for cost-aware boosting budget allocation. We integrate Nitro with state-of-the-art DRL algorithms and frameworks and evaluate them on AWS EC2 and Lambda. Experiments with Mujoco and Atari benchmarks show that Nitro improves the final rewards ( i.e. , training quality) by up to 6× and reduces training costs by up to 42%. Hanfei Yu, Jacob Carter, Hao Wang 0022, Devesh Tiwari, Jian Li 0008, Seung-Jong Park |
Proc. VLDB Endow. | 6 |
| 2024 | Freyr $^+$+: Harvesting Idle Resources in Serverless Computing via Deep Reinforcement LearningabstractServerless computing has revolutionized online service development and deployment with ease-to-use operations, auto-scaling, fine-grained resource allocation, and pay-as-you-go pricing. However, a gap remains in configuring serverless functions—the actual resource consumption may vary due to function types, dependencies, and input data sizes, thus mismatching the static resource configuration by users. Dynamic resource consumption against static configuration may lead to either poor function execution performance or low utilization. This paper proposesFreyr$^+$, a novel resource manager (RM) that dynamically harvests idle resources from over-provisioned functions to accelerate under-provisioned functions for serverless platforms.Freyr$^+$monitors each function's resource utilization in real-time and detects the mismatches between user configuration and actual resource consumption. We design deep reinforcement learning (DRL) algorithms with attention-enhanced embedding, incremental learning, and safeguard mechanism forFreyr$^+$to harvest idle resources safely and accelerate functions efficiently. We have implemented and deployed aFreyr$^+$prototype in a 13-node Apache OpenWhisk cluster using AWS EC2.Freyr$^+$is evaluated on both large-scale simulation and real-world testbed. Experimental results show thatFreyr$^+$harvests 38% of function invocations’ idle resources and accelerates 39% of invocations using harvested resources.Freyr$^+$reduces the 99th-percentile function response latency by 26% compared to the baseline RMs. Hanfei Yu, Hao Wang 0022, Jian Li 0008, Xu Yuan 0001, Seung-Jong Park |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | Libra: Harvesting Idle Resources Safely and Timely in Serverless ClustersabstractServerless computing has been favored by users and infrastructure providers from various industries, including online services and scientific computing. Users enjoy its auto-scaling and ease-of-management, and providers own more control to optimize their service. However, existing serverless platforms still require users to pre-define resource allocations for their functions, leading to frequent misconfiguration by inexperienced users in practice. Besides, functions' varying input data further escalate the gap between their dynamic resource demands and static allocations, leaving functions either over-provisioned or under-provisioned. This paper presents Libra, a safe and timely resource harvesting framework for multi-node serverless clusters. Libra makes precise harvesting decisions to accelerate function invocations with harvested resources and jointly improve resource utilization by profiling dynamic resource demands and availability proactively. Experiments on OpenWhisk clusters with real-world workloads show that Libra reduces response latency by 39% and achieves 3X resource utilization compared to state-of-the-art solutions. Hanfei Yu, Christian Fontenot, Hao Wang 0022, Jian Li 0008, Xu Yuan 0001, Seung-Jong Park |
HPDC | 6 |
| 2022 | Accelerating Serverless Computing by Harvesting Idle ResourcesabstractServerless computing automates fine-grained resource scaling and simplifies the development and deployment of online services with stateless functions. However, it is still non-trivial for users to allocate appropriate resources due to various function types, dependencies, and input sizes. Misconfiguration of resource allocations leaves functions either under-provisioned or over-provisioned and leads to continuous low resource utilization. This paper presents Freyr, a new resource manager (RM) for serverless platforms that maximizes resource efficiency by dynamically harvesting idle resources from over-provisioned functions to under-provisioned functions. Freyr monitors each function’s resource utilization in real-time, detects over-provisioning and under-provisioning, and learns to harvest idle resources safely and accelerates functions efficiently by applying deep reinforcement learning algorithms along with a safeguard mechanism. We have implemented and deployed a Freyr prototype in a 13-node Apache OpenWhisk cluster. Experimental results show that 38.8% of function invocations have idle resources harvested by Freyr, and 39.2% of invocations are accelerated by the harvested resources. Freyr reduces the 99th-percentile function response latency by 32.1% compared to the baseline RMs. Hanfei Yu, Hao Wang 0022, Jian Li 0008, Xu Yuan 0001, Seung-Jong Park |
WWW | 5 |
| 2020 | Distributed de novo assembler for large-scale long-read datasetsabstractThird-generation DNA sequencing technologies such as single-molecule real-time sequencing (SMRT) and nanopore sequencing have the potential to fill the gaps in the existing genome databases since the raw sequences produced by these machines are much longer than those of previous generations and therefore result in more contiguous assemblies. However, these long reads have a high error rate, which makes the assembly process computationally challenging. Moreover, since existing long-read assemblers are designed to run on a single machine, they either take days to complete or run out of memory on even moderate-sized datasets. In this paper, we present a distributed long-read assembler that can assemble large-scale noisy sequence datasets on thousands of cores, resulting in orders of magnitude faster assembly times. By effectively using the map-reduce computation model with a distributed hash-map, both built using a high-performance active messaging middleware, we can assemble a PacBio human genome dataset with 139 billion base-pairs (about 130 GB) in about 33 minutes (using 2,560 cores) compared to more than 38 hours (using 28 cores) with the current state-of-the-art assembler. Sayan Goswami, Kisung Lee, Seung-Jong Park |
IEEE BigData | 3 |
| 2018 | ParLECH: Parallel Long-Read Error Correction with Hadoop
Arghya Kusum Das, Kisung Lee, Seung-Jong Park |
BIBM | 3 |
| 2018 | Towards Distributed Cyberinfrastructure for Smart Cities Using Big Data and Deep Learning TechnologiesabstractRecent advances in big data and deep learning technologies have enabled researchers across many disciplines to gain new insight into large and complex data. For example, deep neural networks are being widely used to analyze various types of data including images, videos, texts, and time-series data. In another example, various disciplines such as sociology, social work, and criminology are analyzing crowd-sourced and online social network data using big data technologies to gain new insight from a plethora of data. Even though many different types of data are being generated and analyzed in various domains, the development of distributed city-level cyberinfrastructure for effectively integrating such data to generate more value and gain insights is still not well-addressed in the research literature. In this paper, we present our current efforts and ultimate vision to build distributed cyberinfrastructure which integrates big data and deep learning technologies with a variety of data for enhancing public safety and livability in cites. We also introduce several methodologies and applications that we are developing on top of the cyberinfrastructure to support diverse community stakeholders in cities. Shayan Shams, Sayan Goswami, Kisung Lee, Seungwon Yang, Seung-Jong Park |
ICDCS | 5 |
| 2018 | GPU-Accelerated Large-Scale Genome AssemblyabstractSpurred by a widening gap between hardware accelerators and traditional processors, numerous bioinformatics applications have harnessed the computing power of GPUs and reported substantial performance improvements compared to their CPU-based counterparts. However, most of these GPU-based applications only focus on the read alignment problem, while the field of de novo assembly still relies mostly on CPU-based solutions. This is primarily due to the nature of the assembly workload which is not only compute-intensive but also extremely data-intensive. Such workloads require large memories, making it difficult to adapt them to use GPUs with their limited memory capacities. To the best of our knowledge, no GPU-based assembler reported in the recent literature has attempted to assemble datasets larger than a few tens of gigabytes, whereas real sequence datasets are often several hundreds of gigabytes in size. In this paper, we present a new GPU-accelerated genome assembler called LaSAGNA, which can assemble large-scale sequence datasets using a single GPU by building string graphs from approximate all-pair overlaps. LaSAGNA can also run on multiple GPUs across multiple compute nodes connected by a high-speed network to expedite the assembly process. To utilize the limited memory on GPUs efficiently, LaSAGNA uses a semi-streaming approach that makes at most a logarithmic number of passes over the input data based on the available memory. Moreover, we propose a two-level streaming model, from disk to host memory and from host memory to device memory, to minimize disk I/O. Using LaSAGNA, we can assemble a 400 GB human genome dataset on a single NVIDIA K40 GPU in 17 hours, and in a little over 5 hours on an 8-node cluster of NVIDIA K20s. Sayan Goswami, Kisung Lee, Shayan Shams, Seung-Jong Park |
IPDPS | 4 |
| 2018 | Deep Generative Breast Cancer Screening and Diagnosis
Shayan Shams, Richard Platania, Jian Zhang 0004, Joohyun Kim 0001, Seung-Jong Park |
MICCAI (2) | 5 |
| 2017 | Minimal Coflow Routing and Scheduling in OpenFlow-Based Cloud Storage Area NetworksabstractResearches affirm that coflow scheduling/routing substantially shortens the average application inner communication time in data center networks(DCNs). The commonly desirable critical features of existing coflow scheduling/routing framework includes (1) coflow scheduling, (2) coflow routing, and (3) per-flow rate-limiting. However, to provide the 3 features, existing frameworks require customized computing frameworks, customized operating systems, or specific external commercial monitoring frameworks on software-defined networking(SDN) switches. These requirements defer or even prohibit the deployment of coflow scheduling/routing in production DCNs. In this paper, we design a coflow scheduling and routing framework, MinCOF which has minimal requirements on hosts and switches for cloud storage area networks(SANs) based on OpenFlow SDN. MinCOF accommodates all critical features of coflow scheduling/routing from previous works. The deployability in production environment is especially taken into consideration. The OpenFlow architecture is capable of processing the traffic load in a cloud SAN. Not necessary requirements for hosts from existing frameworks are migrated to the mature commodity OpenFlow 1.3 Switch and our coflow scheduler. Transfer applications on hosts only need slight enhancements on their existing connection establishment and progress reporting functions. Evaluations reveal that MinCOF decreases the average coflow completion time (CCT) by 12.94% compared to the latest OpenFlow-based coflow scheduling and routing framework. Chui-Hui Chiu, Dipak Kumar Singh, Qingyang Wang 0001, Kisung Lee, Seung-Jong Park |
CLOUD | 5 |
| 2017 | Coflourish: An SDN-Assisted Coflow Scheduling Framework for CloudsabstractExisting coflow scheduling frameworks effectively shorten communication time and completion time of cluster applications. However, existing frameworks only consider available bandwidth on hosts and overlook congestion in the network when making scheduling decisions. Through extensive simulations using the realistic workload probability distribution from Facebook, we observe the performance degradation of the state-of-the-art coflow scheduling framework, Varys, in the cloud environment on a shared data center network (DCN) because of the lack of network congestion information. We propose Coflourish, the first coflow scheduling framework that exploits the congestion feedback assistances from the software-defined-networking(SDN)-enabled switches in the networks for available bandwidth estimation. Our simulation results demonstrate that Coflourish outperforms Varys by up to 75.5% in terms of average coflow completion time under various workload conditions. The proposed work also reveals the potentials of integration with traffic engineering mechanisms in lower levels for further performance optimization. Chui-Hui Chiu, Dipak Kumar Singh, Qingyang Wang 0001, Seung-Jong Park |
CLOUD | 4 |
| 2017 | Augmenting Amdahl's Second Law: A Theoretical Model to Build Cost-Effective Balanced HPC Infrastructure for Data-Driven ScienceabstractHigh-performance analysis of big data demands more computing resources, forcing similar growth in computation cost. So, the challenge to the HPC system designers is providing not only high performance but also high performance at lower cost. For high performance yet cost effective cyberinfrastructure, we propose a new system model augmenting Amdahl's second law for balanced system to optimize price-performance-ratio. We express the optimal balance among CPU-speed, I/O-bandwidth and DRAM-size (i.e., Amdahl's I/O-and memory-number) in terms of application characteristics and hardware cost. Considering Xeon processor and recent hardware prices, we showed that a system needs almost 0.17GBPS I/O-bandwidth and 3GB DRAM per GHz CPU-speed to minimize the price-performance-ratio for data-and compute-intensive applications. To substantiate our claim, we evaluate three different cluster architectures: 1) SupermikeII, a traditional HPC cluster, 2) SwatIII, a regular datacenter, and 3) CeresII, a MicroBrick-based novel hyperscale system. CeresII with 6-Xeon-D1541 cores (2GHz/core), 1-NVMe SSD (2GBPS I/O-bandwidth) and 64GB DRAM per node, closely resembles the optimum produced by our model. Consequently, in terms of price-performance-ratio CeresII outperformed both SupermikeII (by 65-85%) and SwatIII (by 40-50%) for data-and compute-intensive Hadoop benchmarks (TeraSort and WordCount) and our own benchmark genome assembler based on Hadoop and Giraph. Arghya Kusum Das, Jae-Ki Hong, Sayan Goswami, Richard Platania, Kisung Lee, Wooseok Chang, Seung-Jong Park, Ling Liu 0001 |
CLOUD | 7 |
| 2017 | Evaluation of Deep Learning Frameworks Over Different HPC ArchitecturesabstractRecent advances in deep learning have enabled researchers across many disciplines to uncover new insights about large datasets. Deep neural networks have shown applicability to image, time-series, textual, and other data, all of which are available in a plethora of research fields. However, their computational complexity and large memory overhead requires advanced software and hardware technologies to train neural networks in a reasonable amount of time. To make this possible, there has been an influx in development of deep learning software that aim to leverage advanced hardware resources. In order to better understand the performance implications of deep learning frameworks over these different resources, we analyze the performance of three different frameworks, Caffe, TensorFlow, and Apache SINGA, over several hardware environments. This includes scaling up and out with single-and multi-node setups using different CPU and GPU technologies. Notably, we investigate the performance characteristics of NVIDIA's state-of-the-art hardware technology, NVLink, and also Intel's Knights Landing, the most advanced Intel product for deep learning, with respect to training time and utilization. To our best knowledge, this is the first work concerning deep learning bench-marking with NVLink and Knights Landing. Through these experiments, we provide analysis of the frameworks' performance over different hardware environments in terms of speed and scaling. As a result of this work, better insight is given towards both using and developing deep learning tools that cater to current and upcoming hardware technologies. Shayan Shams, Richard Platania, Kisung Lee, Seung-Jong Park |
ICDCS | 4 |
| 2017 | Hadoop-based replica exchange over heterogeneous distributed cyberinfrastructuresabstractSummary We present Hadoop‐based replica exchange (HaRE), a Hadoop‐based implementation of the replica exchange scheme developed primarily for replica exchange statistical temperature molecular dynamics, an example of a large‐scale, advanced sampling molecular dynamics simulation. By using Hadoop as a framework and the MapReduce model for driving replica exchange, an efficient task‐level parallelism is introduced to replica exchange statistical temperature molecular dynamics simulations. In order to demonstrate this, we investigate the performance of our application over various distributed cyberinfrastructures (DCI), including several high‐performance computing systems, our cyberinfrastructure for reconfigurable optical networks testbed, the global environment for network innovations testbed, and the CloudLab testbed. Scalability performance analysis is shown in terms of scale‐out and scale‐up over a single high‐performance computing cluster, EC2, and CloudLab and scale‐across with cyberinfrastructure for reconfigurable optical networks and global environment for network innovations. As a result, we demonstrate that HaRE is capable of efficient execution over both homogeneous and heterogeneous DCI of varying size and configuration. Contributing factors to performance are discussed in order to provide insight towards the effects of computing environment on the execution of HaRE. With these contributions, we propose that similar loosely coupled scientific applications can also take advantage of the scalable, task‐level parallelism Hadoop MapReduce provides over various DCI. Copyright © 2016 John Wiley & Sons, Ltd. Richard Platania, Shayan Shams, Chui-Hui Chiu, Nayong Kim, Joohyun Kim 0001, Seung-Jong Park |
Concurr. Comput. Pract. Exp. | 6 |
| 2016 | Lazer: Distributed memory-efficient assembly of large-scale genomesabstractGenome sequencing technology has witnessed tremendous progress in terms of throughput as well as cost per base pair, resulting in an explosion in the size of data. Consequently, typical sequence assembly tools demand a lot of processing power and memory and are unable to assemble big datasets unless run on hundreds of nodes. In this paper, we present a distributed assembler that achieves both scalability and memory efficiency by using partitioned de Bruijn graphs. By enhancing the memory-to-disk swapping and reducing the network communication in the cluster, we can assemble large sequences such as human genomes (452 GB) on just two nodes in 14.5 hours, and also scale up to 128 nodes in 23 minutes. We also assemble a synthetic wheat genome with 1.1 TB of raw reads on 8 nodes in 18.5 hours and on 128 nodes in 1.25 hours. Sayan Goswami, Arghya Kusum Das, Richard Platania, Kisung Lee, Seung-Jong Park |
IEEE BigData | 5 |
| 2016 | Towards fair and low latency next generation high speed networks: AFCD queuing
Suman Kumar 0001, Cheng Cui, Praveenkumar Kondikoppa, Chui-Hui Chiu, Seung-Jong Park |
J. Netw. Comput. Appl. | 6 |
| 2015 | Evaluating different distributed-cyber-infrastructure for data and compute intensive scientific applicationabstractScientists are increasingly using the current state of the art big data analytic software (e.g., Hadoop, Giraph, etc.) for their data-intensive applications over HPC environment. However, understanding and designing the hardware environment that these data- and compute-intensive applications require for good performance is challenging. With this motivation, we evaluated the performance of big data software over three different distributed-cyber-infrastructures, including a traditional HPC-cluster called SuperMikeII, a regular datacenter called SwatIII, and a novel MicroBrick-based hyperscale system called CeresII, using our own benchmark Parallel Genome Assembler (PGA). PGA is developed atop Hadoop and Giraph and serves as a good real-world example of a data- as well as compute-intensive workload. To evaluate the impact of both individual hardware components as well as overall organization, we changed the configuration of SwatIII in different ways. Comparing the individual impact of different hardware components (e.g., network, storage and memory) over different clusters, we observed 70% improvement in the Hadoop-workload and almost 35% improvement in the Giraph-workload in SwatIII over SuperMikeII by using SSD (thus, increasing the disk I/O rate) and scaling it up in terms of memory (which increases the caching). Then, we provide significant insight on efficient and cost-effective organization of these hardware components. Here, The MicroBrick-based CeresII prototype shows similar performance as SuperMikeII while giving more than 2-times improvement in performance/$ in the entire benchmark test. Arghya Kusum Das, Seung-Jong Park, Jae-Ki Hong, Wooseok Chang |
IEEE BigData | 2 |
| 2015 | Exploring parallelism and desynchronization of TCP over high speed networks with tiny buffers
Cheng Cui, Chui-Hui Chiu, Praveenkumar Kondikoppa, Seung-Jong Park |
Comput. Commun. | 5 |
| 2014 | DMCTCP: Desynchronized Multi-Channel TCP for high speed access networks with tiny buffersabstractThe past few years have witnessed debate on how to improve link utilization of high-speed tiny-size buffer routers. Widely argued proposals for TCP traffic to realize acceptable link capacities mandate: (i) over-provisioned core link bandwidth; and (ii) non-bursty flows; and (iii) tens of thousands of asynchronous flows. However, in high speed access networks where flows are bursty, sparse and synchronous, TCP traffic suffer severely from routers with tiny buffers. We propose a new congestion control algorithm called Desyn-chronized Multi-Channel TCP (DMCTCP) that creates a flow with multiple channels. It avoids TCP loss synchronization by desynchronizing channels, and it avoids sending rate penalties from burst losses. Over a 10 Gb/s large delay network ruled by routers with only a few dozen packets of buffers, our results show that bottleneck link utilization can reach 80% with only 100 flows. Our study is a new step towards the deployment of optical packet switching networks. Cheng Cui, Chui-Hui Chiu, Praveenkumar Kondikoppa, Seung-Jong Park |
ICCCN | 5 |
| 2014 | MapReduce based parallel suffix tree construction for human genomeabstractGenome indexing is the basis for many bioinformatics applications. Read mapping(sequence alignment) is one such application where the goal is to align millions of short reads against reference genome. Several tools are available for read mapping which rely on different indexing techniques to expedite the alignment process. However, many of these contemporary alignment programs are sequential, memory intensive and cannot be easily scaled for larger genomes. Suffix tree is one of the most widely used data structures for indexing strings (genomes). Building a scalable suffix-tree based tool is particularly challenging due to the difficulties involved in parallel construction of the suffix tree. Several suffix tree construction techniques have been proposed till date with focus on space-time tradeoff. Most of these existing works address the construction issue for uniprocessor and cannot be easily extended to utilize modern multi-processor systems. In this paper we investigate and propose a MapReduce based parallel construction of suffix tree. We demonstrate the performance of the algorithm over commodity cluster using up to 32 nodes each having 8GB of primary memory. Umesh Chandra Satish, Praveenkumar Kondikoppa, Seung-Jong Park, Manish Patil, Rahul Shah 0001 |
ICPADS | 3 |
| 2013 | An MPI-enabled MapReduce framework for molecular dynamics simulation applicationsabstractComputational technologies have been extensively investigated to be applied into many application domains. Since the presence of Hadoop, an implementation of MapReduce framework, scientists have applied it to biological sciences, chemistry, medical sciences, and other areas to efficiently process huge data sets. Although Hadoop is fault-tolerant and processes data in parallel, it does not support MPI in computing. The Map/Reduce tasks in Hadoop have to be serial, which results in inefficient scientific computations wrapped in Map/Reduce tasks. In the real world, many applications require MPI techniques due to their nature. Molecular dynamics simulation is one of them. In our research, we proposed a MPI-enabled MapReduce framework for molecular dynamics simulation applications. The MPI module added into Hadoop enables Hadoop to monitor and manage the resources of a Hadoop cluster so that computations incurred in Map tasks can be performed over available resources on the cluster in a parallel manner. We evaluated the proposed framework against a molecular dynamics simulation algorithm, RESTMD, with the application software CHARMM. The experimental results showed that the MPI-enabled framework improves computing efficiency in molecular dynamics simulation. Shuju Bai, Ebrahim Khosravi, Seung-Jong Park |
BIBM | 3 |
| 2013 | A Hadoop approach to advanced sampling algorithms in molecular dynamics simulation on cloud computingabstractCloud computing has emerged as a prevalent computing paradigm with the advantages of virtualization, scalability, fault tolerance, and a usage based pricing model. It has been widely used in various computational research fields. Replica exchange molecular dynamics (REMD) and replica exchange statistical temperature molecular dynamics (RESTMD) are two sampling algorithms in molecular dynamics. Due to the inherent parallel feature of REMD and RESTMD, it is practical to convert them into cloud computing applications. However, the performance of REMD and RESTMD on clouds has not been extensively investigated. In our work, we implemented REMD and RESTMD in cloud computing environment through Hadoop, an open source implementation of MapReduce platform, and evaluated the performance of REMD and RESTMD in cloud computing environment and in high performance computing (HPC) environment as reference. In our research, REMD and RESTMD exhibits good feasibility in cloud computing. Otherwise, cloud computing demonstrates the characteristics of virtualization, elasticity, and scalability, providing a robust environment for REMD and RESTMD. Jin Niu, Shuju Bai, Ebrahim Khosravi, Seung-Jong Park |
BIBM | 4 |
| 2013 | AFCD: An Approximated-Fair and Controlled-Delay Queuing for High Speed NetworksabstractHigh speed networks have characteristics of high bandwidth, long queuing delay, and high burstiness which make it difficult to address issues such as fairness, low queuing delay and high link utilization. Current high speed networks carry heterogeneous TCP flows which makes it even more challenging to address these issues. Since sender centric approaches do not meet these challenges, there have been several proposals to address them at router level via queue management (QM) schemes. These QM schemes have been fairly successful in addressing either fairness issues or large queuing delay but not both at the same time. We propose a new QM scheme called Approximated-Fair and Controlled-Delay (AFCD) queuing for high speed networks that aims to meet following design goals: approximated fairness, controlled low queuing delay, high link utilization and simple implementation. The design of AFCD utilizes a novel synergistic approach by forming an alliance between approximated fair queuing and controlled delay queuing. It uses very small amount of state information in sending rate estimation of flows and makes drop decision based on a target delay of individual flow. Through experimental evaluation in a 10Gbps high speed networking environment, we show AFCD meets our design goals by maintaining approximated fair share of bandwidth among flows and ensuring a controlled very low queuing delay with a comparable link utilization. Cheng Cui, Praveenkumar Kondikoppa, Chui-Hui Chiu, Seung-Jong Park |
ICCCN | 6 |
| 2012 | An evaluation of fairness among heterogeneous TCP variants over 10Gbps high-speed networksabstractSeveral high speed TCP variants are adopted by end users and therefore, heterogeneous congestion control has become the characteristic of newly emerging high-speed networks. In contrast to homogeneous TCP flows, fairness among heterogeneous TCP flows now depends on router parameters such as queue management scheme, buffer size, etc. To the best of our knowledge, this is the first evaluation of fairness among heterogeneous TCP variants over 10Gbps high-speed networks. Our evaluation scenarios for heterogeneous TCP flows consist of TCP variants with substantial presence in current Internet; therefore, TCP-SACK, CUBIC and HSTCP compete for bottleneck bandwidth. Experimental results for fairness are presented for different queue management schemes such as Drop-tail, RED, and CHOKe with varying degree of buffer sizes. We observed heterogeneous TCP flows induce more unfairness than homogeneous ones. RED and CHOKe improve fairness for large buffer sizes as compared to Drop-tail, whereas all three show similar fairness behavior for small buffer sizes. Cheng Cui, Seung-Jong Park |
LCN | 4 |
| 2010 | Congestion-aware topology controls for wireless multi-hop networks
Seung-Jong Park, Raghupathy Sivakumar |
Ad Hoc Networks | 1 |
| 2010 | A loss-event driven scalable fluid simulation method for high-speed networks
Suman Kumar 0001, Seung-Jong Park, S. Sitharama Iyengar |
Comput. Networks | 2 |
| 2010 | Measurement and performance issues of transport protocols over 10 Gbps high-speed optical networks
Suman Kumar 0001, Seung-Jong Park |
Comput. Networks | 3 |
| 2009 | On Transport Protocol Performance Measurement over 10Gbps High Speed Optical NetworksabstractWith the integration of IP and optical technology, fast optical network (of the order of 10 Gbps) has emerged to support international research cooperation such as massive scientific data transfer and next generation Internet related research. Therefore, it is important to analyze the measurement issues to guide the high speed transport protocol research for high speed optical network of the order of 10 Gbps. The objectives of this paper are as follows: i) determine the suitability of TCP parameters such as jumbo frame size, TCP sender and receiver buffer sizes; ii) evaluate TCP performance measurement tools and emulation tools over 10 Gbps high speed optical networks; and iii) compare performance of TCP variants with different metrics, such as throughput and fairness, by varying delays and randomized losses controlled with software emulators. Seung-Jong Park |
ICCCN | 3 |
| 2008 | GARUDA: Achieving Effective Reliability for Downstream Communication in Wireless Sensor NetworksabstractThere exist several applications of sensor networks where the reliability of data delivery can be critical. Although the redundancy inherent in a sensor network might increase the degree of reliability, it by no means can provide any guaranteed reliability semantics. In this paper, we consider the problem of reliable sink-to-sensors data delivery. We first identify several fundamental challenges that need to be addressed and are unique to the environment of wireless sensor networks. We then propose a scalable framework for reliable downstream data delivery that is specifically designed to both address and leverage the characteristics of the wireless sensor networks while achieving the reliability in an efficient manner. Through ns2-based simulations, we evaluate the proposed framework. Seung-Jong Park, Ramanuja Vedantham, Raghupathy Sivakumar, Ian F. Akyildiz |
IEEE Trans. Mob. Comput. | 1 |
| 2007 | Sink-to-sensors congestion control
Ramanuja Vedantham, Raghupathy Sivakumar, Seung-Jong Park |
Ad Hoc Networks | 3 |
| 2005 | Sink-to-sensors congestion controlabstractThe problem of congestion in sensor networks is significantly different from conventional ad-hoc networks and has not been studied to any great extent thus far. In this paper, we focus on providing congestion control from the sink to the sensors in a sensor field. We identify the different reasons for congestion from the sink to the sensors and show the uniqueness of the problem in sensor network environments. We propose a scalable, distributed approach that addresses congestion from the sink to the sensors in a sensor network. Through ns2 based simulations, we evaluate the proposed framework, and show that it performs significantly better over a basic approach which does not provide any congestion control. Ramanuja Vedantham, Raghupathy Sivakumar, Seung-Jong Park |
ICC | 3 |
| 2004 | A scalable approach for reliable downstream data delivery in wireless sensor networksabstractThere exist several applications of sensor networks where reliability of data delivery can be critical. While the redundancy inherent in a sensor network might increase the degree of reliability, it by no means can provide any guaranteed reliability semantics. In this paper, we consider the problem of reliable sink-to-sensors data delivery. We first identify several fundamental challenges that need to be addressed, and are unique to a wireless sensor network environment. We then propose a scalable framework for reliable downstream data delivery that is specifically designed to both address and leverage the characteristics of a wireless sensor network, while achieving the reliability in an efficient manner. Through ns2 based simulations, we evaluate the proposed framework. Seung-Jong Park, Ramanuja Vedantham, Raghupathy Sivakumar, Ian F. Akyildiz |
MobiHoc | 1 |
| 2004 | TCP performance over mobile ad hoc networks: a quantitative studyabstractAbstract In this paper, we study the performance of the transmission control protocol (TCP) over mobile ad‐hoc networks. We present a comprehensive set of simulation results and identify the key factors that impact TCP's performance over ad‐hoc networks. We use a variety of parameters including link failure detection latency, route computation latency, packet level route unavailability index, and flow level route unavailability index to capture the impact of mobility. We relate the impact of mobility on the different parameters to TCP's performance by studying the throughput, loss‐rate and retransmission timeout values at the TCP layer. We conclude from our results that existing approaches to improve TCP performance over mobile ad‐hoc networks have identified and hence focused only on a subset of the affecting factors. In the process, we identify a comprehensive set of factors influencing TCP performance. Finally, using the insights gained through the performance evaluations, we propose a framework calledAtraconsisting of three simple and easily implementable mechanisms at the MAC and routing layers to improve TCP's performance over ad‐hoc networks. We demonstrate thatAtraimproves on the throughput performance of a default protocol stack by 50%–100%. Copyright © 2004 John Wiley & Sons, Ltd. Vaidyanathan Anantharaman, Seung-Jong Park, Karthikeyan Sundaresan, Raghupathy Sivakumar |
Wirel. Commun. Mob. Comput. | 2 |
| 2002 | Load-sensitive transmission power control in wireless ad-hoc networksabstractTransmission power control in ad-hoc networks has hitherto been used only for achieving connectivity of networks. It has been implicitly assumed that the optimal throughput performance in ad-hoc networks can be achieved when using the minimum transmission power required to keep the network connected. However, in this paper we argue that such an assumption remains valid only under high node densities, which is not the characteristic of typical ad-hoc networks. Using both throughput and throughput per unit energy as the optimization criteria, we demonstrate that the optimal transmission power depends on several network characteristics such as the number of stations, the network grid area, and the traffic load. In particular, we show that the optimal power is a function of the network load for typical network scenarios. Finally, we propose two transmission power control algorithms called common power control (CPC), and independent power control (IPC) that adjust the transmission power adaptively, based on the network conditions to optimize throughput performance. Simulation results show that both adaptive power control algorithms achieve better throughput per unit energy than constant power control. Seung-Jong Park, Raghupathy Sivakumar |
GLOBECOM | 1 |
| 1994 | Complexity reduction methods for vector sum excited linear prediction coding
Sung-Joo Kim, Seung-Jong Park, Yung-Hwan Oh |
ICSLP | 2 |