Myungjin Lee

dblp:53/2649 · DBLP profile ↗
← Back
51ranked-venue papers
14as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 28 · 9 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Security and privacy · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 JITServe: SLO-aware LLM Serving with Imprecise Request Information
Wei Zhang 0044, Zhiyu Wu, Yi Mu 0005, Rui Ning, Banruo Liu, Nikhil Sarda, Myungjin Lee, Fan Lai 0001
NSDI7
2025 StitchLLM: Serving LLMs, One Block at a Time
abstract
Bodun Hu, Shuozhe Li, Saurabh Agarwal, Myungjin Lee, Akshay Jajoo, Jiamin Li, Le Xu, Geon-Woo Kim, Donghyun Kim, Hong Xu, Amy Zhang, Aditya Akella. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Bodun Hu, Shuozhe Li, Myungjin Lee, Akshay Jajoo, Jiamin Li 0002, Geon-Woo Kim, Donghyun Kim 0002, Hong Xu 0001, Amy Zhang 0001, Aditya Akella
ACL (1)4
2025 DroEdgeEM: A Drone-Edge Collaborative Emulation Platform for Emerging Situation-aware Applications
abstract
Unmanned Aerial Vehicles (UAVs) are increasingly used for surveillance, mapping, delivery, and search and rescue. These applications require continuous data streams processed through low-latency pipelines to remain situation-aware. Despite advances in hardware and software, UAV uptime is constrained by limited battery or fuel. Adding onboard compute further shortens flight time; where extra payload increases propulsion demand, and the compute itself draws power, so either UAVs carry limited compute or are treated as mobile sensors. Emerging 5G/edge deployments hold promise for supporting such compute limited drones by enabling drone-to-edge task offloading and collaboration for latency-critical processing. However, such infrastructure support is not yet widespread, complicating design and evaluation. Practical and privacy constraints also limit real-world testing, pushing much current research toward theoretical models with uncertain practicality. A key gap is the absence of realistic emulator that couples mobile UAVs and edge nodes within a dynamic virtual world (energy, location, mobility). This work envisions addressing that gap by: 1)adding emulation support on top of existing simulators to emulate necessary infrastructure (mobile drones, edge servers, charging stations) that support data streaming, processing, and control while managing time-varying state; 2)integrating user-defined plug-and-play models (e.g. energy and FPS) to realize realistic emulation behaviors of heterogeneous infrastructure with varying capabilities; and 3)exposing high-level APIs for integrations with application processing pipelines. Using a city-scale vehicle-tracking scenario, we demonstrate the emulator's usability and ease of application integration. We present this as an ongoing step toward an end-to-end platform for developing and assessing drone-edge applications, motivating and easing further research in the domain.
Summit Shrestha, Rit Muliashia, Rohit Rao, Anirudh Sarma, Alan Nussbaum, Myungjin Lee, Umakishore Ramachandran
SEC6
2025 Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
Ziyue Luo, Jia Liu 0002, Myungjin Lee, Ness Shroff
INFOCOM3
2025 SatPipe: Deterministic TCP Adaptation for Highly Dynamic LEO Satellite Networks
Ding Zhao, Xinyu Zhang 0003, Myungjin Lee
INFOCOM3
2025 Always-On, Always-Mine: Federated Recommendation Systems on Personal Home Routers
abstract
As concerns around end-user privacy continues to grow, Federated Recommendation Systems (FRS) have emerged as a privacy-preserving alternative to traditional cloud-based recommendation systems. They keep personal data "always mine"—never leaving the user's device during both training and inference of the recommendation models. However, our empirical study shows that phone-based FRS often suffer from poor recommendation accuracy. This is because, in real-world settings, only a small fraction of phones can contribute data to model updates, as they are often not charging or connected to Wi-Fi when training occurs. In this paper, we propose using "always-on" personal home routers—devices that are continuously powered and consistently network-connected—as a promising alternative to intermittently available mobile phones for FRS. Despite its potential, router-based FRS introduces two key challenges. First, routers have more limited and variable computational resources than phones, leading to slower model convergence during federated training, as the system often waits on slower but data-rich devices. Second, routers introduce an additional layer of distributed infrastructure, increasing the complexity of managing training and inference workflows for application developers. To address these challenges, we present a middleware framework for router-based FRS that includes: (1) a customized federated training strategy that maximizes data utilization from slower routers without significantly increasing training time, and (2) a declarative programming model that simplifies the integration of router-based FRS into applications. Our evaluation shows that the router-based FRS improves recommendation accuracy by up to 53% compared to phone-based systems, while our framework reduces convergence time by up to 43% relative to state-of-the-art federated training methods, with minimal programming effort and runtime overhead.
Myungjin Lee, Zheng Song 0001
Middleware2
2025 Poster: On Harnessing Idle Compute at the Edge for Foundation Model Training
abstract
Foundation model training is increasingly centralized in large cloud data centers because it demands immense compute and memory resources. Training over decentralized edge devices could democratize this ecosystem by harnessing otherwise idle compute, but prior edge-training systems fall short: they scale poorly with model size and device count, exceed per-device memory budgets, incur prohibitive collective communication, and are fragile to heterogeneous and dynamic device availability. We present Cleave, a parameter-server-centric framework that makes tensor-parallel training practical at the edge. Cleave introduces selective hybrid tensor parallelism, which finely shards GEMM-dominated training operations into memory-feasible sub-tasks while avoiding peer-to-peer collectives that become bottlenecks on asymmetric edge links. A cost model guides device selection and shard placement to mitigate stragglers and rapidly adapt to churn. Across OPT and Llama2 models, Cleave matches cloud GPU training efficiency while scaling to thousands of devices. It supports up to 8× more devices than prior edge approaches, reduces per-batch training time by up to 10×, and achieves 100× faster recovery from device failures.
Leyang Xue, Meghana Madhyastha, Myungjin Lee, Amos J. Storkey, Randal C. Burns, Mahesh K. Marina
MobiCom3
2025 Federated Deep Reinforcement Learning-Driven O-RAN for Automatic Multirobot Reconfiguration
abstract
The rapid evolution of Industry 4.0 has led to the emergence of smart factories, where multirobot system autonomously operates to enhance productivity, reduce operational costs, and improve system adaptability. However, maintaining reliable and efficient network operations in these dynamic and complex environments requires advanced automation mechanisms. This study presents a zero-touch network platform that integrates a hierarchical Open Radio Access Network (O-RAN) architecture, enabling the seamless incorporation of advanced machine learning algorithms and dynamic management of communication and computational resources, while ensuring uninterrupted connectivity with multirobot system. Leveraging this adaptability, the platform utilizes federated deep reinforcement learning (FedDRL) to enable distributed decision-making across multiple learning agents, facilitating the adaptive parameter reconfiguration of transmitters (i.e., multirobot system) to optimize long-term system throughput and transmission energy efficiency. Simulation results demonstrate that within the proposed O-RAN-enabled zero-touch network platform, FedDRL achieves a 12% increase in system throughput, a 32% improvement in normalized average transmission energy efficiency, and a 28% reduction in average transmission energy consumption compared to baseline methods such as independent DRL.
Myungjin Lee, Shao-Yu Lien, Suresh Subramaniam 0001, Motoharu Matsuura, Hiroshi Hasegawa, Shih-Chun Lin 0002
NOMS2
2025 MaverIQ: Fingerprint-Guided Extrapolation and Fragmentation-Aware Layering for Intent-Based LLM Serving
abstract
Large Language Models (LLMs) are becoming ubiquitous across industries, where applications demand they fulfill diverse user intents. However, developers currently face the challenge of manually exploring numerous deployment configurations—combinations of parallelism and compression techniques that impact resource usage, latency, cost, and accuracy—to meet these intents. Previous works automate configuration selection and deployment, but, they (a) rely on expensive profiling, and (b) suboptimally utilize the fragmented resource availability in multi-tenant GPU clusters, inflating operational costs for the provider. Moreover, none of these solutions tailors deployment configuration decisions to diverse user intents.
Dimitrios Liakopoulos, Prasoon Sinha, Tianrui Hu, Myungjin Lee, Neeraja J. Yadwadkar
SC4
2025 Optimizing Handover Decisions in Multi-Connectivity Enabled Terrestrial-Satellite Integrated Networks: A Deep Reinforcement Learning Approach
abstract
The integration of 5G terrestrial networks with Low Earth Orbit (LEO) satellites has the potential to provide seamless global connectivity and enhanced service quality, particularly in regions with limited terrestrial infrastructure such as rural areas. Furthermore, the incorporation of Multi-connectivity (MC) enables user equipment (UEs) to maintain simultaneous connections with both terrestrial 5G base stations and LEO satellites, improving system reliability. However, the high mobility of LEO satellites and the dynamic behavior of UEs present significant challenges, particularly in handover decision-making which can adversely impact system throughput, and quality of service (QoS). To address these challenges, we propose a novel deep reinforcement learning-based approach that integrates online Random Ensemble Mixture and Dual Experience Replay into a Dueling Double Deep Q-Network architecture. This proposed scheme intelligently optimizes handover decisions in MC-enabled terrestrial-satellite networks, improving decision accuracy in highly dynamic scenarios. Simulation results demonstrate substantial gains in system throughput, reduced system delay, average handover reduction, and increased transmission success probability, setting new performance benchmarks for integrated terrestrial-satellite networks while adhering to diverse QoS requirements.
Myungjin Lee, Suresh Subramaniam 0001, Motoharu Matsuura, Hiroshi Hasegawa, Shih-Chun Lin 0002
WCNC2
2025 Enhancing Network Traffic Analysis in O-RAN Enabled Next-Generation Networks Through Federated Multi-Task Learning
abstract
The distributed and disaggregated architecture of next-generation (NextG) networks, including 6G has sparked growing interest in federated learning (FL) as a strategy for enabling privacy-preserving collaborative network traffic analysis at the edge. However, FL encounters significant challenges due to data heterogeneity driven by diverse data distributions across edge nodes, and the scarcity of labeled data further worsened by the time-intensive process of data labeling. Although a few studies have addressed these challenges in network traffic analysis tasks using Multi-Task Learning (MTL), existing approaches pre-dominantly focus on single-task FL, centralized model solutions and overlook the integration of MTL in NextG networks. To bridge this gap, we propose O-FedMTL, a novel framework that combines FL with MTL to enable cooperative traffic analysis within an Open Radio Access Network (O-RAN) environment in NextG networks. MTL enhances FL by mitigating the issues of data heterogeneity and labeled data scarcity through shared knowledge derived from multiple interconnected traffic analysis tasks, i.e., traffic classification, flow duration analysis, and bandwidth estimation. Additionally, MTL offers significant benefits by reducing energy consumption and computation costs at the edge through the simultaneous processing of these tasks within a single model. Extensive experimental results demonstrate that O-FedMTL achieves the target global accuracy for traffic classification, flow duration analysis, and bandwidth estimation with 20, 12, and 23 fewer global communication rounds, respectively, compared to the baseline federated averaging. Additionally, O-FedMTL reduces computation costs by 43% compared to the baseline-combined.
Myungjin Lee, Suresh Subramaniam 0001, Motoharu Matsuura, Hiroshi Hasegawa, Shih-Chun Lin 0002
WCNC2
2024 Enhancing Large Language Models through Transforming Reasoning Problems into Classification Tasks
abstract
In this paper, we introduce a novel approach for enhancing the reasoning capabilities of large language models (LLMs) for constraint satisfaction problems (CSPs), by converting reasoning problems into classification tasks. Our method leverages the LLM’s ability to decide when to call a function from a set of logical-linguistic primitives, each of which can interact with a local “scratchpad” memory and logical inference engine. Invocation of these primitives in the correct order writes the constraints to the scratchpad memory and enables the logical engine to verifiably solve the problem. We additionally propose a formal framework for exploring the “linguistic” hardness of CSP reasoning-problems for LLMs. Our experimental results demonstrate that under our proposed method, tasks with significant computational hardness can be converted to a form that is easier for LLMs to solve and yields a 40% improvement over baselines. This opens up new avenues for future research into hybrid cognitive models that integrate symbolic and neural approaches.
Tarun Raheja, Raunak Sinha, Advit Deepak, Will Healy, Jayanth Srinivasa, Myungjin Lee, Ramana Rao Kompella
LREC/COLING6
2024 DεpS: Delayed ε-Shrinking for Faster Once-for-All Training
Aditya Annavajjala, Alind Khare, Animesh Agrawal, Igor Fedorov, Hugo Latapie, Myungjin Lee, Alexey Tumanov
ECCV (89)6
2024 SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-device Inference
Alind Khare, Animesh Agrawal, Aditya Annavajjala, Payman Behnam, Myungjin Lee, Hugo Latapie, Alexey Tumanov
ECCV (79)5
2024 Enabling LEO Satellite Vertical Handover for Massive 6G IoT Random Access
abstract
Integrated terrestrial and satellite networks will allow providing ubiquitous access and global connectivity to remotely deployed battery-activated (machine-to-machine or IoT) sensors or handset devices with messaging/voice capacities via satellite, fostering a series of applications such as intelligent transportation, coastal monitoring, and smart agriculture. In addition, the agile measurement of sensors enables autonomous driving and fast disaster recovery. These applications could generate massive uplink connections in a short period that the network must handle to avoid failures and outages. This paper covers system aspects and refines the 3rd Generation Partnership Project (3GPP)-based solutions to support non-terrestrial networks and make their performance comparable to that of terrestrial networks for providing IoT communications. A general vertical handover framework is introduced for this integrated network to choose the most suitable access technology for a given service. Then, the whole system increases around 100% the successful transmission probability in heavy-loaded scenarios with a slight increment in the number of retransmissions and transmission delays compared to that of single connectivity solutions. Furthermore, we found that by fine-tuning network parameter configurations, the energy consumption and transmission delay can be decreased while providing reliable communications that will benefit small-sized IoT devices with limited resources.
Luis Tello-Oquendo, Shih-Chun Lin 0002, Myungjin Lee
ICC5
2024 A Federated Stochastic Multi-level Compositional Minimax Algorithm for Deep AUC Maximization
abstract
AUC maximization is an effective approach to address the imbalanced data classification problem in federated learning. In the past few years, a couple of federated AUC maximization approaches have been developed based on the minimax optimization. However, directly solving a minimax optimization problem to maximize the AUC score cannot achieve satisfactory performance. To address this issue, we propose to maximize AUC via optimizing a federated multi-level compositional minimax problem. Specifically, we develop a novel federated multi-level compositional minimax algorithm with rigorous theoretical guarantees to solve this new learning paradigm in both algorithmic design and theoretical analysis. To the best of our knowledge, this is the first work studying the multi-level minimax optimization problem. Additionally, extensive empirical evaluations confirm the efficacy of our proposed approach.
Xinwen Zhang, Ali Payani, Myungjin Lee, Richard Souvenir, Hongchang Gao
ICML3
2024 IaC-Eval: A Code Generation Benchmark for Cloud Infrastructure-as-Code Programs
abstract
Infrastructure-as-Code (IaC), an important component of cloud computing, allows the definition of cloud infrastructure in high-level programs. However, developing IaC programs is challenging, complicated by factors that include the burgeoning complexity of the cloud ecosystem (e.g., diversity of cloud services and workloads), and the relative scarcity of IaC-specific code examples and public repositories. While large language models (LLMs) have shown promise in general code generation and could potentially aid in IaC development, no benchmarks currently exist for evaluating their ability to generate IaC code. We present IaC-Eval, a first step in this research direction. IaC-Eval's dataset includes 458 human-curated scenarios covering a wide range of popular AWS services, at varying difficulty levels. Each scenario mainly comprises a natural language IaC problem description and an infrastructure intent specification. The former is fed as user input to the LLM, while the latter is a general notion used to verify if the generated IaC program conforms to the user's intent; by making explicit the problem's requirements that can encompass various cloud services, resources and internal infrastructure details. Our in-depth evaluation shows that contemporary LLMs perform poorly on IaC-Eval, with the top-performing model, GPT-4, obtaining a pass@1 accuracy of 19.36%. In contrast, it scores 86.6% on EvalPlus, a popular Python code generation benchmark, highlighting a need for advancements in this domain. We open-source the IaC-Eval dataset and evaluation framework at https://github.com/autoiac-project/iac-eval to enable future research on LLM-based IaC code generation.
Patrick Tser Jern Kon, Yiming Qiu 0001, Weijun Fan, Owen Park, George Elengikal, Yuxin Kang, Ang Chen 0001, Mosharaf Chowdhury, Myungjin Lee, Xinyu Wang 0006
NeurIPS13
2024 MetaFL: Privacy-preserving User Authentication in Virtual Reality with Federated Learning
abstract
The increasing popularity of virtual reality (VR) has stressed the importance of authenticating VR users while preserving their privacy. Behavioral biometrics, owing to their robustness and ease of collection, compared to traditional modes such as passwords, have become a favored authentication choice. While current approaches that utilize behavioral biometrics to train classifiers for authentication yield promising accuracy, they cause privacy breaches by sharing sensitive data with a server to train a central model. In this paper, we present MetaFL, a first-of-its-kind privacy-preserving VR authentication framework that leverages federated learning (FL) on multi-modal motion data. The design of MetaFL is motivated by our key insight that various modalities of motion data uniquely affect authentication performance for individual users and among different users. It is attributed to the fundamental challenge of privacy-preserving user authentication: users can access only their own data with limited global knowledge. To tackle this issue, MetaFL judiciously selects the most suitable modalities for each user, which is decomposed into within-user ordering and between-user selection to eliminate the complex interplay between various conflicting factors. Moreover, we develop a personalized strategy to initialize FL models, further improving authentication accuracy. Our extensive performance evaluation on six public datasets shows that MetaFL outperforms state-of-the-art FL-based models (e.g., 17--28% higher authentication accuracy), and its accuracy gap with the non-privacy-preserving central model is small (i.e., only <2%).
Ruizhi Cheng, Yuetong Wu, Ashish Kundu, Hugo Latapie, Myungjin Lee, Songqing Chen, Bo Han 0001
SenSys5
2024 Boosting Collaborative Vehicular Perception on the Edge with Vehicle-to-Vehicle Communication
abstract
Collaborative Vehicular Perception (CVP) enables connected and autonomous vehicles (CAVs) to cooperatively extend their views through wirelessly sharing their sensor data. Existing CVP systems employ either a vehicle-to-vehicle (V2V) or vehicle-to-infrastructure (V2I) view exchange paradigm. In this paper, we advocate a hybrid CVP design: our developed system, Harbor, employs V2I as its fundamental underlying framework, and opportunistically employs V2V to boost the performance. In Harbor, vehicles (helpers) may serve as relays to assist other vehicles (helpees) in reaching an edge node, which performs sensor data merging to produce the extended view. We judiciously partition the workload between the edge and vehicles, develop a robust helper-helpee assignment model, and solve it efficiently at runtime. We conduct both real-world tests and large-scale emulation experiments using two prevailing CAV applications: drivable space detection and object detection. Our real-world evaluation conducted at one of the world's first purpose-built autonomous driving testbeds demonstrates that Harbor outperforms state-of-the-art V2V- or V2I-only CVP schemes by up to 36% in detection accuracy, resulting in significantly fewer collisions under dangerous driving scenarios.
Ruiyang Zhu, Xiao Zhu 0001, Anlan Zhang, Xumiao Zhang, Feng Qian 0001, Hang Qiu 0001, Z. Morley Mao, Myungjin Lee
SenSys9
2024 Dissecting Carrier Aggregation in 5G Networks: Measurement, QoE Implications and Prediction
abstract
By aggregating multiple channels, Carrier Aggregation (CA) is an important technology for boosting cellular network bandwidth. Given diverse radio bands made available in 5G networks, CA plays a particularly critical role in achieving the goal of multi-Gbps throughput performance. In this paper, we carry out a timely comprehensive measurement study of CA deployment in commercial 5G networks (as well as 4G networks). We identify the key factors that influence whether CA is deployed and when, as well as which band combinations are used. Thus, we reveal the challenges posed by CA in 5G performance analysis and prediction as well as their implications in application quality-of-experience (QoE). We argue for and develop a novel CA-aware deep learning framework, dubbed Prism5G, which explicitly accounts for the complexity introduced by CA to more effectively predict 5G network throughput performance. Through extensive evaluations, we demonstrate the superiority of Prism5G over existing throughput prediction algorithms. Prism5G improves 5G throughput prediction accuracy by over 14% on average and a maximum of 22%. Using two use cases as examples, we further illustrate how Prism5G can aid applications in optimizing QoE performance.
Wei Ye 0009, Steven Sleder, Anlan Zhang, Udhaya Kumar Dayalan, Ahmad Hassan 0004, Rostand A. K. Fezeu, Akshay Jajoo, Myungjin Lee, Eman Ramadan, Feng Qian 0001, Zhi-Li Zhang
SIGCOMM9
2024 Adaptive Deep Neural Network Inference Optimization with EENet
abstract
Well-trained deep neural networks (DNNs) treat all test samples equally during prediction. Adaptive DNN inference with early exiting leverages the observation that some test examples can be easier to predict than others. This paper presents EENet, a novel early-exiting scheduling framework for multi-exit DNN models. Instead of having every sample go through all DNN layers during prediction, EENet learns an early exit scheduler, which can intelligently terminate the inference earlier for certain predictions, which the model has high confidence of early exit. As opposed to previous early-exiting solutions with heuristics-based methods, our EENet framework optimizes an early-exiting policy to maximize model accuracy while satisfying the given per-sample average inference budget. Extensive experiments are conducted on four computer vision datasets (CIFAR-10, CIFAR-100, ImageNet, Cityscapes) and two NLP datasets (SST-2, AgNews). The results demonstrate that the adaptive inference by EENet can outperform the representative existing early exit techniques. We also perform a detailed visualization analysis of the comparison results to interpret the benefits of EENet.
Fatih Ilhan, Ka-Ho Chow 0001, Sihao Hu, Tiansheng Huang, Selim F. Tekin, Wenqi Wei 0001, Yanzhao Wu 0001, Myungjin Lee, Ramana Rao Kompella, Hugo Latapie, Gaowen Liu, Ling Liu 0001
WACV8
2023 Flame: Simplifying Topology Extension in Federated Learning
abstract
Distributed machine learning approaches, including a broad class of federated learning (FL) techniques, present a number of benefits when deploying machine learning applications over widely distributed infrastructures. The benefits are highly dependent on the details of the underlying machine learning topology, which specifies the functionality executed by the participating nodes, their dependencies and interconnections. Current systems lack the flexibility and extensibility necessary to customize the topology of a machine learning deployment. We present Flame, a new system that provides flexibility of the topology configuration of distributed FL applications around the specifics of a particular deployment context, and is easily extensible to support new FL architectures. Flame achieves this via a new high-level abstraction Topology Abstraction Graphs (TAGs). TAGs decouple the ML application logic from the underlying deployment details, making it possible to specialize the application deployment with reduced development effort. Flame is released as an open source project, and its flexibility and extensibility support a variety of topologies and mechanisms, and can facilitate the development of new FL methodologies.
Harshit Daga, Jaemin Shin 0005, Dhruv Garg, Ada Gavrilovska, Myungjin Lee, Ramana Rao Kompella
SoCC5
2023 An In-Depth Measurement Analysis of 5G mmWave PHY Latency and Its Impact on End-to-End Delay
Rostand A. K. Fezeu, Eman Ramadan, Wei Ye 0009, Benjamin Minneci, Jack Xie, Arvind Narayanan, Ahmad Hassan 0004, Feng Qian 0001, Zhi-Li Zhang, Jaideep Chandrashekar, Myungjin Lee
PAM11
2022 GLYCO: a tool to quantify glycan shielding of glycosylated proteins
abstract
MOTIVATION: Glycans play important roles in protein folding and cell-cell interactions-and, furthermore, glycosylation of protein antigens can dramatically impact immune responses. While there have been attempts to quantify the glycan shielding or coverage of a protein surface, none of the publicly available tools analyzes glycan shielding computationally at an atomistic level. RESULTS: Here, we developed an in silico approach, GLYCO (GLYcan COverage), to quantify the glycan shielding of a protein surface. The software provides insights into glycan-dense/sparse regions of the entire protein surface or a subset of the protein surface. GLYCO calculates glycan shielding from a single coordinate file or from multiple coordinate files, for instance, as obtained from molecular dynamics simulations or by nuclear magnetic resonance spectroscopy structure determination, enabling analysis of glycan dynamics. Overall, GLYCO provides fundamental insights into the glycan shielding of glycosylated proteins. AVAILABILITY AND IMPLEMENTATION: GLYCO is freely available at GitHub (https://github.com/myungjinlee/GLYCO). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Myungjin Lee, Mateo Reveiz, Reda Rawi, Peter D. Kwong, Gwo-Yu Chuang
Bioinform.1
2020 A solution to MPTCP's inefficiencies under the incast problem for Data Center Networks
Morteza Kheirkhah, Myungjin Lee
Comput. Commun.2
2019 AMP: An Adaptive Multipath TCP for Data Center Networks
abstract
MPTCP and its ECN-capable variants such as XMP and DCM have recently been introduced to effectively exploit the path diversity of modern data center networks (DCNs). Although these multipath schemes improve overall network throughput compared to single-path schemes due to their fast, host-based, load balancing ability, they failed to address the following two problems: TCP incast and last hop unfairness. Firstly, these mechanisms cause frequent TCP incast collapses when used for workloads with a many-to-one communication pattern, commonly found in DCNs. Secondly, the last hop unfairness problem severely violates network fairness as single-path flows achieve 2-5 times less throughput than multipath flows. To effectively tackle these problems, we propose the Adaptive MultiPath (AMP) congestion control mechanism that quickly detects the onset of these problems and transforms its multipath flow into a single-path flow. Once these problems disappear, AMP safely reverses this transformation and continues data transmission via multiple paths. Our evaluation results under a diverse set of scenarios in a large-scale fat-tree topology demonstrate that AMP is robust to the TCP incast problem and improves network fairness between multipath and single-path flows significantly with no performance loss.
Morteza Kheirkhah, Myungjin Lee
Networking2
2018 Fault Localization in Large-Scale Network Policy Deployment
abstract
The recent advances in network management automation and Software-Defined Networking (SDN) facilitate network policy management tasks. At the same time, these new technologies create a new mode of failure in the management cycle itself. Network policies are presented in an abstract model at a centralized controller and deployed as low-level rules across network devices. Thus, any software and hardware element in that cycle can be a potential cause of underlying network problems. In this paper, we present and solve a network policy fault localization problem that arises in operating policy management frameworks for a production network. We formulate our problem via risk modeling and propose a greedy algorithm that quickly localizes faulty policy objects in the network policy. We then design and develop SCOUT-a fully-automated system that produces faulty policy objects and further pinpoints physical-level failures which made the objects faulty. Evaluation results using a real testbed and extensive simulations demonstrate that SCOUT detects faulty objects with small false positives and false negatives.
Praveen Tammana, Chandra Nagarajan, Pavan Mamillapalli, Ramana Rao Kompella, Myungjin Lee
ICDCS5
2018 ApproxIoT: Approximate Analytics for Edge Computing
abstract
IoT-enabled devices continue to generate a massive amount of data. Transforming this continuously arriving raw data into timely insights is critical for many modern online services. For such settings, the traditional form of data analytics over the entire dataset would be prohibitively limiting and expensive for supporting real-time stream analytics. In this work, we make a case for approximate computing for data analytics in IoT settings. Approximate computing aims for efficient execution of workflows where an approximate output is sufficient instead of the exact output. The idea behind approximate computing is to compute over a representative sample instead of the entire input dataset. Thus, approximate computing- based on the chosen sample size - can make a systematic tradeoff between the output accuracy and computation efficiency. This motivated the design of APPROXIOT- a data analytics system for approximate computing in IoT. To realize this idea, we designed an online hierarchical stratified reservoir sampling algorithm that uses edge computing resources to produce approximate output with rigorous error bounds. To showcase the effectiveness of our algorithm, we implemented APPROXIOT based on Apache Kafka and evaluated its effectiveness using a set of microbenchmarks and real-world case studies. Our results show that APPROXIOT achieves a speedup 1:3×-9:9× with varying sampling fraction of 80% to 10% compared to simple random sampling.
Zhenyu Wen, Do Le Quoc, Pramod Bhatotia, Ruichuan Chen, Myungjin Lee
ICDCS5
2018 Distributed Network Monitoring and Debugging with SwitchPointer
Praveen Tammana, Rachit Agarwal 0001, Myungjin Lee
NSDI3
2018 Formal Analysis of Sneak-Peek: A Data Centre Attack and Its Mitigations
Wei Chen 0023, Yuhui Lin, Vashti Galpin, Vivek Nigam, Myungjin Lee, David Aspinall 0001
SEC5
2017 Randomizing TCP payload size for TCP fairness in data center networks
Soojeon Lee, Dongman Lee, Myungjin Lee, Hyungsoo Jung 0001, Byoung-Sun Lee
Comput. Networks3
2016 Simplifying Datacenter Network Debugging with PathDump
Praveen Tammana, Rachit Agarwal 0001, Myungjin Lee
OSDI3
2015 TCPRand: Randomizing TCP payload size for TCP fairness in data center networks
abstract
As many-to-one traffic patterns prevail in data center networks, TCP flows often suffer from severe unfairness in sharing bottleneck bandwidth, which is known as the TCP outcast problem. The cause of the TCP outcast problem is the bursty packet losses by a drop-tail queue that triggers TCP timeouts and leads to decreasing the congestion window. This paper proposes TCPRand, a transport layer solution to TCP outcast. The main idea of TCPRand is the randomization of TCP payload size, which breaks synchronized packet arrivals between flows from different input ports. We investigate how TCPRand reduces consecutive packet drops and demonstrate various benefits of TCPRand with extensive experiments and ns-3 simulation. Our evaluation results show that TCPRand guarantees the superior enhancement of TCP fairness with negligible overheads in all of our test cases.
Soojeon Lee, Myungjin Lee, Dongman Lee, Hyungsoo Jung 0001, Byoung-Sun Lee
INFOCOM2
2015 A flow measurement architecture to preserve application structure
Myungjin Lee, Mohammad Y. Hajjat, Ramana Rao Kompella, Sanjay G. Rao
Comput. Networks1
2014 On the impact of 802.11n frame aggregation on end-to-end available bandwidth estimation
abstract
We consider for the first time available bandwidth estimation (ABE) in the context of 802.11n, which is fast replacing the legacy 802.11a/b/g networks. We experimentally show that the frame aggregation (FA) feature of 802.11n is the dominant one among 802.11n features affecting the ABE. Using an indoor 802.11n wireless testbed, we compare three ABE tools (WBest, DietTopp and pathChirp) in various cross-traffic scenarios. We find that FA significantly hurts the accuracy of all ABE tools; DietTopp and pathChirp are relatively more robust than WBest. Because faster available bandwidth estimation and less intrusiveness are desirable properties of any ABE tool and WBest satisfies them relatively better than the other two tools, we conduct an in-depth investigation into the harmful effect of FA on ABE using WBest. This in turn led us to come up with two key design principles to counter FA effects: (1) treating aggregated probes as one jumbo probe; and (2) generating a larger number of probes. We then develop an enhanced version of WBest termed WBest+ that incorporates these principles. Our evaluation shows that the new version is effective in achieving accurate ABE in the presence of FA.
Arsham Farshad, Myungjin Lee, Mahesh K. Marina
SECON2
2014 Application-driven bandwidth guarantees in datacenters
abstract
Providing bandwidth guarantees to specific applications is becoming increasingly important as applications compete for shared cloud network resources. We present CloudMirror, a solution that provides bandwidth guarantees to cloud applications based on a new network abstraction and workload placement algorithm. An effective network abstraction should enable applications to easily and accurately specify their requirements, while simultaneously enabling the infrastructure to provision resources efficiently for deployed applications. Prior research has approached the bandwidth guarantee specification by using abstractions that resemble physical network topologies. We present a contrasting approach of deriving a network abstraction based on application communication structure, called Tenant Application Graph or TAG. CloudMirror also incorporates a new workload placement algorithm that efficiently meets bandwidth requirements specified by TAGs while factoring in high availability considerations. Extensive simulations using real application traces and datacenter topologies show that CloudMirror can handle 40% more bandwidth demand than the state of the art (e.g., the Oktopus system), while improving high availability from 20% to 70%.
Jeongkeun Lee, Yoshio Turner, Myungjin Lee, Lucian Popa 0002, Sujata Banerjee, Joon-Myung Kang, Puneet Sharma 0001
SIGCOMM3
2014 FineComb: Measuring Microscopic Latency and Loss in the Presence of Reordering
abstract
Modern stock trading and cluster applications require microsecond latencies and almost no losses in data centers. This paper introduces an algorithm called FineComb that can obtain fine-grain end-to-end loss and latency measurements between edge routers in these networks. Such a mechanism can allow managers to distinguish between latencies and loss singularities caused by servers and those caused by the network. Compared to prior work, such as Lossy Difference Aggregator (LDA), which focused on switch-level latency measurements, the requirement of end-to-end latency measurements introduces the challenge of reordering that occurs commonly in IP networks due to churn. The problem is even more acute in switches across data center networks that employ multipath routing algorithms to exploit the inherent path diversity. Without proper care, a loss estimation algorithm can confound loss and reordering; furthermore, any attempt to aggregate delay estimates in the presence of reordering results in severe errors. FineComb deals with these problems using order-agnostic packet digests and a simple new idea we call stash recovery. Our evaluation demonstrates that FineComb is orders of magnitude more accurate than LDA in loss and delay estimates in the presence of reordering.
Myungjin Lee, Sharon Goldberg, Ramana Rao Kompella, George Varghese
IEEE/ACM Trans. Netw.1
2013 High-Fidelity Per-Flow Delay Measurements With Reference Latency Interpolation
abstract
New applications such as soft real-time data center applications, algorithmic trading, and high-performance computing require extremely low latency (in microseconds) from networks. Network operators today lack sufficient fine-grain measurement tools to detect, localize, and repair delay spikes that cause application service level agreement (SLA) violations. A recently proposed solution called LDA provides a scalable way to obtain latency, but only provides aggregate measurements. However, debugging application-specific problems requires per-flow measurements since different flows may exhibit significantly different characteristics even when they are traversing the same link. To enable fine-grained per-flow measurements in routers, we propose a new scalable architecture called reference latency interpolation (RLI) that is based on our observation that packets potentially belonging to different flows that are closely spaced to each other exhibit similar delay properties. In our evaluation using simulations over real traces, we show that while having small overhead, RLI achieves a median relative error of 12% and one to two orders of magnitude higher accuracy than previous per-flow measurement solutions. We also observe RLI achieves as high accuracy as LDA in aggregate latency estimation, and RLI outperforms LDA in standard deviation estimation.
Myungjin Lee, Nick G. Duffield, Ramana Rao Kompella
IEEE/ACM Trans. Netw.1
2012 MAPLE: a scalable architecture for maintaining packet latency measurements
abstract
Latency has become an important metric for network monitoring since the emergence of new latency-sensitive applications (e.g., algorithmic trading and high-performance computing). To satisfy the need, researchers have proposed new architectures such as LDA and RLI that can provide fine-grained latency measurements. However, these architectures are fundamentally ossified in their design as they are designed to provide only a specific pre-configured aggregate measurement---either average latency across all packets (LDA) or per-flow latency measurements (RLI). Network operators, however, need latency measurements at both finer (e.g., packet) as well as flexible (e.g., flow subsets) levels of granularity. To bridge this gap, we propose an architecture called MAPLE that essentially stores packet-level latencies in routers and allows network operators to query the latency of arbitrary traffic sub-populations. MAPLE is built using scalable data structures with small storage needs (uses only 12.8 bits/packet), and uses a novel mechanism to reduce the query bandwidth significantly (by a factor of 17 compared to the naive method of sending packet queries individually).
Myungjin Lee, Nick G. Duffield, Ramana Rao Kompella
Internet Measurement Conference1
2012 A scalable architecture for maintaining packet latency measurements
abstract
Latency has become an important metric for network monitoring since the emergence of new latency-sensitive applications (e.g., algorithmic trading and high-performance computing). In this paper, to provide latency measurements at both finer (e.g., packet) as well as flexible (e.g., flow subsets) levels of granularity, we propose an architecture called MAPLE that essentially stores packet-level latencies in routers and allows network operators to query the latency of arbitrary traffic sub-populations. MAPLE is built using a scalable data structure called SVBF with small storage needs.
Myungjin Lee, Nick G. Duffield, Ramana Rao Kompella
SIGMETRICS1
2012 Searching and ranking method of relevant resources by user intention on the Semantic Web
Myungjin Lee, Wooju Kim, Sangun Park 0001
Expert Syst. Appl.1
2012 Opportunistic Flow-Level Latency Estimation Using Consistent NetFlow
abstract
The inherent measurement support in routers (SNMP counters or NetFlow) is not sufficient to diagnose performance problems in IP networks, especially for flow-specific problems where the aggregate behavior within a router appears normal. Tomographic approaches to detect the location of such problems are not feasible in such cases as active probes can only catch aggregate characteristics. To address this problem, in this paper, we propose a Consistent NetFlow (CNF) architecture for measuring per-flow delay measurements within routers. CNF utilizes the existing NetFlow architecture that already reports the first and last timestamps per flow, and it proposes hash-based sampling to ensure that two adjacent routers record the same flows. We devise a novel Multiflow estimator that approximates the intermediate delay samples from other background flows to significantly improve the per-flow latency estimates compared to the naive estimator that only uses actual flow samples. In our experiments using real backbone traces and realistic delay models, we show that the Multiflow estimator is accurate with a median relative error of less than 20% for flows of size greater than 100 packets. We also show that Multiflow estimator performs two to three times better than a prior approach based on trajectory sampling at an equivalent packet sampling rate.
Myungjin Lee, Nick G. Duffield, Ramana Rao Kompella
IEEE/ACM Trans. Netw.1
2011 Scheduling in mapreduce-like systems for fast completion time
abstract
Large-scale data processing needs of enterprises today are primarily met with distributed and parallel computing in data centers. MapReduce has emerged as an important programming model for these environments. Since today's data centers run many MapReduce jobs in parallel, it is important to find a good scheduling algorithm that can optimize the completion times of these jobs. While several recent papers focused on optimizing the scheduler, there exists very little theoretical understanding of the scheduling problem in the context of MapReduce. In this paper, we seek to address this problem by first presenting a simplified abstraction of the MapReduce scheduling problem, and then formulate the scheduling problem as an optimization problem.We devise various online and offline algorithms to arrive at a good ordering of jobs to minimize the overall job completion times. Since optimal solutions are hard to compute (NP-hard), we propose approximation algorithms that work within a factor of 3 of the optimal. Using simulations, we also compare our online algorithm with standard scheduling strategies such as FIFO, Shortest Job First and show that our algorithm consistently outperforms these across different job distributions.
Hyunseok Chang, Murali S. Kodialam, Ramana Rao Kompella, T. V. Lakshman, Myungjin Lee, Sarit Mukherjee
INFOCOM5
2011 RelSamp: Preserving application structure in sampled flow measurements
abstract
The Internet has significantly evolved in the number and variety of applications. Network operators need mechanisms to constantly monitor and study these applications. Given modern applications routinely consist of several flows, potentially to many different destinations, existing measurement approaches such as Sampled NetFlow sample only a few flows per application session. To address this issue, in this paper, we introduce RelSamp architecture that implements the notion of related sampling where flows that are part of the same application session are given higher probability. In our evaluation using real traces, we show that RelSamp achieves 5-10x more flows per application session compared to Sampled NetFlow for the same effective number of sampled packets. We also show that behavioral and statistical classification approaches such as BLINC, SVM and C4.5 achieve up to 50% better classification accuracy compared to Sampled NetFlow, while not breaking existing management tasks such as volume estimation.
Myungjin Lee, Mohammad Y. Hajjat, Ramana Rao Kompella, Sanjay G. Rao
INFOCOM1
2011 Using NetMagic to observe fine-grained per-flow latency measurements
abstract
We introduce NetMagic to demonstrate the efficacy of RLI architecture RLI for the fine-grained per-flow latency measurements. In this demo, the main function of RLI is implemented in NetMagic, which is the key component of our experimental network comprising several computers and switches. We are going to show how NetMagic can provide rapid implementation and evaluation of RLI architecture that is difficult with commercial switch or router platforms. In the demo, the estimated fine-grained per-flow latency by RLI is monitored and dynamically presented. Further, the true latency with a resolution of 8ns is also provided by NetMagic for the evaluation. The efficacy of RLI architecture can be observed in a real-time fashion by the difference between estimated latencies and true ones.
Tao Li 0008, Zhigang Sun 0002, Chunbo Jia, Myungjin Lee
SIGCOMM5
2011 Fine-grained latency and loss measurements in the presence of reordering
abstract
Modern trading and cluster applications require microsecond latencies and almost no losses in data centers. This paper introduces an algorithm called FineComb that can estimate fine-grain end-to-end loss and latency measurements between edge routers in these data center networks. Such a mechanism can allow managers to distinguish between latencies and loss singularities caused by servers and those caused by the network. Compared to prior work, such as Lossy Difference Aggregator (LDA), that focused on switch-level latency measurements, the requirement of end-to-end latency measurements introduces the challenge of reordering that occurs commonly in IP networks due to churn. The problem is even more acute in switches across data center networks that employ multipath routing algorithms to exploit the inherent path diversity. Without proper care, a loss estimation algorithm can confound loss and reordering; further, any attempt to aggregate delay estimates in the presence of reordering results in severe errors. FineComb deals with these problems using order-agnostic packet digests and a simple new idea we call stash recovery. Our evaluation demonstrates that FineComb can provide orders of magnitude better accuracy in loss and delay estimates in the presence of reordering compared to LDA.
Myungjin Lee, Sharon Goldberg, Ramana Rao Kompella, George Varghese
SIGMETRICS1
2010 Two Samples are Enough: Opportunistic Flow-level Latency Estimation using NetFlow
abstract
The inherent support in routers (SNMP counters or NetFlow) is not sufficient to diagnose performance problems in IP networks, especially for flow-specific problems and hence, the aggregate behavior within a router appears normal. To address this problem, in this paper, we propose a Consistent NetFlow (CNF) architecture for measuring per-flow performance measurements within routers. CNF utilizes NetFlow architecture that already reports the first and last timestamps per-flow, and hash-based sampling for ensuring that two routers record same flows. We devise a novel Multiflow estimator that approximates the intermediate delay samples from other background flows to improve the per-flow latency estimates significantly compared to the naive estimator that only uses actual flow samples. In our experiments using real backbone traces and realistic delay models, we show that Multiflow estimator is accurate with a median relative error of less than 20% for flows of size greater than 100 packets. We also show that prior approach based on trajectory sampling performs about 2-3× worse.
Myungjin Lee, Nick G. Duffield, Ramana Rao Kompella
INFOCOM1
2010 Not all microseconds are equal: fine-grained per-flow measurements with reference latency interpolation
abstract
New applications such as algorithmic trading and high-performance computing require extremely low latency (in microseconds). Network operators today lack sufficient fine-grain measurement tools to detect, localize and repair performance anomalies and delay spikes that cause application SLA violations. A recently proposed solution called LDA provides a scalable way to obtain latency, but only provides aggregate measurements. However, debugging application-specific problems requires per-flow measurements, since different flows may exhibit significantly different characteristics even when they are traversing the same link. To enable fine-grained per-flow measurements in routers, we propose a new scalable architecture called reference latency interpolation (RLI) that is based on our observation that packets potentially belonging to different flows that are closely spaced to each other exhibit similar delay properties. In our evaluation using simulations over real traces, we show that RLI achieves a median relative error of 12% and one to two orders of magnitude higher accuracy than previous per-flow measurement solutions with small overhead.
Myungjin Lee, Nick G. Duffield, Ramana Rao Kompella
SIGCOMM1
2009 Semantic Web Constraint Language and its application to an intelligent shopping agent
Hak-Jin Kim, Wooju Kim, Myungjin Lee
Decis. Support Syst.3
2008 A cross-layer approach for TCP optimization over wireless and mobile networks
Myungjin Lee, Moonsoo Kang, Jeonghoon Mo
Comput. Commun.1
2006 A New Key Management Scheme for Distributed Encrypted Storage Systems
Myungjin Lee, Hyokyung Bahn, Kijoon Chae
ICCSA (1)1