Kyunghan Lee

dblp:49/6532 · DBLP profile ↗
← Back
70ranked-venue papers
11as first author
29since 2021 · last 2026
0000-0001-8647-1476ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 59 · 11 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Decento: A New Scalable Interactive Live Streaming System via Control Plane Decentralization
Jongyun Lee, Sangtae Ha, Kyunghan Lee
INFOCOM5
2026 PAVE: Mitigating Non-Congestive Delay for Seamless Video Calls over NextG Mobile Networks
Goodsol Lee, Seyeon Kim 0001, Juheon Yi, Junhong Min, Sangtae Ha, Kyunghan Lee, Saewoong Bahk
INFOCOM6
2026 SAIL: Redesigning Collaborative Language Inference with a Single Server-to-Mobile Handoff
abstract
While on-device small language models (SLMs) enable responsive inference, they often lack the accuracy required for complex reasoning tasks; conversely, server-based large language models (LLMs) deliver high accuracy but incur prohibitive latency for real-time mobile applications. Collaborative inference approaches, such as split computing and speculative decoding, promise to bridge this gap, yet both require repeated network synchronization during generation, accumulating delays that negate their computational benefits. We present SAIL, a collaborative inference framework that achieves both high accuracy and low latency through a strategy we term Prefix Handoff Inference (PHI). In PHI, a server LLM generates the difficult early tokens and hands off completion to the mobile SLM, consolidating communication into a single handoff while providing sufficient context for the SLM to maintain accuracy. SAIL operationalizes PHI through three core modules: the handoff decision module determining the optimal prefix length tailored to each query's difficulty, the branch prediction module speculatively generating candidate continuations on-device while awaiting the server prefix, and the adaptive control module coordinating hand-off and branch prediction under fluctuating network and server conditions to meet service level objectives. Together, these modules enable SAIL to deliver up to 76% accuracy improvement over mobile SLMs with 98.9% server LLM accuracy retention, while strictly meeting latency constraints across translation, math, and expert level QA.
Gibum Park, Yonghwa Cho, Chanjeong Park, Kyunghan Lee
MobiSys5
2026 Ouroboros: Instilling Motion Awareness in ViTs for Efficient Video Analytics on the Edge
abstract
While Vision Transformers (ViTs) have emerged as foundation models for visual recognition, their high computational demands hinder deployment on edge platforms. Temporal redundancy across video frames offers a natural opportunity to reuse prior computations; however, existing methods remain far from ideal, often relying on simple frame-difference signals. To address this, we propose Ouroboros, a framework that encompasses geometric redundancy from spatial displacement of content. We achieve this by aligning invariant content to consistent coordinates across frames, enabled by warping each frame into a global coordinate system via motion vectors from a hardware-accelerated encoder. Yet, this design raises two key challenges: (i) preserving content that drifts out of the limited coordinate system and (ii) maintaining spatial continuity at frame borders. Ouroboros resolves these challenges by introducing a toroidal (i.e., wrap-around) input space and reassigning positional encodings to track displaced content. Leveraging the significant patch reduction via a system-efficient partial computation scheme, our approach accelerates inference by up to 2.61× and reduces energy consumption by 64.5% on NVIDIA Jet-son Orin devices, with <1% accuracy loss on object detection and instance segmentation, outperforming prior methods. Designed to process only non-redundant patches, Ouroboros also excels as an offloading system, yielding higher accuracy at lower bandwidth compared to prior schemes. The source code is available at https://github.com/ckswjd99-lab/Ouroboros.
Chanjeong Park, Donggyu Yang, Sooyoung Kwon, Gibum Park, Carlee Joe-Wong, Kyunghan Lee
MobiSys6
2026 QCON: Seamless QoE-Aware 5G Streaming via Multi-Connectivity
Goodsol Lee, Junhong Min, Seyeon Kim 0001, Juheon Yi, Kwang Taik Kim, Mung Chiang, Sangtae Ha, Kyunghan Lee, Saewoong Bahk
NSDI8
2026 eXpressSFU: Toward Super-Scalable Video Conferencing with SmartNICs
S. M. H. Hosseini, Seyeon Kim 0001, Kyunghan Lee, Nam Bui, Dirk Grunwald, Sangtae Ha
NSDI4
2026 DeepSFU: Scalable Deepfake Detection for Video Conferencing
abstract
Deepfakes have emerged as a significant threat to online communications, enabling nearly indistinguishable impersonation of executives, public figures, and trusted contacts during video calls. While state-of-the-art deepfake detection models can achieve high accuracy offline, deploying them in real-time video conferencing systems remains challenging: the added computation quickly violates interactive latency budgets and greatly limits scalability. Our empirical analysis reveals that video decoding and frame movement dominate the detection pipeline, together accounting for approximately 86.6% of per-frame processing time.
Shirin Ebadi, S. M. H. Hosseini, Woongsub Shin, Evan Ram, Youngwook Son, Seyeon Kim 0001, Nam Bui, Kyunghan Lee, Eric Keller, Sangtae Ha
SIGCOMM9
2025 ADQ: Application-Aware Socket Buffer Dequeueing for Mobile Devices
abstract
The end-to-end latency requirement in mobile applications is not limited to the time taken from the server to the user device. The perceived latency by the users also includes the time taken for the data to traverse from the mobile operating system's kernel to the application layer. The kernels of modern mobile devices are designed to reduce the network packet processing delay by enqueuing incoming data in socket buffer and quickly dequeueing socket buffer for immediate data delivery to the application layer. However, this approach results in unnecessary CPU load on the mobile device due to the frequent execution of socket buffer dequeueing. We propose ADQ, which utilizes Application Data Unit (ADU) information—the smallest data unit interpretable by the application—in the mobile kernel to conduct minimal socket buffer dequeueing, aiming to reduce unnecessary CPU load without degrading delay performance. Our evaluation in a real-time video streaming scenario shows that ADQ can maintain a minimal level of latency using a restricted amount of CPU resources. Furthermore, in the presence of high CPU loads from competing tasks, we demonstrate that ADQ can improve delay performance compared to the default mobile kernel.
Dongwook Choi, Shinik Park, Donggyu Yang, Kyunghan Lee
CCNC4
2025 Empilo: Realizing Immersive Mobile 3D Video Conferencing through Parameterized Communication
abstract
In this work, we explore a new communication paradigm for immersive 3D video conferencing, termed parameterized communication, which dramatically reduces bandwidth usage by eliminating the need to exchange excessive volumetric data. Instead, this approach extracts a compact set of informative parameters representing key elements in the 3D space, transmits only these parameters, and reconstructs the scene on the receiving end. Translating this concept into practice, we present Empilo, a mobile 3D conferencing system composed of a face parameter extractor and a neural rendering-based scene generator. However, while neural rendering excels at synthesizing arbitrary views of objects without explicit 3D models, its heavy computational demands present a major obstacle for mobile deployment. To overcome this challenge, we propose a novel technique called truncated ray marching, which significantly reduces computational overhead by replacing iterative MLP inferences with a single-pass of a shallow neural network. Furthermore, to ensure a consistently immersive experience, we structure the neural-free lightweight renderer as a decoupled component, dedicated to delivering rapid responsiveness to dynamic viewpoint changes. These breakthroughs on computation together enable Empilo to rely entirely on mobile resources, achieving real-time performance with a frame generation time of 30.3 ms and a re-rendering latency of just 6.6 ms—all while operating at an exceptionally low bitrate of 24 kbps. Our approach provides valuable guidance for the practical deployment of 3D conferencing, envisioning accessibility on par with platforms like FaceTime and Zoom.
Donggyu Yang, Wooseung Nam, Byunggu Kang, Kyunghan Lee
MobiSys4
2025 AoRA: AI-on-RAN for Backhaul-free Edge Inference
abstract
In cellular networks, edge intelligence is often enabled by MultiAccess Edge Computing (MEC), which aims to bring AI services closer to end users. Although MEC reduces latency by placing computation near the network edge, it remains external to the Radio Access Network (RAN) and its native execution environment, thereby introducing additional transport and buffering delays. The emerging AI-on-RAN paradigm proposes to overcome these limitations but remains largely conceptual, lacking practical implementation and feasibility validation. In this paper, we present AoRA, the first framework realizing the AI-on-RAN vision by dynamically utilizing available computational headroom in GPU- and NPU-accelerated RAN platforms to deliver AI services directly from within the base stations. AoRA leverages containerized AI workloads in the 5G RAN stack to enable inlined and opportunistic AI service provisioning without degrading core telecom operations. The framework is fully compliant with O-RAN interfaces and can operate seamlessly alongside existing edge computing infrastructures. Evaluations show that AoRA reduces transport latency by over 30% compared to MEC and 70% compared to cloud-based setups.
Siyavushkhon Kholmatov, Seongsik Cho, Kyunghan Lee, Song Chong
SIGCOMM3
2025 NeuroBalancer: Balancing System Frequencies With Punctual Laziness for Timely and Energy-Efficient DNN Inferences
abstract
On-device deep neural network (DNN) inference is often desirable for user experience and privacy. Existing solutions have fully utilized resources to minimize inference latency. However, they result in severe energy inefficiency by completing DNN inference much earlier than the required service interval. It poses a new challenge of how to make DNN inferences in a punctual and energy-efficient manner. To tackle this challenge, we propose a new resource allocation strategy for DNN processing, namelypunctual lazinessthat disperses its workload as efficiently as possible over time within its strict delay constraint. This strategy is particularly beneficial for neural workloads since a DNN comprises a set of popular operators whose latency and energy consumption are predictable. Through this understanding, we propose NeuroBalancer, an operator-aware core and memory frequency scaling framework that balances those frequencies as efficiently as possible while making timely inferences. We implement and evaluate NeuroBalancer on off-the-shelf Android devices with various state-of-the-art DNN models. Our results show that NeuroBalancer successfully meets a given inference latency requirements while saving energy consumption up to 43.9% and 21.1% compared to the Android's default governor and up to 42.1% and 18.6% compared to SysScale, the state-of-the-art mobile governor on CPU and GPU, respectively.
Kyungmin Bin, Seyeon Kim 0001, Sangtae Ha, Song Chong, Kyunghan Lee
IEEE Trans. Mob. Comput.5
2025 nCTX: A Neural Network-Powered Lossless Compressive Transmission Using Shared Information
abstract
In this work, we explore the possibility of a new delivery method for lossless data, namely compressive transmission. It aims at minimizing the transmission data volume at runtime by exploiting the tailored information shared between the sender and the receiver. There are two approaches to leverage shared information for compression: 1) using a DNN-based codec as a proxy for shared information and 2) applying redundancy elimination using deduplication. However, these approaches have not been studied in depth to utilize the trade-off between the compression rate and the amount of shared information. Compared to these approaches, compressive transmission is unique as it fully leverages the abundance of information available on both sides, which is chosen and placed purposely. To bring the concept to reality, we propose nCTX, a neural network-powered Compressive Transmission System that adaptively exploits a generative model and matching blocks. nCTX extracts the optimal semantic data from the input data, exploiting shared information to closely imitate the original and compensate it with the offset (i.e., difference). Extensive evaluations in mobile platforms confirm that nCTX reduces the transmission volume significantly by 25.8% and 23.3% compared to FLIF and RC, the state-of-the-art image codecs, respectively, in comparable or shorter computation times.
Wooseung Nam, Sungyong Lee, Kyunghan Lee
IEEE Trans. Mob. Comput.4
2025 N-Epitomizer: A Semantic Offloading Framework Leveraging Essential Information for Timely Neural Network Inferences
abstract
Offloading neural network inferences from resource-constrained mobile devices to an edge server over wireless networks is becoming more crucial as the neural networks get heavier. To this end, recent studies have tried to make this offloading process more efficient. However, the most fundamental question on extracting and offloading the minimal amount of necessary information that does not degrade the inference accuracy has remained unanswered. We call such an ideal offloading semantic offloading and propose N-epitomizer, a new offloading framework that enables semantic offloading, thus achieving more reliable and timely inferences in highly-fluctuated or even low-bandwidth wireless networks. To realize N-epitomizer, we design an autoencoder-based scalable encoder trained to extract the most informative data and scale its output size to meet the latency and accuracy requirements of inferences over a network. We also accelerate N-epitomizer by exploiting light-weight knowledge distillation for the encoder design and decoder slimming for the decoder design, reducing its overall computation time significantly. Moreover, we extend our N-epitomizer to support multiple DNNs by extracting and offloading the union of the essential information required for each DNN. Our evaluation shows that N-epitomizer achieves exceptionally high compression for images without compromising inference accuracy, which is 21$\times$, 77$\times$, and 192$\times$higher than JPEG compression, and 20$\times$, 55$\times$, and 86$\times$higher than the state-of-the-art DNN-aware image compression GRACE for semantic segmentation, depth estimation, and classification, respectively. Our results show N-epitomizer’s strong potential as the first semantic offloading system to guarantee end-to-end latency even under highly varying cellular networks.
Wooseung Nam, Sungyong Lee, Jinsung Lee, Huijeong Choe, Sangtae Ha, Kyunghan Lee
IEEE Trans. Netw.6
2024 Exstream: A Delay-minimized Streaming System with Explicit Frame Queueing Delay Measurement
abstract
Network fluctuations can cause unpredictable degradation of the user’s quality of experience (QoE) on real-time video streaming. The intrinsic property of real-time video streaming, which generates delay-sensitive and chunk-based video frames, makes the situation even more complicated. Although previous approaches have tried to alleviate this problem by controlling the video bitrate based on the current network capacity estimate, they do not take into account the explicit queueing delay experienced by the video frame in determining the bitrate of upcoming video frames. To tackle this problem, we propose a new real-time video streaming system, Exstream, that can adapt to dynamic network conditions with the help of video bitrate control method and bandwidth estimation method designed to support real-time video streaming environments. Exstream explicitly estimates the queueing delay experienced by the video frame based on the transmission time budget that each frame can maximally utilize, which depends on the frame generation interval, and adjusts the bitrate of newly generated video frames to suppress the queueing delay level close to zero. Our comprehensive experiments demonstrate that Exstream achieves lower frame delay than four existing systems, Salsify, WebRTC, Skype, and Hangouts without frequent video frame skip.
Shinik Park, Junseon Kim, Jongyun Lee, Sangtae Ha, Kyunghan Lee
INFOCOM6
2024 CoActo: CoActive Neural Network Inference Offloading with Fine-grained and Concurrent Execution
abstract
Collaborative inference is the current state-of-the-art solution for mobile-server neural network inference offloading. However, we find that existing collaborative inference solutions only focus on partitioning the DNN computation, which is only a small part of achieving an efficient DNN offloading system. What ultimately determines the performance of DNN offloading is how the execution system utilizes the characteristics of the given DNN offloading task on the mobile, network, and server resources of the offloading environment. To this end, we design CoActo, a DNN execution system built from the ground up for mobile-server inference offloading. Our key design philosophy is Coactive Inference Offloading, which is a new, improved concept of DNN offloading that adds two properties, 1) fine-grained expression of DNNs and 2) concurrency of runtime resources, to existing collaborative inference. In CoActo, system components go beyond simple model splitting of existing approaches and operate more proactively to achieve the coactive execution of inference workloads. CoActo dynamically schedules concurrent interleaving of the mobile, server, and network operations to actively increase resource utilization, enabling lower end-to-end latency. We implement CoActo for various mobile devices and server environments and evaluate our system with distinct environment settings and DNN models. The experimental results show that our system achieves up to 2.1 times speed-up compared to the state-of-the-art collaborative inference solutions.
Kyungmin Bin, Jongseok Park 0002, Chanjeong Park, Seyeon Kim 0001, Kyunghan Lee
MobiSys5
2024 The investigation of a digitalized projective psychological assessment: Comparison to human expert on bender gestalt test
Won-Du Chang, Byeongjun Kim, Bogeum Kim, Kyunghan Lee, Yeonji Kim, Jueun Hwang, Seong-Jin Choi
Multim. Tools Appl.4
2024 An Empirical Study of 5G: Effect of Edge on Transport Protocol and Application Performance
abstract
In this paper, we conduct a measurement study on operational 5G networks deployed across different frequency bands (mmWave and sub-6GHz) and server locations (mobile edge and Internet cloud). Specifically, we assess 5G performance in both uplink and downlink across multiple operators’ networks. We then carry out extensive comparisons of transport-layer protocols using ten different algorithms in full-fledged 5G networks, including an edge computing environment. Finally, we evaluate representative mobile applications over the 5G network with and without edge servers. Our comprehensive measurements provide several insights that affect the experience of 5G users: (i) With a 5G edge server, existing TCP congestion control algorithms can achieve throughput up to 1.8Gbps with only a single flow. (ii) The maximum TCP receive buffer size, which is set by off-the-shelf 5G phones, can limit the throughput performance of 5G networks, which is not observed in 4G LTE-A networks. (iii) Despite significant latency gains in download-centric applications, the 5G edge service provides limited benefits to CPU-intensive tasks or those that use significant uplink bandwidth. To our knowledge, this is the first measurement-driven understanding of 5G edge computing “in the wild,” which can provide an answer to how edge computing would perform in real 5G networks.
Hyoyoung Lim, Jinsung Lee, Jongyun Lee, Sandesh Dhawaskar Sathyanarayana, Junseon Kim, Kwang Taik Kim, Youngbin Im, Mung Chiang, Dirk Grunwald, Kyunghan Lee, Sangtae Ha
IEEE Trans. Mob. Comput.11
2024 Enabling Delay-Guaranteed Congestion Control With One-Bit Feedback in Cellular Networks
abstract
Unexpected large packet delays are often observed in cellular networks due to huge network queuing caused by excessive traffic coming into the network. To deal with the large queue problem, many congestion control algorithms try to find out how much traffic the network can accommodate, either by measuring network performance or by directly providing explicit information. However, due to the nature of the control in which queue growth should be observed or the necessity to modify the overall network architecture, existing algorithms are experiencing difficulties in keeping queues within a strict bound. In this paper, we propose a novel congestion control algorithm based on simple feedback, ECLAT which can provide bounded queuing delay using only one-bit signaling already available in traditional network architecture. To do so, a base station or a router running ECLAT 1) calculates how many packets each flow should transmit and 2) analyzes when congestion feedback needs to be forwarded to adjust the flow’s packet transmission to the desired rate. Our extensive experiments in our testbed demonstrate that ECLAT achieves strict queuing delay bounds, even in the dynamic cellular network environment.
Junseon Kim, Youngbin Im, Kyunghan Lee
IEEE/ACM Trans. Netw.3
2023 ENTRO: Tackling the Encoding and Networking Trade-off in Offloaded Video Analytics
abstract
With the rapid advances of deep learning and the commercialization of high-definition cameras in mobile and embedded devices, the demands from latency-critical applications such as AR and XR for high-quality video analytics (HVA) are soaring. By the nature of HVA aiming at enabling detailed analytics even for small objects, its on-device implementation is suffering from thermal and battery issues, which makes offloaded HVA an attractive solution. This work provides unique observations on the tradeoff pertaining to offloaded HVA: the frame encoding time, the frame transmission time, and the HVA accuracy. Our observations pose a fundamental question: given a latency budget, how to choose the encoding option that properly combines between the encoding time and the transmission time to maximize the HVA accuracy. To answer this question, we propose an offloaded HVA system, ENTRO, which exploits this tradeoff in real-time to maximize the HVA accuracy under the latency budget. Our extensive evaluations with ENTRO implemented on Nvidia AGX Xavier and Samsung Galaxy S20 Ultra over WiFi networks show 8.8× improvement in latency without accuracy loss compared to DDS, the state-of-the-art offloaded video analytics. Our evaluation over commercial 5G and LTE networks also indicates that ENTRO flexibly adapts its encoding option under the tradeoff and enables the latency-bounded HVA with 4K frames.
Seyeon Kim 0001, Kyungmin Bin, Donggyu Yang, Sangtae Ha, Song Chong, Kyunghan Lee
ACM Multimedia6
2023 ASPEN: Breaking Operator Barriers for Efficient Parallelization of Deep Neural Networks
abstract
Modern Deep Neural Network (DNN) frameworks use tensor operators as the main building blocks of DNNs. However, we observe that operator-based construction of DNNs incurs significant drawbacks in parallelism in the form of synchronization barriers. Synchronization barriers of operators confine the scope of parallel computation to each operator and obscure the rich parallel computation opportunities that exist across operators. To this end, we present ASPEN, a novel parallel computation solution for DNNs that achieves fine-grained dynamic execution of DNNs, which (1) removes the operator barriers and expresses DNNs in dataflow graphs of fine-grained tiles to expose the parallel computation opportunities across operators, and (2) exploits these opportunities by dynamically locating and scheduling them in runtime. This novel approach of ASPEN enables opportunistic parallelism, a new class of parallelism for DNNs that is unavailable in the existing operator-based approaches. ASPEN also achieves high resource utilization and memory reuse by letting each resource asynchronously traverse depthwise in the DNN graph to its full computing potential. We provide challenges and solutions to our approach and show that our proof-of-concept implementation of ASPEN on CPU shows exceptional performance, outperforming state-of-the-art inference systems of TorchScript and TVM by up to 3.2$\times$ and 4.3$\times$, respectively.
Jongseok Park 0002, Kyungmin Bin, Gibum Park, Sangtae Ha, Kyunghan Lee
NeurIPS5
2023 Converge: QoE-driven Multipath Video Conferencing over WebRTC
abstract
Video conferencing has become a daily necessity, but protocols to support video conferencing have yet to keep pace despite the innovation in next-generation networks. As video resolutions increase and mobile applications using multiple cameras for photos and videos become popular, the need to meet the Quality of Experience (QoE) requirements is growing. Multipath protocols could be a possible solution.
Sandesh Dhawaskar Sathyanarayana, Kyunghan Lee, Dirk Grunwald, Sangtae Ha
SIGCOMM2
2023 DeepVehicleSense: An Energy-Efficient Transportation Mode Recognition Leveraging Staged Deep Learning Over Sound Samples
abstract
In this paper, we present a new transportation mode recognition system for smartphones called DeepVehicleSense, which is widely applicable to mobile context-aware services. DeepVehicleSense aims at achieving three performance objectives: high accuracy, low latency, and low power consumption at once by exploiting sound characteristics captured from the built-in microphone while being on candidate transportations. To attain high energy efficiency, DeepVehicleSense adopts hierarchical accelerometer-based triggers that minimize the activation of the microphone of smartphones. Further, to achieve high accuracy and low latency, DeepVehicleSense makes use of non-linear filters that can best extract the transportation sound samples. For recognition of five different transportation modes, we design a deep learning based sound classifier using a novel deep neural network architecture with multiple branches. Our staged inference technique can significantly reduce runtime and energy consumption while maintaining high accuracy for the majority of samples. Through 263-hour datasets collected by seven different Android phone models, we demonstrate that DeepVehicleSense achieves the recognition accuracy of 97.44% with only sound samples of 2 seconds at the power consumption of 35.08 mW on average for all-day monitoring.
Sungyong Lee, Jinsung Lee, Kyunghan Lee
IEEE Trans. Mob. Comput.3
2022 R-FEC: RL-based FEC Adjustment for Better QoE in WebRTC
abstract
The demand for video conferencing applications has seen explosive growth while users still often face unsatisfactory quality of experience (QoE). Video conferencing applications adopt Forward Error Correction (FEC) as a recovery mechanism to meet tight latency requirements and overcome packet losses prevalent in the network. However, many studies mainly focused on video rate control by neglecting the complex interactions of this video recovery mechanism on the rate control and its impact on the user QoE. Deciding the right amount of FEC for the current video rate under a dynamically changing network environment is not straightforward. For instance, the higher FEC may enhance the tolerance to packet losses, but it may increase latency due to FEC processing overhead and hurt the video quality due to the additional bandwidth used for FEC. To address this issue, we propose R-FEC which is a reinforcement learning (RL) based framework for video and FEC bitrate decisions in video conferencing. R-FEC aims to improve overall QoE by automatically learning through the results of past decisions and adjusting video and FEC bitrates to maximize the user QoE while minimizing the congestion in the network. Our experiments show that R-FEC outperforms the state-of-the-art solutions in video conferencing, with up to 27% improvement in its video rate and 6dB PSNR improvement in video quality over the default WebRTC.
Insoo Lee, Seyeon Kim 0001, Sandesh Dhawaskar Sathyanarayana, Kyungmin Bin, Song Chong, Kyunghan Lee, Dirk Grunwald, Sangtae Ha
ACM Multimedia6
2022 mGEMM: low-latency convolution with minimal memory overhead optimized for mobile devices
abstract
The convolution layer is the key building block in many neural network designs. Most high-performance implementations of the convolution operation rely on GEMM (General Matrix Multiplication) to achieve high computational throughput with a large workload size. However, in mobile environments, the user experience priority puts focus on low-latency inferences over a single or limited batch size. This signifies two major problems of current GEMM-based solutions: 1) GEMM-based solutions require mapping the convolution operation to GEMM, causing overheads in both computation and memory, 2) GEMM-based solutions lose large opportunities of data reuse while mapping, leading to under-utilization of the given hardware. Through an in-depth analysis of current GEMM-based solutions, we identify the root cause of these problems, and we propose mGEMM, a convolution solution that overcomes the aforementioned problems, without changes in accuracy. mGEMM expands the structure of GEMM in such a way that it can accommodate the convolution operation without any overhead, while the existing algorithms suffer from inefficiencies in converting the convolution operation to a static GEMM algorithm. Our extensive evaluations done over various neural networks and test devices show that mGEMM outperforms the existing solutions in the aspects of latency, memory overhead, and energy consumption. In running a real-world application, YoloV3-Tiny object detection, mGEMM achieves up to 1.29× and 1.58× speedup in total latency and convolution latency compared to the state-of-the-art, resulting in 15.5% reduction in energy consumption while using only near-minimum heap memory.
Jongseok Park 0002, Kyungmin Bin, Kyunghan Lee
MobiSys3
2021 ECLAT: An ECN Marking System for Latency Guarantee in Cellular Networks
abstract
As the importance of latency performance increases, a number of multi-bit feedback-based congestion control mechanisms have been proposed for explicit latency control in cellular networks. However, due to their reactive nature and limited access to the network queue, while latency reduction was possible, latency guarantee has not been achieved. Also, due to the need for end-host modifications, it was hard to commonly provide latency benefit to all connected devices. To this end, we propose a novel network-assisted congestion control, ECLAT, which can always bound the queuing delay within a delay-budget through ECN-based single-bit feedback while maintaining high link utilization for any device. To do so, ECLAT 1) calculates its target operating point for each flow, which is related to the maximum allowable cwnd to meet the delay-budget under time-varying cellular networks, and 2) determines its single-bit feedback policy to limit cwnd within the target operating point. Our extensive experiments in our testbed demonstrate that ECLAT is able to bound the queuing delays of multiple flows within their delay-budget and achieve high utilization even in the dynamic cellular network environment.
Junseon Kim, Youngbin Im, Kyunghan Lee
INFOCOM3
2021 Demystifying Commercial Video Conferencing Applications
abstract
Video conferencing applications have seen explosive growth both in the number of available applications and their use. However, there have been few studies on the detailed analysis of video conferencing applications with respect to network dynamics, yet understanding these dynamics is essential for network design and improving these applications. In this paper, we carry out an in-depth measurement and modeling study on the rate control algorithms used in six popular commercial video conferencing applications. Based on macroscopic behaviors commonly observed across these applications in our extensive measurements, we construct a unified architecture to model the rate control mechanisms of individual applications. We then reconstruct each application's rate control by inferring key parameters that closely follow its rate control and quality adaptation behaviors. To our knowledge, this is the first work that reverse-engineers rate control algorithms of popular video conferencing applications, which are often unknown or hidden as they are proprietary software. We confirm our analysis and models using an end-to-end testbed that can capture the dynamics of each application under a variety of network conditions. We also show how we can use these models to gain insights into the particular behaviors of an application in two practical scenarios.
Insoo Lee, Jinsung Lee, Kyunghan Lee, Dirk Grunwald, Sangtae Ha
ACM Multimedia3
2021 zTT: learning-based DVFS with zero thermal throttling for mobile devices
abstract
DVFS (dynamic voltage and frequency scaling) is a system-level technique that adjusts voltage and frequency levels of CPU/GPU at runtime to balance energy efficiency and high performance. DVFS has been studied for many years, but it is considered still challenging to realize a DVFS that performs ideally for mobile devices for two main reasons: i) an optimal power budget distribution between CPU and GPU in a power-constrained platform can only be defined by the application performance, but conventional DVFS implementations are mostly application-agnostic; ii) mobile platforms experience dynamic thermal environments for many reasons such as mobility and holding methods, but conventional implementations are not adaptive enough to such environmental changes. In this work, we propose a deep reinforcement learning-based frequency scaling technique, zTT. zTT learns thermal environmental characteristics and jointly scales CPU and GPU frequencies to maximize the application performance in an energy-efficient manner while achieving zero thermal throttling. Our evaluations for zTT implemented on Google Pixel 3a and NVIDIA JETSON TX2 platform with various applications show that zTT can adapt quickly to changing thermal environments, consistently resulting in high application performance with energy efficiency. In a high-temperature environment where a rendering application with the default mobile DVFS fails to keep producing more than a target frame rate, zTT successfully manages to do so even with 23.9% less average power consumption.
Seyeon Kim 0001, Kyungmin Bin, Sangtae Ha, Kyunghan Lee, Song Chong
MobiSys4
2021 An Inter-Data Encoding Technique that Exploits Synchronized Data for Network Applications
abstract
In a variety of network applications, there exists a significant amount of shared data between two end hosts. Examples include data synchronization services that replicate data from one node to another. Given that shared data may have a high correlation with new data to transmit, we question how such shared data can be best utilized to improve the efficiency of data transmission. To answer this, we develop an inter-data encoding technique, SyncCoding, that effectively replaces bit sequences of the data to be transmitted with the pointers to their matching bit sequences in the shared data so called references. By doing so, SyncCoding can reduce data traffic, speed up data transmission, and save energy consumption for transmission. Our evaluations of SyncCoding implemented in Linux show that it outperforms existing popular encoding techniques, Brotli, LZMA, Deflate, and Deduplication. The gains of SyncCoding over those techniques in the perspective of data size after compression in a cloud storage scenario are about 12.5, 20.8, 30.1, and 66.1 percent, and are about 78.4, 80.3, 84.3, and 94.3 percent in a web browsing scenario, respectively.
Wooseung Nam, Ness Shroff, Kyunghan Lee
IEEE Trans. Mob. Comput.4
2021 Toward Programmable DOCSIS 4.0 Networks: Adaptive Modulation in OFDM Channels
abstract
The sixth generation of DOCSIS standard is currently under development for the provisioning of multi-Gbps services over cable networks. Building upon DOCSIS 3.1 (D3.1), DOCSIS 4.0 (D4) introduces several features including full-duplex transmission and extended-spectrum, which benefit from subcarrier-level OFDM modulation configurations to adapt to varying channel conditions. To exploit the full potential of D4, we propose a softwarized adaptive subcarrier modulation management framework. The optimization system consists of (i) a clustering mechanism that classifies CMs (Cable Modems) with a similar channel condition into the same group using a sparsified K-means algorithm and (ii) an efficient profile generation mechanism to balance achieved channel throughput and packet error rate within the same group. Then, we implement key elements of the softwarized system using a virtualized network function in our DOCSIS experimental testbed that enables the programmatic control of OFDM channels using D4 performance parameters. Our experimental results show that the proposed optimization function offers significant improvements in OFDM channel throughput over current industry management practices. Furthermore, we confirm via simulations that using a novel clustering algorithm for the classification of CM populations and a new bit-loading method measurably enhances channel performance in large-scale distributed deployment scenarios.
Jason Schnitzer, Prasanth Prahladan, Parisa Rahimzadeh, Chad Humble, Jinsung Lee, Kyunghan Lee, Sangtae Ha
IEEE Trans. Netw. Serv. Manag.7
2020 PERCEIVE: deep learning-based cellular uplink prediction using real-time scheduling patterns
abstract
As video calls and personal broadcasting become popular, the demand for mobile live streaming over cellular uplink channels is growing fast. However, current live streaming solutions are known to suffer from frequent uplink throughput fluctuations causing unnecessary video stalls and quality drops. As a remedy to this problem, we propose PERCEIVE, a deep learning-based uplink throughput prediction framework. PERCEIVE exploits a 2-stage LSTM (Long Short Term Memory) design and makes throughput predictions for the next 100ms. Our extensive evaluations show that PERCEIVE, trained with LTE network traces from three major operators in the U.S., achieves high accuracy in the uplink throughput prediction with only 7.67% mean absolute error and outperforms existing prediction techniques. We integrate PERCEIVE with WebRTC, a popular video streaming platform from Google, as a rate adaptation module. Our implementation on the Android phone demonstrates that it can improve PSNR by up to 6dB (4x) over the default WebRTC while providing less streaming latency.
Jinsung Lee, Sungyong Lee, Jongyun Lee, Sandesh Dhawaskar Sathyanarayana, Hyoyoung Lim, Sangeeta Ramakrishnan, Dirk Grunwald, Kyunghan Lee, Sangtae Ha
MobiSys10
2019 I Sent It: Where Does Slow Data Go to Wait?
abstract
Emerging applications like virtual reality (VR), augmented reality (AR), and 360-degree video aim to exploit the unprecedentedly low latencies promised by technologies like the tactile Internet and mobile 5G networks. Yet these promises are still unrealized. In order to fulfill them, it is crucial to understand where packet delays happen, which impacts protocol performance such as throughput and latency. In this work, we empirically find that sender-side protocol stack delays can cause high end-to-end latencies, though existing solutions primarily address network delays. Unfortunately, however, current latency diagnosis tools cannot even distinguish between delays on network links and delays in the end hosts. To close this gap, we present ELEMENT, a latency diagnosis framework that decomposes end-to-end TCP latency into endhost and network delays, without requiring admin privileges at the sender or receiver.
Youngbin Im, Parisa Rahimzadeh, Brett Shouse, Shinik Park, Carlee Joe-Wong, Kyunghan Lee, Sangtae Ha
EuroSys6
2018 ExLL: an extremely low-latency congestion control for mobile cellular networks
abstract
Since the diagnosis of severe bufferbloat in mobile cellular networks, a number of low-latency congestion control algorithms have been proposed. However, due to the need for continuous bandwidth probing in dynamic cellular channels, existing mechanisms are designed to cyclically overload the network. As a result, it is inevitable that their latency deviates from the smallest possible level (i.e., minimum RTT). To tackle this problem, we propose a new low-latency congestion control, ExLL, which can adapt to dynamic cellular channels without overloading the network. To do so, we develop two novel techniques that run on the cellular receiver: 1) cellular bandwidth inference from the downlink packet reception pattern and 2) minimum RTT calibration from the inference on the uplink scheduling interval. Furthermore, we incorporate the control framework of FAST into ExLL's cellular specific inference techniques. Hence, ExLL can precisely control its congestion window to not overload the network unnecessarily. Our implementation of ExLL on Android smartphones demonstrates that ExLL reduces latency much closer to the minimum RTT compared to other low-latency congestion control algorithms in both static and dynamic channels of LTE networks.
Shinik Park, Jinsung Lee, Junseon Kim, Ji Hoon Lee, Sangtae Ha, Kyunghan Lee
CoNEXT6
2017 iMUTE: Energy-optimal update policy for perishable mobile contents
abstract
Mobile applications that provide ever-changing information such as social media and news feeds applications are designed to consistently update their contents in the background. This operation, often called “prefetching”, provides the users with immediate access to up-to-date contents. However, such updates often result in the unwanted side-effect of draining the battery of mobile devices. It is considered as pure waste when updated contents are not accessed before being renewed. In this paper, we develop an optimal strategy to update the contents in the background under a given energy constraint. The key challenge is to predict when the user will access the contents in a probabilistic manner from the statistics of the accessed patterns in the past. We model our problem as a constrained Markov decision process (C-MDP) and propose to tackle its high complexity with a two-step solution that combines: (1) a threshold-based backward induction algorithm for the Lagrangian relaxation of our C-MDP, and (2) an iterative root finding algorithm, iMUTE (iterative Method for optimal UpdaTe policy with Energy constraint). We prove that iMUTE converges superlinearly to the optimal solution of the original C-MDP under a mild condition. We also experimentally verify that iMUTE outperforms the periodic policy as well as the additive and multiplicative increase policies that are adopted in the Doze mode of Android systems and HUSH, in terms of user experience and energy saving.
Fang Liu 0020, Kyunghan Lee, Ness Shroff
ICNP3
2017 SyncCoding: A compression technique exploiting references for data synchronization services
abstract
In this work, we raise a question on why the abundant information previously shared between a server and its client is not effectively utilized in the exchange of a new data which may be highly correlated with the shared data. We formulate this question as an encoding problem that is applicable to general data synchronization services including a wide range of Internet services such as cloud data synchronization, web browsing, messaging, and even data streaming. To this problem, we propose a new encoding technique, SyncCoding that maximally replaces subsets of the data to be transmitted with the coordinates pointing to the matching subsets included in the set of relevant shared data, called references. SyncCoding can be easily integrated into a transport layer protocol such as HTTP and enables significant reduction of network traffic. Our experimental evaluations of SyncCoding implemented in Linux shows that it outperforms existing popular encoding techniques, Brotli, LZMA, Deflate, and Deduplication in two practical use networking applications: cloud data sharing and web browsing. The gains of SyncCoding over Brotli, LZMA, Deflate, and Deduplication in the encoded size to be transmitted are shown to be about 12.4%, 20.1%, 29.9%, and 61.2% in the cloud data sharing and about 78.3%, 79.6%, 86.1%, and 92.9% in the web browsing, respectively. The gains of SyncCoding over Brotli, LZMA, and Deflate when Deduplication is applied in advance are about 7.4%, 10.6%, and 17.4% in the cloud data sharing and about 79.4%, 82.0%, and 83.2% in the web browsing, respectively.
Wooseung Nam, Kyunghan Lee
ICNP3
2017 VehicleSense: A reliable sound-based transportation mode recognition system for smartphones
abstract
A new transportation mode recognition system for smartphones, VehicleSense that is widely applicable to mobile context-aware services is proposed. VehicleSense aims at achieving three performance objectives: high accuracy, low latency, and low power consumption at once by exploiting sound characteristics captured from the built-in microphone while being on candidate transportations. To attain high energy efficiency, VehicleSense adopts hierarchical accelerometer-based triggers that minimize the activation of the microphone of smartphones. Further, to attain high accuracy and low latency, VehicleSense makes use of non-linear filters that can best extract the transportation sound samples. Our 186-hour log of sound and accelerometer data collected by seven different Android smartphone models confirms that VehicleSense achieves the recognition accuracy of 98.2% with only 0.5 seconds of sound sampling at the power consumption of 26.1 mW on average for all day monitoring.
Sungyong Lee, Jinsung Lee, Kyunghan Lee
WoWMoM3
2017 CAS: Context-Aware Background Application Scheduling in Interactive Mobile Systems
abstract
Each individual's usage behavior on mobile devices depends on a variety of factors, such as time, location, and previous actions. Hence, context-awareness provides great opportunities to make the networking and computing capabilities of mobile systems more personalized and more efficient in managing their resources. To this end, we first reveal new findings from our own Android user experiment: 1) the launching probabilities of applications follow Zipf's law and 2) inter-running and running times of applications conform to log-normal distributions. We also find contextual dependencies between application usage patterns, for which we classify contexts autonomously with unsupervised learning methods. Using the knowledge acquired, we develop a context-aware application scheduling framework, context-aware application scheduler (CAS), that adaptively unloads and preloads background applications for a joint optimization in which the energy saving is maximized and the user discomfort from the scheduling is minimized. Our trace-driven simulations with 96 user traces demonstrate that the context-aware design of the CAS enables it to outperform existing process scheduling algorithms. Our implementation of the CAS over Android platforms and its end-to-end evaluations verify that its human-involved design indeed provides substantial user-experience gains in both energy and application launching latency.
Kyunghan Lee, Euijin Jeong, Jaemin Jo, Ness Shroff
IEEE J. Sel. Areas Commun.2
2017 CarrierMix: How Much Can User-side Carrier Mixing Help?
abstract
Energy consumption for cellular communication is increasingly gaining importance in smartphone battery lifetime as the bandwidth of wireless communication and the demand for mobile traffic increase. For energy-efficient cellular communication, we tackle two energy characteristics of cellular networks: (1) transmission energy highly varies upon channel condition, and (2) transmission of a packet accompanies unnecessary tail energy waste. Under the objective of transmitting packets when the best channel is provided as well as a number of packets are accumulated, we propose a new mobile collaboration framework “CarrierMix” that aggregates smart devices across multiple heterogeneous cellular carriers. Compared to the standalone operation, even without a buffering delay, CarrierMix allows better channel and reduces more tail energy in a statistical point of view. To maximize the energy benefit while maintaining the fairness among the nodes in collaboration, we further develop a dynamic programming framework providing the optimal algorithm of CarrierMix and its approximated heuristic. Trace-driven simulations on our experimental HSPA/EVDO/LTE network traces show that CarrierMix of five devices achieves up to 42 percent of energy reduction.
Kyunghan Lee, Yeongjin Kim, Song Chong
IEEE Trans. Mob. Comput.2
2016 Context-aware application scheduling in mobile systems: what will users do and not do next?
abstract
Usage patterns of mobile devices depend on a variety of factors such as time, location, and previous actions. Hence, context-awareness can be the key to make mobile systems to become personalized and situation dependent in managing their resources. We first reveal new findings from our own Android user experiment: (i) the launching probabilities of applications follow Zipf's law, and (ii) inter-running and running times of applications conform to log-normal distributions. We also find context-dependency in application usage patterns, for which we classify contexts in a personalized manner with unsupervised learning methods. Using the knowledge acquired, we develop a novel context-aware application scheduling framework, CAS that adaptively unloads and preloads background applications in a timely manner. Our trace-driven simulations with 96 user traces demonstrate the benefits of CAS over existing algorithms. We also verify the practicality of CAS by implementing it on the Android platform.
Kyunghan Lee, Euijin Jeong, Jaemin Jo, Ness Shroff
UbiComp2
2016 TravelMiner: On the Benefit of Path-Based Mobility Prediction
abstract
Mobility predictions are becoming more valuable in various applications with the rise of mobile devices. Given that existing prediction techniques are composed of two key procedures: 1) profiling past mobility trajectories as sequences of discrete atomic states (e.g., grid locations, semantic locations) and capturing them with an appropriate statistical model, 2) making a prediction on the next state using the statistical model, TravelMiner tackles the former with paths utilized as the atomic states for the first time, where the paths are defined as sub-trajectories with no branches. Comparing to available location-based predictors, TravelMiner makes a fundamental difference in that it is able to predict the sequence of paths rather than locations, which is far more detailed in the perspective of knowing the exact route to follow. TravelMiner enables this benefit by extracting disjoint paths from GPS trajectories via a similarity metric for curves, called Frechet distance and keeping the sequences of such paths in a statistical model, called probabilistic radix tree. Our extensive simulations over the GPS trajectories of 124 users reveal that TravelMiner outperforms other predictors in diverse popular performance metrics including predictability, prediction accuracy and prediction resolution.
Jaeseong Jeong, Kyunghan Lee, Beknazar Abdikamalov, Kimin Lee, Song Chong
SECON2
2016 DRWA: A Receiver-Centric Solution to Bufferbloat in Cellular Networks
abstract
The problem of overbuffering in the current Internet (termed as bufferbloat) has drawn the attention of the research community in recent years. Cellular networks keep large buffers at base stations to smooth out the bursty data traffic over the time-varying channels and are hence apt to bufferbloat. However, despite their growing importance due to the boom of smart phones, we still lack a comprehensive study of bufferbloat in cellular networks and its impact on TCP performance. In this paper, we conducted extensive measurement of the 3G/4G networks of the four major U.S. carriers and the largest carrier in Korea. We revealed the severity of bufferbloat in current cellular networks and discovered some ad-hoc tricks adopted by smart phone vendors, which mitigate the impact of bufferbloat but result in performance degradation under various practical scenarios. To address the problem, we propose dynamic receiver window adjustment (DRWA) that requires slight TCP modification only in smart phones, thus guarantees quick deployment via over-the-air updates. Our extensive real-world tests confirm that DRWA reduces the latency of TCP flows by 25-49 percent and increase TCP throughput by up to 51 percent in certain scenarios. It is further verified that DRWA has significant effect on the latency of Voice-over-IP traffic and on the traffic going through TCP split scenarios with performance enhancing proxy.
Haiqing Jiang, Yaogong Wang, Kyunghan Lee, Injong Rhee
IEEE Trans. Mob. Comput.3
2016 On Stochastic Confidence of Information Spread in Opportunistic Networks
abstract
Predicting spreading patterns of information or virus has been a popular research topic for which various mathematical tools have been developed. These tools have mainly focused on estimating the average time of spread to a fraction (e.g.,$\alpha$) of the agents, i.e., so-called average$\alpha$-completion time$E(T_{\alpha})$. We claim that understanding stochastic confidence on the time$T_{\alpha}$rather than only its average gives more comprehensive knowledge on the spread behavior and wider engineering choices. Obviously, the knowledge also enables us to effectively accelerate or decelerate a spread. To demonstrate the benefits of understanding the distribution of spread time, we introduce a new metric$G_{\alpha, \beta}$that denotes the time required to guarantee$\alpha$completion (i.e., penetration) with probability$\beta$. Also, we develop a new framework characterizing$G_{\alpha, \beta}$for various spread parameters such as number of seeders, contact rates between agents, and heterogeneity in contact rates. We apply our technique to a large-scale experimental vehicular trace and show that it is possible to allocate resources for acceleration of spread in a far more elaborated way compared to conventional average-based mathematical tools.
Yoora Kim, Kyunghan Lee, Ness Shroff
IEEE Trans. Mob. Comput.2
2016 ACMI: FM-Based Indoor Localization via Autonomous Fingerprinting
abstract
We present ACMI, an FM-based indoor localization system that does not require proactive site profiling. ACMI constructs the fingerprint database based on pure estimation of indoor received signal strength (RSS) distribution, where only the signals transmitted from commercial FM radio stations are used. Based on extensive field measurement study, we established our own signal propagation model that harnesses FM radio characteristics and open information of FM transmission towers in combination with the floor-plan of a building. Output of the model is an RSS fingerprint database. Using the fingerprint database as a knowledge base, ACMI refines a positioning result via the two-step process; parameter calibration and path matching, during its runtime. Without site profiling, our evaluation indicates that ACMI in seven campus locations and three downtown buildings using eight distinguished FM stations finds positions with only about 6 and 10 meters of errors on average, respectively.
Sungro Yoon, Kyunghan Lee, YeoCheon Yun, Injong Rhee
IEEE Trans. Mob. Comput.2
2016 Resource-Efficient Mobile Multimedia Streaming With Adaptive Network Selection
abstract
From the advancements of mobile display and network infrastructure, mobile users can enjoy high quality mobile video streaming anywhere, anytime. However, most mobile users are still reluctant to use high quality video streaming when they are mobile due to costly cellular data and high energy consumption. In this work, we develop scheduling algorithms for resource-efficient mobile video streaming, which minimize the weighted sum objective of cellular cost and energy consumption. We first model the scheduling problem as a Markov decision process and propose an optimal scheduling algorithm based on dynamic programming. Then, we derive a heuristic algorithm that approximates the optimal algorithm. To evaluate the performance of proposed algorithms, we run simulation over YouTube video traces with audience retention graphs and mobility/connectivity traces in public transportation (e.g., commuting). Through extensive simulations, we show that our proposed scheduling algorithm has negligible performance loss compared to the optimal scheduling algorithm, where it saves 59% of cellular cost and 41% of energy compared to the YouTube default scheduler. We also implement our scheduling algorithm on an Android platform, and experimentally evaluate the performance compared to existing streaming policies.
Kyunghan Lee, Choongwoo Han, Song Chong
IEEE Trans. Multim.2
2015 Max Contribution: An Online Approximation of Optimal Resource Allocation in Delay Tolerant Networks
abstract
In this paper, a joint optimization of link scheduling, routing and replication for delay-tolerant networks (DTNs) has been studied. The optimization problems for resource allocation in DTNs are typically solved using dynamic programming which requires knowledge of future events such as meeting schedules and durations. This paper defines a new notion of approximation to the optimality for DTNs, called snapshot approximation where nodes are not clairvoyant, i.e., not looking ahead into future events, and thus decisions are made using only contemporarily available knowledges. Unfortunately, the snapshot approximation still requires solving an NP-hard problem of maximum weighted independent set (MWIS) and a global knowledge of who currently owns a copy and what their delivery probabilities are. This paper proposes an algorithm, Max-Contribution (MC) that approximates MWIS problem with a greedy method and its distributed online approximation algorithm, Distributed Max-Contribution (DMC) that performs scheduling, routing and replication based only on locally and contemporarily available information. Through extensive simulations based on real GPS traces tracking over 4,000 taxies and 500 taxies for about 30 days and 25 days in two different large cities, DMC is verified to perform closely to MC and outperform existing heuristically engineered resource allocation algorithms for DTNs.
Kyunghan Lee, Jaeseong Jeong, Yung Yi, Hyungsuk Won, Injong Rhee, Song Chong
IEEE Trans. Mob. Comput.1
2014 An analytical framework to characterize the efficiency and delay in a mobile data offloading system
abstract
Smart mobile devices are generating a tremendous amount of data traffic that is putting stress on even the most advanced cellular networks. Delayed offloading has recently been proposed as an efficient mechanism to substantially alleviate this stress. The idea is simple. It allows a mobile device to delay transmission of data packets for a certain amount of time, while it searches WiFi (or similarly femtocell) networks to offload the data during the time. When the time expires, it completes the remaining portion of the delayed transmission through the cellular network that is available at the moment. In this paper, we develop an analytical framework using an embedded Markov process for the delayed offloading system. We provide a closed-form expression for estimating how much data generated by the users can be offloaded to WiFi networks from cellular networks even when there are non-Markovian data arrivals and service interruptions. We conduct extensive numerical studies with various ranges of delay, service interruption time, arrived data, and service rate. These numerical studies show that the current deployment of WiFi networks measured from a metropolitan city is capable of offloading about 80% of the generated data with 30 minutes of delay and 1 Mbps of WiFi data rate, but increasing the data rate does not help improve the amount of offloading. Further studies using this framework on two new deployment strategies of WiFi networks give guidance on how to upgrade WiFi networks by revealing that the amount of offloading for 30 minutes of delay and 1 Mbps of data rate can be drastically improved to about 90% or 98% according to the strategy.
Yoora Kim, Kyunghan Lee, Ness Shroff
MobiHoc2
2014 PhonePool: On energy-efficient mobile network collaboration with provider aggregation
abstract
Energy consumption for cellular communication is increasingly gaining importance in smartphone battery lifetime as the bandwidth of wireless communication and the demand for mobile traffic increase. For energy-efficient cellular communication, we tackle two energy characteristics of cellular networks: (1) transmission energy highly varies upon channel condition, and (2) transmission of a packet accompanies unnecessary tail energy waste. Under the objective of transmitting packets when the best channel is provided as well as a number of packets are accumulated, we propose a new mobile collaboration framework “PhonePool” that aggregates smart devices across multiple cellular providers. Compared to the standalone operation, even without a buffering delay, PhonePool allows better channel and reduces more tail energy in a statistical point of view. To maximize the energy benefit while maintaining the fairness among the nodes in collaboration, we further develop a dynamic programming framework providing the optimal algorithm of PhonePool and its approximated heuristic. Trace-driven simulations on our experimental HSPA/EVDO/LTE network traces show that PhonePool of 5 devices achieves up to 42% of energy reduction.
Kyunghan Lee, Yeongjin Kim, Song Chong
SECON2
2014 ExMin: A routing metric for novel opportunity gain in Delay Tolerant Networks
Jaeseong Jeong, Kyunghan Lee, Yung Yi, Injong Rhee, Song Chong
Comput. Networks2
2014 Opportunistic networks
Chiara Boldrini, Kyunghan Lee, Melek Önen, Jörg Ott, Elena Pagani
Comput. Commun.2
2014 Depth manipulation using disparity histogram analysis for stereoscopic 3D
Younghui Kim, Jungjin Lee, Kyehyun Kim, Kyunghan Lee, Jun-yong Noh
Vis. Comput.5
2013 Providing probabilistic guarantees on the time of information spread in opportunistic networks
abstract
A variety of mathematical tools have been developed for predicting the spreading patterns in a number of varied environments including infectious diseases, computer viruses, and urgent messages broadcast to mobile agents (e.g., humans, vehicles, and mobile devices). These tools have mainly focused on estimating the average time for the spread to reach a fraction (e.g., α) of the agents, i.e., the so-called average completion time E(Tα). We claim that providing probabilistic guarantee on the time for the spread Tαrather than only its average gives a much better understanding of the spread, and hence could be used to design improved methods to prevent epidemics or devise accelerated methods for distributing data. To demonstrate the benefits, we introduce a new metric Gα,βthat denotes the time required to guarantee α completion with probability β, and develop a new framework to characterize the distribution of Tαfor various spread parameters such as number of seeds, level of contact rates, and heterogeneity in contact rates. We apply our technique to an experimental mobility trace of taxies in Shanghai and show that our framework enables us to allocate resources (i.e., to control spread parameters) for acceleration of spread in a far more efficient way than the state-of-the-art.
Yoora Kim, Kyunghan Lee, Ness Shroff, Injong Rhee
INFOCOM2
2013 FM-based indoor localization via automatic fingerprint DB construction and matching
abstract
We present ACMI, an FM-based indoor localization that does not require proactive site profiling. ACMI constructs the fingerprint database based on the pure estimation of indoor RSS distribution, where the signals transmitted from commercial FM radio stations are used. For this, ACMI makes use of our signal model harnessing public transmission information of FM stations in a combination with a floorplan of a building. Using the fingerprint database as the knowledge base, ACMI actively performs multi-level online signal matching to infer the current location of a mobile user. ACMI achieves good indoor localization accuracy even without site profiling efforts. We evaluate ACMI with extensive indoor experiments in 7 different locations with over 1,100 indoor spots. The results show that ACMI achieves up to 89% room identification and accuracy of 6m localization error on average using 8 FM broadcast signals.
Sungro Yoon, Kyunghan Lee, Injong Rhee
MobiSys2
2013 On the Critical Delays of Mobile Networks Under Lévy Walks and Lévy Flights
abstract
Delay-capacity tradeoffs for mobile networks have been analyzed through a number of research works. However, Lévy mobility known to closely capture human movement patterns has not been adopted in such work. Understanding the delay-capacity tradeoff for a network with Lévy mobility can provide important insights into understanding the performance of real mobile networks governed by human mobility. This paper analytically derives an important point in the delay-capacity tradeoff for Lévy mobility, known as the critical delay. The critical delay is the minimum delay required to achieve greater throughput than what conventional static networks can possibly achieve (i.e., O(1/√n) per node in a network with n nodes). The Lévy mobility includes Lévy flight and Lévy walk whose step-size distributions parametrized by α ∈ (0,2] are both heavy-tailed while their times taken for the same step size are different. Our proposed technique involves: 1) analyzing the joint spatio-temporal probability density function of a time-varying location of a node for Lévy flight, and 2) characterizing an embedded Markov process in Lévy walk, which is a semi-Markov process. The results indicate that in Lévy walk, there is a phase transition such that for α ∈ (0,1), the critical delay is always Θ(n[1/2]), and for α ∈ [1,2] it is Θ(n[(α)/2]). In contrast, Lévy flight has the critical delay Θ(n[(α)/2]) for α ∈ (0,2].
Kyunghan Lee, Yoora Kim, Song Chong, Injong Rhee, Yung Yi, Ness Shroff
IEEE/ACM Trans. Netw.1
2013 Mobile Data Offloading: How Much Can WiFi Deliver?
abstract
This paper presents a quantitative study on the performance of 3G mobile data offloading through WiFi networks. We recruited 97 iPhone users from metropolitan areas and collected statistics on their WiFi connectivity during a two-and-a-half-week period in February 2010. Our trace-driven simulation using the acquired whole-day traces indicates that WiFi already offloads about 65% of the total mobile data traffic and saves 55% of battery power without using any delayed transmission. If data transfers can be delayed with some deadline until users enter a WiFi zone, substantial gains can be achieved only when the deadline is fairly larger than tens of minutes. With 100-s delays, the achievable gain is less than only 2%-3%, whereas with 1 h or longer deadlines, traffic and energy saving gains increase beyond 29% and 20%, respectively. These results are in contrast to the substantial gain (20%-33%) reported by the existing work even for 100-s delayed transmission using traces taken from transit buses or war-driving. In addition, a distribution model-based simulator and a theoretical framework that enable analytical studies of the average performance of offloading are proposed. These tools are useful for network providers to obtain a rough estimate on the average performance of offloading for a given WiFi deployment condition.
Kyunghan Lee, Yung Yi, Injong Rhee, Song Chong
IEEE/ACM Trans. Netw.1
2012 Reducing redundant cross-ISP traffic in peer-to-peer systems via explicit coordination
abstract
Locality-aware P2P file sharing systems have drawn the attention of the research community as a promising technique to alleviate the tussle between P2P traffic and ISPs. However, existing locality-based schemes mainly focus on how to distinguish whether a peer is local or not and pay less attention to how to utilize the locality information. Typically, they simply bias the peer selection towards local peers in the hope that it will reduce cross-ISP traffic. In this paper, we argue that such coarse-grained control is inadequate and propose a fine-grained mechanism called Swarm-over-Swarm (SOS). Through explicit coordination on piece selection among local peers, SOS is much more effective in reducing redundant cross-ISP traffic than existing schemes. We implement SOS in both NS-2 and a real BitTorrent client and demonstrate the effectiveness of our solution via both simulations and Internet experiments.
Alphonse H. A. Selvanayagam, Yaogong Wang, Haiqing Jiang, Kyunghan Lee, Injong Rhee
CCNC4
2012 Tackling bufferbloat in 3G/4G networks
abstract
The problem of overbuffering in the current Internet (termed as bufferbloat) has drawn the attention of the research community in recent years. Cellular networks keep large buffers at base stations to smooth out the bursty data traffic over the time-varying channels and are hence apt to bufferbloat. However, despite their growing importance due to the boom of smart phones, we still lack a comprehensive study of bufferbloat in cellular networks and its impact on TCP performance. In this paper, we conducted extensive measurement of the 3G/4G networks of the four major U.S. carriers and the largest carrier in Korea. We revealed the severity of bufferbloat in current cellular networks and discovered some ad-hoc tricks adopted by smart phone vendors to mitigate its impact. Our experiments show that, due to their static nature, these ad-hoc solutions may result in performance degradation under various scenarios. Hence, a dynamic scheme which requires only receiver-side modification and can be easily deployed via over-the-air (OTA) updates is proposed. According to our extensive real-world tests, our proposal may reduce the latency experienced by TCP flows by 25% ~ 49% and increase TCP throughput by up to 51% in certain scenarios.
Haiqing Jiang, Yaogong Wang, Kyunghan Lee, Injong Rhee
Internet Measurement Conference3
2012 Revisiting delay-capacity tradeoffs for mobile networks: The delay is overestimated
abstract
In the literature, one of the key assumptions in characterizing the scaling laws for wireless mobile networks, is to assume that nodes do not communicate while being mobile. In other words, contact opportunities are not considered during the mobility process itself. However, we find that this assumption leads to an inflated estimate of the delay, even in an order sense. To address this issue, a new framework that allows nodes to communicate while being mobile is proposed in this paper. Under this framework, it is shown that delays to obtain various levels of throughput for i.i.d. mobility model are overestimated and a new tighter delay-capacity tradeoff is suggested. Also, the framework is used to analytically derive the delay-capacity tradeoff of Lévy flight model for various levels of throughput, where Lévy flight is a random walk of a power-law flight distribution with an exponent α ∈ (0, 2]. It is known as a mobility model which closely captures human movement patterns. The tradeoffs from the proposed framework between the delay (D̅) and per-node throughput (λ) indicate that D̅ = O(√(max(1,nλ3))) holds for i.i.d. mobility and D̅ = O(√(min(n1+αλ,n2))) holds for Lévy flight.
Yoora Kim, Kyunghan Lee, Ness Shroff, Injong Rhee
INFOCOM2
2012 Demo: VisualComm - new robust channel of file transfer in mobile communications
abstract
There is a lack of method to communicate wirelessly without underlying networking infrastructure. We adopt the concept of camera-LCD communication and proposed this VisualComm system, a channel to communicate wirelessly through transmission and reception of a QR-code stream. VisualComm implements a series of techniques to overcome big challenges of frame missing, perspective distortion and viewing angle issue. In this demo, we show our VisualComm system can achieve real-time wireless file transfer with a reasonable bandwidth (of hundreds of Kbps) and good robustness.
Khiem Lam, Sirote Chaiwattanapong, Kyunghan Lee, Injong Rhee
MobiSys4
2012 SLAW: Self-Similar Least-Action Human Walk
abstract
Many empirical studies of human walks have reported that there exist fundamental statistical features commonly appearing in mobility traces taken in various mobility settings. These include: 1) heavy-tail flight and pause-time distributions; 2) heterogeneously bounded mobility areas of individuals; and 3) truncated power-law intercontact times. This paper reports two additional such features: a) The destinations of people (or we say waypoints) are dispersed in a self-similar manner; and b) people are more likely to choose a destination closer to its current waypoint. These features are known to be influential to the performance of human-assisted mobility networks. The main contribution of this paper is to present a mobility model called Self-similar Least-Action Walk (SLAW) that can produce synthetic mobility traces containing all the five statistical features in various mobility settings including user-created virtual ones for which no empirical information is available. Creating synthetic traces for virtual environments is important for the performance evaluation of mobile networks as network designers test their networks in many diverse network settings. A performance study of mobile routing protocols on top of synthetic traces created by SLAW shows that SLAW brings out the unique performance features of various routing protocols.
Kyunghan Lee, Seongik Hong, Seong Joon Kim, Injong Rhee, Song Chong
IEEE/ACM Trans. Netw.1
2011 Delay-capacity tradeoffs for mobile networks with Lévy walks and Lévy flights
abstract
This paper analytically derives the delay-capacity tradeoffs for Lévy mobility: Lévy walks and Lévy flights. Lévy mobility is a random walk with a power-law flight distribution. α is the power-law slope of the distribution and 01/2) and for 1 ≤ α ≤ 2, is Θ(nα/2). In contrast, Lévy flight has critical delay Θ(nα/2) for 0 <; α ≤ 2.
Kyunghan Lee, Yoora Kim, Song Chong, Injong Rhee, Yung Yi
INFOCOM1
2011 On the levy-walk nature of human mobility
abstract
We report that human walk patterns contain statistically similar features observed in Levy walks. These features include heavy-tail flight and pause-time distributions and the super-diffusive nature of mobility. Human walks are not random walks, but it is surprising that the patterns of human walks and Levy walks contain some statistical similarity. Our study is based on 226 daily GPS traces collected from 101 volunteers in five different outdoor sites. The heavy-tail flight distribution of human mobility induces the super-diffusivity of travel, but up to 30 min to 1 h due to the boundary effect of people's daily movement, which is caused by the tendency of people to move within a predefined (also confined) area of daily activities. These tendencies are not captured in common mobility models such as random way point (RWP). To evaluate the impact of these tendencies on the performance of mobile networks, we construct a simple truncated Levy walk mobility (TLW) model that emulates the statistical features observed in our analysis and under which we measure the performance of routing protocols in delay-tolerant networks (DTNs) and mobile ad hoc networks (MANETs). The results indicate the following. Higher diffusivity induces shorter intercontact times in DTN and shorter path durations with higher success probability in MANET. The diffusivity of TLW is in between those of RWP and Brownian motion (BM). Therefore, the routing performance under RWP as commonly used in mobile network studies and tends to be overestimated for DTNs and underestimated for MANETs compared to the performance under TLW.
Injong Rhee, Seongik Hong, Kyunghan Lee, Seong Joon Kim, Song Chong
IEEE/ACM Trans. Netw.4
2010 Mobile data offloading: how much can WiFi deliver?
abstract
This paper presents a quantitative study on the performance of 3G mobile data offloading through WiFi networks. We recruited about 100 iPhone users from metropolitan areas and collected statistics on their WiFi connectivity during about a two and half week period in February 2010. Our trace-driven simulation using the acquired traces indicates that WiFi already offloads about 65% of the total mobile data traffic and saves 55% of battery power without using any delayed transmission. If data transfers can be delayed with some deadline until users enter a WiFi zone, substantial gains can be achieved only when the deadline is fairly larger than tens of minutes. With 100 second delays, the achievable gain is less than only 2--3%. But with 1 hour or longer deadline, traffic and energy saving gains increase beyond 29% and 20%, respectively. These results are in stark contrast to the substantial gain (20 to 33%) reported by the existing work even for 100 second delayed transmission using traces taken from transit buses or war-driving. The major performance difference comes from traces: while bus and war-driving traces contain much shorter connection and inter-connection times, our traces reflects the daily mobility patterns of average users more accurately.
Kyunghan Lee, Injong Rhee, Song Chong, Yung Yi
CoNEXT1
2010 Max-Contribution: On Optimal Resource Allocation in Delay Tolerant Networks
abstract
This is by far the first paper considering joint optimization of link scheduling, routing and replication for disruption-tolerant networks (DTNs). The optimization problems for resource allocation in DTNs are typically solved using dynamic programming which requires knowledge of future events such as meeting schedules and durations. This paper defines a new notion of optimality for DTNs, called snapshot optimality where nodes are not clairvoyant, i.e., cannot look ahead into future events, and thus decisions are made using only contemporarily available knowledge. Unfortunately, the optimal solution for snapshot optimality still requires solving an NP-hard problem of maximum weight independent set and a global knowledge of who currently owns a copy and what their delivery probabilities are. This paper presents a new efficient approximation algorithm, called Distributed Max-Contribution (DMC) that performs greedy scheduling, routing and replication based only on locally and contemporarily available information. Through a simulation study based on real GPS traces tracking over 4000 taxies for about 30 days in a large city, DMC outperforms existing heuristically engineered resource allocation algorithms for DTNs.
Kyunghan Lee, Yung Yi, Jaeseong Jeong, Hyungsuk Won, Injong Rhee, Song Chong
INFOCOM1
2010 STEP: A spatio-temporal mobility model for humans walks
abstract
The movement of people is by-products of spatial and temporal correlations. People go to a place at a certain time with a purpose and they meet because they are in the same place at the same time. Extracting and representing the statistical features of spatio-temporal correlations inherent in human mobility is the goal of this paper. Most of the existing human mobility models focus on representing only the spatial features of human mobility (e.g., where and how they visit) and devote less attention to representing the temporal features (when they visit). This paper shows through GPS experiments simultaneously tracking the movement of about 200 students in two university campuses that many existing models cannot capture the temporal features and their correlations with the spatial features, and proposes a new mobility model called spatio-temporal mobility model (STEP) that aims to remedy this deficiency. The paper reports a work-in-progress for evaluating the performance of STEP and the other mobility models in representing inherent spatio-temporal features of human mobility such as the inter-contact times and diffusion speeds which are captured in the GPS experiments.
Seongik Hong, Kyunghan Lee, Injong Rhee
MASS2
2010 Mobile data offloading: how much can WiFi deliver?
abstract
This is a quantitative study on the performance of 3G mobile data offloading through WiFi networks. We recruited about 100 iPhone users from a metropolitan area and collected statistics on their WiFi connectivity during about a two and half week period in February 2010. We find that a user is in WiFi coverage for 70% of the time on average and the distributions of WiFi connection and disconnection times have a strong heavy-tail tendency with means around 2 hours and 40 minutes, respectively. Using the acquired traces, we run trace-driven simulation to measure offloading efficiency under diverse conditions e.g. traffic types, deadlines and WiFi deployment scenarios. The results indicate that if users can tolerate a two hour delay in data transfer (e.g, video and image up-loads), the network can offload 70% of the total 3G data traffic on average. We also develop a theoretical framework that permits an analytical study of the average performance of offloading. This tool is useful for network providers to obtain a rough estimate on the average performance of offloading for a given inputWiFi deployment condition.
Kyunghan Lee, Injong Rhee, Yung Yi, Song Chong
SIGCOMM1
2009 SLAW: A New Mobility Model for Human Walks
abstract
Simulating human mobility is important in mobile networks because many mobile devices are either attached to or controlled by humans and it is very hard to deploy real mobile networks whose size is controllably scalable for performance evaluation. Lately various measurement studies of human walk traces have discovered several significant statistical patterns of human mobility. Namely these include truncated power-law distributions of flights, pause-times and inter-contact times, fractal way-points, and heterogeneously defined areas of individual mobility. Unfortunately, none of existing mobility models effectively captures all of these features. This paper presents a new mobility model called SLAW (self-similar least action walk) that can produce synthetic walk traces containing all these features. This is by far the first such model. Our performance study using using SLAW generated traces indicates that SLAW is effective in representing social contexts present among people sharing common interests or those in a single community such as university campus, companies and theme parks. The social contexts are typically common gathering places where most people visit during their daily lives such as student unions, dormitory, street malls and restaurants. SLAW expresses the mobility patterns involving these contexts by fractal way points and heavy-tail flights on top of the way points. We verify through simulation that SLAW brings out the unique performance features of various mobile network routing protocols.
Kyunghan Lee, Seongik Hong, Seong Joon Kim, Injong Rhee, Song Chong
INFOCOM1
2009 Cross-Layer Survivability in WDM-Based Networks
abstract
In layered networks, a single failure at a lower layer may cause multiple failures in the upper layers. As a result, traditional schemes that protect against single failures may not be effective in cross-layer networks. In this paper, we introduce the problem of maximizing the connectivity of layered networks. We show that connectivity metrics in layered networks have significantly different meaning than their single-layer counterparts. Results that are fundamental to survivable single-layer network design, such as the Max-Flow Min-Cut theorem, are no longer applicable to the layered setting. We propose new metrics to measure connectivity in layered networks and analyze their properties. We use one of the metrics, Min Cross Layer Cut, as the objective for the survivable lightpath routing problem, and develop several algorithms to produce lightpath routings with high survivability. This allows the resulting cross-layer architecture to be resilient to failures.
Kyunghan Lee, Eytan H. Modiano
INFOCOM1
2008 On the Levy-Walk Nature of Human Mobility
abstract
We report that human walks performed in outdoor settings of tens of kilometers resemble a truncated form of Levy walks commonly observed in animals such as monkeys, birds and jackals. Our study is based on about one thousand hours of GPS traces involving 44 volunteers in various outdoor settings including two different college campuses, a metropolitan area, a theme park and a state fair. This paper shows that many statistical features of human walks follow truncated power-law, showing evidence of scale-freedom and do not conform to the central limit theorem. These traits are similar to those of Levy walks. It is conjectured that the truncation, which makes the mobility deviate from pure Levy walks, comes from geographical constraints including walk boundary, physical obstructions and traffic. None of commonly used mobility models for mobile networks captures these properties. Based on these findings, we construct a simple Levy walk mobility model which is versatile enough in emulating diverse statistical patterns of human walks observed in our traces. The model is also used to recreate similar power-law inter-contact time distributions observed in previous human mobility studies. Our network simulation indicates that the Levy walk features are important in characterizing the performance of mobile network routing performance.
Injong Rhee, Seongik Hong, Kyunghan Lee, Song Chong
INFOCOM4
2007 A Group of People Acts like a Black Body in a Wireless Mesh Network
abstract
A wireless mesh network (WMN) is being considered for commercial use in spite of several unaddressed issues. In this paper we focus on one of the most critical issues: the impact of ambient motion of entities like people on the channel characteristics and on the WMN performance. A human body in an electro-magnetic (EM) field acts as an scatterer that absorbs 60% of incident EM energy, thereby shadowing the receiver. This human body model along with the human mobility behavior gives rise to a black body (a group movement) effect that traps the incident EM wave with repetitive internal reflections. The black body theory is verified by simulating the WiSEMesh testbed in picoKAIST, a tool based on deterministic ray tube method. Experimental results show each link exhibiting a unique channel variation pattern in presence of the black body. Based on the pattern we provide several insights in WMN deployment and protocol design.
Sachin Lal Shrestha, Anseok Lee, Jinsung Lee, Dong-Wook Seo, Kyunghan Lee, Junhee Lee 0002, Song Chong, NohHoon Myung
GLOBECOM5
2007 Human Mobility Patterns and Their Impact on Delay Tolerant Networks
Injong Rhee, Seongik Hong, Kyunghan Lee, Song Chong
HotNets4
2005 Currency boosts content dissemination in noncooperative ad-hoc networks
abstract
In most of research works on the wireless ad-hoc network, it is often assumed that all nodes in the network are cooperative to relay packets. However, it is natural for nodes to be reluctant to cooperate by force due to the consumption of resources. Thus, a concept of noncooperative ad-hoc networks is being widely accepted in the latest research works. For the noncooperative ad-hoc networks, a framework which stimulates nodes to mutually cooperate was proposed by W. Yuen. It was shown that the framework utilizes more net capacity of the network than the multi-hop cooperative ad-hoc network does by exploiting data diversity and eliminating redundant bandwidth wastes for the multi-hop packet relays. In this paper, we suggest a content dissemination protocol which adopts currency and we show that adopting currency outperforms the previously proposed one through extensive simulations
Kyunghan Lee, Song Chong
GLOBECOM1