EDBT 2026 Demo / reviewers in the wild / expert
Junchen Jiang
dblp:49/8398
· DBLP profile ↗
78ranked-venue papers
13as first author
36since 2021 · last 2026
0000-0002-6877-1683ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 57 · 12 first-author · 19 since 2021Systems, architecture and hardware · 12 · 9 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeNC++: Efficient Diffusion-Enhanced Neural Codec for End-to-end Semantic Streaming at the EdgeabstractThe neural-enhanced video streaming (NeVS) has been an emerging technique to integrate neural models into video codecs for higher streaming efficiency. The state-of-the-art methods, e.g., DeNC and Gemino, typically compress videos in RGB space and restore video quality via a neural enhancement model hosted on the external media server. However, these methods are not always accessible in resource-constrained edge environments due to their heavy reliance on the media server's computation, which undermines end-to-end performance and restricts NeVS's usage boundary. This limitation raises an interesting question: is it possible to make NeVS lightweight so that all neural codec operations can be handled directly by clients' edge devices? In this paper, we present the answer yes and develop a new plug-and-play module called DeNC++, which significantly improves the compression-restoration-overhead trade-off over existing methods. Our core design philosophy is to wrap all the codec operations within a latent semantic space, in which the original high-dimensional visual signals are efficiently embedded into low-dimensional semantic representations. With this fundamental transformation, DeNC++'s neural encoder introduces the triple semantic-bitwidth-resolution compression to effectively lower the streaming traffic. Meanwhile, we make DeNC++'s neural decoder aware of the perceptual loss caused by its encoder and design tiny generative models to guarantee high restoration quality. We also strictly restrict the runtime computational overhead and accelerate the neural enhancement process, making DeNC++ compatible with commodity edge devices. Real-world evaluations reveal that DeNC++ consistently provides higher restoration quality while achieving 24-55 times higher compression ratio and 5-7 times end-to-end speedup over the latest NeVS solutions. Qihua Zhou, Wangjiang Gong, Zili Meng, Yaxiong Xie, Yaodong Huang, Junchen Jiang, Laizhong Cui |
AAAI | 6 |
| 2026 | DroidSpeak: KV Cache Sharing Across Fine-tuned Model Variants
Yuhan Liu 0004, Shaoting Feng, Zhuohan Gu, Kuntai Du, Hanchen Li, Yihua Cheng, Junchen Jiang, Shan Lu 0001, Madan Musuvathi, Esha Choukse |
NSDI | 9 |
| 2025 | Earth+: On-Board Satellite Imagery Compression Leveraging Historical Earth ObservationsabstractDue to limited downlink (satellite-to-ground) capacity, over 90% of the images captured by the earth-observation satellites are not downloaded to the ground. To overcome the downlink limitation, we present Earth+, a new on-board satellite imagery compression system that identifies and downloads only changed areas in each image compared to latest on-board reference images of the same location. The key of Earth+ is that it obtains latest on-board reference images by letting the ground stations upload images recently captured by all satellites in the constellation. To our best knowledge, Earth+ is the first system that leverages images across an entire satellite constellation to enable more images to be downloaded to the ground (by better satellite imagery compression). Our evaluation shows that to download images of the same area, Earth+ can reduce the downlink usage by 3.3× compared to state-of-the-art on-board image compression techniques without sacrificing imagery quality or using more resources (downlink, computation or storage). Kuntai Du, Yihua Cheng, Peder A. Olsen, Shadi A. Noghabi, Junchen Jiang |
ASPLOS (1) | 5 |
| 2025 | CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge FusionabstractLarge language models (LLMs) often incorporate multiple text chunks in their inputs to provide the necessary contexts. To speed up the prefill of the long LLM inputs, one can pre-compute the KV cache of a text and re-use the KV cache when the context is reused as the prefix of another LLM input. However, the reused text chunks are not always the input prefix, which makes precomputed KV caches not directly usable since they ignore the text's cross-attention with the preceding texts. Thus, the benefits of reusing KV caches remain largely unrealized. Hanchen Li, Yuhan Liu 0004, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu 0001, Junchen Jiang |
EuroSys | 9 |
| 2025 | Holmes: Localizing Irregularities in LLM Training with Mega-scale GPU Clusters
Zhiyi Yao, Pengbo Hu, Congcong Miao, Xuya Jia, Zuning Liang, Yuedong Xu 0001, Chunzhi He, Mingzhuo Chen, Xiang Li 0010, Zekun He, Yachen Wang, Xianneng Zou, Junchen Jiang |
NSDI | 14 |
| 2025 | PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model ApplicationsabstractBesides typical generative applications, like ChatGPT, GitHub Copilot, and Cursor, we observe an emerging trend that LLMs are increasingly used in traditional discriminative tasks, such as recommendation, credit verification, and data labeling. The key characteristic of these emerging use cases is that the LLM generates only a single output token, rather than an arbitrarily long sequence of tokens. We refer to this as a prefill-only workload. However, since existing LLM engines assume arbitrary output lengths, they fail to leverage the unique properties of prefill-only workloads. In this paper, we present PrefillOnly, the first LLM inference engine that improves the inference throughput and latency by fully embracing the properties of prefill-only workloads. First, since it generates only one token, PrefillOnly only needs to store the KV cache of only the last computed layer, rather than of all layers. This drastically reduces the GPU memory footprint of LLM inference and allows handling long inputs without using solutions that reduce throughput, such as cross-GPU KV cache parallelization. Second, because the output length is fixed, rather than arbitrary, PrefillOnly can precisely determine the job completion time (JCT) of each prefill-only request before it starts. This enables efficient JCT-aware scheduling policies such as shortest prefill first. PrefillOnly can process up to 4× larger queries per second without inflating the average and P99 latency. Kuntai Du, Bowen Wang 0016, Chen Zhang 0001, Qing Lan, Hejian Sang, Yihua Cheng, Yifan Qiao 0002, Ion Stoica, Junchen Jiang |
SOSP | 12 |
| 2025 | METIS: Fast Quality-Aware RAG Systems with Configuration AdaptationabstractRAG (Retrieval Augmented Generation) allows LLMs (large language models) to generate better responses with external knowledge, but using more external knowledge causes higher response delay. Prior work focuses either on reducing the response delay (e.g., better scheduling of RAG queries) or on maximizing quality (e.g., tuning the RAG workflow), but they fall short in systematically balancing the tradeoff between the delay and quality of RAG responses. To balance both quality and response delay, this paper presents METIS, the first RAG system that jointly schedules queries and adapts the key RAG configurations of each query, such as the number of retrieved text chunks and synthesis methods. Using four popular RAG-QA datasets, we show that compared to the state-of-the-art RAG optimization schemes, METIS reduces the generation latency by 1.64 – 2.54× without sacrificing generation quality. Siddhant Ray, Rui Pan 0003, Zhuohan Gu, Kuntai Du, Shaoting Feng, Ganesh Ananthanarayanan, Ravi Netravali, Junchen Jiang |
SOSP | 8 |
| 2024 | Fed2PKD: Bridging Model Diversity in Federated Learning via Two-Pronged Knowledge DistillationabstractHeterogeneous federated learning (HFL) enables collaborative learning across clients with diverse model architectures and data distributions while preserving privacy. However, existing HFL approaches often struggle to effectively address the challenges posed by model diversity, leading to suboptimal performance and limited generalization ability. This paper pro-poses Fed2PKD, a novel HFL framework that tackles these challenges through a two-pronged knowledge distillation approach. Fed2PKD combines prototypical contrastive knowledge distillation to align client embeddings with global class prototypes and semi-supervised global knowledge distillation to capture global data characteristics. Experimental results on three benchmarks (MNIST, CIFAR10, and CIFAR100) demonstrate that Fed2PKD significantly outperforms existing state-of-the-art HFL methods, achieving average improvements of up to 30.53%, 13.89%, and 5.80 % in global model accuracy, respectively. Furthermore, Fed2PKD enables personalized models for each client, adapting to their specific data distributions and model architectures while benefiting from global knowledge sharing. Theoretical analysis provides convergence guarantees for Fed2PKD under realistic assumptions. Fed2PKD represents a significant step forward in HFL, unlocking the potential for privacy-preserving collaborative learning in real-world scenarios with model and data diversity. Zaipeng Xie, Junchen Jiang, Ruiqian Han |
CLOUD | 4 |
| 2024 | Towards Domain-Specific Network Transport for Distributed DNN Training
Hao Wang 0116, Han Tian, Jingrong Chen 0004, Xinchen Wan, Jiacheng Xia, Gaoxiong Zeng, Wei Bai 0001, Junchen Jiang, Yong Wang 0046, Kai Chen 0005 |
NSDI | 8 |
| 2024 | GRACE: Loss-Resilient Real-Time Video through Neural Codecs
Yihua Cheng, Hanchen Li, Anton Arapin, Qizheng Zhang, Yuhan Liu 0004, Kuntai Du, Francis Y. Yan, Amrita Mazumdar, Nick Feamster, Junchen Jiang |
NSDI | 13 |
| 2024 | ARTEMIS: Adaptive Bitrate Ladder Optimization for Live Video Streaming
Farzad Tashtarian, Abdelhak Bentaleb, Hadi Amirpour, Sergey Gorinsky, Junchen Jiang, Hermann Hellwagner, Christian Timmerer |
NSDI | 5 |
| 2024 | ChameleonAPI: Automatic and Efficient Customization of Neural Networks for ML Applications
Yuhan Liu 0004, Chengcheng Wan 0001, Kuntai Du, Henry Hoffmann, Junchen Jiang, Shan Lu 0001, Michael Maire |
OSDI | 5 |
| 2024 | CacheGen: KV Cache Compression and Streaming for Fast Large Language Model ServingabstractAs large language models (LLMs) take on complex tasks, their inputs are supplemented with longer contexts that incorporate domain knowledge. Yet using long contexts is challenging as nothing can be generated until the whole context is processed by the LLM. While the context-processing delay can be reduced by reusing the KV cache of a context across different inputs, fetching the KV cache, which contains large tensors, over the network can cause high extra network delays. Yuhan Liu 0004, Hanchen Li, Yihua Cheng, Siddhant Ray, Qizheng Zhang, Kuntai Du, Shan Lu 0001, Ganesh Ananthanarayanan, Michael Maire, Henry Hoffmann, Ari Holtzman, Junchen Jiang |
SIGCOMM | 14 |
| 2024 | NetLLM: Adapting Large Language Models for NetworkingabstractMany networking tasks now employ deep learning (DL) to solve complex prediction and optimization problems. However, current design philosophy of DL-based algorithms entails intensive engineering overhead due to the manual design of deep neural networks (DNNs) for different networking tasks. Besides, DNNs tend to achieve poor generalization performance on unseen data distributions/environments. Duo Wu, Xianda Wang, Yaqi Qiao, Zhi Wang 0001, Junchen Jiang, Shuguang Cui, Fangxin Wang 0001 |
SIGCOMM | 5 |
| 2023 | Raising the Level of Abstraction for Time-State Analytics With the Timeline Framework
Henry Milner, Yihua Cheng, Jibin Zhan, Hui Zhang 0001, Vyas Sekar, Junchen Jiang, Ion Stoica |
CIDR | 6 |
| 2023 | Online Profiling and Adaptation of Quality Sensitivity for Internet VideoabstractA key to video streaming systems is knowing how sensitive quality of experience (QoE) is to quality metrics (e.g., buffering ratio and average bitrate). In the conventional wisdom, such quality sensitivity should be profiled by offline user studies because QoE is equally sensitive to quality metrics everywhere for an entire genre of videos. However, recent studies show that quality sensitivity varies substantially both across videos and within a video, giving rise to a new potential for improving QoE and serving more users without using more bandwidth. Unfortunately, offline profiling cannot capture the variability of quality sensitivity within a new video (e.g., a new TV show episode or live sports event), if users join to watch it within a short time window. Yihua Cheng, Hui Zhang 0001, Junchen Jiang |
SoCC | 3 |
| 2023 | OneAdapt: Fast Adaptation for Deep Learning Applications via BackpropagationabstractDeep learning inference on streaming media data, such as object detection in video or LiDAR feeds and text extraction from audio waves, is now ubiquitous. To achieve high inference accuracy, these applications typically require significant network bandwidth to gather high-fidelity data and extensive GPU resources to run deep neural networks (DNNs). While the high demand for network bandwidth and GPU resources could be substantially reduced by optimally adapting the configuration knobs, such as video resolution and frame rate, current adaptation techniques fail to meet three requirements simultaneously: adapt configurations (i) with minimum extra GPU or bandwidth overhead (ii) to reach near-optimal decisions based on how the data affects the final DNN's accuracy, and (iii) do so for a range of configuration knobs. This paper presents OneAdapt, which meets these requirements by leveraging a gradient-ascent strategy to adapt configuration knobs. The key idea is to embrace DNNs' differentiability to quickly estimate the accuracy's gradient to each configuration knob, called AccGrad. Specifically, OneAdapt estimates AccGrad by multiplying two gradients: InputGrad (i.e., how each configuration knob affects the input to the DNN) and DNNGrad (i.e., how the DNN input affects the DNN inference output). We evaluate OneAdapt across five types of configurations, four analytic tasks, and five types of input data. Compared to state-of-the-art adaptation schemes, OneAdapt cuts bandwidth usage and GPU usage by 15-59% while maintaining comparable accuracy or improves accuracy by 1-5% while using equal or fewer resources. Kuntai Du, Yuhan Liu 0004, Yitian Hao, Qizheng Zhang, Ganesh Ananthanarayanan, Junchen Jiang |
SoCC | 8 |
| 2023 | Estimating WebRTC Video QoE Metrics Without Using Application HeadersabstractThe increased use of video conferencing applications (VCAs) has made it critical to understand and support end-user quality of experience (QoE) by all stakeholders in the VCA ecosystem, especially network operators, who typically do not have direct access to client software. Existing VCA QoE estimation methods use passive measurements of application-level Real-time Transport Protocol (RTP) headers. However, a network operator does not always have access to RTP headers in all cases, particularly when VCAs use custom RTP protocols (e.g., Zoom) or due to system constraints (e.g., legacy measurement systems). Given this challenge, this paper considers the use of more standard features in the network traffic, namely, IP and UDP headers, to provide per-second estimates of key VCA QoE metrics such as frames rate and video resolution. We develop a method that uses machine learning with a combination of flow statistics (e.g., throughput) and features derived based on the mechanisms used by the VCAs to fragment video frames into packets. We evaluate our method for three prevalent VCAs running over WebRTC: Google Meet, Microsoft Teams, and Cisco Webex. Our evaluation consists of 54,696 seconds of VCA data collected from both (1), controlled in-lab network conditions, and (2) real-world networks from 15 households. We show that the ML-based approach yields similar accuracy compared to the RTP-based methods, despite using only IP/UDP data. For instance, we can estimate FPS within 2 FPS for up to 83.05% of one-second intervals in the real-world data, which is only 1.76% lower than using the application-level RTP headers. Taveesh Sharma, Tarun Mangla, Arpit Gupta, Junchen Jiang, Nick Feamster |
IMC | 4 |
| 2023 | Gemini: Divide-and-Conquer for Practical Learning-Based Internet Congestion ControlabstractLearning-based Internet congestion control algorithms have attracted much attention due to their potential performance improvement over traditional algorithms. However, such performance improvement is usually at the expense of black-box design and high computational overhead, which prevent them from large-scale deployment over production networks. To address this problem, we propose a novel Internet congestion control algorithm called Gemini. It contains a parameterized congestion control module, which is white-box designed with low computational overhead, and an online parameter optimization module, which serves to adapt the parameterized congestion control module to different networks for higher transmission performance. Extensive trace-driven emulations reveal Gemini achieves better balances between delay and throughput than state-of-the-art algorithms. Moreover, we successfully deploy Gemini over production networks. The evaluation results show that the average throughput of Gemini is 5% higher than that of Cubic (4% higher than that of BBR) over a mobile application downloading service and 61% higher than that of Cubic (33% higher than that of BBR) over a commercial network speed-test benchmarking service. Wenzheng Yang, Yan Liu 0047, Chen Tian 0001, Junchen Jiang, Lingfeng Guo |
INFOCOM | 4 |
| 2023 | RECL: Responsive Resource-Efficient Continuous Learning for Video Analytics
Mehrdad Khani Shirkoohi, Ganesh Ananthanarayanan, Kevin Hsieh, Junchen Jiang, Ravi Netravali, Yuanchao Shu, Mohammad Alizadeh, Paramvir Bahl |
NSDI | 4 |
| 2023 | Run-Time Prevention of Software Integration Failures of Machine Learning APIsabstractDue to the under-specified interfaces, developers face challenges in correctly integrating machine learning (ML) APIs in software. Even when the ML API and the software are well designed on their own, the resulting application misbehaves when the API output is incompatible with the software. It is desirable to have an adapter that converts ML API output at runtime to better fit the software need and prevent integration failures. In this paper, we conduct an empirical study to understand ML API integration problems in real-world applications. Guided by this study, we present SmartGear, a tool that automatically detects and converts mismatching or incorrect ML API output at run time, serving as a middle layer between ML API and software. Our evaluation on a variety of open-source applications shows that SmartGear detects 70% incompatible API outputs and prevents 67% potential integration failures, outperforming alternative solutions. Chengcheng Wan 0001, Yuhan Liu 0004, Kuntai Du, Henry Hoffmann, Junchen Jiang, Michael Maire, Shan Lu 0001 |
Proc. ACM Program. Lang. | 5 |
| 2023 | Enabling Edge-Cloud Video Analytics for Robotics ApplicationsabstractEmerging deep learning-based video analytics tasks demand computation-intensive neural networks and powerful computing resources on the cloud to achieve high accuracy. Due to the latency requirement and limited network bandwidth, edge-cloud systems adaptively compress the data to strike a balance between overall analytics accuracy and bandwidth consumption. However, the degraded data leads to another issue of poortail accuracy, which means the extremely low accuracy of a few semantic classes and video frames. Autonomous robotics applications especially value the tail accuracy performance but suffer using the prior edge-cloud systems. We present Runespoor, an edge-cloud video analytics system to manage the tail accuracy and enable emerging robotics applications. We train and deploy a super-resolution model tailored for the tail accuracy of analytics tasks on the server to significantly improves the performance on hard-to-detect classes and sophisticated frames. During online operation, we use an adaptive data rate controller to further improve the tail performance by instantly adjusting the data rate policy according to the video content. Our evaluation shows that Runespoor improves class-wise tail accuracy by up to 300%, frame-wise 90%/99% tail accuracy by up to 22%/54%, and greatly improves the overall accuracy and bandwidth trade-off. Weiyan Wang, Duowen Liu, Xin Jin 0008, Junchen Jiang, Kai Chen 0005 |
IEEE Trans. Cloud Comput. | 5 |
| 2023 | CocoSketch: High-Performance Sketch-Based Measurement Over Arbitrary Partial Key QueryabstractSketch-based measurement has emerged as a promising solutions due to its high accuracy and resource efficiency. Prior sketches focus on measuring single flow keys and cannot support measurement on multiple keys. This work takes a significant step towards supporting arbitrary partial key queries, which aims to provide information for any key in the predefined range of possible flow keys. The designed system, CocoSketch, casts arbitrary partial key queries to the subset sum estimation problem and makes the theoretical tools for subset sum estimation practical. CocoSketch utilizes two techniques: (1) stochastic variance minimization to significantly reduce per-packet update delay, and (2) removing circular dependencies in the per-packet update logic to make the implementation hardware-friendly. This paper extends the conference version by discussing how CocoSketch adapts to new measurement requirements, including: (1) collecting the exact information of specified flow keys, and (2) distributed measurement. CocoSketch is implemented on five popular platforms (CPU, Open vSwitch, Redis, P4, and FPGA). Experiment results show that compared to baselines that use traditional single-key sketches, CocoSketch improves average packet processing throughput by$27.2\times $and accuracy by$10.4\times $when measuring six flow keys. Ruijie Miao, Yinda Zhang 0002, Ruwen Zhang, Tong Yang 0003, Zaoxing Liu, Junchen Jiang |
IEEE/ACM Trans. Netw. | 8 |
| 2022 | Minimizing packet retransmission for real-time video analyticsabstractIn smart-city and video analytics (VA) applications, high-quality data streams (video frames) must be accurately analyzed with a low delay. Since maintaining high accuracy requires compute-intensive deep neural nets (DNNs), these applications often stream massive video data to remote, more powerful cloud servers, giving rise to a strong need for low streaming delay between video sensors and cloud servers while still delivering enough data for accurate DNN inference. In response, many recent efforts have proposed distributed VA systems that aggressively compress/prune video frames deemed less important to DNN inference, with the underlying assumptions being that (1) without increasing available bandwidth, reducing delays means sending fewer bits, and (2) the most important frames can be precisely determined before streaming. This short paper challenges both views. First, in high-bandwidth networks, the delay of real-time videos is primarily bounded by packet losses and delay jitters, so reducing bitrate is not always as effective as reducing packet retransmissions. Second, for many DNNs, the impact of missing a video frame depends not only on itself but also on which other frames have been received or lost. We argue that some changes must be made in the transport layer, to determine whether to resend a packet based on the packet's impact on DNN's inference dependent on which packets have been received. While much research is needed toward an optimal design of DNN-driven transport layer, we believe that we have taken the first step in reducing streaming delay while maintaining a high inference accuracy. Kuntai Du, Junchen Jiang |
SoCC | 3 |
| 2022 | FedDGIC: Reliable and Efficient Asynchronous Federated Learning with Gradient CompensationabstractAsynchronous federated learning is a distributed machine learning paradigm that may alleviate the impact of straggler nodes and improve the efficiency of federated training. However, some nodes can become sluggish, and node dropout may frequently happen for various reasons, such as network connection constraints, energy deficits, and system faults. Consequently, the global model may deviate from the desired convergence direction and lead to suboptimal results. This work proposes an asynchronous federated learning framework, FedDGIC, to mitigate the impact of the node dropout problem. The proposed framework can improve training efficiency by utilizing a dynamic grouping algorithm with gradient compensation. Experiments are performed in a real federated learning environment using two datasets, i.e., MNIST and CIFAR-10. Compared with three state-of-the-art methods, the proposed FedDGIC can significantly improve training efficiency and provide reliable asynchronous federated learning. Zaipeng Xie, Junchen Jiang, Zhihao Qu, Hanxiang Liu |
ICPADS | 2 |
| 2022 | Bandwidth-Efficient Multi-video Prefetching for Short Video StreamingabstractApplications that allow sharing of user-created short videos exploded in popularity in recent years. A typical short video application allows a user to swipe away the current video being watched and start watching the next video in a video queue. Such user interface causes significant bandwidth waste if users frequently swipe a video away before finishing watching. Solutions to reduce bandwidth waste without impairing the Quality of Experience (QoE) are needed. Solving the problem requires adaptively prefetching of short video chunks, which is challenging as the download strategy needs to match unknown user viewing behavior and network conditions. In our work, we first formulate the problem of adaptive multi-video prefetching in short video streaming. Then, to facilitate the integration and comparison of researchers' algorithms towards solving the problem, we design and implement a discrete-event simulator, which we release as open source. Finally, based on the organization of the Short Video Streaming Grand Challenge at ACM Multimedia 2022, we analyze and summarize the algorithms of the contestants, with the hope of promoting the research community towards addressing this problem. Xutong Zuo, Yishu Li, Mohan Xu, Wei Tsang Ooi, Jiangchuan Liu, Junchen Jiang, Xinggong Zhang, Kai Zheng 0003, Yong Cui 0001 |
ACM Multimedia | 6 |
| 2022 | Ekya: Continuous Learning of Video Analytics Models on Edge Compute Servers
Romil Bhardwaj, Zhengxu Xia, Ganesh Ananthanarayanan, Junchen Jiang, Yuanchao Shu, Nikolaos Karianakis, Kevin Hsieh, Paramvir Bahl, Ion Stoica |
NSDI | 4 |
| 2022 | Privid: Practical, Privacy-Preserving Video Analytics Queries
Frank Cangialosi, Neil Agarwal, Venkat Arun, Junchen Jiang, Srinivas Narayana, Anand D. Sarwate, Ravi Netravali |
NSDI | 4 |
| 2022 | Genet: automatic curriculum generation for learning adaptation in networkingabstractAs deep reinforcement learning (RL) showcases its strengths in networking, its pitfalls are also coming to the public's attention. Training on a wide range of network environments leads to suboptimal performance, whereas training on a narrow distribution of environments results in poor generalization. Zhengxu Xia, Francis Y. Yan, Junchen Jiang |
SIGCOMM | 4 |
| 2021 | Sayer: Using Implicit Feedback to Optimize System PoliciesabstractWe observe that many system policies that make threshold decisions involving a resource (e.g., time, memory, cores) naturally reveal additional, or implicit feedback. For example, if a system waits X min for an event to occur, then it automatically learns what would have happened if it waited < X min, because time has a cumulative property. This feedback tells us about alternative decisions, and can be used to improve the system policy. However, leveraging implicit feedback is difficult because it tends to be one-sided or incomplete, and may depend on the outcome of the event. As a result, existing practices for using feedback, such as simply incorporating it into a data-driven model, suffer from bias. Mathias Lécuyer, Sang Hoon Kim, Mihir Nanavati, Junchen Jiang, Siddhartha Sen 0001, Aleksandrs Slivkins, Amit Sharma 0007 |
SoCC | 4 |
| 2021 | Towards Performance Clarity of Edge Video Analytics
Zhujun Xiao, Zhengxu Xia, Haitao Zheng 0001, Ben Y. Zhao, Junchen Jiang |
SEC | 5 |
| 2021 | Precise error estimation for sketch-based flow measurementabstractAs a class of approximate measurement approaches, sketching algorithms have significantly improved the estimation of network flow information using limited resources. While these algorithms enjoy sound error-bound analysis under worst-case scenarios, their actual errors can vary significantly with the incoming flow distribution, making their traditional error bounds too "loose" to be useful in practice. In this paper, we propose a simple yet rigorous error estimation method to more precisely analyze the errors for posterior sketch queries by leveraging the knowledge from the sketch counters. This approach will enable network operators to understand how accurate the current measurements are and make appropriate decisions accordingly (e.g., identify potential heavy users or answer "what-if" questions to better provision resources). Theoretical analysis and trace-driven experiments show that our estimated bounds on sketch errors are much tighter than previous ones and match the actual error bounds in most cases. Peiqing Chen, Yuhan Wu 0001, Tong Yang 0003, Junchen Jiang, Zaoxing Liu |
Internet Measurement Conference | 4 |
| 2021 | Enabling Edge-Cloud Video Analytics for Robotics ApplicationsabstractEmerging deep learning-based video analytics tasks demand computation-intensive neural networks and powerful computing resources on the cloud to achieve high accuracy. Due to the latency requirement and limited network bandwidth, edge-cloud systems adaptively compress the data to strike a balance between overall analytics accuracy and bandwidth consumption. However, the degraded data leads to another issue of poor tail accuracy, which means the extremely low accuracy of a few semantic classes and video frames. Autonomous robotics applications especially value the tail accuracy performance but suffer using the prior edge-cloud systems.We present Runespoor, an edge-cloud video analytics system to manage the tail accuracy and enable emerging robotics applications. We train and deploy a super-resolution model tailored for the tail accuracy of analytics tasks on the server to significantly improves the performance on hard-to-detect classes and sophisticated frames. During online operation, we use an adaptive data rate controller to further improve the tail performance by instantly adjusting the data rate policy according to the video content. Our evaluation shows that Runespoor improves class-wise tail accuracy by up to 300%, frame-wise 90%/99% tail accuracy by up to 22%/54%, and greatly improves the overall accuracy and bandwidth trade-off. Weiyan Wang, Duowen Liu, Xin Jin 0008, Junchen Jiang, Kai Chen 0005 |
INFOCOM | 5 |
| 2021 | SENSEI: Aligning Video Streaming Quality with Dynamic User Sensitivity
Yiyang Ou, Siddhartha Sen 0001, Junchen Jiang |
NSDI | 4 |
| 2021 | CocoSketch: high-performance sketch-based measurement over arbitrary partial key queryabstractSketch-based measurement has emerged as a promising alternative to the traditional sampling-based network measurement approaches due to its high accuracy and resource efficiency. While there have been various designs around sketches, they focus on measuring one particular flow key, and it is infeasible to support many keys based on these sketches. In this work, we take a significant step towards supporting arbitrary partial key queries, where we only need to specify a full range of possible flow keys that are of interest before measurement starts, and in query time, we can extract the information of any key in that range. We design CocoSketch, which casts arbitrary partial key queries to the subset sum estimation problem and makes the theoretical tools for subset sum estimation practical. To realize desirable resource-accuracy tradeoffs in software and hardware platforms, we propose two techniques: (1) stochastic variance minimization to significantly reduce per-packet update delay, and (2) removing circular dependencies in the per-packet update logic to make the implementation hardware-friendly. We implement CocoSketch on four popular platforms (CPU, Open vSwitch, P4, and FPGA) and show that compared to baselines that use traditional single-key sketches, CocoSketch improves average packet processing throughput by 27.2x and accuracy by 10.4x when measuring six flow keys. Yinda Zhang 0002, Zaoxing Liu, Tong Yang 0003, Jizhou Li, Ruijie Miao, Peng Liu 0047, Ruwen Zhang, Junchen Jiang |
SIGCOMM | 9 |
| 2021 | BDS+: An Inter-Datacenter Data Replication System With Dynamic Bandwidth SeparationabstractMany important cloud services require replicating massive data from one datacenter (DC) to multiple DCs. While the performance of pair-wise inter-DC data transfers has been much improved, prior solutions are insufficient to optimize bulk-data multicast, as they fail to explore the rich inter-DC overlay paths that exist in geo-distributed DCs, as well as the remaining bandwidth reserved for online traffic under fixed bandwidth separation scheme. To take advantage of these opportunities, we present BDS+, a near-optimal network system for large-scale inter-DC data replication. BDS+ is an application-level multicast overlay network with a fully centralized architecture, allowing a central controller to maintain an up-to-date global view of data delivery status of intermediate servers, in order to fully utilize the available overlay paths. Furthermore, in each overlay path, it leverages dynamic bandwidth separation to make use of the remaining available bandwidth reserved for online traffic. By constantly estimating online traffic demand and rescheduling bulk-data transfers accordingly, BDS+ can further speed up the massive data multicast. Through a pilot deployment in one of the largest online service providers and large-scale real-trace simulations, we show that BDS+ can achieve 3- 5× speedup over the provider's existing system and several well-known overlay routing baselines of static bandwidth separation. Moreover, dynamic bandwidth separation can further reduce the completion time of bulk data transfers by 1.2 to 1.3 times. Yuchao Zhang 0004, Xiaohui Nie, Junchen Jiang, Wendong Wang 0003, Ke Xu 0002, Youjian Zhao, Martin J. Reed, Kai Chen 0005, Guang Yao |
IEEE/ACM Trans. Netw. | 3 |
| 2020 | Spatula: Efficient cross-camera video analytics on large camera networksabstractCameras are deployed at scale with the purpose of searching and tracking objects of interest (e.g., a suspected person) through the camera network on live videos. Such cross-camera analytics is data and compute intensive, whose costs grow with the number of cameras and time. We present Spatula, a cost-efficient system that enables scaling cross-camera analytics on edge compute boxes to large camera networks by leveraging the spatial and temporal cross-camera correlations. While such correlations have been used in computer vision community, Spatula uses them to drastically reduce the communication and computation costs by pruning search space of a query identity (e.g., ignoring frames not correlated with the query identity’s current position). Spatula provides the first system substrate on which cross-camera analytics applications can be built to efficiently harness the cross-camera correlations that are abundant in large camera deployments. Spatula reduces compute load by $8.3\times$ on an 8-camera dataset, and by $23\times-86\times$ on two datasets with hundreds of cameras (simulated from real vehicle/pedestrian traces). We have also implemented Spatula on a testbed of 5 AWS DeepLens cameras. Samvit Jain, Ganesh Ananthanarayanan, Junchen Jiang, Yuanchao Shu, Paramvir Bahl, Joseph Gonzalez 0001 |
SEC | 5 |
| 2020 | Server-Driven Video Streaming for Deep Learning InferenceabstractVideo streaming is crucial for AI applications that gather videos from sources to servers for inference by deep neural nets (DNNs). Unlike traditional video streaming that optimizes visual quality, this new type of video streaming permits aggressive compression/pruning of pixels not relevant to achieving high DNN inference accuracy. However, much of this potential is left unrealized, because current video streaming protocols are driven by the video source (camera) where the compute is rather limited. We advocate that the video streaming protocol should be driven by real-time feedback from the server-side DNN. Our insight is two-fold: (1) server-side DNN has more context about the pixels that maximize its inference accuracy; and (2) the DNN's output contains rich information useful to guide video streaming. We present DDS (DNN-Driven Streaming), a concrete design of this approach. DDS continuously sends a low-quality video stream to the server; the server runs the DNN to determine where to re-send with higher quality to increase the inference accuracy. We find that compared to several recent baselines on multiple video genres and vision tasks, DDS maintains higher accuracy while reducing bandwidth usage by upto 59% or improves accuracy by upto 9% with no additional bandwidth usage. Kuntai Du, Ahsan Pervaiz, Aakanksha Chowdhery, Qizheng Zhang, Henry Hoffmann, Junchen Jiang |
SIGCOMM | 7 |
| 2019 | Rethinking Transport Layer Design for Distributed Machine LearningabstractMotivated by the increasing scale of data, we see a growing need of high performance distributed machine learning systems. Many research works are being proposed to improve distributed machine learning performance. Jiacheng Xia, Gaoxiong Zeng, Junxue Zhang 0001, Weiyan Wang, Wei Bai 0001, Junchen Jiang, Kai Chen 0005 |
APNet | 6 |
| 2019 | Pano: optimizing 360° video streaming with a better understanding of quality perceptionabstractStreaming 360° videos requires more bandwidth than non-360° videos. This is because current solutions assume that users perceive the quality of 360° videos in the same way they perceive the quality of non-360° videos. This means the bandwidth demand must be proportional to the size of the user's field of view. However, we found several quality-determining factors unique to 360° videos, which can help reduce the bandwidth demand. They include the moving speed of a user's viewpoint (center of the user's field of view), the recent change of video luminance, and the difference in depth-of-fields of visual objects around the viewpoint. Yu Guan 0005, Chengyuan Zheng, Xinggong Zhang, Zongming Guo, Junchen Jiang |
SIGCOMM | 5 |
| 2019 | Zooming in on wide-area latencies to a global cloud providerabstractThe network communications between the cloud and the client have become the weak link for global cloud services that aim to provide low latency services to their clients. In this paper, we first characterize WAN latency from the viewpoint of a large cloud provider Azure, whose network edges serve hundreds of billions of TCP connections a day across hundreds of locations worldwide. In particular, we focus on instances of latency degradation and design a tool, BlameIt, that enables cloud operators to localize the cause (i.e., faulty AS) of such degradation. BlameIt uses passive diagnosis, using measurements of existing connections between clients and the cloud locations, to localize the cause to one of cloud, middle, or client segments. Then it invokes selective active probing (within a probing budget) to localize the cause more precisely. We validate BlameIt by comparing its automatic fault localization results with that arrived at by network engineers manually, and observe that BlameIt correctly localized the problem in all the 88 incidents. Further, BlameIt issues 72X fewer active probes than a solution relying on active probing alone, and is deployed in production at Azure. Sundararajan Renganathan, Ganesh Ananthanarayanan, Junchen Jiang, Venkat N. Padmanabhan, Manuel Schröder, Matt Calder, Arvind Krishnamurthy |
SIGCOMM | 4 |
| 2019 | E2E: embracing user heterogeneity to improve quality of experience on the webabstractConventional wisdom states that to improve quality of experience (QoE), web service providers should reduce the median or other percentiles of server-side delays. This work shows that doing so can be inefficient due to user heterogeneity in how the delays impact QoE. From the perspective of QoE, the sensitivity of a request to delays can vary greatly even among identical requests arriving at the service, because they differ in the wide-area network latency experienced prior to arriving at the service. In other words, saving 50ms of server-side delay affects different users differently. Siddhartha Sen 0001, Daniar Heri Kurniawan, Haryadi S. Gunawi, Junchen Jiang |
SIGCOMM | 5 |
| 2018 | Demystifying Deep Learning in NetworkingabstractWe are witnessing a surge of efforts in networking community to develop deep neural networks (DNNs) based approaches to networking problems. Most results so far have been remarkably promising, which is arguably surprising given how intensively these problems have been studied before. Despite these promises, there has not been much systematic work to understand the inner workings of these DNNs trained in networking settings, their generalizability in different workloads, and their potential synergy with domain-specific knowledge. The problem of model opacity would eventually impede the adoption of DNN-based solutions in practice. This position paper marks the first attempt to shed light on the interpretability of DNNs used in networking problems. Inspired by recent research in ML towards interpretable ML models, we call upon this community to similarly develop techniques and leverage domain-specific insights to demystify the DNNs trained in networking settings, and ultimately unleash the potential of DNNs in an explainable and reliable way. Ying Zheng 0004, Xinyu You, Yuedong Xu 0001, Junchen Jiang |
APNet | 5 |
| 2018 | BDS: a centralized near-optimal overlay network for inter-datacenter data replicationabstractMany important cloud services require replicating massive data from one datacenter (DC) to multiple DCs. While the performance of pair-wise inter-DC data transfers has been much improved, prior solutions are insufficient to optimize bulk-data multicast, as they fail to explore the capability of servers to store-and-forward data, as well as the rich inter-DC overlay paths that exist in geo-distributed DCs. To take advantage of these opportunities, we present BDS, an application-level multicast overlay network for large-scale inter-DC data replication. At the core of BDS is a fully centralized architecture, allowing a central controller to maintain an up-to-date global view of data delivery status of intermediate servers, in order to fully utilize the available overlay paths. To quickly react to network dynamics and workload churns, BDS speeds up the control algorithm by decoupling it into selection of overlay paths and scheduling of data transfers, each can be optimized efficiently. This enables BDS to update overlay routing decisions in near realtime (e.g., every other second) at the scale of multicasting hundreds of TB data over tens of thousands of overlay paths. A pilot deployment in one of the largest online service providers shows that BDS can achieve 3-5 x speedup over the provider's existing system and several well-known overlay routing baselines. Yuchao Zhang 0004, Junchen Jiang, Ke Xu 0002, Xiaohui Nie, Martin J. Reed, Guang Yao, Kai Chen 0005 |
EuroSys | 2 |
| 2018 | Chameleon: scalable adaptation of video analyticsabstractApplying deep convolutional neural networks (NN) to video data at scale poses a substantial systems challenge, as improving inference accuracy often requires a prohibitive cost in computational resources. While it is promising to balance resource and accuracy by selecting a suitable NN configuration (e.g., the resolution and frame rate of the input video), one must also address the significant dynamics of the NN configuration's impact on video analytics accuracy. We present Chameleon, a controller that dynamically picks the best configurations for existing NN-based video analytics pipelines. The key challenge in Chameleon is that in theory, adapting configurations frequently can reduce resource consumption with little degradation in accuracy, but searching a large space of configurations periodically incurs an overwhelming resource overhead that negates the gains of adaptation. The insight behind Chameleon is that the underlying characteristics (e.g., the velocity and sizes of objects) that affect the best configuration have enough temporal and spatial correlation to allow the search cost to be amortized over time and across multiple video feeds. For example, using the video feeds of five traffic cameras, we demonstrate that compared to a baseline that picks a single optimal configuration offline, Chameleon can achieve 20-50% higher accuracy with the same amount of resources, or achieve the same accuracy with only 30--50% of the resources (a 2-3X speedup). Junchen Jiang, Ganesh Ananthanarayanan, Peter Bodík, Siddhartha Sen 0001, Ion Stoica |
SIGCOMM | 1 |
| 2017 | Biases in Data-Driven Networking, and What to Do About ThemabstractRecent efforts highlight the promise of data-driven approaches to optimize network decisions. Many such efforts use trace-driven evaluation; i.e., running offline analysis on network traces to estimate the potential benefits of different policies before running them in practice. Unfortunately, such frameworks can have fundamental pitfalls (e.g., skews due to previous policies that were used in the data collection phase and insufficient data for specific subpopulations) that could lead to misleading estimates and ultimately suboptimal decisions. In this paper, we shed light on such pitfalls and identify a promising roadmap to address these pitfalls by leveraging parallels in causal inference, namely the Doubly Robust estimator. Mihovil Bartulovic, Junchen Jiang, Sivaraman Balakrishnan, Vyas Sekar, Bruno Sinopoli |
HotNets | 2 |
| 2017 | Pytheas: Enabling Data-Driven Quality of Experience Optimization Using Group-Based Exploration-Exploitation
Junchen Jiang, Vyas Sekar, Hui Zhang 0001 |
NSDI | 1 |
| 2016 | CFA: A Practical Prediction System for Video QoE Optimization
Junchen Jiang, Vyas Sekar, Henry Milner, Davis Shepherd, Ion Stoica, Hui Zhang 0001 |
NSDI | 1 |
| 2016 | Via: Improving Internet Telephony Call Quality Using Predictive Relay SelectionabstractInteractive real-time streaming applications such as audio-video conferencing, online gaming and app streaming, place stringent requirements on the network in terms of delay, jitter, and packet loss. Many of these applications inherently involve client-to-client communication, which is particularly challenging since the performance requirements need to be met while traversing the public wide-area network (WAN). This is different from the typical situation of cloud-to-client communication, where the WAN can often be bypassed by moving a communication end-point to a cloud “edge”, close to the client. Can we nevertheless take advantage of cloud resources to improve the performance of real-time client-to-client streaming over the WAN? Junchen Jiang, Rajdeep Das, Ganesh Ananthanarayanan, Philip A. Chou, Venkat N. Padmanabhan, Vyas Sekar, Esbjorn Dominique, Marcin Goliszewski, Dalibor Kukoleca, Renat Vafin, Hui Zhang 0001 |
SIGCOMM | 1 |
| 2016 | CS2P: Improving Video Bitrate Selection and Adaptation with Data-Driven Throughput PredictionabstractBitrate adaptation is critical in ensuring good users’ quality-of-experience (QoE) in Internet video delivery system. Several efforts have argued that accurate throughput prediction can dramatically improve (1) initial bitrate selection for low startup delay and high initial resolution; (2) midstream bitrate adaptation for high QoE. However, prior ef- forts did not systematically quantify real-world throughput predictability or develop good prediction algorithms. To bridge this gap, this paper makes three key technical contributions: First, we analyze the throughput characteristics in a dataset with 20M+ sessions. We find: (a) Sessions sharing similar key features (e.g., ISP, region) present similar initial values and dynamical patterns; (b) There is a natural “stateful” dynamical behavior within a given session. Second, building on these insights, we develop CS2P, a better throughput prediction system. CS2P leverages data-driven approach to learn (a) clusters of similar sessions, (b) an initial throughput predictor, and (c) a Hidden-Markov-Model based midstream predictor modeling the stateful evolution of throughput. Third, we develop a prototype system and show by trace-driven simulation and real-world experiments that CS2P outperforms state-of-art by 40% and 50% median pre- diction error respectively for initial and midstream through- put and improves QoE by 14% over buffer-based adaptation algorithm. Yi Sun 0004, Xiaoqi Yin, Junchen Jiang, Vyas Sekar, Fuyuan Lin, Nanshu Wang, Bruno Sinopoli |
SIGCOMM | 3 |
| 2015 | C3: Internet-Scale Control Plane for Video Quality Optimization
Aditya Ganjam, Faisal Siddiqui, Jibin Zhan, Ion Stoica, Junchen Jiang, Vyas Sekar, Hui Zhang 0001 |
NSDI | 6 |
| 2015 | Practical, Real-time Centralized Control for CDN-based Live Video DeliveryabstractLive video delivery is expected to reach a peak of 50 Tbps this year. This surging popularity is fundamentally changing the Internet video delivery landscape. CDNs must meet users' demands for fast join times, high bitrates, and low buffering ratios, while minimizing their own cost of delivery and responding to issues in real-time. Wide-area latency, loss, and failures, as well as varied workloads ("mega-events" to long-tail), make meeting these demands challenging. Matthew K. Mukerjee, David Naylor, Junchen Jiang, Dongsu Han, Srinivasan Seshan, Hui Zhang 0001 |
SIGCOMM | 3 |
| 2014 | EONA: Experience-Oriented Network ArchitectureabstractThere is a growing recognition among researchers, industry practitioners, and service providers of the need to optimize user-perceived application experience. Network infrastructure owners (i.e., ISPs) have traditionally been left out of this equation, leading to repeated tussles between content providers and ISPs. In parallel, application providers have to deploy complex workarounds that reverse engineer the network's impact on application-level metrics. In this work, we make the case for EONA, a new network paradigm where application providers and network providers can collaborate meaningfully to improve application experience. We observe a confluence of technology trends that are enablers for EONA: the ability to collect large volumes of client-side application measurements, the emergence of novel "big data" platforms for real-time analytics, and new control plane capabilities for ISPs (e.g., SDN, IXPs, NFV). We highlight the challenges and opportunities in designing suitable EONA interfaces between infrastructure and application providers and EONA-enhanced control loops that leverage these interfaces to optimize user experience. Junchen Jiang, Vyas Sekar, Ion Stoica, Hui Zhang 0001 |
HotNets | 1 |
| 2014 | Using Video-Based Measurements to Generate a Real-Time Network Traffic MapabstractWe envision a real-time network traffic map for the Internet, where each network link is annotated with its capacity and its current utilization, with an interface that networked applications can query to inform their control decisions. While this goal is simple to state, it has been out of our reach due to concerns over measurement overhead and coverage. Our insight is that the rise of Internet video and the availability of measurements from video players present an unprecedented opportunity to address these issues. We outline a preliminary roadmap to build on this opportunity to realize a global traffic map. Yi Sun 0004, Junchen Jiang, Vyas Sekar, Hui Zhang 0001, Fuyuan Lin, Nanshu Wang |
HotNets | 2 |
| 2014 | Enabling near real-time central control for live video delivery in CDNsabstractUser-created live video streaming is marking a fundamental shift in the workload of live video delivery. However, live-video-specific challenges and the viral nature of user-created content makes it difficult for current CDNs to deliver 1) high-quality, 2) highly-scalable, and 3) highly-responsive service. We present the design and implementation of VDN, a new control plane for CDNs designed to optimize the delivery of live streams within the CDN. VDN satisfies these requirements by using two approaches: 1) optimizing directly for video quality (not just throughput) and 2) combining centralized control with local control, allowing VDN to adapt to traffic dynamics and network failures at fine timescales. Matthew K. Mukerjee, JungAh Hong, Junchen Jiang, David Naylor, Dongsu Han, Srinivasan Seshan, Hui Zhang 0001 |
SIGCOMM | 3 |
| 2014 | Kangaroo: Accelerating String Matching by Running Multiple Collaborative Finite State MachinesabstractString matching is a key technique for network security applications such as network intrusion detection systems and antivirus scanners, where the payload of every packet is inspected against thousands of patterns in real time. As the transmission rate of Internet links is getting higher and higher, the speed of matching engines is required to be faster and faster. Existing deterministic finite automaton (DFA)-based approaches achieve high throughput at the expense of extremely expensive memory cost; therefore, they are not suitable for the scenarios where only limited on-chip memory resources are available. To achieve fast matching speed while controlling memory expense, in this paper, we propose Kangaroo, a compact string matching scheme that scans multiple characters each time by running multiple small-sized finite state machines in parallel. Specifically, Kangaroo processes k consecutive characters mostly in one cycle by accessing k different memories in parallel, where k is a predefined factor that can be tuned based on the requirement of applications. Kangaroo is memory efficient. Experimental evaluations on Snort and ClamAV rule sets show that a tenfold increase in speed can be practically achieved by a single Kangaroo matching engine with a reduced memory cost comparing with the state-of-the-art DFA-based approaches. Xiaofei Wang 0006, Bin Liu 0001, Junchen Jiang, Yang Xu 0010, Yi Wang 0004, Xiaojun Wang 0001 |
IEEE J. Sel. Areas Commun. | 3 |
| 2014 | TFA: A Tunable Finite Automaton for Pattern Matching in Network Intrusion Detection SystemsabstractDeterministic finite automatons (DFAs) and nondeterministic finite automatons (NFAs) are two typical automatons used in the network intrusion detection system. Although they both perform regular expression matching, they have quite different performance and memory usage properties. DFAs provide fast and deterministic matching performance but suffer from the well-known state explosion problem. NFAs are compact, but their matching performance is unpredictable and with no worst case guarantee. In this paper, we propose a new automaton representation of regular expressions, called tunable finite automaton (TFA), to deal with the DFAs' state explosion problem and the NFAs' unpredictable performance problem. Different from a DFA, which has only one active state, a TFA allows multiple concurrent active states. Thus, the total number of states required by the TFA to track the matching status is much smaller than that required by the DFA. Different from an NFA, a TFA guarantees that the number of concurrent active states is bounded by a bound factor b that can be tuned during the construction of the TFA according to the needs of the application for speed and storage. Simulation results based on regular expression rule sets from Snort and Bro show that, with only two concurrent active states, a TFA can achieve significant reductions in the number of states and memory usage, e.g., a 98% reduction in the number of states and a 95% reduction in memory space. Yang Xu 0010, Junchen Jiang, Rihua Wei, Yang Song 0031, H. Jonathan Chao |
IEEE J. Sel. Areas Commun. | 2 |
| 2014 | Improving Fairness, Efficiency, and Stability in HTTP-Based Adaptive Video Streaming With FestiveabstractModern video players today rely on bit-rate adaptation in order to respond to changing network conditions. Past measurement studies have identified issues with today's commercial players when multiple bit-rate-adaptive players share a bottleneck link with respect to three metrics: fairness, efficiency, and stability. Unfortunately, our current understanding of why these effects occur and how they can be mitigated is quite limited. In this paper, we present a principled understanding of bit-rate adaptation and analyze several commercial players through the lens of an abstract player model consisting of three main components: bandwidth estimation, bit-rate selection, and chunk scheduling. Using framework, we identify the root causes of several undesirable interactions that arise as a consequence of overlaying the video bit-rate adaptation over HTTP. Building on these insights, we develop a suite of techniques that can systematically guide the tradeoffs between stability, fairness, and efficiency and thus lead to a general framework for robust video adaptation. We pick one concrete instance from this design space and show that it significantly outperforms today's commercial players on all three key metrics across a range of experimental scenarios. Junchen Jiang, Vyas Sekar, Hui Zhang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2013 | Shedding light on the structure of internet video quality problems in the wildabstractThe key role that video quality plays in impacting user engagement, and consequently providers' revenues, has motivated recent efforts in improving the quality of Internet video. This includes work on adaptive bitrate selection, multi-CDN optimization, and global control plane architectures. Before we embark on deploying these designs, we need to first understand the nature of video of quality problems to see if this complexity is necessary, and if simpler approaches can yield comparable benefits. Junchen Jiang, Vyas Sekar, Ion Stoica, Hui Zhang 0001 |
CoNEXT | 1 |
| 2012 | Improving fairness, efficiency, and stability in HTTP-based adaptive video streaming with FESTIVEabstractMany commercial video players rely on bitrate adaptation logic to adapt the bitrate in response to changing network conditions. Past measurement studies have identified issues with today's commercial players with respect to three key metrics---efficiency, fairness, and stability---when multiple bitrate-adaptive players share a bottleneck link. Unfortunately, our current understanding of why these effects occur and how they can be mitigated is quite limited. Junchen Jiang, Vyas Sekar, Hui Zhang 0001 |
CoNEXT | 1 |
| 2012 | Scalable Name Lookup in NDN Using Effective Name Component EncodingabstractName-based route lookup is a key function for Named Data Networking (NDN). The NDN names are hierarchical and have variable and unbounded lengths, which are much longer than IPv4/6 address, making fast name lookup a challenging issue. In this paper, we propose an effective Name Component Encoding (NCE) solution with the following two techniques: (1) A code allocation mechanism is developed to achieve memory-efficient encoding for name components, (2) We apply an improved State Transition Arrays to accelerate the longest name prefix matching and design a fast and incremental update mechanism which satisfies the special requirements of NDN forwarding process, namely to insert, modify, and delete name prefixes frequently. Furthermore, we analyze the memory consumption and time complexity of NCE. Experimental results on a name set containing 3,000,000 names demonstrate that compared with the character trie NCE reduces overall 30% memory. Besides, NCE performs a few millions lookups per second (on an Intel 2.8 GHz CPU), a speedup of over 7 times compared with the character trie. Our evaluation results also show that NCE can scale up to accommodate the potential future growth of the name sets. Yi Wang 0004, Keqiang He, Huichen Dai, Wei Meng 0001, Junchen Jiang, Bin Liu 0001, Yan Chen 0004 |
ICDCS | 5 |
| 2012 | Reducing power of traffic manager in routers via dynamic on/off-chip schedulingabstractGreen networking in the Internet becomes increasingly important. In a high-performance router, the dominant power consumer on the Internet, half of its total power usage goes into the line-cards, where the traffic managers inside consume most of it. In this paper, we propose an energy-efficient design on the traffic manager architecture for packet buffering and storage. Unlike traditional routers where packets are always kept in off-chip memory, we propose a dynamic on-chip and off-chip scheduling mechanism, called Dynamic Packet Manager (DPM), to reduce both peak and average power consumption caused by the traffic manager. DPM buffers packets in a small on-chip memory in the light-traffic period, and activates the off-chip memory on when the on-chip memory is to overflow. In this design, when the traffic is light, the off-chip memory is put into power saving state by clock gating so that the average power consumption is reduced. With an on-chip flow based and off-chip class-based design, DPM can save one off-chip memory otherwise used for the per-flow index information storage, therefore further reduce the peak power usage. We present the theoretic analysis guiding the implementation of the DPM mechanism. Experiments on three prototypes implemented on different hardware show that the peak and average power consumptions can be reduced by 27.9% and 37.5% respectively, along with less on-chip memory cost. Besides, the traffic manger with DPM shows better performance on average packet scheduling delay than the one without DPM. Jindou Fan, Chengchen Hu, Keqiang He, Junchen Jiang, Bin Liu 0001 |
INFOCOM | 4 |
| 2012 | Tracking millions of flows in high speed networks for application identificationabstractToday's Internet applications exhibit increased diversity, while the Internet routers are still oblivious to this trend. To improve the end-to-end application QoS, one solution is to embed the application information explicitly in packet headers, but it will bring global changes. Another local solution is router-assisted traffic differentiation. To achieve this, the functionalities including packet identification and flow tracking inside the router are required. While most existing studies focus on the former, fewer efforts are put on the later. Given a large flow table is involved, how to track millions of concurrent flows in a cost-effective manner on a router's line card raises a great space-time challenge. To address this, we design an on-chip/off-chip flow tracking system to accommodate millions of flows and achieve the throughput at tens of Gigabits. By exploiting temporal locality and heavy-tailedness of Layer-4 traffic, we design the Adaptive Least Frequently Evicted (ALFE) replacement policy to catch elephant flows, therefore maintain a high cache hit rate. To alleviate performance penalty due to the cache misses, we organize the flow table in a fixed-allocated manner to fully utilize modern DRAM's burst feature. We have implemented a research prototype using FPGA for performance evaluation. The experiment results show that our system can reach 80% hit rate with a small-sized cache of 16K entries, while achieving 70Mpps throughput. This enables backbone line rate processing. Further, more than 40% power saving can be achieved by our system, which is fast and accurate with only 3% FPGA resource usage. Tian Pan 0001, Xiaoyu Guo 0008, Junchen Jiang, Hao Wu 0023, Bin Liu 0001 |
INFOCOM | 4 |
| 2012 | A case for a coordinated internet video control planeabstractVideo traffic already represents a significant fraction of today's traffic and is projected to exceed 90% in the next five years. In parallel, user expectations for a high quality viewing experience (e.g., low startup delays, low buffering, and high bitrates) are continuously increasing. Unlike traditional workloads that either require low latency (e.g., short web transfers) or high average throughput (e.g., large file transfers), a high quality video viewing experience requires sustained performance over extended periods of time (e.g., tens of minutes). This imposes fundamentally different demands on content delivery infrastructures than those envisioned for traditional traffic patterns. Our large-scale measurements over 200 million video sessions show that today's delivery infrastructure fails to meet these requirements: more than 20% of sessions have a rebuffering ratio ≥ 10% and more than 14% of sessions have a video startup delay ≥ 10s. Using measurement-driven insights, we make a case for a video control plane that can use a global view of client and network conditions to dynamically optimize the video delivery in order to provide a high quality viewing experience despite an unreliable delivery infrastructure. Our analysis shows that such a control plane can potentially improve the rebuffering ratio by up to 2× in the average case and by more than one order of magnitude under stress. Florin Dobrian, Henry Milner, Junchen Jiang, Vyas Sekar, Ion Stoica, Hui Zhang 0001 |
SIGCOMM | 4 |
| 2012 | MOIST: A Scalable and Parallel Moving Object Indexer with School Tracking abstractLocation-Based Service (LBS) is rapidly becoming the next ubiquitous technology for a wide range of mobile applications. To support applications that demand nearest-neighbor and history queries, an LBS spatial indexer must be able to efficiently update, query, archive and mine location records, which can be in contention with each other. In this work, we propose MOIST, whose baseline is a recursive spatial partitioning indexer built upon BigTable. To reduce update and query contention, MOIST groups nearby objects of similar trajectory into the same school, and keeps track of only the history of school leaders. This dynamic clustering scheme can eliminate redundant updates and hence reduce update latency. To improve history query processing, MOIST keeps some history data in memory, while it flushes aged data onto parallel disks in a locality-preserving way. Through experimental studies, we show that MOIST can support highly efficient nearest-neighbor and history queries and can scale well with an increasing number of users and update frequency. Junchen Jiang, Hongji Bao, Edward Y. Chang |
Proc. VLDB Endow. | 1 |
| 2011 | Parallel Name Lookup for Named Data NetworkingabstractName-based route lookup is a key function for Named Data Networking (NDN). The NDN names are hierarchical and have variable and unbounded lengths, which are much longer than IPv4/6 address, making fast name lookup a challenging issue. In this paper, we propose a parallel architecture for NDN name lookup called Parallel Name Lookup (PNL) which leverages hardware parallelism to achieve high lookup speedup while keeping a low and controllable memory redundancy. The core of PNL is an allocation algorithm that maps the logically tree-based structure to physically parallel modules, with low computational complexity. We evaluate the PNL's performance and show that PNL dramatically accelerates the name lookup process. Furthermore, with certain knowledge of prior probability, the speedup can be significantly improved. Yi Wang 0004, Huichen Dai, Junchen Jiang, Keqiang He, Wei Meng 0001, Bin Liu 0001 |
GLOBECOM | 3 |
| 2011 | StriD²FA: Scalable Regular Expression Matching for Deep Packet InspectionabstractDeep packet inspection (DPI) has become one of the key components of a Network Intrusion Detection System (NIDS) and it compares packet content to a set of rules written in regular expression. The need to keep up with ever-increasing line speed has forced NIDS designers to move to hardware or high-speed memory where memory resources are limited. In this paper, we present LBM, a novel accelerating scheme for regular expression matching which converts the original byte stream into much shorter integer stream and then matches it with a variant of DFA, called StriD2FA. In the instance of LBM that we realize, 10 to 15 speedup is reasonable while the memory is much smaller than traditional DFA. Xiaofei Wang 0006, Junchen Jiang, Yi Tang 0002, Bin Liu 0001, Xiaojun Wang 0001 |
ICC | 2 |
| 2011 | Measurements on movie distribution behavior in Peer-to-Peer networksabstractPeer-to-Peer (P2P) mode dominates the way that files are shared over the Internet today. A measurement study on the user behavior during the P2P file sharing is important and helpful to better understand and design P2P networks. In this paper, we developed a method to collect information about peers and connections in movie sharing at the BitTorrent client side. Based on the collected data, we have derived 5 observations in the influence upon peers and connections distributions over geographic areas (at different levels of continents, countries, cities) by differences of population, GDP (Gross Domestic Product), time zone and life style. Xiaofei Wang 0006, Xiaojun Wang 0001, Chengchen Hu, Keqiang He, Junchen Jiang, Bin Liu 0001 |
Integrated Network Management | 5 |
| 2011 | S3: Smart selection of sampling function for passive network measurementabstractFlow size statistics is a fundamental task of passive measurement. In order to bound the estimation error of passive measurement for both small and large flows, previous probabilistic counter updating algorithms used linear or nonlinear sampling function to automatically adjust the sampling rate. However, each of these methods employed a pre-set and fixed sampling function during the measurement period. As a result, the performance would vary for different flow distributions. In this paper, we propose a Smart Selection Sampling (S3) approach, which can tune the sampling function to reach a comparatively lower relative error. The key component of S3 is a heuristic algorithm leveraging the flow distribution information to determine a better sampling function so as to achieve better measurement accuracy. Experiments under real trace and synthetic traces demonstrate that S3 is more accurate than the previous work if given the same memory sizes to accommodate flow statistics counters. Chengchen Hu, Junchen Jiang |
LCN | 3 |
| 2010 | A2C: Anti-Attack Counters for Traffic MeasurementabstractFlow-level sampling methods have been widely studied and extensively employed in network traffic measurement systems. However, traffic anomalies are becoming more prevalent and severe in the Internet, which pose great challenges to the traffic measurement. Existing solutions targeted at such scenario have either low accuracy or high memory usage. In this paper, we propose a two-stage sampling approach-Anti Attack Counters (A2C) and an efficient parameter adapting method to solve the problem. The proposed sampling mechanism can adapt to the network condition automatically and collect more information even under severe traffic attacks. Theoretical analysis on accuracy and resource requirement is presented in our work. Furthermore, we validate our approach using both synthetic and real traces. The experimental results demonstrate that A2C is of high resilience while providing significantly improved measurement accuracy with reduced memory occupation comparing with other existing anti-attack countermeasures. Keqiang He, Chengchen Hu, Junchen Jiang, Yachao Zhou, Bin Liu 0001 |
GLOBECOM | 3 |
| 2010 | Skip Finite Automaton: A Content Scanning Engine to Secure Enterprise NetworksabstractToday's file sharing networks are creating potential security problems to enterprise networks, i.e., the leakage of confidential documents. In order to prevent such leakage, we propose the Data Leakage Prevention System (DLPS) which is applied at the entrance of the enterprise network to filter out the outgoing sensitive information. The DLPS is based on a content scanning engine which defines a new type of matching problem, called longest overlap matching which also exits in many other applications as a basic problem where contents are delivered by small blocks. We study the problem by comparing it with the traditional pattern matching problem in Deep Packet Inspection (DPI) of Network Intrusion Detection Systems (NIDS) whose solutions are based on finite automata. We develop a new finite automata representation called Skip-Finite Automata (Skip-FA) which detects the packets carrying sensitive information by using default transitions to implicitly track the overlapping parts between packets' payloads and sensitive files. The simulation results shows that our system achieves a matching speed of about 10B+ per memory access for small file set (>;20KB) and 100B+ per memory access for large file set (>;2500KB). We also find that the memory consumption of Skip-FA is almost the same to that of the original files. Junchen Jiang, Yi Tang 0002, Bin Liu 0001, Yang Xu 0010, Xiaofei Wang 0006 |
GLOBECOM | 1 |
| 2010 | Independent Parallel Compact Finite Automatons for Accelerating Multi-String MatchingabstractMulti-string matching is a key technique for implementing network security applications like Network Intrusion Detection Systems (NIDS). Existing DFA-based approaches always tradeoff between memory and throughput, and fail to has the best of both worlds. This paper extends the classic longest prefix principle from single-character to multi-character string matching and proposes a multi-string matching acceleration scheme named Independent Parallel Compact Finite Automata (PC-FA). In the scheme, DFA is divided into k PC-FAs, each of which can process one character from the input stream, achieving a speedup up to k with reduced memory occupation. Theoretical proof is given for the equivalency between traditional DFA and PC-FA approach. Experimental evaluations show that seven times of speedup can be practically achieved with a reduced memory size than up-to-date DFA-based compression approaches. Yi Tang 0002, Junchen Jiang, Xiaofei Wang 0006, Bin Liu 0001, Yang Xu 0010 |
GLOBECOM | 2 |
| 2010 | Cache-Based Scalable Deep Packet Inspection with Predictive AutomatonabstractRegular expression (Regex) becomes the standard signature language for security and application detection. Deterministic finite automata (DFAs) are widely used to perform regex matching in linear time. Previously researches mostly focus on how to compress DFA to reduce memory requirements in recent years. However, memory requirement is not the only problem caused by DFA explosion when implementation DFA matching system. In this paper, we propose a new issue in DFA matching procedure. We notice that the DFA produced from regex never considers the physical locality of logical neighbor, which results in a low cache hit rate when using cache as matching accelerator. This problem becomes severe for current increasingly complex security regex which producing huge DFA with nearly no locality in physical location. We propose to solve this problem through reordering the state number of existing DFA and further put forward two methods on reordering DFA from different viewpoints. In our algorithms, we achieve more than twice cache hit rate compared with traditional method. Moreover, our methods will not affect the existing matching system. Hence, all the cache hit rate improvement is achieved without any cost in wire speed matching. Yi Tang 0002, Junchen Jiang, Xiaofei Wang 0006, Yi Wang 0004, Bin Liu 0001 |
GLOBECOM | 2 |
| 2010 | Parallel Architecture for High Throughput DFA-Based Deep Packet InspectionabstractMulti-pattern matching is a key technique for implementing network security applications such as Network Intrusion Detection/Protection Systems (NIDS/NIPSes) where every packet is inspected against predefined attack signatures written in regular expressions (regexes). To this end, Deterministic Finite Automaton (DFA) is widely used for multi-regex matching, but existing DFAbased researches have claimed high throughput at an expenses of extremely high memory cost. In this paper, we propose a parallel architecture of DFA called Parallel DFA (PDFA), using multiple flow aggregations to increase the throughput with nearly no extra memory cost. The basic idea is to selectively store the DFA in multiple memory modules which can be accessed in parallel and to explore the potential parallelism. The memory cost of our system in both the average cases and the worst cases is analyzed, optimized and evaluated by numerical results. The evaluation shows that we obtain an average speedup of about 0.5k to 0.7k where k is the number of parallel memory modules under our synthetic trace and compressed real trace in a statistical average case, compared with the traditional DFA-based matching approaches. Junchen Jiang, Xiaofei Wang 0006, Keqiang He, Bin Liu 0001 |
ICC | 1 |
| 2010 | Pattern-Based DFA for Memory-Efficient and Scalable Multiple Regular Expression MatchingabstractIn Network Intrusion Detection System, De-terministic Finite Automaton (DFA) is widely used to compare packet content at a constant speed against a set of patterns specified in regular expressions (regex patterns). However, combining many regex patterns into a single DFA causes a serious state explosion. Partitioning the pat-tern set into several subsets, each of which produces a small DFA, is a practical way to deflate the state explosion. In this paper, we propose a regex pattern grouping scheme based on a new DFA model called Pattern-Based DFA (P-DFA) which supports efficient pattern-based op-erations, such as insertion, deletion, and etc. By using these basic operations, one can easily measure the state explo-sion when combining a set of regex patterns into a single DFA. Based on the privilege, we develop regex grouping algorithms for mitigating the state explosion in parallel and sequential matching environments, respectively. The evaluation shows that under the same constraints, our ap-proach requires only half the number of groups compared with the most well-known algorithms. Junchen Jiang, Yang Xu 0010, Tian Pan 0001, Yi Tang 0002, Bin Liu 0001 |
ICC | 1 |
| 2010 | Deflation DFA: Remembering History is AdequateabstractThere is an increasing demand for network devices to perform deep packet inspection (DPI) to enhance network security. In DPI the packet payload is compared against a set of predefined patterns which can be specified using regular expressions (regexes). It is well-known that mapping regexes to deterministic finite automata (DFA) will suffer from the state explosion problem. Through observation, we attribute DFA explosion to the necessity of remembering matching history. In this paper, we investigate how to record the matching history efficiently and propose an extended DFA approach for regex matching called fcq-FA, which can make a memory size reduction of about 1000 times with a fully automated approach. In fcq-FA, we use pipeline queues and counters to help recording the matching history. Hence, state explosion caused by Kleene closure and repetitions can be definitely avoided. Further, it achieves a fully automated signature compilation with polynomial running time and space. Yi Tang 0002, Tianfan Xue, Junchen Jiang, Bin Liu 0001 |
ICC | 3 |
| 2010 | NetShield: massive semantics-based vulnerability signature matching for high-speed networksabstractAccuracy and speed are the two most important metrics for Network Intrusion Detection/Prevention Systems (NIDS/NIPSes). Due to emerging polymorphic attacks and the fact that in many cases regular expressions (regexes) cannot capture the vulnerability conditions accurately, the accuracy of existing regex-based NIDS/NIPS systems has become a serious problem. In contrast, the recently-proposed vulnerability signatures (a.k.a data patches) can exactly describe the vulnerability conditions and achieve better accuracy. However, how to efficiently apply vulnerability signatures to high speed NIDS/NIPS with a large ruleset remains an untouched but challenging issue. Zhichun Li, Gao Xia, Yi Tang 0002, Yan Chen 0004, Bin Liu 0001, Junchen Jiang, Yuezhou Lv |
SIGCOMM | 7 |
| 2009 | SPC-FA: synergic parallel compact finite automaton to accelerate multi-string matching with low memoryabstractDeterministic Finite Automaton (DFA) is well-known for its constant matching speed in worst case, and widely used in multi-string matching, which is a critical technique in high performance Network Intrusion Detection System (NIDS) design. Existing DFA-based researches achieve high throughput at the expense of extremely high memory cost, so they fail to be used in situations like embedded systems where very tight memory resource is available. In this paper, we propose a memory-efficient multi-string matching acceleration scheme named Synergic Parallel Compact (SPC) Match Engine, which can provide a high matching speedup with no extra memory cost than the traditional DFA. Our scheme can be understood as consisting of k SPC-FAs, each of which can process one character from the input stream, causing achieving a constant speedup factor k with reduced memory occupation. Experimental evaluations with Snort and ClamAV rulesets show that a speedup of 9X can be practically achieved by a single SPC Match Engine instance with a reduced memory size than the up-to-date DFA-based compression approaches. Junchen Jiang, Yi Tang 0002, Bin Liu 0001, Xiaofei Wang 0006, Yang Xu 0010 |
ANCS | 1 |