F. Donelson Smith

dblp:s/FDonelsonSmith · also Frank Donelson Smith · DBLP profile ↗
← Back
44ranked-venue papers
0as first author
7since 2021 · last 2024
0000-0002-8080-7589ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 18Systems, architecture and hardware · 15 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Statistical verification of autonomous system controllers under timing uncertainties
Bineet Ghosh, Clara Hobbs, Shengjie Xu 0005, F. Donelson Smith, James H. Anderson, P. S. Thiagarajan, Benjamin Berg, Parasara Sridhar Duggirala, Samarjit Chakraborty
Real Time Syst.4
2022 Making Powerful Enemies on NVIDIA GPUs
abstract
Graphics Processing Units (GPUs) are widely used in safety-critical real-time systems such as autonomous vehicles due to their high performance on artificial intelligence (AI) work-loads. As the computing power of recent GPUs keeps growing, it becomes increasingly possible to allow multiple independent programs to access the GPU concurrently. This complicates timing analysis, as contention for shared GPU resources renders execution times less predictable and worst-case execution times (WCETs) difficult to estimate. This paper provides a method for producing enemy programs that intentionally contend for GPU resources in order to enable more confident measurement-based WCET estimations. This paper provides an experiment-driven method to design effective enemy programs for several different interference channels—specific shared resources within the GPU through which concurrent computations may impact others' execution times. The method is flexible and can be applied to different GPU sharing mechanisms. The enemies are evaluated against a large number of real GPU applications, and the results indicate that these enemies cause higher slowdowns for GPU tasks than other baseline resource-stressing methods.
Tyler Yandrofski, Nathan Otterness, James H. Anderson, F. Donelson Smith
RTSS5
2021 Timing-Predictable Vision Processing for Autonomous Systems
abstract
Vision processing for autonomous systems today involves implementing machine learning algorithms and vision processing libraries on embedded platforms consisting of CPUs, GPUs and FPGAs. Because many of these use closed-source proprietary components, it is very difficult to perform any timing analysis on them. Even measuring or tracing their timing behavior is challenging, although it is the first step towards reasoning about the impact of different algorithmic and implementation choices on the end-to-end timing of the vision processing pipeline. In this paper we discuss some recent progress in developing tracing, measurement and analysis infrastructure for determining the timing behavior of vision processing pipelines implemented on state-of-the-art FPGA and GPU platforms.
Tanya Amert, Michael Balszun, Martin Geier 0001, F. Donelson Smith, James H. Anderson, Samarjit Chakraborty
DATE4
2021 Perception Computing-Aware Controller Synthesis for Autonomous Systems
abstract
Feedback control loops are ubiquitous in any autonomous system. The design flow for any controller starts by determining a control strategy, while abstracting away all implementation details. However, when designing controllers for autonomous systems, there is significant computation associated with the perception modules. For example, this involves vision processing using deep neural networks on multicore CPU+accelerator platforms. Such computation can be organized in many different ways, with each choice resulting in very different sensor-to-actuator delays and tradeoffs between cost, delay, and accuracy. Further, each of these choices requires the control strategy to be designed accordingly. It is not possible for a control designer to enumerate and account for all of these choices manually, or abstract them away as “implementation details” as done in traditional controller design. In this paper we outline this problem and discuss how automated controller-synthesis techniques could help in addressing it.
Clara Hobbs, Debayan Roy, Parasara Sridhar Duggirala, F. Donelson Smith, Soheil Samii, James H. Anderson, Samarjit Chakraborty
DATE4
2021 Simultaneous Multithreading in Mixed-Criticality Real-Time Systems
abstract
Simultaneous multithreading (SMT) enables enhanced computing capacity by allowing multiple tasks to execute concurrently on the same computing core. Despite its benefits, its use has been largely eschewed in work on real-time systems due to concerns that tasks running on the same core may adversely interfere with each other. In this paper, the safety of using SMT in a mixed-criticality multicore context is considered in detail. To this end, a prior open-source framework called MC2(mixedcriticality on multicore), which provides features for mitigating cache and memory interference, was re-implemented to support SMT on an SMT-capable multicore platform. The creation of this new, configurable MC2variant entailed producing the first operating-system implementations of several recently proposed real-time SMT schedulers and tying them together within a mixed-criticality context. These schedulers introduce new spatialisolation challenges, which required introducing isolation at both the L2 and L3 cache levels. The efficacy of the resulting MC2variant is demonstrated via three experimental efforts. The first involved obtaining execution data using a wide range of benchmark suites, including TACLeBench, DIS, SD-VBS, and synthetic microbenchmarks. The second involved conducting a large-scale overhead-aware schedulability study, parameterized by the collected benchmark data, to elucidate schedulability tradeoffs. The third involved experiments involving case-study task systems. In the schedulability study, the use of SMT proved capable of increasing platform capacity by an average factor of 1.32. In the case-study experiments, deadline misses of highly critical tasks were never observed.
Joshua Bakita, Shareef Ahmed, Sims Osborne, F. Donelson Smith, James H. Anderson
RTAS6
2021 TimeWall: Enabling Time Partitioning for Real-Time Multicore+Accelerator Platforms
abstract
Across a range of safety-critical domains, an evolution is underway to endow embedded systems with "thinking" capabilities by using artificial-intelligence (AI) techniques. This evolution is being fueled by the availability of high-performance embedded hardware, typically multicore machines augmented with accelerators. Unfortunately, existing software certification processes rely on time partitioning to isolate system components, and this sense of isolation can be broken by accelerator usage. To address this issue, this paper presents TimeWall, a time-partitioning framework for multicore+accelerator platforms. When applied alongside existing methods for alleviating spatial interference, TimeWall can help enable component-wise certification on multicore+accelerator platforms. The challenges in realizing a TimeWall implementation are discussed in detail in this paper. Additionally, the temporal isolation TimeWall affords is examined experimentally, including via a case study of a computer-vision perception application, on a real platform.
Tanya Amert, Zelin Tong, Sergey Voronov, Joshua Bakita, F. Donelson Smith, James H. Anderson
RTSS5
2021 The price of schedulability in cyclic workloads: The history-vs.-response-time-vs.-accuracy trade-off
Tanya Amert, Ming Yang 0036, Sergey Voronov, Saujas Nandi, Thanh Vu 0001, James H. Anderson, F. Donelson Smith
J. Syst. Archit.7
2020 The Price of Schedulability in Multi-Object Tracking: The History-vs.-Accuracy Trade-Off
abstract
Autonomous vehicles often employ computer-vision (CV) algorithms that track the movements of pedestrians and other vehicles to maintain safe distances from them. These algorithms are usually expressed as real-time processing graphs that have cycles due to back edges that provide history information. If immediate back history is required, then such a cycle must execute sequentially. Due to this requirement, any graph that contains a cycle with utilization exceeding 1.0 is categorically unschedulable, i.e., bounded graph response times cannot be guaranteed. Unfortunately, such cycles can occur in practice, particularly if conservative execution-time assumptions are made, as befits a safety-critical system. This dilemma can be obviated by allowing older back history, which enables parallelism in cycle execution at the expense of possibly affecting the accuracy of tracking. However, the efficacy of this solution hinges on the resulting history-vs.-accuracy trade-off that it exposes. In this paper, this trade-off is explored in depth through an experimental study conducted using the open-source CARLA autonomous-driving simulator. Somewhat surprisingly, easing away from always requiring immediate back history proved to have only a marginal impact on accuracy in this study.
Tanya Amert, Ming Yang 0036, Saujas Nandi, Thanh Vu 0001, James H. Anderson, F. Donelson Smith
ISORC6
2020 Supporting I/O and IPC via fine-grained OS isolation for mixed-criticality real-time tasks
Namhoon Kim, Nathan Otterness, James H. Anderson, F. Donelson Smith, Donald E. Porter
Real Time Syst.5
2019 Re-Thinking CNN Frameworks for Time-Sensitive Autonomous-Driving Applications: Addressing an Industrial Challenge
abstract
Vision-based perception systems are crucial for profitable autonomous-driving vehicle products. High accuracy in such perception systems is being enabled by rapidly evolving convolution neural networks (CNNs). To achieve a better understanding of its surrounding environment, a vehicle must be provided with full coverage via multiple cameras. However, when processing multiple video streams, existing CNN frameworks often fail to provide enough inference performance, particularly on embedded hardware constrained by size, weight, and power limits. This paper presents the results of an industrial case study that was conducted to re-think the design of CNN software to better utilize available hardware resources. In this study, techniques such as parallelism, pipelining, and the merging of per-camera images into a single composite image were considered in the context of a Drive PX2 embedded hardware platform. The study identifies a combination of techniques that can be applied to increase throughput (number of simultaneous camera streams) without significantly increasing per-frame latency (camera to CNN output) or reducing per-stream accuracy.
Ming Yang 0036, Shige Wang, Joshua Bakita, Thanh Vu 0001, F. Donelson Smith, James H. Anderson, Jan-Michael Frahm
RTAS5
2018 Avoiding Pitfalls when Using NVIDIA GPUs for Real-Time Tasks in Autonomous Systems
abstract
NVIDIA's CUDA API has enabled GPUs to be used as computing accelerators across a wide range of applications. This has resulted in performance gains in many application domains, but the underlying GPU hardware and software are subject to many non-obvious pitfalls. This is particularly problematic for safety-critical systems, where worst-case behaviors must be taken into account. While such behaviors were not a key concern for earlier CUDA users, the usage of GPUs in autonomous vehicles has taken CUDA programs out of the sole domain of computer-vision and machine-learning experts and into safety-critical processing pipelines. Certification is necessary in this new domain, which is problematic because GPU software may have been developed without any regard for worst-case behaviors. Pitfalls when using CUDA in real-time autonomous systems can result from the lack of specifics in official documentation, and developers of GPU software not being aware of the implications of their design choices with regards to real-time requirements. This paper focuses on the particular challenges facing the real-time community when utilizing CUDA-enabled GPUs for autonomous applications, and best practices for applying real-time safety-critical principles.
Ming Yang 0036, Nathan Otterness, Tanya Amert, Joshua Bakita, James H. Anderson, F. Donelson Smith
ECRTS6
2018 Making OpenVX Really "Real Time"
abstract
OpenVX is a recently ratified standard that was expressly proposed to facilitate the design of computer-vision (CV) applications used in real-time embedded systems. Despite its real-time focus, OpenVX presents several challenges when validating real-time constraints. Many of these challenges are rooted in the fact that OpenVX only implicitly defines any notion of a schedulable entity. Under OpenVX, CV applications are specified in the form of processing graphs that are inherently considered to execute monolithically end-to-end. This monolithic execution hinders parallelism and can lead to significant processing-capacity loss. Prior work partially addressed this problem by treating graph nodes as schedulable entities, but under OpenVX, these nodes represent rather coarse-grained CV functions, so the available parallelism that can be obtained in this way is quite limited. In this paper, a much more fine-grained approach for scheduling OpenVX graphs is proposed. This approach was designed to enable additional parallelism and to eliminate schedulability-related processing-capacity loss that arises when programs execute on both CPUs and graphics processing units (GPUs). Response-time analysis for this new approach is presented and its efficacy is evaluated via a case study involving an actual CV application.
Ming Yang 0036, Tanya Amert, Kecheng Yang 0001, Nathan Otterness, James H. Anderson, F. Donelson Smith, Shige Wang
RTSS6
2017 TCP Rapid: From theory to practice
abstract
Delay and rate-based alternatives to TCP congestion-control have been around for nearly three decades and have seen a recent surge in interest. However, such designs have faced significant resistance in being deployed on a wide-scale across the Internet - this has been mostly due to serious concerns about noise in delay measurements, pacing inter-packet gaps, and/or required changes to the standard TCP stack/headers. With the advent of high-speed networking, some of these concerns become even more significant. In this paper, we consider Rapid, a recent proposal for ultra-high speed congestion control, which perhaps stretches each of these challenges to the greatest extent. Rapid adopts a framework of continuous fine-scale bandwidth probing, which requires a potentially different and finely-controlled gap for every packet, high-precision timestamping of received packets, and reliance on fine-scale changes in inter-packet gaps. While simulation-based evaluations of Rapid show that it has outstanding performance gains along several important dimensions, these will not translate to the real-world unless the above challenges are addressed. We design a Linux implementation of Rapid after carefully considering each of these challenges. Our evaluations on a 10Gbps testbed confirm that the implementation can indeed achieve the claimed performance gains, and that it would not have been possible unless each of the above challenges was addressed.
Qianwen Yin, Jasleen Kaur 0001, F. Donelson Smith
INFOCOM3
2017 Allowing Shared Libraries While Supporting Hardware Isolation in Multicore Real-Time Systems
abstract
The desire to support real-time applications on multicore platforms has led to intense recent interest in techniques for reducing memory-related hardware interference. These techniques typically rely on mechanisms that ensure per-task isolation properties with respect to cache and memory accesses. In most prior work on such techniques, any sharing of memory pages by different tasks is defined away, as sharing breaks isolation. In reality, however, sharing is common. In this paper, one source of sharing is considered, namely, the usage of shared libraries. Such sharing can be obviated by statically linking libraries, but this solution can degrade schedulability by exhausting memory capacity. An alternative approach is proposed herein that allows library pages to be shared while preserving isolation properties. This approach is presented in the context of the MC2 framework and a schedulability-based evaluation of it is presented. Such an evaluation must necessarily consider memory-capacity limits. As a secondary contribution, this paper considers such limits for the first time in the context of MC2.
Namhoon Kim, Micaiah Chisholm, Nathan Otterness, James H. Anderson, F. Donelson Smith
RTAS5
2017 An Evaluation of the NVIDIA TX1 for Supporting Real-Time Computer-Vision Workloads
abstract
Autonomous vehicles are an exemplar for forward-looking safety-critical real-time systems where significant computing capacity must be provided within strict size, weight, and power (SWaP) limits. A promising way forward in meeting these needs is to leverage multicore platforms augmented with graphics processing units (GPUs) as accelerators. Such an approach is being strongly advocated by NVIDIA, whose Jetson TX1 board is currently a leading multicore+GPU solution marketed for autonomous systems. Unfortunately, no study has ever been published that expressly evaluates the effectiveness of the TX1, or any other comparable platform, in hosting safety-critical real-time workloads. In this paper, such a study is presented. Specifically, the TX1 is evaluated via benchmarking efforts, blackbox evaluations of GPU behavior, and case-study evaluations involving computer-vision workloads inspired by autonomousdriving use cases. Autonomous vehicles are an exemplar for forward-looking safety-critical real-time systems where significant computing capacity must be provided within strict size, weight, and power (SWaP) limits. A promising way forward in meeting these needs is to leverage multicore platforms augmented with graphics processing units (GPUs) as accelerators. Such an approach is being strongly advocated by NVIDIA, whose Jetson TX1 board is currently a leading multicore+GPU solution marketed for autonomous systems. Unfortunately, no study has ever been published that expressly evaluates the effectiveness of the TX1, or any other comparable platform, in hosting safety-critical real-time workloads. In this paper, such a study is presented. Specifically, the TX1 is evaluated via benchmarking efforts, blackbox evaluations of GPU behavior, and case-study evaluations involving computer-vision workloads inspired by autonomous-driving use cases.
Nathan Otterness, Ming Yang 0036, Sarah Rust, Eunbyung Park, James H. Anderson, F. Donelson Smith, Alexander C. Berg, Shige Wang
RTAS6
2017 GPU Scheduling on the NVIDIA TX2: Hidden Details Revealed
abstract
The push towards fielding autonomous-driving capabilities in vehicles is happening at breakneck speed. Semi-autonomous features are becoming increasingly common, and fully autonomous vehicles are optimistically forecast to be widely available in just a few years. Today, graphics processing units (GPUs) are seen as a key technology in this push towards greater autonomy. However, realizing full autonomy in mass-production vehicles will necessitate the use of stringent certification processes. Currently available GPUs pose challenges in this regard, as they tend to be closed-source “black boxes” that have features that are not publicly disclosed. For certification to be tenable, such features must be documented. This paper reports on such a documentation effort. This effort was directed at the NVIDIA TX2, which is one of the most prominent GPU-enabled platforms marketed today for autonomous systems. In this paper, important aspects of the TX2's GPU scheduler are revealed as discerned through experimental testing and validation.
Tanya Amert, Nathan Otterness, Ming Yang 0036, James H. Anderson, F. Donelson Smith
RTSS5
2017 Attacking the one-out-of-m multicore problem by combining hardware management with mixed-criticality provisioning
Namhoon Kim, Bryan C. Ward, Micaiah Chisholm, James H. Anderson, F. Donelson Smith
Real Time Syst.5
2017 Packet-Scale Congestion Control Paradigm
abstract
This paper presents the packet-scale paradigm for designing end-to-end congestion control protocols for ultra-high speed networks. The paradigm discards the legacy framework of RTT-scale protocols, and instead builds upon two revolutionary foundations-that of continually probing for available bandwidth at short timescales, and that of adapting the data sending rate so as to avoid overloading the network. Through experimental evaluations with a prototype, we report high performance gains along several dimensions in high-speed networks-the steady-state throughput, adaptability to dynamic cross-traffic, RTT-fairness, and co-existence with the conventional TCP traffic mixes. The paradigm also opens up several issues that are less of a concern for traditional protocols-we summarize our approaches for addressing these.
Rebecca Lovewell, Qianwen Yin, Tianrong Zhang, Jasleen Kaur 0001, F. Donelson Smith
IEEE/ACM Trans. Netw.5
2016 Attacking the One-Out-Of-m Multicore Problem by Combining Hardware Management with Mixed-Criticality Provisioning
abstract
The multicore revolution is having limited impact in safety-critical application domains. A key reason is the "one-out-of-m" problem: when validating real-time constraints on an m-core platform, excessive analysis pessimism can effectively negate the processing capacity of the additional m-1 cores so that only "one core's worth" of capacity is available. Two approaches have been investigated previously to address this problem: mixed-criticality allocation techniques, which provision less-critical software components less pessimistically, and hardware-management techniques, which make the underlying platform itself more predictable. A better way forward may be to combine both approaches, but to show this, fundamentally new criticality-cognizant hardware-management tradeoffs must be explored. Such tradeoffs are investigated herein in the context of a large-scale, overhead-aware schedulability study. This study was guided by extensive trace data obtained by executing benchmark tasks on a new variant of the MC^2 framework that supports configurable criticality-based hardware management. This study shows that the two approaches mentioned above can be much more effective when applied together instead of alone.
Namhoon Kim, Bryan C. Ward, Micaiah Chisholm, Cheng-Yang Fu, James H. Anderson, F. Donelson Smith
RTAS6
2016 Reconciling the Tension Between Hardware Isolation and Data Sharing in Mixed-Criticality, Multicore Systems
abstract
Recent work involving a mixed-criticality framework called MC2 has shown that, by combining hardware-management techniques and criticality-aware task provisioning, capacity loss can be significantly reduced when supporting real-time workloads on multicore platforms. However, as in most other prior research on multicore hardware management, tasks were assumed in that work to not share data. Data sharing is problematic in the context of hardware management because it can violate the isolation properties hardware-management techniques seek to ensure. Clearly, for research on such techniques to have any practical impact, data sharing must be permitted. Towards this goal, this paper presents a new version of MC2 that permits tasks to share data within and across criticality levels through shared memory. Several techniques are presented for mitigating capacity loss due to data sharing. The effectiveness of these techniques is demonstrated by means of a large-scale, overhead-aware schedulability study driven by micro-benchmark data.
Micaiah Chisholm, Namhoon Kim, Bryan C. Ward, Nathan Otterness, James H. Anderson, F. Donelson Smith
RTSS6
2014 Can Bandwidth Estimation Tackle Noise at Ultra-high Speeds?
abstract
While existing bandwidth estimation tools have been shown to perform well on 100Mbps networks, they fail to do so at gigabit and higher network speeds. This is because finer inter-packet gaps are needed to probe for higher rates -- fine gaps are more susceptible to be disturbed by small-scale buffering-related noise. In this paper, we evaluate existing noise reduction techniques for tackling the issue, and show that they are ineffective on 10Gbps links. We propose a novel smoothing strategy, Buffering-aware Spike Smoothing (BASS), which can be applied effectively to both single-rate and multi-rate probing frameworks and help significantly in scaling bandwidth estimation to ultra-high speed networks. Besides, we provide first evidence that accurate bandwidth estimation using our strategy can help improve the performance of congestion-control protocols on real 10Gbps networks.
Qianwen Yin, Jasleen Kaur 0001, F. Donelson Smith
ICNP3
2014 Scaling Bandwidth Estimation to High Speed Networks
Qianwen Yin, Jasleen Kaur 0001, F. Donelson Smith
PAM3
2012 Towards Traffic Benchmarks for Empirical Networking Research: The Role of Connection Structure in Traffic Workload Modeling
abstract
Networking research would be well served by the adoption of a set of traffic benchmarks to model network applications for empirical evaluations; such benchmarks are common in many other areas of computing. While it has long been known that certain aspects of modeling traffic, such as round trip time, can dramatically affect application and network performance, there is still no agreement as to how such components should be controlled within an experiment. In this paper we advance the discussion of standards for empirical networking research by demonstrating how certain components of network traffic, such as the structure of application data exchanges within a TCP connection, can have a larger impact on the results obtained through experimentation than other dimensions of traffic such as round-trip time. Such findings point to the pressing need for traffic benchmarks in networking research. Through testbed experiments performed with synthetically generated network traffic from two very different traffic sources, and using several models of TCP connection structure, we demonstrate the strong effects of connection structure in traffic workload modeling on performance measures such as queue length at routers, number of active connections in the network, user response times, and connection durations.
Jay Aikat, Shaddi Hasan, Kevin Jeffay, F. Donelson Smith
MASCOTS4
2007 Modeling and generating TCP application workloads
abstract
In order to perform valid experiments, traffic generators used in network simulators and testbeds require contemporary models of traffic as it exists on real network links. Ideally one would like a model of the workload created by the full range of applications running on the Internet today. Unfortunately, at best, all that is available to the research community are a small number of models for single applications or application classes such as the web or peer-to-peer. We present a method for creating a model of the full TCP application workload that generates the traffic flowing on a network link. From this model, synthetic workload traffic can be generated in a simulation that is statistically similar to the traffic observed on the real link. The model is generated automatically using only a simple packet-header trace and requires no knowledge of the actual identity or mix of TCP applications on the network. We present the modeling method and a traffic generator that will enable researchers to conduct network experiments with realistic, easy-to-update TCP application workloads. An extensive validation study is performed using Abilene and university traces. The method is validated by comparing traces of synthetically generated traffic to the original traces for a set of important measures of realism. We also show how workload models can be re-sampled to generate statistically valid randomized and rescaled variations.
Félix Hernández-Campos, Kevin Jeffay, F. Donelson Smith
BROADNETS3
2007 A Performance Study of Loss Detection/Recovery in Real-world TCP Implementations
abstract
TCP is the dominant transport protocol used in the Internet and its performance fundamentally governs the performance of Internet applications. It is well-known that packet losses can adversely affect the connection duration of TCP connections - however, what is not fully understood is how well does the TCP design deal with losses. In this paper, we systematically evaluate the impact of design parameters associated with TCP's loss detection/recovery mechanisms on the performance of real-world TCP connections. For this, we rely on an analysis tool that partially emulates the sender-side TCP implementations of 5 prominent OSes for passively analyzing the traces of TCP connections. Our study conducts passive analysis of more than 2.8 million real Internet TCP connections. We find that the recommended as well as widely-implemented settings of TCP parameters are not optimal for a significant fraction of Internet connections.
Sushant Rewaskar, Jasleen Kaur 0001, F. Donelson Smith
ICNP3
2007 The effects of active queue management and explicit congestion notification on web performance
Long Le, Jay Aikat, Kevin Jeffay, F. Donelson Smith
IEEE/ACM Trans. Netw.4
2006 A Loss and Queuing-Delay Controller for Router Buffer Management
abstract
Active queue management (AQM) in routers has been proposed as a solution to some of the scalability issues associated with TCP’s pure end-to-end approach to congestion control. However, beyond congestion control, controlling queues in routers is important because unstable router queues can cause poor application performance. Existing AQM schemes explicitly try to control router queues by probabilistically dropping (or marking) packets. We argue that while controlling router queues is important, this control needs to be tempered by a consideration of the overall lossrate at the router. Solely attempting to control queue length can induce loss-rates that have as negative an effect on application and network performance as the large queues that existing AQM schemes were trying to avoid. Thus controlling queue length without regard to loss-rate can be counterproductive. In this work we demonstrate that by jointly controlling queue length and loss-rate, both network and application performance are improved. We present a novel AQM design that attempts to simultaneously optimize queue length and loss-rate. Our algorithm, called loss and queuing delay control (LQD), is a control theoretic scheme that explicitly treats loss-rate as a control parameter. LQD is shown to provide stable control analytically and is evaluated empirically by comparing its performance against other control theoretic AQM designs (PI and REM). The results of evaluation in a laboratory testbed under realistic traffic mixes and loads show that LQD results in lower overall loss rates and that applications see lower average queue lengths than with PI or REM.
Long Le, Kevin Jeffay, F. Donelson Smith
ICDCS3
2006 Quantifying the effects of recent protocol improvements to TCP: Impact on Web performance
Michele C. Weigle, Kevin Jeffay, F. Donelson Smith
Comput. Commun.3
2005 Understanding Patterns of TCP Connection Usage with Statistical Clustering
abstract
We describe a new methodology for understanding how applications use TCP to exchange data. The method is useful for characterizing TCP workloads and synthetic traffic generation. Given a packet header trace, the method automatically constructs a source-level model of the applications using TCP in a network without any a priori knowledge of which applications are actually present in a network. From this source-level model, statistical feature vectors can be defined for each TCP connection in the trace. Hierarchical cluster analysis can then be performed to identify connections that are statistically homogeneous and that are likely exerting similar demands on a network. We apply the methods to packet header traces taken from the UNC and Abilene networks and show how classes of similar connections can be automatically detected and modeled.
Félix Hernández-Campos, Andrew B. Nobel, F. Donelson Smith, Kevin Jeffay
MASCOTS3
2005 Long-range dependence in a changing Internet traffic mix
Cheolwoo Park, Félix Hernández-Campos, J. S. Marron, F. Donelson Smith
Comput. Networks4
2005 Delay-based early congestion detection and adaptation in TCP: impact on web performance
Michele C. Weigle, Kevin Jeffay, F. Donelson Smith
Comput. Commun.3
2004 Differential Congestion Notification: Taming the Elephants
abstract
Active queue management (AQM) in routers has been proposed as a solution to some of the scalability issues associated with TCP's pure end-to-end approach to congestion control. A recent study of AQM demonstrated its effectiveness in reducing the response times of Web request/response exchanges as well as increasing link throughput and reducing loss rates [L. Le et al., 2003]. However, use of the ECN (explicit congestion notification) signaling protocol was required to outperform drop-tail queuing. Since ECN is not currently widely deployed on end-systems, we investigate an alternative to ECN, namely applying AQM differentially to flows based on a heuristic classification of the flow's transmission rate. Our approach, called differential congestion notification (DCN), distinguishes between "small" flows and "large" high-bandwidth flows and only provides congestion notification to large high-bandwidth flows. We compare DCN to other prominent AQM schemes and demonstrate that for Web and general TCP traffic, DCN outperforms all the other AQM designs, including those previously designed to differentiate between flows based on their size and rate.
Long Le, Jay Aikat, Kevin Jeffay, F. Donelson Smith
ICNP4
2004 Stochastic Models for Generating Synthetic HTTP Source Traffic
abstract
New source-level models for aggregated HTTP traffic and a design for their integration with the TCP transport layer are built and validated using two large-scale collections of TCP/IP packet header traces. An implementation of the models and the design in the ns network simulator can be used to generate web traffic in network simulations
William S. Cleveland, Kevin Jeffay, F. Donelson Smith, Michele C. Weigle
INFOCOM5
2004 Variable heavy tails in Internet traffic
Félix Hernández-Campos, J. S. Marron, Gennady Samorodnitsky, F. Donelson Smith
Perform. Evaluation4
2003 Variability in TCP round-trip times
abstract
We measured and analyzed the variability in round trip times (RTTs) within TCP connections using passive measurement techniques. We collected eight hours of bidirectional traces containing over 22 million TCP connections between end-points at a large university campus and almost $1$ million remote locations. Of these, we used over 1 million TCP connections that yield 10 or more valid RTT samples, to examine RTT variability within a TCP connection. Our results indicate that contrary to observations in several previous studies, RTT values within a connection vary widely. Our results have implications for designing better simulation models, and understanding how round trip times affect the dynamic behavior and throughput of TCP connections.
Jay Aikat, Jasleen Kaur 0001, F. Donelson Smith, Kevin Jeffay
Internet Measurement Conference3
2003 The effects of active queue management on web performance
abstract
We present an empirical study of the effects of active queue management (AQM) on the distribution of response times experienced by a population of web users. Three prominent AQM schemes are considered: the Proportional Integrator (PI) controller, the Random Exponential Marking (REM) controller, and Adaptive Random Early Detection (ARED). The effects of these AQM schemes were studied alone and in combination with Explicit Congestion Notification (ECN). Our major results are:
Long Le, Jay Aikat, Kevin Jeffay, F. Donelson Smith
SIGCOMM4
2001 Tuning RED for Web traffic
abstract
We study the effects of RED on the performance of Web browsing with a novel aspect of our work being the use of a user-centric measure of performance: response time for HTTP request-response pairs. We empirically evaluate RED across a range of parameter settings and offered loads. Our results show that: (1) contrary to expectations, compared to an FIFO queue, RED has a minimal effect on HTTP response times for offered loads up to 90% of link capacity; (2) response times at loads in this range are not substantially affected by RED parameters; (3) between 90% and 100% load, RED can be carefully tuned to yield performance somewhat superior to FIFO, however, response times are quite sensitive to the actual RED parameter values selected; and (4) in such heavily congested networks, RED parameters that provide the best link utilization produce poorer response times. We conclude that for links carrying only Web traffic, RED queue management appears to provide no clear advantage over tail-drop FIFO for end-user response times.
Mikkel Christiansen, Kevin Jeffay, David E. Ott, F. Donelson Smith
IEEE/ACM Trans. Netw.4
2000 Tuning RED for web traffic
abstract
We study the effects of RED on the performance of Web browsing with a novel aspect of our work being the use of a user-centric measure of performance - response time for HTTP request-response pairs. We empirically evaluate RED across a range of parameter settings and offered loads. Our results show that: (1) contrary to expectations, compared to a FIFO queue, RED has a minimal effect on HTTP response times for offered loads up to 90% of link ca?pacity, (2) response times at loads in this range are not substantially effected by RED pa?rameters, (3) between 90% and 100% load, RED can be carefully tuned to yield performance somewhat superior to FIFO, however, response times are quite sensitive to the actual RED pa?rameter values selected, and (4) in such heavily congested networks, RED parameters that provide the best link utilization produce poorer response times. We conclude that for links carrying only web traf?fic, RED queue management appears to provide no clear advantage over tail-drop FIFO for end-user response times.
Mikkel Christiansen, Kevin Jeffay, David E. Ott, F. Donelson Smith
SIGCOMM4
1998 Proportional Share Scheduling of Operating System Services for Real-Time Applications
abstract
While there is currently great interest in the problem of providing real time services in general purpose operating systems, the issue of real time scheduling of internal operating system activities has received relatively little attention. Without such real time scheduling, the system is susceptible to conditions such as receive livelock-a situation in which an operating system spends all its time processing arriving network packets, and application processes, even if scheduled with a real time scheduler, are starved. We investigate the problem of scheduling operating system activities such as network protocol processing in a proportional share manner. We describe a proportional share implementation of the FreeBSD operating system and demonstrate that it solves the receive livelock problem. Packets are processed within the operating system only at the cumulative rate at which the destination applications are prepared to receive them. If packets arrive at a faster rate then they are discarded after consuming minimal system resources. In this manner the performance of "well behaved" applications is unaffected by "misbehaving" applications. We demonstrate this effect by running a set of multimedia applications under a variety of network conditions on a set of increasingly sophisticated proportional share implementations of FreeBSD and comparing their performance. This work contributes to our knowledge of the engineering of proportional share real time systems.
Kevin Jeffay, F. Donelson Smith, A. Moorthy, James H. Anderson
RTSS2
1994 Transport and Display Mechanisms for Multimedia Conferencing Across Packet-Switched Networks
Kevin Jeffay, Donald L. Stone, F. Donelson Smith
Comput. Networks ISDN Syst.3
1992 Architecture of the Artifact-Based Collaboration System Matrix
abstract
ABSTRACT The UNC Collaboratory project is concerned with both the process of collaboration and with computer systems to support that process. Here, we describe a component of the Artifact-Based Collaboration (ABC) system, called the Matrix, that provides an infrastructure in which existing single-user applications can be incorporated with few, if any, changes and used collaboratively. We take the position that what is needed is not new tools but better infrastructure for using familiar single-user tools collectively. The paper discusses the Matrix architecture, a Virtual Screen component, and generic functions that provide conferencing, hyperlinking, and recording of users' actions for all applications.
Kevin Jeffay, Jin-Kun Lin, John Menges, F. Donelson Smith, John B. Smith
CSCW4
1992 Adaptive, Best-Effort Delivery of Digital Audio and Video Across Packet-Switched Networks
Kevin Jeffay, Donald L. Stone, Terry Talley, F. Donelson Smith
NOSSDAV4
1992 Kernel support for live digital audio and video
Kevin Jeffay, Donald L. Stone, F. Donelson Smith
Comput. Commun.3
1991 Kernel Support for Live Digital Audio and Video
Kevin Jeffay, Donald L. Stone, F. Donelson Smith
NOSSDAV3