Martin F. Arlitt

dblp:76/6339 · DBLP profile ↗
← Back
66ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0001-6167-2255ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 19 · 5 first-author · 1 since 2021Systems, architecture and hardware · 14 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 14 · 3 first-author · 1 since 2021Security and privacy · 8 · 5 since 2021Databases, data management, data science and information retrieval · 8 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Bitcoin Flows from Sanctioned Sources
abstract
Cryptocurrencies’ pseudonymity property poses regulatory challenges and has attracted illicit actors that try to avoid oversight. To counter this, the U.S. Treasury’s Office of Foreign Assets Control (OFAC) sanctions individuals and entities using Bitcoin for cybercrime, terrorism financing, and other illicit activities. However, the effectiveness of these measures remains uncertain. This study analyzes over 13 million Bitcoin transactions linked to sanctioned entities, tracing fund flows and exchange interactions. We find that ~175,000 BTC was moved before sanctions took effect, with only 50 BTC remaining post-sanction, indicating preemptive fund displacement. Cybercrime-linked addresses accounted for the largest transfers—sometimes exceeding $1 billion—while sanctioned entities favored large, direct transactions to exchanges. Despite activity dropping immediately after sanctions took effect, some entities continued transacting for up to 1,500 days, exposing enforcement gaps. Our findings highlight key challenges in sanction enforcement, including delayed restrictions, exchange compliance gaps, and strategic fund movements. These insights inform policymakers and regulators seeking to strengthen cryptocurrency financial controls.
Axel Flodmark, Rasmus Samuelson, David Hasselquist, Martin F. Arlitt, Niklas Carlsson
LCN4
2024 Trust Issue(r)s: Certificate Revocation and Replacement Practices in the Wild
David Cerenius, Martin Kaller, Carl Magnus Bruhner, Martin F. Arlitt, Niklas Carlsson
PAM (2)4
2024 On the Dark Side of the Coin: Characterizing Bitcoin Use for Illicit Activities
Hampus Rosenquist, David Hasselquist, Martin F. Arlitt, Niklas Carlsson
PAM (2)3
2023 Packet-Level Analysis of Zoom Performance Anomalies
abstract
In this paper, we use Wireshark packet-level traces to study the performance of the Zoom network application. Our work is motivated by several anecdotal reports of Zoom performance problems on our campus network during the Fall 2021 semester. Through the collection and analysis of Wireshark traces from different vantage points, we are able to pinpoint the root cause of the Zoom performance problems, which is a congested external Internet link for our campus network. We also identify several characteristics of the Zoom application that exacerbate its performance issues on congested and lossy networks, due to multi-layer protocol interactions.
Mehdi Karamollahi, Carey L. Williamson, Martin F. Arlitt
ICPE3
2023 Simulation modeling of Zoom traffic on a campus network: A case study
Mehdi Karamollahi, Carey L. Williamson, Martin F. Arlitt
Perform. Evaluation3
2022 Deep Learning for Network Traffic Data
abstract
Network traffic data is key in addressing several important cybersecurity problems, such as intrusion and malware detection, and network management problems, such as application and device identification. However, it poses several challenges to building machine learning models. Two main challenges are manual feature engineering and scarcity of training data due to privacy and security concerns. In this tutorial we provide a comprehensive review of recent advances to address these challenges through use of deep learning. Network traffic data can be cast as a multivariate time-series (sequential) data, attributed graph data, or image data to leverage representation learning architectures available in deep learning. To preserve data privacy, generative methods, such as GANs and autoregressive neural architectures can be used to synthesize realistic network traffic data. In particular, our tutorial is organized into three parts: 1) we describe network traffic data, applications to security and network management, and challenges; 2) we present different deep learning architectures used for representation learning instead of feature engineering of network traffic data; and, 3) we describe use of generative neural models for synthetic generation of network traffic data.
Manish Marwah, Martin F. Arlitt
KDD2
2022 Changing of the Guards: Certificate and Public Key Management on the Internet
Carl Magnus Bruhner, Oscar Linnarsson, Matús Nemec, Martin F. Arlitt, Niklas Carlsson
PAM4
2022 Zoom Session Quality: A Network-Level View
Albert Choi, Mehdi Karamollahi, Carey L. Williamson, Martin F. Arlitt
PAM4
2022 Zoomiversity: A Case Study of Pandemic Effects on Post-secondary Teaching and Learning
Mehdi Karamollahi, Carey L. Williamson, Martin F. Arlitt
PAM3
2021 Fast and Efficient Performance Tuning of Microservices
abstract
The microservice architecture is being increasingly adopted. Microservices often rely on containerization technology, facilitating agile development and permitting flexible deployment on cloud platforms. Many microservice applications are interactive. Consequently, there is a need for pre-deployment performance tuning techniques to ensure that an application will meet its end user response time requirements post-deployment. Additionally, the tuning process should be efficient, i.e., allocate just enough resources to minimize costs in cloud-based deployments. Furthermore, the tuning process needs to be fast to facilitate agile deployments. We design and evaluate a technique called MOAT (Microservice Application Performance Tuner) that embodies these requiremenis. MOAT conducts iterative performance tests to determine resource allocations for the individual microservices in an application for any given workload. It exploits a novel optimization technique that identifies resource allocations while requiring only a limited number of performance tests to explore the tuning space. Validation using an experimental system shows that MOAT outperforms a competing approach based on Bayesian optimization in terms of both solution speed and resource allocation efficiency.
Vahid MirzaEbrahim Mostofi, Diwakar Krishnamurthy, Martin F. Arlitt
CLOUD3
2020 Contention Aware Web of Things Emulation Testbed
abstract
Since the advent of the Web, new Web benchmarking tools have frequently been introduced to keep up with evolving workloads and environments. The introduction of Web of Things (WoT) marks the beginning of another important paradigm that requires new benchmarking tools and testbeds. Such a WoT benchmarking testbed can enable the comparison of different WoT application configurations and workload scenarios under assumptions regarding WoT application resource demands and WoT device network characteristics. The powerful computational capabilities of modern commodity multicore servers along with the limited resource consumption footprints of WoT devices suggest the feasibility of a benchmarking testbed that can emulate the application behaviour of a large number of WoT devices on just a single multicore server. However, to obtain test results that reflect the true performance of the system being emulated, care must be exercised to detect and consider the impact of testbed bottlenecks on performance results. For example, if too many WoT devices are emulated then performance metrics obtained from a test run, e.g., WoT device response times, would only reflect contention among emulated devices for shared multicore server resources instead of providing a true indication of the performance of the WoT system being emulated. We develop a testbed that helps a user emulate a system consisting of multiple WoT devices on a single multicore server by exploiting Docker containers. Furthermore, we devise a novel mechanism for the user to check whether shared resource contention in the testbed has impacted the integrity of test results. Our solution allows for careful scaling of experiments and enables resource efficient evaluation of a wide range of WoT systems, architectures, application characteristics, workload scenarios, and network conditions.
Raoufehsadat Hashemian, Niklas Carlsson, Diwakar Krishnamurthy, Martin F. Arlitt
ICPE4
2019 ACE - An Anomaly Contribution Explainer for Cyber-Security Applications
abstract
In this paper we introduce Anomaly Contribution Explainer or ACE, a tool to explain security anomaly detection models in terms of the model features through a regression framework, and its variant, ACE-KL, which highlights the important anomaly contributors. ACE and ACE-KL provide insights in diagnosing which attributes significantly contribute to an anomaly by building a specialized linear model to locally approximate the anomaly score that a black-box model generates. We conducted experiments with these anomaly detection models to detect security anomalies on both synthetic data and real data. In particular, we evaluate performance on three public data sets: CERT insider threat, netflow logs, and Android malware. The experimental results are encouraging: our methods consistently identify the correct contributing feature in the synthetic data where ground truth is available; similarly, for real data sets, our methods point a security analyst in the direction of the underlying causes of an anomaly, including in one case leading to the discovery of previously overlooked network scanning activity. We have made our source code publicly available.
Xiao Zhang 0017, Manish Marwah, I-Ta Lee, Martin F. Arlitt, Dan Goldwasser
IEEE BigData4
2019 Campus-Level Instagram Traffic: A Case Study
abstract
Instagram is a popular network application for photo sharing, video streaming, and online social media interaction. In this paper, we present results from an initial characterization study of Instagram network traffic, as viewed from a large campus edge network. Despite the challenges of NAT, DHCP, end-to-end encryption, and high traffic volume, we are able to identify key characteristics of Instagram traffic, which exceeds 1 TB per day. The main highlights from our study include classic observations such as diurnal usage patterns, Zipf-like distributions for IP frequency-rank profile, and heavy-tailed transfer size distributions.
Steffen Berg Klenow, Carey L. Williamson, Martin F. Arlitt, Sina Keshvadi
MASCOTS3
2018 Modeling, Analysis, and Characterization of Periodic Traffic on a Campus Edge Network
abstract
Traffic in today's edge networks is diverse, exhibiting many different patterns. This paper focuses on periodic network traffic, which is often used by known network services (e.g., Network Time Protocol, Akamai CDN) as well as by malicious applications (e.g., botnets, vulnerability scanning). We use a simple and flexible SQL-based approach as our computational model for detecting periodic traffic, and apply it to the analysis of seven weeks of Bro connection logs from a campus edge network. Our results show that periodic traffic analysis is effective for detecting P2P, gaming, cloud, scanning, and botnet traffic flows, which often exhibit periodic network communications. We present a classication taxonomy for periodic traffic, and provide an in-depth characterization of this traffic on our campus edge network.
Mackenzie Haffey, Martin F. Arlitt, Carey L. Williamson
MASCOTS2
2017 A First Look at the CT Landscape: Certificate Transparency Logs in Practice
Josef Gustafsson, Gustaf Overier, Martin F. Arlitt, Niklas Carlsson
PAM3
2017 IRIS: Iterative and Intelligent Experiment Selection
abstract
Benchmarking is a widely-used technique to quantify the performance of software systems. However, the design and implementation of a benchmarking study can face several challenges. In particular, the time required to perform a benchmarking study can quickly spiral out of control, owing to the number of distinct variables to systematically examine. In this paper, we propose IRIS, an IteRative and Intelligent Experiment Selection methodology, to maximize the information gain while minimizing the duration of the benchmarking process. IRIS selects the region to place the next experiment point based on the variability of both dependent, i.e., response, and independent variables in that region. It aims to identify a performance function that minimizes the response variable prediction error for a constant and limited experimentation budget. We evaluate IRIS for a wide selection of experimental, simulated and synthetic systems with one, two and three independent variables. Considering a limited experimentation budget, the results show IRIS is able to reduce the performance function prediction error up to 4.3 times compared to equal distance experiment point selection. Moreover, we show that the error reduction can further improve through system-specific parameter tuning. Analysis of the error distributions obtained with IRIS reveals that the technique is particularly effective in regions where the response variable is sensitive to changes in the independent variables.
Raoufehsadat Hashemian, Niklas Carlsson, Diwakar Krishnamurthy, Martin F. Arlitt
ICPE4
2017 Data Analytics for Managing Power in Commercial Buildings
abstract
Commercial buildings are significant consumers of electricity. We propose a number of methods for managing power in commercial buildings. The first step toward better energy management in commercial buildings is monitoring consumption. However, instrumenting every electrical panel in a large commercial building is an expensive proposition. In this article, we demonstrate that it is also unnecessary. Specifically, we propose a greedy meter (sensor) placement algorithm based on maximization of information gain subject to a cost constraint. The algorithm provides a near-optimal solution guarantee, and our empirical results demonstrate a 15% improvement in prediction power over conventional methods. Next, to identify power-saving opportunities, we use an unsupervised anomaly detection technique based on a low-dimensional embedding. Furthermore, to enable a building manager to effectively plan for demand response programs, we evaluate several solutions for fine-grained, short-term load forecasting. Our investigation reveals that support vector regression and an ensemble model work best overall. Finally, to better manage resources such as lighting and HVAC, we propose a semisupervised approach combining hidden Markov models (HMMs) and a standard classifier to model occupancy based on readily available port-level network statistics. We show that the proposed two-step approach simplifies the occupancy model while achieving good accuracy. The experimental results demonstrate an average occupancy estimation error of 9.3% with a potential reduction of 9.5% in lighting load using our occupancy models.
Gowtham Bellala, Manish Marwah, Martin F. Arlitt, Geoff Lyon, Cullen E. Bash, Amip Shah
ACM Trans. Cyber Phys. Syst.3
2015 IoTAbench: an Internet of Things Analytics Benchmark
abstract
The commoditization of sensors and communication networks is enabling vast quantities of data to be generated by and collected from cyber-physical systems. This ``Internet-of-Things" (IoT) makes possible new business opportunities, from usage-based insurance to proactive equipment maintenance. While many technology vendors now offer ``Big Data" solutions, a challenge for potential customers is understanding quantitatively how these solutions will work for IoT use cases. This paper describes a benchmark toolkit called IoTAbench for IoT Big Data scenarios. This toolset facilitates repeatable testing that can be easily extended to multiple IoT use cases, including a user's specific needs, interests or dataset. We demonstrate the benchmark via a smart metering use case involving an eight-node cluster running the HP Vertica analytics platform. The use case involves generating, loading, repairing and analyzing synthetic meter readings. The intent of IoTAbench is to provide the means to perform ``apples-to-apples" comparisons between different sensor data and analytics platforms. We illustrate the capabilities of IoTAbench via a large experimental study, where we store 22.8 trillion smart meter readings totaling 727 TB of data in our eight-node cluster.
Martin F. Arlitt, Manish Marwah, Gowtham Bellala, Amip Shah, Jeff Healey, Ben Vandiver
ICPE1
2014 Characterizing the scalability of a Web application on a multi-core server
abstract
SUMMARY The advent of multi‒core technology motivates new studies to understand how efficiently Web servers utilize such hardware. This paper presents a detailed performance study of a Web server application deployed on a modern eight‒core server. Our study shows that default Web server configurations result in poor scalability with increasing core counts. We study two different types of workloads, namely, a workload with intense TCP/IP related OS activity and the SPECweb2009 Support workload with more application‒level processing. We observe that the scaling behaviour is markedly different for these workloads, mainly because of the difference in the performance of static and dynamic requests. While static requests perform poorly when moving from using one socket to both sockets in the system, the converse is true for dynamic requests. We show that, contrary to what was suggested by previous work, Web server scalability improvement policies need to be adapted based on the type of workload experienced by the server. The results of our experiments reveal that with workload‒specific Web server configuration strategies, a multi‒core server can be utilized up to 80% while still serving requests without significant queuing delays; utilizations beyond 90% are also possible, while still serving requests with ‘acceptable’ response times. Copyright © 2014 John Wiley & Sons, Ltd.
Raoufehsadat Hashemian, Diwakar Krishnamurthy, Martin F. Arlitt, Niklas Carlsson
Concurr. Comput. Pract. Exp.3
2013 Maximizing server utilization while meeting critical SLAs via weight-based collocation management
Sergey Blagodurov, Daniel Gmach, Martin F. Arlitt, Yuan Chen 0001, Chris Hyser, Alexandra Fedorova
IM3
2013 Maximizing server utilization while meeting critical SLAs via weight-based collocation management
Sergey Blagodurov, Daniel Gmach, Martin F. Arlitt, Yuan Chen 0001, Chris Hyser, Alexandra Fedorova
IM3
2013 Improving the scalability of a multi-core web server
abstract
Improving the performance and scalability of Web servers enhances user experiences and reduces the costs of providing Web-based services. The advent of Multi-core technology motivates new studies to understand how efficiently Web servers utilize such hardware. This paper presents a detailed performance study of a Web server application deployed on a modern 2 socket, 4-cores per socket server. Our study show that default, "out-of-the-box" Web server configurations can cause the system to scale poorly with increasing core counts. We study two different types of workloads, namely a workload that imposes intense TCP/IP related OS activity and the SPECweb2009 Support workload, which incurs more application-level processing. We observe that the scaling behaviour is markedly different for these two types of workloads, mainly due to the difference in the performance characteristics of static and dynamic requests. The results of our experiments reveal that with workload-specific Web server configuration strategies a modern Multi-core server can be utilized up to 80% while still serving requests without significant queuing delays; utilizations beyond 90% are also possible, while still serving requests with acceptable response times.
Raoufehsadat Hashemian, Diwakar Krishnamurthy, Martin F. Arlitt, Niklas Carlsson
ICPE3
2012 Fine-Grained Photovoltaic Output Prediction Using a Bayesian Ensemble
abstract
Local and distributed power generation is increasingly relianton renewable power sources, e.g., solar (photovoltaic or PV) andwind energy. The integration of such sources into the power grid ischallenging, however, due to their variable and intermittent energyoutput. To effectively use them on alarge scale, it is essential to be able to predict power generation at afine-grained level. We describe a novel Bayesian ensemble methodologyinvolving three diverse predictors. Each predictor estimates mixingcoefficients for integrating PV generation output profiles but capturesfundamentally different characteristics. Two of them employ classicalparameterized (naive Bayes) and non-parametric (nearest neighbor) methods tomodel the relationship between weather forecasts and PV output. The thirdpredictor captures the sequentiality implicit in PV generation and uses motifsmined from historical data to estimate the most likely mixture weights usinga stream prediction methodology. We demonstrate the success and superiority of ourmethods on real PV data from two locations that exhibit diverse weatherconditions. Predictions from our model can be harnessed to optimize schedulingof delay tolerant workloads, e.g., in a data center.
Prithwish Chakraborty, Manish Marwah, Martin F. Arlitt, Naren Ramakrishnan
AAAI3
2012 Passive crowd-based monitoring of World Wide Web infrastructure and its performance
abstract
The World Wide Web and the services it provides are continually evolving. Even for a single time instant, it is a complex task to methodologically determine the infrastructure over which these services are provided and the corresponding effect on user perceived performance. For such tasks, researchers typically rely on active measurements or large numbers of volunteer users. In this paper, we consider an alternative approach, which we refer to as passive crowd-based monitoring. More specifically, we use passively collected proxy logs from a global enterprise to observe differences in the quality of service (QoS) experienced by users on different continents. We also show how this technique can measure properties of the underlying infrastructures of different Web content providers. While some of these properties have been observed using active measurements, we are the first to show that many of these properties (such as location of servers) can be obtained using passive measurements of actual user activity. Passive crowd-based monitoring has the advantages that it does not add any overhead on Web infrastructure, it does not require any specific software on the clients, but still captures the performance and infrastructure observed by actual Web usage.
Martin F. Arlitt, Niklas Carlsson, Carey L. Williamson, Jerome A. Rolia
ICC1
2012 Overcoming Web Server Benchmarking Challenges in the Multi-core Era
abstract
Web-based services are used by many organizations to support their customers and employees. An important consideration in developing such services is ensuring the Quality of Service (QoS) that users experience is acceptable. Recent years have seen a shift toward deploying Web service son multi-core hardware. Leveraging the performance benefits of multi-core hardware is a non-trivial task. In particular, systematic Web server benchmarking techniques are needed so organizations can verify their ability to meet customer QoS objectives while effectively utilizing such hardware. However, our recent experiences suggest that the multi-core era imposes significant challenges to Web server benchmarking. In particular, due to limitations of current hardware monitoring tools, we found that a large number of experiments are needed to detect complex bottlenecks that can arise in a multi-core system due to contention for shared resources such as cache hierarchy, memory controllers and processor inter-connects. Furthermore, multiple load generator instances are needed to adequately stress multi-core hardware. This leads to practical challenges in validating and managing the test results. This paper describes the automation strategies we employed to overcome these challenges. We make our test harness available for other researchers and practitioners working on similar studies.
Raoufehsadat Hashemian, Diwakar Krishnamurthy, Martin F. Arlitt
ICST3
2012 Following the electrons: methods for power management in commercial buildings
abstract
Commercial buildings are significant consumers of electricity. The first step towards better energy management in commercial buildings is monitoring consumption. However, instrumenting every electrical panel in a large commercial building is expensive and wasteful. In this paper, we propose a greedy meter (sensor) placement algorithm based on maximization of information gained, subject to a cost constraint. The algorithm provides a near-optimal solution guarantee. Furthermore, to identify power saving opportunities, we use an unsupervised anomaly detection technique based on a low-dimensional embedding. Further, to better manage resources such as lighting and HVAC, we propose a semi-supervised approach combining hidden Markov models (HMM) and a standard classifier to model occupancy based on readily available port-level network statistics.
Gowtham Bellala, Manish Marwah, Martin F. Arlitt, Geoff Lyon, Cullen E. Bash
KDD3
2012 Characterizing cyberlocker traffic flows
abstract
Cyberlockers have recently become a very popular means of distributing content. Today, cyberlocker traffic accounts for a non-negligible fraction of the total Internet traffic volume, and is forecasted to grow significantly in the future. The underlying protocol used in cyberlockers is HTTP, and increased usage of these services could drastically alter the characteristics of Web traffic. In light of the evolving nature of Web traffic, updated traffic models are required to capture this change. Despite their popularity, there has been limited work on understanding the characteristics of traffic flows originating from cyberlockers. Using a year-long trace collected from a large campus network, we present a comprehensive characterization study of cyberlocker traffic at the transport layer. We use a combination of flow-level and host-level characteristics to provide insights into the behavior of cyberlockers and their impact on networks. We also develop statistical models that capture the salient features of cyberlocker traffic. Studying the transport-layer interaction is important for analyzing reliability, congestion, flow control, and impact on other layers as well as Internet hosts. Our results can be used in developing improved traffic simulation models that can aid in capacity planning and network traffic management.
Aniket Mahanti, Niklas Carlsson, Martin F. Arlitt, Carey L. Williamson
LCN3
2012 Policy and mechanism for carbon-aware cloud applications
abstract
This position paper explores research challenges facing carbon-aware cloud applications. These applications run inside of a renewable-energy datacenter, provision resources on demand, and seek to minimize their use of carbon-heavy, grid energy. First, we argue that carbon-aware applications need new provisioning policies that address the uncertainty of renewable energy. Second, we argue that renewable-energy datacenters need mechanisms to determine the contribution of grid energy to specific application workloads. We propose first-cut solutions to these problems and present preliminary results for carbon-aware Web server running on a small renewable-energy cluster.
Christopher Stewart, Daniel Gmach, Martin F. Arlitt
NOMS4
2012 A Longitudinal Characterization of Local and Global BitTorrent Workload Dynamics
Niklas Carlsson, György Dán, Anirban Mahanti, Martin F. Arlitt
PAM4
2012 Towards a net-zero data center
abstract
A world consisting of billions of service-oriented client devices and thousands of data centers can deliver a diverse range of services, from social networking to management of our natural resources. However, these services must scale in order to meet the fundamental needs of society. To enable such scaling, the total cost of ownership of the data centers that host the services and comprise the vast majority of service delivery costs will need to be reduced. As energy drives the total cost of ownership of data centers, there is a need for a new paradigm in design and management of data centers that minimizes energy used across their lifetimes, from “cradle to cradle”. This tutorial article presents a blueprint for a “net-zero data center”: one that offsets any electricity used from the grid via adequate on-site power generation that gets fed back to the grid at a later time. We discuss how such a data center addresses the total cost of ownership, illustrating that contrary to the oft-held view of sustainability as “paying more to be green”, sustainable data centers—built on a framework that focuses on integrating supply and demand management from end-to-end—can concurrently lead to lowest cost and lowest environmental impact.
Prithviraj Banerjee, Chandrakant D. Patel, Cullen E. Bash, Amip Shah, Martin F. Arlitt
ACM J. Emerg. Technol. Comput. Syst.5
2012 Web workload generation challenges - an empirical investigation
abstract
SUMMARY Workload generators are widely used for testing the performance of Web‐based systems. Typically, these tools are also used to collect measurements such as throughput and end‐user response times that are often used to characterize the QoS provided by a system to its users. However, our study finds that Web workload generation is more difficult than it seems. In examining the popular RUBiS client generator, we found that reported response times could be grossly inaccurate, and that the generated workloads were less realistic than expected, causing server scalability to be incorrectly estimated. Using experimentation, we demonstrate how the Java virtual machine and the Java network library are the root causes of these issues. Our work serves as an example of how to verify the behavior of a Web workload generator. Copyright © 2011 John Wiley & Sons, Ltd.
Raoufehsadat Hashemian, Diwakar Krishnamurthy, Martin F. Arlitt
Softw. Pract. Exp.3
2011 Unsupervised Disaggregation of Low Frequency Power Measurements
abstract
Fear of increasing prices and concern about climate change are motivating residential power conservation efforts. We investigate the effectiveness of several unsupervised disaggregation methods on low frequency power measurements collected in real homes. Specifically, we consider variants of the factorial hidden Markov model. Our results indicate that a conditional factorial hidden semi-Markov model, which integrates additional features related to when and how appliances are used in the home and more accurately represents the power use of individual appliances, outperforms the other unsupervised disaggregation methods. Our results show that unsupervised techniques can provide perappliance power usage information in a non-invasive manner, which is ideal for enabling power conservation efforts.
Hyungsul Kim, Manish Marwah, Martin F. Arlitt, Geoff Lyon, Jiawei Han 0001
SDM3
2011 Improving the efficiency of information collection and analysis in widely-used it applications
abstract
Modern IT environments collect and analyze increasingly large volumes of data for a growing number of purposes (e.g., automated management, security, regulatory compliance, etc.). Simultaneously, such environments are challenged by the need to minimize their environmental footprints. A general solution to this problem is to utilize IT resources more efficiently.
Sergey Blagodurov, Martin F. Arlitt
ICPE2
2011 Towards more effective utilization of computer systems
abstract
Globally, vast infrastructures of Information Technology (IT) equipment are deployed. Much of this infrastructure is under utilized to ensure acceptable response times. This results in less than ideal use of the capital investment used to purchase the IT equipment. To improve the sustainability of IT, we focus on increasing the effective utilization of computer systems. Our results show that computer systems running delay-sensitive (e.g., Web) workloads can be more effectively utilized while still maintaining adequate (e.g., mean or upper percentile) response times. In particular, these computer systems can simultaneously support delay-tolerant workloads, to increase the value of work done by a computer system over time.
Niklas Carlsson, Martin F. Arlitt
ICPE2
2011 A Socratic method for validation of measurement-based networking research
Balachander Krishnamurthy, Walter Willinger, Phillipa Gill, Martin F. Arlitt
Comput. Commun.4
2011 Characterizing the file hosting ecosystem: A view from the edge
Aniket Mahanti, Carey L. Williamson, Niklas Carlsson, Martin F. Arlitt, Anirban Mahanti
Perform. Evaluation4
2011 Characterizing Intelligence Gathering and Control on an Edge Network
abstract
There is a continuous struggle for control of resources at every organization that is connected to the Internet. The local organization wishes to use its resources to achieve strategic goals. Some external entities seek direct control of these resources, for purposes such as spamming or launching denial-of-service attacks. Other external entities seek indirect control of assets (e.g., users, finances), but provide services in exchange for them. Using a year-long trace from an edge network, we examine what various external organizations know about one organization. We compare the types of information exposed by or to external organizations using either active ( reconnaissance ) or passive ( surveillance ) techniques. We also explore the direct and indirect control external entities have on local IT resources.
Martin F. Arlitt, Niklas Carlsson, Phillipa Gill, Aniket Mahanti, Carey L. Williamson
ACM Trans. Internet Techn.1
2011 Characterizing Organizational Use of Web-Based Services: Methodology, Challenges, Observations, and Insights
abstract
Today’s Web provides many different functionalities, including communication, entertainment, social networking, and information retrieval. In this article, we analyze traces of HTTP activity from a large enterprise and from a large university to identify and characterize Web-based service usage. Our work provides an initial methodology for the analysis of Web-based services. While it is nontrivial to identify the classes, instances, and providers for each transaction, our results show that most of the traffic comes from a small subset of providers, which can be classified manually. Furthermore, we assess both qualitatively and quantitatively how the Web has evolved over the past decade, and discuss the implications of these changes.
Phillipa Gill, Martin F. Arlitt, Niklas Carlsson, Anirban Mahanti, Carey L. Williamson
ACM Trans. Web2
2010 Leveraging Organizational Etiquette to Improve Internet Security
abstract
As more and more organizations rely on the Internet for their daily operation, Internet security becomes increasingly critical. Unfortunately, the vast resources available on the Internet are attracting many malicious users and organizations, including organized crime syndicates. With such organizations disguising their activity by operating from the machines owned and operated by legitimate organizations, we argue that responsible organizations could improve overall Internet security by strengthening their own. By improving their organizational etiquette, legitimate organizations will make it more difficult for malicious users and organizations to hide. Towards this goal, we propose a system to identify and eliminate malicious activity on edge networks. We use a year-long trace of activity from an edge network to characterize the malicious activity at an edge network and demonstrate the potential effectiveness of our system.
Niklas Carlsson, Martin F. Arlitt
ICCCN2
2010 Ambient Interference Effects in Wi-Fi Networks
Aniket Mahanti, Niklas Carlsson, Carey L. Williamson, Martin F. Arlitt
Networking4
2009 Characterization of FriendFeed - A Web-based Social Aggregation Service
Trinabh Gupta, Sanchit Garg, Anirban Mahanti, Niklas Carlsson, Martin F. Arlitt
ICWSM5
2008 Facebook Meets the Virtualized Enterprise
abstract
ldquoWeb 2.0rdquo and ldquocloud computingrdquo are revolutionizing the way IT infrastructure is accessed and managed. Web 2.0 technologies such as blogs, wikis and social networking platforms provide Internet users with easier mechanisms to produce Web content and to interact with each other. Cloud computing technologies are aimed at running applications as services over the Internet on a scalable infrastructure. In this paper we explore the advantages of using Web 2.0 and cloud computing technologies in an enterprise setting to provide employees with a comprehensive and transparent environment for utilizing applications. To demonstrate the effectiveness of this approach we have developed an environment that uses a social networking platform to provide access to a legacy application. The application is hosted on an internal cloud computing infrastructure that adapts dynamically to user demands. Initial feedback suggests this approach provides an improved user experience while simplifying management and increasing effective utilization of the underlying IT resources.
Roger Curry, Cameron Kiddle, Nayden Markatchev, Rob Simmonds, Tingxi Tan, Martin F. Arlitt, Bruce Walker
EDOC6
2008 The Flattening Internet Topology: Natural Evolution, Unsightly Barnacles or Contrived Collapse?
Phillipa Gill, Martin F. Arlitt, Zongpeng Li, Anirban Mahanti
PAM2
2008 A comparative analysis of web and peer-to-peer traffic
abstract
Peer-to-Peer (P2P) applications continue to grow in popularity, and have reportedly overtaken Web applications as the single largest contributor to Internet traffic. Using traces collected from a large edge network, we conduct an extensive analysis of P2P traffic, compare P2P traffic with Web traffic, and discuss the implications of increased P2P traffic. In addition to studying the aggregate P2P traffic, we also analyze and compare the two main constituents of P2P traffic in our data, namely BitTorrent and Gnutella. The results presented in the paper may be used for generating synthetic workloads, gaining insights into the functioning of P2P applications, and developing network management strategies. For example, our results suggest that new models are necessary for Internet traffic. As a first step, we present flow-level distributional models for Web and P2P traffic that may be used in network simulation and emulation experiments.
Naimul Basher, Aniket Mahanti, Anirban Mahanti, Carey L. Williamson, Martin F. Arlitt
WWW5
2007 Youtube traffic characterization: a view from the edge
abstract
This paper presents a traffic characterization study of the popular video sharing service, YouTube. Over a three month period we observed almost 25 million transactions between users on an edge network and YouTube, including more than 600,000 video downloads. We also monitored the globally popular videos over this period of time.
Phillipa Gill, Martin F. Arlitt, Zongpeng Li, Anirban Mahanti
Internet Measurement Conference2
2007 Comparing Wired-side and Wireless-side WLAN Monitoring Techniques: A Case Study
abstract
Wireless local area networks (WLANs) have become omnipresent: WLANs are available at airports, coffee shops, university campuses, corporate environments, and homes. This surge in the popularity of WLANs motivates the study of how these networks are used. Characterizing WLANs, however, is complicated by a number of factors including the geographic diversity of WLAN deployments and the need for capturing activity in the wireless environment instead of the wired environment. In this paper, we describe our experiences with the deployment and use of a remote passive wireless-side measurement infrastructure for monitoring usage of WLANs, and compare our results with a commonly used wired-side measurement technique.
Aniket Mahanti, Carey L. Williamson, Martin F. Arlitt, Anirban Mahanti
LCN3
2007 Semi-supervised network traffic classification
abstract
No abstract available.
Jeffrey Erman, Anirban Mahanti, Martin F. Arlitt, Ira Cohen, Carey L. Williamson
SIGMETRICS3
2007 Assessing the Completeness of Wireless-side Tracing Mechanisms
abstract
Analyzing traces of wireless network activity has many pragmatic purposes, from capacity planning to network design. Unfortunately, capturing complete traces of wireless traffic is difficult, and using incomplete traces can degrade the quality of the aforementioned analyses. In this paper we examine three different methods for estimating the completeness of wireless traces. We find that a method that examines MAC-layer sequence numbers provides the most accurate results. We also examine the effect of the placement of wireless sensors on the completeness of wireless-side traces. We determine that locating sensors such that the signal strengths between clients and access points is over 40% results in low miss rates at the sensor, and few CRC errors.
Aniket Mahanti, Martin F. Arlitt, Carey L. Williamson
WOWMOM2
2007 Identifying and discriminating between web and peer-to-peer traffic in the network core
abstract
Traffic classification is the ability to identify and categorize network traffic by application type. In this paper, we consider the problem of traffic classification in the network core.Classification at the core is challenging because only partial information about the flows and their contributors is available. We address this problem by developing a framework that can classify a flow using only unidirectional flow information. We evaluated this approach using recent packet traces that we collected and pre-classified to establish a "base truth". From our evaluation, we find that flow statistics for the server-to-client direction of a TCP connection provide greater classification accuracy than the flow statistics for the client-to-server direction. Because collection of the server-to-client flow statistics may not always be feasible, we developed and validated an algorithm that can estimate the missing statistics froma unidirectional packet trace.
Jeffrey Erman, Anirban Mahanti, Martin F. Arlitt, Carey L. Williamson
WWW3
2007 Offline/realtime traffic classification using semi-supervised learning
Jeffrey Erman, Anirban Mahanti, Martin F. Arlitt, Ira Cohen, Carey L. Williamson
Perform. Evaluation3
2007 Remote analysis of a distributed WLAN using passive wireless-side measurement
Aniket Mahanti, Carey L. Williamson, Martin F. Arlitt
Perform. Evaluation3
2006 Internet Traffic Identification using Machine Learning
abstract
We apply an unsupervised machine learning approach for Internet traffic identification and compare the results with that of a previously applied supervised machine learning approach. Our unsupervised approach uses an expectation maximization (EM) based clustering algorithm and the supervised approach uses the naive Bayes classifier. We find the unsupervised clustering technique has an accuracy up to 91% and outperform the supervised technique by up to 9%. We also find that the unsupervised technique can be used to discover traffic from previously unknown applications and has the potential to become an excellent tool for exploring Internet traffic.
Jeffrey Erman, Anirban Mahanti, Martin F. Arlitt
GLOBECOM3
2005 Dealing with Scale and Adaptation of Global Web Services Management
abstract
Service oriented architectures (SOA) are becoming the prevalent approach for realizing modern services and systems. SOA offers superior support for autonomy (decoupling) and heterogeneity compared to previous generation middleware systems, resulting in more scalable and adaptive solutions. However, SOA have not adequately addressed management, while traditional management solutions do not sufficiently scale to address the needs of (global) Web services. We propose scalable management based on models and industry standards. We discuss a use case for global service management, we present its design, implementation and preliminary evaluation. We retain all the benefits of SOA while also enabling global scale manageability. Our approach provides manageability that is comprehensible for administrators yet automated enough for integration into autonomous systems.
William Vambenepe, Carol Thompson, Vanish Talwar, Sandro Rafaeli, Bryan S. Murray, Dejan S. Milojicic, Subu Iyer, Keith I. Farkas, Martin F. Arlitt
ICWS9
2005 Adaptive entitlement control of resource containers on shared servers
abstract
In this paper, we describe the design of online feedback control algorithms to dynamically adjust entitlement values for a resource container on a server shared by multiple applications. The goal is to determine the minimum level of entitlement for the container such that its hosted application achieves desired performance levels. Classic control theory is used for both model identification and controller design. Specific implementation issues that affect the closed-loop system performance are discussed. A self-tuning adaptive controller is also presented to handle limited variations in the workload. The controllers were implemented and evaluated on a testbed using the HP-UX PRM as the resource container and the Apache Web server as the hosted application in the container. In all experiments, our controller was able to quickly converge to the proper level of CPU entitlement for the Web server to track its performance target. By using our entitlement control system, shared servers can potentially reach much higher resource utilization while meeting service level objectives for the hosted applications under changing operating conditions.
Xiaoyun Zhu, Sharad Singhal, Martin F. Arlitt
Integrated Network Management4
2005 Quartermaster - a resource utility system
abstract
Utility computing is envisioned as the future of enterprise IT environments. Achieving utility computing is a daunting task, because enterprise users have diverse and complex needs. In this paper we describe quartermaster, an integrated set of tools that addresses some of these needs. Quartermaster supports the entire lifecycle of computing tasks - including design, deployment, operation, and decommissioning of each task. Although individual components of this lifecycle have been addressed in earlier work, quartermaster integrates them in a unified framework using model-based automation. All tools within quartermaster are integrated using models based on the common information model (CIM), an industry-standard model from the distributed management task force (DMTF). The paper discusses the quartermaster implementation, and describes two case studies using quartermaster.
Sharad Singhal, Martin F. Arlitt, Dirk Beyer 0002, Sven Graupner, Vijay Machiraju, Jim Pruyne, Jerome A. Rolia, Akhil Sahai, Cipriano A. Santos, Julie Ward, Xiaoyun Zhu
Integrated Network Management2
2005 Predicting Short-Transfer Latency from TCP Arcana: A Trace-based Validation
Martin F. Arlitt, Balachander Krishnamurthy, Jeffrey C. Mogul
Internet Measurement Conference1
2004 Statistical service assurances for applications in utility grid environments
Jerome A. Rolia, Xiaoyun Zhu, Martin F. Arlitt, Artur Andrzejak 0001
Perform. Evaluation3
2004 Understanding Web server configuration issues
abstract
Abstract This paper proposes a methodological approach to the evaluation of Web server performance in a simple local area network test environment. The paper examines how different system and application configuration parameters can, over a range of workloads, impact the performance of a Web server. Our approach relies on relatively fine‐grain reporting of performance data for a broad set of system‐level metrics. Graphical visualization of these performance indices helps to identify the primary system bottleneck in each configuration studied. The Apache Web server is used as a case study to demonstrate the methodology. Our experiments quantify the performance implications of several configuration decisions common to any Web server implementation, and also serve to illustrate several performance anomalies specific to the Apache Web server (if misconfigured). Copyright © 2004 John Wiley & Sons, Ltd.
Martin F. Arlitt, Carey L. Williamson
Softw. Pract. Exp.1
2003 Resource Access Management for a Utility Hosting Enterprise Applications
Jerome A. Rolia, Xiaoyun Zhu, Martin F. Arlitt
Integrated Network Management3
2003 Grids for Enterprise Applications
Jerome A. Rolia, Jim Pruyne, Xiaoyun Zhu, Martin F. Arlitt
JSSPP4
2002 Web server benchmarking using parallel WAN emulation
abstract
This paper discusses the use of a parallel discrete-event network emulator called the Internet Protocol Traffic and Network Emulator (IP-TNE) for Web server benchmarking. The experiments in this paper demonstrate the feasibility of high-performance WAN emulation using parallel discrete-event simulation techniques on shared-memory multiprocessors. Our experiments with the Apache Web server achieve 3400 HTTP transactions per second for simple Web workloads, and 1000 HTTP transactions per second for realistic Web workloads, for static document retrieval across emulated WAN topologies of up to 4096 concurrent Web/TCP clients. The results show that WAN characteristics, including round-trip delays, link speeds, packet losses, packet sizes, and bandwidth asymmetry, all have significant impacts on Web server performance. WAN emulation enables stress testing and benchmarking of Web server performance in ways that may not be possible in simple LAN test scenarios.
Rob Simmonds, Carey L. Williamson, Russell J. Bradford, Martin F. Arlitt, Brian W. Unger
SIGMETRICS4
2002 A case study of Web server benchmarking using parallel WAN emulation
Carey L. Williamson, Rob Simmonds, Martin F. Arlitt
Perform. Evaluation3
2001 Characterizing the scalability of a large web-based shopping system
abstract
This article presents an analysis of five days of workload data from a large Web-based shopping system. The multitier environment of this Web-based shopping system includes Web servers, application servers, database servers, and an assortment of load-balancing and firewall appliances. We characterize user requests and sessions and determine their impact on system performance scalability. The purpose of our study is to assess scalability and support capacity planning exercises for the multitier system. We find that horizontal scalability is not always an adequate mechanism for supporting increased workloads and that personalization and robots can have a significant impact on system scalability.
Martin F. Arlitt, Diwakar Krishnamurthy, Jerome A. Rolia
ACM Trans. Internet Techn.1
2000 Performance evaluation of Web proxy cache replacement policies
Martin F. Arlitt, Rich Friedrich, Tai Jin
Perform. Evaluation1
1997 Internet Web servers: workload characterization and performance implications
abstract
This paper presents a workload characterization study for Internet Web servers. Six different data sets are used in the study: three from academic environments, two from scientific research organizations, and one from a commercial Internet provider. These data sets represent three different orders of magnitude in server activity, and two different orders of magnitude in time duration, ranging from one week of activity to one year. The workload characterization focuses on the document type distribution, the document size distribution, the document referencing behavior, and the geographic distribution of server requests. Throughout the study, emphasis is placed on finding workload characteristics that are common to all the data sets studied. Ten such characteristics are identified. The paper concludes with a discussion of caching and performance issues, using the observed workload characteristics to suggest performance enhancements that seem promising for Internet Web servers.
Martin F. Arlitt, Carey L. Williamson
IEEE/ACM Trans. Netw.1
1996 Web Server Workload Characterization: The Search for Invariants
abstract
The phenomenal growth in popularity of the World Wide Web (WWW, or the Web) has made WWW traffic the largest contributor to packet and byte traffic on the NSFNET backbone. This growth has triggered recent research aimed at reducing the volume of network traffic produced by Web clients and servers, by using caching, and reducing the latency for WWW users, by using improved protocols for Web interaction.Fundamental to the goal of improving WWW performance is an understanding of WWW workloads. This paper presents a workload characterization study for Internet Web servers. Six different data sets are used in this study: three from academic (i.e., university) environments, two from scientific research organizations, and one from a commercial Internet provider. These data sets represent three different orders of magnitude in server activity, and two different orders of magnitude in time duration, ranging from one week of activity to one year of activity.Throughout the study, emphasis is placed on finding workload invariants: observations that apply across all the data sets studied. Ten invariants are identified. These invariants are deemed important since they (potentially) represent universal truths for all Internet Web servers. The paper concludes with a discussion of caching and performance issues, using the invariants to suggest performance enhancements that seem most promising for Internet Web servers.
Martin F. Arlitt, Carey L. Williamson
SIGMETRICS1