VLDB 2026 Research / reviewers in the wild / expert
Aleksandar Kuzmanovic
dblp:k/AleksandarKuzmanovic
· DBLP profile ↗
69ranked-venue papers
9as first author
14since 2021 · last 2026
0000-0003-2622-6019ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 44 · 8 first-author · 7 since 2021Databases, data management, data science and information retrieval · 12 · 3 since 2021Systems, architecture and hardware · 10 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3Security and privacy · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RAG on the Decentralized Web: An Empirical Study of Akash and GolemabstractDecentralized compute markets claim to be an alternative to the cloud, yet the systems community lacks hard evidence on whether they can run modern workloads end to end. We put this claim to the test by deploying a Retrieval-Augmented Generation (RAG) pipeline on the Akash and Golem networks, two popular decentralized computing markets. Our analysis reveals a striking paradox: supply is abundant, with over 10,000 CPU cores and 50 TiB of memory available, yet utilization remains minimal (about 0.3% on Akash). When mapped carefully, several RAG pipeline stages execute correctly and up to 3 × cheaper than AWS; others degrade under node churn and network bottlenecks. We also observe that decentralization in practice is already hybrid, with centralized gateways and offchain services masking blockchain complexity. Taken together, our findings suggest that decentralized markets are neither hype nor drop-in cloud replacements, but an underutilized substrate whose viability depends on better scheduling, reliability mechanisms, and trust primitives. Suting Chen, Matteo Varvello, Aleksandar Kuzmanovic, Yunming Xiao |
APNet | 3 |
| 2025 | Leveraging Cross-Directional Dependency in Realtime Interactive StreamingabstractRealtime interactive streaming, an emerging paradigm that enables instantaneous two-way interactions with diverse streaming traffic, introduces new network challenges. This paper identifies a unique characteristic of this paradigm — cross-directional dependency. By leveraging this dependency, we propose a novel prioritization-based approach to optimize realtime interactive streaming. Our design aligns application-layer demands with transport-layer behavior, prioritizes critical traffic to mitigate buffer overflows caused by streaming microbursts, and transforms chaotic congestion into a predictable process. We extend the QUIC priority system to accommodate unreliable QUIC datagrams and develop an end-to-end framework for evaluating this emerging streaming paradigm. Extensive Internet-scale experiments validate the effectiveness of our system, demonstrating up to 9.4 × reduction in motion-to-photon latency and 82 × reduction in freeze frame rates. Sen Lin 0009, Andre Chen, Kevin Zhikai Chen, Aleksandar Kuzmanovic |
MMAsia | 4 |
| 2025 | PreAcher: Secure and Practical Password Pre-Authentication by Content Delivery Networks
Shihan Lin, Suting Chen, Yunming Xiao, Yanqi Gu, Aleksandar Kuzmanovic, Xiaowei Yang 0001 |
NSDI | 5 |
| 2025 | Unlocking ECMP Programmability for Precise Traffic Control
Yunming Xiao, Weizhen Dang, Xiang Li 0223, Zekun He, Jilong Wang 0001, Aleksandar Kuzmanovic, Ang Chen 0001, Congcong Miao |
NSDI | 9 |
| 2025 | Enabling Anonymous Online Streaming Analytics at the Network EdgeabstractIn recent years, content hyper-giants have increasingly deployed server infrastructure and services close to end-users within “eyeball” networks. Still, online streaming analytics has largely remained unaffected by this trend. This is despite the fact that most of the “big data” is received in real-time and is most valuable at the time of arrival. The inability to process data at the network edge is caused by a common setting where user profiles, necessary for analytics, are stored deep in the data center backends. This setting also carries privacy concerns as such user profiles are individually identifiable, yet the users are almost blind to what data is associated with their identities and how the data is analyzed. In this article, we revise this arrangement, and plant encrypted semantic cookies at the user end. By redesigning the cookie content without altering existing protocols, semantic cookies enable the capture and pre-processing of user data at edge ISPs or CDNs while preserving user anonymity. Additionally, lightweight cryptographic algorithms like partially homomorphic encryption can protect web providers’ proprietary data from CDNs. We present Snatch , a QUIC-based streaming analytics prototype that achieves up to 200x faster user analytics, with common-case improvements of 10-30x. Yunming Xiao, Yanqi Gu, Sen Lin 0009, Aleksandar Kuzmanovic |
ACM Trans. Comput. Syst. | 5 |
| 2024 | Snatch: Online Streaming Analytics at the Network EdgeabstractIn recent years, we have witnessed a growing trend of content hyper-giants deploying server infrastructure and services close to end-users, in "eyeball" networks. Still, one of the services that remained largely unaffected by this trend is online streaming analytics. This is despite the fact that most of the "big data" is received in real time and is most valuable at the time of arrival. The inability to process requests at the network edge is caused by a common setting where user profiles, necessary for analytics, are stored deep in the data center back-ends. This setting also carries privacy concerns as such user profiles are individually identifiable, yet the users are almost blind to what data is associated with their identities and how the data is analyzed. In this paper, we revise this arrangement, and plant encrypted semantic cookies at the user end. Without altering any of the existing protocols, this enables capturing and analytically pre-processing user requests soon after they are generated, at edge ISPs or content providers' off-nets. In addition, it ensures user anonymity perseverance during the analytics. We design and implement Snatch, a QUIC-based streaming analytics prototype, and demonstrate that it speeds up user analytics by up to 200x, and by 10-30x in the common case. Yunming Xiao, Sen Lin 0009, Aleksandar Kuzmanovic |
EuroSys | 4 |
| 2024 | Optimizing Traffic in Public-Facing Data Centers Amid Internet ProtocolsabstractRapid development has been witnessed in optimizing the performance of data centers over the past decade. However, such advances thriving in private data centers are rarely deployed in public-facing data centers. A major challenge is synchronizing optimization signals—such as flow sizes, server assignments, and load information—with the traffic they are intended to optimize, especially across networks controlled by different entities. In this paper, we propose CloudCookie, a versatile signal carrier within Internet protocols that ensures bidirectional signal presence without client-side cooperation. To exemplify CloudCookie's benefits on public-facing data center traffic, we design a set of easy-to-deploy data center infrastructures, including load balancers and switches, to leverage application layer awareness and enable efficient flow packet scheduling and load balancing. Our evaluation shows that these advances synergistically optimize the 99thpercentile of flow completion time by up to 20 x for the majority of flows. Sen Lin 0009, Aleksandar Kuzmanovic |
ICNP | 3 |
| 2024 | Conspirator: SmartNIC-Aided Control Plane for Distributed ML Workloads
Yunming Xiao, Diman Zad Tootaghaj, Aditya Dhakal, Lianjie Cao, Puneet Sharma 0001, Aleksandar Kuzmanovic |
USENIX ATC | 6 |
| 2023 | TENSOR: Lightweight BGP Non-Stop RoutingabstractAs the solitary inter-domain protocol, BGP plays an important role in today's Internet. Its failures threaten network stability and will usually result in large-scale packet losses. Thus, the non-stop routing (NSR) capability that protects inter-domain connectivity from being disrupted by various failures, is critical to any Autonomous System (AS) operator. Replicating the BGP and underlying TCP connection status is key to realizing NSR. But existing NSR solutions, which heavily rely on OS kernel modifications, have become impractical due to providers' adoption of virtualized network gateways for better scalability and manageability. Congcong Miao, Yunming Xiao, Marco Canini, Ruiqiang Dai, Shengli Zheng, Jilong Wang 0001, Jiwu Bu, Aleksandar Kuzmanovic, Yachen Wang |
SIGCOMM | 8 |
| 2023 | Demo: PDNS: A Fully Privacy-Preserving DNSabstractThe Domain Name System (DNS) is a key component of Internet-based communication and its privacy has been neglected for years. Recently, DNS over HTTPS has improved the situation by fixing the issue of in-path middleboxes. Further progress has been made with proxy-based solutions such as Oblivious DoH, which separate a user's identity from their DNS queries. However, these solutions rely on non-collusion between DNS resolvers and proxy networks. This paper instead proposes PDNS, a new DNS extension that uses Private Information Retrieval to allow DNS resolvers to operate on blind queries, thereby eliminating any privacy leaks. Yunming Xiao, Chenkai Weng, Ruijie Yu, Peizhi Liu, Matteo Varvello, Aleksandar Kuzmanovic |
SIGCOMM | 6 |
| 2023 | Decoding the Kodi EcosystemabstractFree and open-source media centers are experiencing a boom in popularity for the convenience they offer users seeking to remotely consume digital content. Kodi is today’s most popular home media center, with millions of users worldwide. Kodi’s popularity derives from its ability to centralize the sheer amount of media content available on the Web, both free and copyrighted . Researchers have been hinting at potential security concerns around Kodi, due to add-ons injecting unwanted content as well as user settings linked with security holes. Motivated by these observations, this article conducts the first comprehensive analysis of the Kodi ecosystem: 15,000 Kodi users from 104 countries, 11,000 unique add-ons, and data collected over 9 months. Our work makes three important contributions. Our first contribution is that we build “crawling” software ( de-Kodi ) which can automatically install a Kodi add-on, explore its menu, and locate (video) content. This is challenging for two main reasons. First, Kodi largely relies on visual information and user input which intrinsically complicates automation. Second, the potential sheer size of this ecosystem (i.e., the number of available add-ons) requires a highly scalable crawling solution. Our second contribution is that we develop a solution to discover Kodi add-ons. Our solution combines Web crawling of popular websites where Kodi add-ons are published (LazyKodi and GitHub) and SafeKodi , a Kodi add-on we have developed which leverages the help of Kodi users to learn which add-ons are used in the wild and, in return, offers information about how safe these add-ons are, e.g., do they track user activity or contact sketchy URLs/IP addresses. Our third contribution is a classifier to passively detect Kodi traffic and add-on usage in the wild. Our analysis of the Kodi ecosystem reveals the following findings. We find that most installed add-ons are unofficial but safe to use. Still, 78% of the users have installed at least one unsafe add-on, and even worse, such add-ons are among the most popular. In response to the information offered by SafeKodi, one-third of the users reacted by disabling some of their add-ons. However, the majority of users ignored our warnings for several months attracted by the content such unsafe add-ons have to offer. Last but not least, we show that Kodi’s auto-update, a feature active for 97.6% of SafeKodi users, makes Kodi users easily identifiable by their ISPs. While passively identifying which Kodi add-on is in use is, as expected, much harder, we also find that many unofficial add-ons do not use HTTPS yet, making their passive detection straightforward. 1 Yunming Xiao, Matteo Varvello, Marc Anthony Warrior, Aleksandar Kuzmanovic |
ACM Trans. Web | 4 |
| 2022 | Blockchain Mining: Optimal Resource AllocationabstractHaving enabled numerous applications, blockchains have attracted not only much attention, in the past decade, but also huge amount of resources: talent, capital, energy, etc. Focusing on the mining side of the market, in this paper, we aim at understanding how to efficiently use the resources mining and staking pools attract. We start with developing predictions about factors that increase the efficient allocation of pools' resources. We then test our predictions based on a general model for optimal resource allocation that we develop, as well as data we collected on pools' actual resource allocations. We find that pools can increase resource efficiency by mining for more blockchains as well as by increasing the frequency of resource re-allocation. Further, we enroll to mining pools as a miner to understand and comment on how pools can encourage their miners to increase the efficiency of their allocation. While our empirical investigation mostly focuses on the BTC family, we show that our theory and results are general and applicable to the Ethereum family as well as other proof-of-work (PoW) and proof-of-stake (PoS) chains. Yunming Xiao, Sarit Markovich, Aleksandar Kuzmanovic |
AFT | 3 |
| 2021 | Web-LEGO: Trading Content Strictness for Faster WebpagesabstractThe current Internet content delivery model assumes strict mapping between a resource and its descriptor, e.g., a JPEG file and its URL. Content Distribution Networks (CDNs) extend it by replicating the same resources across multiple locations, and introducing multiple descriptors. The goal of this work is to build Web-LEGO, an opt-in service, to speedup webpages at client side. Our rationale is to replace the slow original content with fast similar or equal content. Further, we perform a reality check of this idea both in term of the prevalence of CDN-less websites, availability of similar content, and user perception of similar webpages via millions of scale automated tests and thousands of real users. Then, we devise Web-LEGO, and address natural concerns on content inconsistency and copyright infringements. The final evaluation shows that Web-LEGO brings significant improvements both in term of reduced Page Load Time (PLT) and user-perceived PLT. Specifically, CDN-less websites provide more room for speedup than CDN-hosted ones, i.e., 7x more in the median case. Besides, Web-LEGO achieves high visual accuracy (94.2%) and high scores from a paid survey: 92% of the feedback collected from 1,000 people confirm Web-LEGO's accuracy as well as positive interest in the service. Pengfei Wang 0013, Matteo Varvello, Chunhe Ni, Ruiyun Yu, Aleksandar Kuzmanovic |
INFOCOM | 5 |
| 2021 | Utilizing Web Trackers for Sybil DefenseabstractUser tracking has become ubiquitous practice on the Web, allowing services to recommend behaviorally targeted content to users. In this article, we design Alibi, a system that utilizes such readily available personalized content, generated by recommendation engines in real time, as a means to tame Sybil attacks. In particular, by using ads and other tracker-generated recommendations as implicit user “certificates,” Alibi is capable of creating meta-profiles that allow for rapid and inexpensive validation of users’ uniqueness, thereby enabling an Internet-wide Sybil defense service. We demonstrate the feasibility of such a system, exploring the aggregate behavior of recommendation engines on the Web and demonstrating the richness of the meta-profile space defined by such inputs. We further explore the fundamental properties of such meta-profiles, i.e., their construction, uniqueness, persistence, and resilience to attacks. By conducting a user study, we show that the user meta-profiles are robust and show important scaling effects. We demonstrate that utilizing even a moderate number of popular Web sites empowers Alibi to tame large-scale Sybil attacks. Marcel Flores, Andrew Kahn, Marc Anthony Warrior, Alan Mislove, Aleksandar Kuzmanovic |
ACM Trans. Web | 5 |
| 2020 | De-Kodi: Understanding the Kodi EcosystemabstractFree and open source media centers are currently experiencing a boom in popularity for the convenience and flexibility they offer users seeking to remotely consume digital content. This newfound fame is matched by increasing notoriety—for their potential to serve as hubs for illegal content—and a presumably ever-increasing network footprint. It is fair to say that a complex ecosystem has developed around Kodi, composed of millions of users, thousands of “add-ons”—Kodi extensions from 3rd-party developers—and content providers. Motivated by these observations, this paper conducts the first analysis of the Kodi ecosystem. Our approach is to build “crawling” software around Kodi which can automatically install an addon, explore its menu, and locate (video) content. This is challenging for many reasons. First, Kodi largely relies on visual information and user input which intrinsically complicates automation. Second, no central aggregators for Kodi addons exist. Third, the potential sheer size of this ecosystem requires a highly scalable crawling solution. We address these challenges with de-Kodi, a full fledged crawling system capable of discovering and crawling large cross-sections of Kodi’s decentralized ecosystem. With de-Kodi, we discovered and tested over 9,000 distinct Kodi addons. Our results demonstrate de-Kodi, which we make available to the general public, to be an essential asset in studying one of the largest multimedia platforms in the world. Our work further serves as the first ever transparent and repeatable analysis of the Kodi ecosystem at large. Marc Anthony Warrior, Yunming Xiao, Matteo Varvello, Aleksandar Kuzmanovic |
WWW | 4 |
| 2019 | Kaleidoscope: A Crowdsourcing Testing Tool for Web Quality of ExperienceabstractToday's webpages development cycle consists of constant iterations with the goal to improve user retention, time spent on site, and overall quality of experience. Big companies like Google, Facebook, Amazon, etc. invest a lot of time and money to perform online testing. The prohibitive costs of these approaches are an entry barrier for smaller players. Further, the lack of a substantial user-base can be problematic to ensure statistical significance within a reasonable duration. In this paper we propose Kaleidoscope, an automated tool to evaluate Web features at a large scale, quickly, accurately, and at a reasonable price. Kaleidoscope can test two crucial user-perceived Web features - the style and page loading. As far as we know, it is the first testing tool to replay page loading by controlling visual changes on a webpage. Kaleidoscope allows to concurrently load a webpage in two versions (e.g., different fonts, with vs without ads) that are shown to a participant side-by-side. Further, Kaleidoscope also allows a participant to interact with each webpage version and provide feedback, e.g., respond to a questionnaire previously prepared by an "experimenter". Kaleidoscope supports both voluntary and paid testers from FigureEight, a popular crowdsourcing platform. Using hundreds of FigureEight testers, we validate that Kaleidoscope matches the accuracy of trusted in-lab tests while providing results about 12x faster (and arguably at a lower cost) than A/B testing. Finally, we showcase how to use Kaleidoscope's page loading feature to study the user-perceived page load time (uPLT) of a webpage. Pengfei Wang 0013, Matteo Varvello, Aleksandar Kuzmanovic |
ICDCS | 3 |
| 2019 | Perceiving Internet Anomalies via CDN Replica ShiftsabstractAnomalies are a ubiquitous and inevitable phenomenon associated with a complex and large-scale system such as the Internet. While measuring and analyzing network anomalies is as old as the Internet itself, comprehensively detecting anomalies at a global scale is a challenging task that requires a significant measurement infrastructure. In this paper, we demonstrate that the production Content Distribution Networks (CDNs), and their pervasive network infrastructure, could be effectively utilized to detect Internet anomalies. Our approach avoids direct network measurements and instead relies on “abnormal” spatial and temporal CDN replica shifts to indirectly sense anomalies. We measure replica shifts for five CDNs (Google, Amazon, Akamai, Fastly, and Incapsula) for two months. Contrary to our expectations, we find that (i) Google's and Amazon's CDNs, which are characterized by rich connectivity and infrastructure, are not best suited for our method because they effectively mask anomalies; (ii) Akamai is the most “sophisticated” of all evaluated CDNs, yet again not best suited to detect anomalies because it reacts exceptionally to much smaller network performance variations; (iii) Fastly's and Incapsula's replica shifts strongly correlate with network anomalies, making them viable anomaly predictors. Yihao Jia, Aleksandar Kuzmanovic |
INFOCOM | 2 |
| 2018 | Mining the web with webcoinabstractFour major search engines, Google in particular, hold a unique position in enabling the use of the Internet, as they alone direct over 98% of Internet users to the content they seek, using proprietary indices. While the contribution of these companies is undeniable, their design is necessarily affected by their economic interests, which may or may not align with those of the users, raising concerns regarding their effect on the availability of information around the globe. While multiple academic and commercial projects aimed to distribute and democratize the Web search, they failed to gain much traction, mostly due to inferior results and lack of incentives for participation. In this paper, we show how complex networking-intensive tasks can be crowdsourced using Bitcoin's incentive model. We present Webcoin, a novel distributed digital-currency which utilizes networking resources rather then computational, and can only be mined through Web indexing. Webcoin provides both the incentives and the means to create Google-scale indices, freely available to competing services and the public. Webcoin's design overcomes numerous unique challenges, such as index verification, scalability, and nodes' ability to actively manipulate webpages. We deploy 200 fully-functioning Webcoin nodes and demonstrate their low bandwidth requirements. Uri Klarman, Marcel Flores, Aleksandar Kuzmanovic |
CoNEXT | 3 |
| 2018 | Fury Route: Leveraging CDNs to Remotely Measure Network Distance
Marcel Flores, Alexander Wenzel, Aleksandar Kuzmanovic |
PAM | 4 |
| 2017 | Drongo: Speeding Up CDNs with Subnet Assimilation from the ClientabstractCurrently, the attempt to choose the "best" content replica server for a client is carried out solely by CDNs. While CDNs have a decent view of load distribution and content placement, they receive little input from the clients themselves. We propose a hybrid solution, subnet assimilation, where the client participates in the server selection process while still leaving the final say to the CDN. Subnet assimilation allows clients to declare their own "network location," different from the actual one, which in turn drives a CDN towards making better decisions. To demonstrate, we introduce Drongo, a client-side system, readily deployable on existing clients without any changes to the CDNs, that employs subnet assimilation to dramatically improve replica server selection. We implemented and extensively evaluated Drongo on a set of 429 clients spread across 177 countries and 6 major CDNs. We show that Drongo affects 69.93% of all clients, prompting better CDN replica choices which reduce the latency of affected requests by up to an order of magnitude and by 24.89% on average across six major providers, with Google's performance improving by 50% in the median case. Our results indicate that client participation holds great opportunities for the advancement of CDN performance. Marc Anthony Warrior, Uri Klarman, Marcel Flores, Aleksandar Kuzmanovic |
CoNEXT | 4 |
| 2017 | Oak: User-Targeted Web PerformanceabstractWeb performance has long proved to be one of the most sought after and difficult to achieve components for the web. Since the inception of the modern web infrastructure, the situation has been growing in complexity, adding remote hosts and objects, providing everything from computation infrastructure, content distribution capability, and targeted advertising. While many of these components provide improvements for some users, the complexity of the Internet often leaves other users suffering from poor performance. We propose Oak, a system which addresses client performance on the individual level, hence addressing challenges which may be unique to the user. Oak measures a user's performance for objects loading on a page, and determines which components are under-performing. Oak further provides an automated mechanism by which sites are able to replace resources with those provided by a better performing alternative service for a particular user. In this work, we demonstrate the prevalence of under-performing services on the web, finding that over 60% of the Alexa Top 500 have at least one under-preforming server. We further evaluate Oak on experimental and popular existing webpages, and demonstrate its effectiveness in making decisions in existing environments and with a distributed user base. Marcel Flores, Alexander Wenzel, Aleksandar Kuzmanovic |
ICDCS | 3 |
| 2016 | Enabling router-assisted congestion control on the InternetabstractEnabling communication between routers and endpoints has long been sought after as an approach to congestion control in the Internet. However, the narrow-waist of TCP/IP has complicated the deployment of such communication. In this paper, we present Kick-Ass1, a congestion control mechanism that enables explicit rate congestion control protocols to be deployed within the TCP/IP stack. The key idea is to utilize packet lengths as a vehicle to communicate fine-grained explicit rate and other information from routers to endpoints and vice versa. Given that our approach (i) requires no explicit coordination among Kick-Ass routers, (ii) no explicit coordination among Kick-Ass routers and endpoints, and (iii) is effective on paths that include legacy routers, it provides a practical road towards a faster Internet, today. Using large-scale simulations, testbed experiments, and wide-area Internet evaluations, we demonstrate that (i) a basic explicit-rate protocol using the Kick-Ass mechanism improves flow completion times by up to an order of magnitude and outperforms endpoint-based approaches, including CUBIC and PCC. (ii) Kick-Ass is incrementally deployable on the Internet. (iii) Deploying Kick-Ass at end-hosts and edge routers can enable the above performance benefits, without waiting for universal adoption. (iv) Our packet-fragmentation mechanism is well behaved on the Internet. Marcel Flores, Alexander Wenzel, Aleksandar Kuzmanovic |
ICNP | 3 |
| 2016 | Understanding Factors That Affect Web Traffic via Twitter
Chunjing Xiao, Zhiguang Qin, Xucheng Luo, Aleksandar Kuzmanovic |
WISE (2) | 4 |
| 2015 | Wi-FM: Resolving Neighborhood Wireless Network Affairs by Listening to MusicabstractFM radio, typically broadcast in the 87.5 to 108.0Mhz range, is widely available in urban areas and beyond. Contrary to GPS, it effectively penetrates buildings, contrary to 3G/4G or TV, FM radio receivers are becoming freely available in mobile devices. Indeed, nearly every smart phone and many other consumer electronics today have a built-in FM chip. In this paper, we demonstrate that this ubiquitous in-the-air and on-device FM radio availability presents a unique opportunity to address some of the fundamental wireless networking problems. In particular, we focus on the problem commonly arising in home networks where devices from neighboring, yet autonomous and non-collaborative, Wi-Fi networks systematically "step on each other's feet", i.e., interfere and degrade each other's performance. We show that the digital signal that accompanies broadcast FM radio has sufficient structure to enable effective scheduling relative to it. It thus provides a common reference for neighboring devices to harmonize their transmissions, yet without requiring any explicit communication among them. To the best of our knowledge, our system is the first to enable such mutually-beneficial, autonomous, and implicit harmonization among Wi-Fi devices across administrative network bounds. Marcel Flores, Uri Klarman, Aleksandar Kuzmanovic |
ICNP | 3 |
| 2014 | A CDN-based Domain Name System
Chunjing Xiao, Qiyao Wang, Yuehui Jin, Aleksandar Kuzmanovic |
Comput. Commun. | 5 |
| 2014 | How to Improve Your Search Engine Ranking: Myths and RealityabstractSearch engines have greatly influenced the way people access information on the Internet, as such engines provide the preferred entry point to billions of pages on the Web. Therefore, highly ranked Web pages generally have higher visibility to people and pushing the ranking higher has become the top priority for Web masters. As a matter of fact, Search Engine Optimization (SEO) has became a sizeable business that attempts to improve their clients’ ranking. Still, the lack of ways to validate SEO’s methods has created numerous myths and fallacies associated with ranking algorithms. In this article, we focus on two ranking algorithms, Google’s and Bing’s, and design, implement, and evaluate a ranking system to systematically validate assumptions others have made about these popular ranking algorithms. We demonstrate that linear learning models, coupled with a recursive partitioning ranking scheme, are capable of predicting ranking results with high accuracy. As an example, we manage to correctly predict 7 out of the top 10 pages for 78% of evaluated keywords. Moreover, for content-only ranking, our system can correctly predict 9 or more pages out of the top 10 ones for 77% of search terms. We show how our ranking system can be used to reveal the relative importance of ranking features in a search engine’s ranking function, provide guidelines for SEOs and Web masters to optimize their Web pages, validate or disprove new ranking features, and evaluate search engine ranking results for possible ranking bias. Ao-Jan Su, Y. Charlie Hu, Aleksandar Kuzmanovic, Cheng-Kok Koh |
ACM Trans. Web | 3 |
| 2013 | Rayleigh-normalized Gaussian noise in blind signal fusion
Aaron Ballew, Aleksandar Kuzmanovic, Chung-Chieh Lee |
FUSION | 2 |
| 2013 | Searching for Spam: Detecting Fraudulent Accounts via Web Search
Marcel Flores, Aleksandar Kuzmanovic |
PAM | 2 |
| 2013 | Mosaic: quantifying privacy leakage in mobile networksabstractWith the proliferation of online social networking (OSN) and mobile devices, preserving user privacy has become a great challenge. While prior studies have directly focused on OSN services, we call attention to the privacy leakage in mobile network data. This concern is motivated by two factors. First, the prevalence of OSN usage leaves identifiable digital footprints that can be traced back to users in the real-world. Second, the association between users and their mobile devices makes it easier to associate traffic to its owners. These pose a serious threat to user privacy as they enable an adversary to attribute significant portions of data traffic including the ones with NO identity leaks to network users' true identities. To demonstrate its feasibility, we develop the Tessellation methodology. By applying Tessellation on traffic from a cellular service provider (CSP), we show that up to 50% of the traffic can be attributed to the names of users. In addition to revealing the user identity, the reconstructed profile, dubbed as "mosaic," associates personal information such as political views, browsing habits, and favorite apps to the users. We conclude by discussing approaches for preventing and mitigating the alarming leakage of sensitive user information. Ning Xia, Han Hee Song, Marios Iliofotou, Antonio Nucci, Zhi-Li Zhang, Aleksandar Kuzmanovic |
SIGCOMM | 7 |
| 2012 | Selective Behavior in Online Social NetworksabstractAccording to the classical communication theories, known as Gate keeping and Selective Exposure, individuals tend to have selective behavior when they disseminate and receive information based on their psychological preferences. Selective behavior related to these two theories have been broadly studied separately. While, thanks to the advent of Online Social Networks (OSNs), larger-scale feedback and user information can be collected. In this paper, based on these data, We analyze the correlation among users' properties (such as age, gender, and cultural background) and analyze their selective behavior by tagging users as disseminators and/or audiences in YouTube, Flickr, and Twitter. We find that despite enormous amount of content available in OSNs, users have a comparatively small selective range and do exhibit selective behavior properties. In particular, they pay the most attention to the content published by disseminators that share similar properties, i.e., gender, age, and country. Nonetheless, we also find significant differences and commonalities among the three OSNs with respect to selective behavior. In particular, (i) the proportion and properties of disseminators, audiences, and dual-role users are quite different for the three networks, (ii) the global level of information spread in Flickr is almost two times than that in Twitter and YouTube is approximately the median one, (iii) For a given country, the global level of information spread is different for different OSNs. For a given OSN, it is different for different countries, (iv) despite ubiquitous presence of dual-role users in OSNs, most of such users are very active as either disseminators or audiences, but not both. Our findings are not only useful for understanding these two theories, but also have applications ranging from advertising and recommendation systems to developing predicting models. Chunjing Xiao, Ling Su, Juan Bi, Yuxia Xue, Aleksandar Kuzmanovic |
Web Intelligence | 5 |
| 2012 | Extracting user web browsing patterns from non-content network traces: The online advertising case study
Gabriel Maciá-Fernández, Rafael Rodríguez-Gómez, Aleksandar Kuzmanovic |
Comput. Networks | 4 |
| 2012 | P2P as a CDN: A new service model for file sharing
Amit Mondal, Ionut Trestian, Aleksandar Kuzmanovic |
Comput. Networks | 4 |
| 2012 | Taming the Mobile Data Deluge With Drop ZonesabstractHuman communication has changed by the advent of smartphones. Using commonplace mobile device features, they started uploading large amounts of content that increases. This increase in demand will overwhelm capacity and limits the providers' ability to provide the quality of service demanded by their users. In the absence of technical solutions, cellular network providers are considering changing billing plans to address this. Our contributions are twofold. First, by analyzing user content upload behavior, we find that the user-generated content problem is a user behavioral problem. Particularly, by analyzing user mobility and data logs of 2 million users of one of the largest US cellular providers, we find that: 1) users upload content from a small number of locations; 2) because such locations are different for users, we find that the problem appears ubiquitous. However, we find that: 3) there exists a significant lag between content generation and uploading times, and 4) with respect to users, it is always the same users to delay. Second, we propose a cellular network architecture. Our approach proposes capacity upgrades at a select number of locations called Drop Zones. Although not particularly popular for uploads originally, Drop Zones seamlessly fall within the natural movement patterns of a large number of users. They are therefore suited for uploading larger quantities of content in a postponed manner. We design infrastructure placement algorithms and demonstrate that by upgrading infrastructure in only 963 base stations across the entire US, it is possible to deliver 50% of content via Drop Zones. Ionut Trestian, Supranamaya Ranjan, Aleksandar Kuzmanovic, Antonio Nucci |
IEEE/ACM Trans. Netw. | 3 |
| 2011 | Fusion of live audio recordings for blind noise reduction
Aaron Ballew, Aleksandar Kuzmanovic, Chung-Chieh Lee |
FUSION | 2 |
| 2011 | Understanding the Network and User-Targeting Properties of Web Advertising NetworksabstractAdvertising has become an integral and inseparable part of the World Wide Web. However, neither public auditing nor monitoring mechanisms still exist in this emerging area. In this paper, we present our initial efforts on building a network and content-level auditing service for Web-based ad networks. Our network-level measurements -- charting the network infrastructure and quantifying the ad platforms' delay performance -- can help commissioners to evaluate their networks from end users' perspective, and let advertisers choose commissioners that better fit their needs. Our content-level measurements -- understanding the ad distribution mechanisms and evaluating location-based and behavioral targeting approaches -- bring useful auditing information to all entities involved in the on line advertising business. We extensively evaluate Google's, AOL's, and Ad blade's ad networks and demonstrate how their different design philosophies dominantly affect their performance at both network and content levels. Daniel Burgener, Aleksandar Kuzmanovic, Gabriel Maciá-Fernández |
ICDCS | 3 |
| 2011 | Taming user-generated content in mobile networks via Drop ZonesabstractSmartphones have changed the way people communicate. Most prominently, using commonplace mobile device features (e.g., high resolution cameras), they started producing and uploading large amounts of content that increases at an exponential pace. In the absence of viable technical solutions, some cellular network providers are considering to start charging special usage fees to address the problem. Our contributions are twofold. First, we find that the user-generated content problem is a user-behavioral problem. By analyzing user mobility and data logs of close to 2 million users of a cellular network, we find that (i) users upload content from a small number of locations, typically corresponding to their home or work locations; (ii) because such locations are different for different users, we find that the problem appears ubiquitous, since user-generated content uploads grow exponentially at most locations. However, we also find that (Hi) there exists a significant lag between content generation and uploading times. For example, we find that 55% of content that is uploaded via mobile phones is at least 1 day old. Second, based on the above insights, we propose a new cellular network architecture. Our approach proposes capacity upgrades at a select number of locations called Drop Zones. Although not particularly popular for uploads originally, Drop Zones seamlessly fall within the natural movement patterns of a large number of users. They are therefore better suited for uploading larger quantities of content in a postponed manner. We design infrastructure placement algorithms and demonstrate that by upgrading infrastructure in only 963 base-stations across the entire United States, it is possible to deliver 50% of total content via the Drop Zones. Ionut Trestian, Supranamaya Ranjan, Aleksandar Kuzmanovic, Antonio Nucci |
INFOCOM | 3 |
| 2011 | Towards Street-Level Client-Independent IP Geolocation
Daniel Burgener, Marcel Flores, Aleksandar Kuzmanovic, Cheng Huang 0002 |
NSDI | 4 |
| 2011 | Understanding Crowds' Migration on the WebabstractConsider a network where nodes are websites and the weight of a link that connects two nodes corresponds to the average number of users that visits both of the two websites over longer timescales. Such user-driven Web network is not only invaluable for understanding how crowds' interests collectively spread on the Web, but also useful for applications such as advertising or search. In this paper, we manage to construct such a network by 'putting together' pieces of information publicly available from the popular analytics websites. Our contributions are threefold. First, we design a crawler and a normalization methodology that enable us to construct a user-driven Web network based on limited publicly-available information, and validate the high accuracy of our approach. Second, we evaluate the unique properties of our network, and demonstrate that it exhibits small-world, seed-free, and scale-free phenomena. Finally, we build an application, website selector, on top of the user-driven network. The core concept utilized in the website selector is that by exploiting the knowledge that a number of websites share a number of common users, an advertiser might prefer displaying his ads only on a subset of these websites to optimize the budget allocation, and in turn increase the visibility of his ads on other websites. Our websites elector system is tailored for ad commissioners and it could be easily embedded in their ad selection algorithms. Komal Pal, Aleksandar Kuzmanovic |
Web Intelligence | 3 |
| 2010 | A Case for WiFi Relay: Improving VoIP Quality for WiFi UsersabstractVoice over Internet (VoIP) has been experiencing enormous growth in recent years. While posed to replace traditional PSTN for both enterprise and residential customers, VoIP has yet to achieve the same level of quality and reliability as PSTN. One key challenge is that a growing segment of customers is increasingly relying on WiFi connections. VoIP over WiFi (VoWiFi) experiences significant degradation in quality because of packet losses, mostly due to WiFi's low capacity, varying signal strength, interference, etc. To understand this problem, we have developed and deployed a comprehensive measurement platform in a global enterprise network. From large-scale real-world traces, we quantitatively analyze the impact of WiFi connections and study measures to mitigate such impact. Our results confirm that WiFi connections incur significantly more packet losses than wirelines, but these losses can be effectively concealed by sending each packet up to five times (heavy replication). Due to WiFi's inherent overhead, heavy replication only marginally increases WiFi airtime. To avoid the overhead on wirelines, we further propose a relay-based solution, where heavy replication only occurs between endpoints and nearby relays, and is removed before packets are transmitted on inter-branch long haul links or the public Internet. The solution has been implemented and deployed in the global enterprise network, and measurement results confirm that it can indeed greatly improve the performance of VoIP for WiFi users. In particular, it reduces the percentage of poor calls from 35% to 10%; and increases the percentage of acceptable ones from 45% to 70%. Amit Mondal, Cheng Huang 0002, Jin Li 0001, Aleksandar Kuzmanovic |
ICC | 5 |
| 2010 | Measurement and Diagnosis of Address Misconfigured P2P TrafficabstractMisconfigured P2P traffic caused by bugs in volunteer-developed P2P software or by attackers is prevalent. It influences both end users and ISPs. In this paper, we discover and study address-misconfigured P2P traffic, a major class of such misconfiguration. P2P address misconfiguration is a phenomenon in which a large number of peers send P2P file downloading requests to a ``random'' target on the Internet. On measuring three Honeynet datasets spanning four years and across five different /8 networks, we find address-misconfigured P2P traffic on average contributes 38.9% of Internet background radiation, increasing by more than 100% every year. In this paper, we design the P2PScope, a measurement tool, to detect and diagnose such unwanted traffic. We find, in all the P2P systems, address misconfiguration is caused by resource mapping contamination, i.e., the sources returned for a given file ID through P2P indexing are not valid. Different P2P systems have different reasons for such contamination. For eMule, we find that the root cause is mainly a network byte ordering problem in the eMule Source Exchange protocol. For BitTorrent misconfiguration, one reason is that anti-P2P companies actively inject bogus peers into the P2P system. Another reason is that the KTorrent implementation has a byte order problem. We also design approaches to detect anti-P2P peers without false positives. Zhichun Li, Anup Goyal, Yan Chen 0004, Aleksandar Kuzmanovic |
INFOCOM | 4 |
| 2010 | ISP-Enabled Behavioral Ad Targeting without Deep Packet InspectionabstractOnline advertising is a rapidly growing industry currently dominated by the search engine 'giant' Google. In an attempt to tap into this huge market, Internet Service Providers (ISPs) started deploying deep packet inspection techniques to track and collect user browsing behavior. However, such techniques violate wiretap laws that explicitly prevent intercepting the contents of communication without gaining consent from consumers. In this paper, we show that it is possible for ISPs to extract user browsing patterns without inspecting contents of communication. Our contributions are threefold. First, we develop a methodology and implement a system that is capable of extracting web browsing features from stored non-content based records of online communication, which could be legally shared. When such browsing features are correlated with information collected by independently crawling the Web, it becomes possible to recover the actual web pages accessed by clients. Second, we systematically evaluate our system on the Internet and demonstrate that it can successfully recover user browsing patterns with high accuracy. Finally, our findings call for a comprehensive legislative reform that would not only enable fair competition in the online advertising business, but more importantly, protect the consumer rights in a more effective way. Gabriel Maciá-Fernández, Rafael Rodríguez-Gómez, Aleksandar Kuzmanovic |
INFOCOM | 4 |
| 2010 | SureCall: Towards glitch-free real-time audio/video conferencingabstractGlobal enterprises are increasingly adopting unified communication solutions over traditional telephone systems. Such solutions provide integrated audio/video conferencing and messaging services, and enable flexible working environments by allowing mobile and dispersed users to communicate and collaborate easily and efficiently. The ultimate goal of unified communications is to ensure a smooth and best possible user experience across all scenarios. To address this challenge and understand the impact of various network scenarios on unified audio/video conferencing, we have developed a distributed experimental platform - SureCall - and deployed it on over 80 machines across a global enterprise and many residential networks. SureCall has collected worth of more than 6 months of packet-level audio/video conferencing traces. Through in-depth analysis of these traces, we have quantitatively compared how key performance metrics, such as packet loss and jitter, as well as the correlation between them, are affected by the enterprise and residential networks, by WiFi connections and VPN links, etc. In addition, we show how SureCall can serve as an ideal platform to design, experiment and validate new schemes and algorithms. We have developed a new audio quality classifier using the SureCall platform, which is being experimented with the recent release of Office Communicator solution for large-scale validation. Amit Mondal, Ross Cutler, Cheng Huang 0002, Jin Li 0001, Aleksandar Kuzmanovic |
IWQoS | 5 |
| 2010 | How to Improve Your Google Ranking: Myths and RealityabstractSearch engines have greatly influenced the way people access information on the Internet as such engines provide the preferred entry point to billions of pages on the Web. Therefore, highly ranked web pages generally have higher visibility to people and pushing the ranking higher has become the top priority for webmasters. As a matter of fact, search engine optimization (SEO) has became a sizeable business that attempts to improve their clients' ranking. Still, the natural reluctance of search engine companies to reveal their internal mechanisms and the lack of ways to validate SEO's methods have created numerous myths and fallacies associated with ranking algorithms; Google'sin particular. In this paper, we focus on the Google ranking algorithm and design, implement, and evaluate a ranking system to systematically validate assumptions others have made about this popular ranking algorithm. We demonstrate that linear learning models, coupled with a recursive partitioning ranking scheme, are capable of reverse engineering Google's ranking algorithm with high accuracy. As an example, we manage to correctly predict 7 out of the top 10 pages for 78% of evaluated keywords. Moreover, for content-only ranking, our system can correctly predict 9 or more pages out of the top 10 ones for 77% of search terms. We show how our ranking system can be used to reveal the relative importance of ranking features in Google's ranking function, provide guidelines for SEOs and webmasters to optimize their web pages, validate or disapprove new ranking features, and evaluate search engine ranking results for possible ranking bias. Ao-Jan Su, Y. Charlie Hu, Aleksandar Kuzmanovic, Cheng-Kok Koh |
Web Intelligence | 3 |
| 2010 | Analyzing content-level properties of the web adversphereabstractAdvertising has become an integral and inseparable part of the World Wide Web. However, neither public auditing nor monitoring mechanisms still exist in this emerging area. In this paper, we present our initial efforts on building a content-level auditing service for web-based ad networks. Our content-level measurements - understanding the ad distribution mechanisms and evaluating location-based and behavioral targeting approaches - bring useful auditing information to all entities involved in the online advertising business. We extensively evaluate Google's, AOL's, and Adblade's ad networks and demonstrate how their different design philosophies dominantly affect their performance at the content level. Daniel Burgener, Aleksandar Kuzmanovic, Gabriel Maciá-Fernández |
WWW | 3 |
| 2010 | Upgrading mice to elephants: effects and end-point solutions
Amit Mondal, Aleksandar Kuzmanovic |
IEEE/ACM Trans. Netw. | 2 |
| 2010 | Googling the internet: profiling internet endpoints via the world wide web
Ionut Trestian, Supranamaya Ranjan, Aleksandar Kuzmanovic, Antonio Nucci |
IEEE/ACM Trans. Netw. | 3 |
| 2009 | Measuring serendipity: connecting people, locations and interests in a mobile 3G networkabstractCharacterizing the relationship that exists between people's application interests and mobility properties is the core question relevant for location-based services, in particular those that facilitate serendipitous discovery of people, businesses and objects. In this paper, we apply rule mining and spectral clustering to study this relationship for a population of over 280,000 users of a 3G mobile network in a large metropolitan area. Our analysis reveals that (i) People's movement patterns are correlated with the applications they access, e.g., stationary users and those who move more often and visit more locations tend to access different applications. (ii) Location affects the applications accessed by users, i.e., at certain locations, users are more likely to evince interest in a particular class of applications than others irrespective of the time of day. (iii) Finally, the number of serendipitous meetings between users of similar cyber interest is larger in regions with higher density of hotspots. Our analysis demonstrates how cellular network providers and location-based services can benefit from knowledge of the inter-play between users and their locations and interests. Ionut Trestian, Supranamaya Ranjan, Aleksandar Kuzmanovic, Antonio Nucci |
Internet Measurement Conference | 3 |
| 2009 | Supporting application network flows with multiple QoS constraintsabstractThere is a growing need to support real-time applications over the Internet. Real-time interactive applications often have multiple quality-of-service (QoS) requirements which are application specific. Traditional provisioning of QoS in the Internet through IP routing - Intserv or Diffserv - faces many technical challenges, and is also deterred by the huge deployment issues. As an alternative, application providers often build their own application-specific overlay networks to meet their QoS requirements. In this paper, we present a unified framework which can serve diverse applications with multiple QoS constraints. Our scalable flow route management architecture, called MCQoS, employs a hybrid approach using a path vector protocol to disseminate aggregated path information combined with on-demand path discovery to find paths that match the diverse QoS requirements. It uses a distributed algorithm to dynamically adapt to an alternate path when the current path fails to satisfy the required QoS constraints. We do large-scale simulation and analysis to show that our approach is both efficient and scalable, and that it substantially outperforms the state of the art protocols in accuracy. Our simulation results show that MCQoS can reduce the false negative percentage to less than 1% compared with 5-10% in other approaches, and eliminates false positives, whereas other schemes have false positive rates of 10-20% with minimal increase in protocol overhead. Finally, we implemented and deployed our system on the Planetlab testbed for evaluation in a real network environment. Amit Mondal, Puneet Sharma 0001, Sujata Banerjee, Aleksandar Kuzmanovic |
IWQoS | 4 |
| 2009 | Drafting behind Akamai: inferring network conditions based on CDN redirections
Ao-Jan Su, David R. Choffnes, Aleksandar Kuzmanovic, Fabián E. Bustamante |
IEEE/ACM Trans. Netw. | 3 |
| 2008 | Relative Network Positioning via CDN RedirectionsabstractMany large-scale distributed systems can benefit from a service that allows them to select among alternative nodes based on their relative network positions. A variety of approaches propose new measurement infrastructures that attempt to scale this service to large numbers of nodes by reducing the amount of direct measurements to end hosts. In this paper, we introduce a new approach to relative network positioning that eliminates direct probing by leveraging pre-existing infrastructure. Specifically, we exploit the dynamic association of nodes with replica servers from large content distribution networks (CDNs) to determine relative position information - we call this approach CDN-based relative network positioning (CRP). We demonstrate how CRP can support two common examples of location information used by distributed applications: server selection and dynamic node clustering. After describing CRP in detail, we present results from an extensive wide-area evaluation that demonstrates its effectiveness. Ao-Jan Su, David R. Choffnes, Fabián E. Bustamante, Aleksandar Kuzmanovic |
ICDCS | 4 |
| 2008 | Monitoring persistently congested Internet linksabstractMeasurement tools that can accurately locate and monitor congested Internet links would significantly help us understand how the Internet operates. However, developing such tools is challenging, especially when our concerned target is congestion on the core Internet links rather than that on the relatively easily measured access links. Congestion on core links — and persistent congestion in particular — can reveal systematic problems such as routing pathologies, poorly-engineered network policies, or non-cooperative inter-AS relationships. In this paper, we present Pong, a novel tool capable of accurately locating and monitoring a subset of non-access Internet links that exhibit persistent congestion over longer time scales. Pong takes advantage of the persistently congested link property to overcome the long-lasting challenges common for delay-based inference tools. In addition, it exploits the same property to (i) infer otherwise unknown underlying path conditions, (ii) determine appropriate queuing delay thresholds to reveal congestion, (iii) achieve high accuracy with low probing rate, and (iv) detect moments of its own inaccuracy. Finally, Pong can quantify measurement results’ accuracy comprehensively, allowing us to further select vantage points that maximize the observability of the underlying congestion. Leiwen Deng, Aleksandar Kuzmanovic |
ICNP | 2 |
| 2008 | Thinning akamaiabstractGlobal-scale Content Distribution Networks (CDNs), such as Akamai, distribute thousands of servers worldwide providing a highly reliable service to their customers. Not only has reliability been one of the main design goals for such systems - they are engineered to operate under severe and constantly changing number of server failures occurring at all times. Consequently, in addition to being resilient to component or network outages, CDNs are inherently considered resilient to denial-of-service (DoS) attacks as well.In this paper, we focus on Akamai's (audio and video) streaming service and demonstrate that the current system design is highly vulnerable to intentional service degradations. We show that (i) the discrepancy among streaming flows' lifetimes and DNS redirection timescales, (ii) the lack of isolation among customers and services, (e.g., video on demand vs. live streaming), (iii) a highly transparent system design, (iv) a strong bias in the stream popularity, and (v) minimal clients' tolerance for low-quality viewing experiences, are all factors that make intentional service degradations highly feasible. We demonstrate that it is possible to impact arbitrary customers' streams in arbitrary network regions: not only by targeting appropriate points at the streaming network's edge, but by effectively provoking resource bottlenecks at a much higher level in Akamai's multicast hierarchy. We provide countermeasures to help avoid such vulnerabilities and discuss how lessons learned from this research could be applied to improve DoS-resiliency of large-scale distributed and networked systems in general. Ao-Jan Su, Aleksandar Kuzmanovic |
Internet Measurement Conference | 2 |
| 2008 | Unconstrained endpoint profiling (googling the internet)abstractUnderstanding Internet access trends at a global scale, i.e., what do people do on the Internet, is a challenging problem that is typically addressed by analyzing network traces. However, obtaining such traces presents its own set of challenges owing to either privacy concerns or to other operational difficulties. The key hypothesis of our work here is that most of the information needed to profile the Internet endpoints is already available around us - on the web. Ionut Trestian, Supranamaya Ranjan, Aleksandar Kuzmanovic, Antonio Nucci |
SIGCOMM | 3 |
| 2008 | Pollution attacks and defenses for Internet caching systems
Leiwen Deng, Yan Gao 0003, Yan Chen 0004, Aleksandar Kuzmanovic |
Comput. Networks | 4 |
| 2007 | A Poisoning-Resilient TCP StackabstractWe treat the problem of large-scale TCP poisoning: an attacker, who is able to monitor TCP packet headers in the network, can deny service to all flows traversing the monitoring point simply by injecting a single spoofed data or control packet into each of the flows. One of the entities responsible for this severe vulnerability is certainly the TCP protocol itself: it behaves as a "dummy" state machine that can more-than-easily become desynchronized by an attacker. In this paper, we explore ways for upgrading TCP endpoints into viable DoS-resilient protocol entities, capable of mitigating large-scale poisoning attacks. We show, by means of analytical modeling, simulations, and Internet experiments, how small upgrades implemented by the endpoints can dramatically improve resilience to attacks. The key mechanisms unique to our approach are (i) deferred protocol reaction, used to accurately detect poisoning attacks; (ii) forward nonces, applied to distinguish among different traffic sources during the attack; and (iii) self-clocking-based correlation, utilized for successfully detecting legitimate packet streams. Our solution solely relies on the protocol design, it is incrementally deployable, and TCP friendly. Amit Mondal, Aleksandar Kuzmanovic |
ICNP | 2 |
| 2007 | When TCP Friendliness Becomes HarmfulabstractShort TCP flows may suffer significant response-time performance degradations during network congestion. Unfortunately, this creates an incentive for misbehavior by clients ofinteractiveapplications(e.g., gaming, telnet, Web): to send "dummy" packets into the network at aTCP-fairrate even when they have no data to send, thus improving their performance in moments when they do have data to send. Even though no "law" is violated in this way, a large-scale deployment of such an approach has the potential to seriously jeopardize one of the core Internet's principles -statisticalmultiplexing. We quantify, by means of analytical modeling and simulation, gains achievable by the above misbehavior. Further, we explore techniques that both misbehaving and regular clients can apply to optimize their performance. Our research indicates that easy-to-implement application-level techniques are capable of dramatically reducing incentives for conducting the above transgressions, still without compromising the idea of statistical multiplexing. Amit Mondal, Aleksandar Kuzmanovic |
INFOCOM | 2 |
| 2007 | Pong: diagnosing spatio-temporal internet congestion propertiesabstractThe ability to accurately detect congestion events in the Internet and reveal their spatial (i.e., where they happen?) and temporal (i.e., how frequently they occur and how long they last?) properties would significantly improve our understanding of how the Internet operates. In this paper we present Pong, a novel measurement tool capable of effectively diagnosing congestion events over short (e.g., ~100ms or longer) time-scales, and simultaneously locating congested points within a single hop on an end-to-end path at the granularity of a single link. Leiwen Deng, Aleksandar Kuzmanovic |
SIGMETRICS | 2 |
| 2007 | Receiver-centric congestion control with a misbehaving receiver: Vulnerabilities and end-point solutions
Aleksandar Kuzmanovic, Edward W. Knightly |
Comput. Networks | 1 |
| 2006 | Internet Cache Pollution Attacks and CountermeasuresabstractProxy caching servers are widely deployed in today's Internet. While cooperation among proxy caches can significantly improve a network's resilience to denial-of-service (DoS) attacks, lack of cooperation can transform such servers into viable DoS targets. In this paper, we investigate a class of pollution attacks that aim to degrade a proxy's caching capabilities, either by ruining the cache file locality, or by inducing false file locality. Using simulations, we propose and evaluate the effects of pollution attacks both in Web and peer-to- peer (p2p) scenarios, and reveal dramatic variability in resilience to pollution among several cache replacement policies. We develop efficient methods to detect both false-locality and locality-disruption attacks, as well as a combination of the two. To achieve high scalability for a large number of clients/requests without sacrificing the detection accuracy, we leverage streaming computation techniques, Le., bloom filters. Evaluation results from large-scale simulations show that these mechanisms are effective and efficient in detecting and mitigating such attacks. Furthermore, a squid-based implementation demonstrates that our protection mechanism forces the attacker to launch extremely large distributed attacks in order to succeed. Yan Gao 0003, Leiwen Deng, Aleksandar Kuzmanovic, Yan Chen 0004 |
ICNP | 3 |
| 2006 | Drafting behind Akamai (travelocity-based detouring)abstractTo enhance web browsing experiences, content distribution networks (CDNs) move web content "closer" to clients by caching copies of web objects on thousands of servers worldwide. Additionally, to minimize client download times, such systems perform extensive network and server measurements, and use them to redirect clients to different servers over short time scales. In this paper, we explore techniques for inferring and exploiting network measurements performed by the largest CDN, Akamai; our objective is to locate and utilize quality Internet paths without performing extensive path probing or monitoring.Our contributions are threefold. First, we conduct a broad measurement study of Akamai's CDN. We probe Akamai's network from 140 PlanetLab vantage points for two months. We find that Akamai redirection times, while slightly higher than advertised, are sufficiently low to be useful for network control. Second, we empirically show that Akamai redirections overwhelmingly correlate with network latencies on the paths between clients and the Akamai servers. Finally, we illustrate how large-scale overlay networks can exploit Akamai redirections to identify the best detouring nodes for one-hop source routing. Our research shows that in more than 50% of investigated scenarios, it is better to route through the nodes "recommended" by Akamai, than to use the direct paths. Because this is not the case for the rest of the scenarios, we develop lowoverhead pruning algorithms that avoid Akamai-driven paths when they are not beneficial. Ao-Jan Su, David R. Choffnes, Aleksandar Kuzmanovic, Fabián E. Bustamante |
SIGCOMM | 3 |
| 2006 | Low-rate TCP-targeted denial of service attacks and counter strategies
Aleksandar Kuzmanovic, Edward W. Knightly |
IEEE/ACM Trans. Netw. | 1 |
| 2006 | TCP-LP: low-priority service via end-point congestion control
Aleksandar Kuzmanovic, Edward W. Knightly |
IEEE/ACM Trans. Netw. | 1 |
| 2005 | The power of explicit congestion notificationabstractDespite the fact that Explicit Congestion Notification (ECN) demonstrated a clear potential to substantially improve network performance, recent network measurements reveal an extremely poor usage of this option in today's Internet. In this paper, we analyze the roots of this phenomenon and develop a set of novel incentives to encourage network providers, end-hosts, and web servers to apply ECN.Initially, we examine a fundamental drawback of the current ECN specification, and demonstrate that the absence of ECN indications in TCP control packets can dramatically hinder system performance. While security reasons primarily prevent the usage of ECN bits in TCP SYN packets, we show that applying ECN to TCP SYN ACK packets can significantly improve system performance without introducing any novel security or stability side-effects. Our network experiments on a cluster of web servers show a dramatic performance improvement over the existing ECN specification: throughput increases by more than 40%, while the average web response-time simultaneously decreases by nearly an order of magnitude.In light of the above finding, using large-scale simulations, modeling, and network experiments, we re-investigate the relevance of ECN, and provide a set of practical recommendations and insights: (i) ECN systematically improves the performance of all investigated AQM schemes; contrary to common belief, this particularly holds for RED. (ii) The impact of ECN is highest for web-only traffic mixes such that even a generic AQM algorithm with ECN support outperforms all non-ECN-enabled AQM schemes that we investigated. (iii) Primarily due to moderate queuing levels, the superiority of ECN over other AQM mechanisms largely holds for high-speed backbone routers, even in more general traffic scenarios. (iv) End-hosts that apply ECN can exercise the above performance benefits instantly, without waiting for the entire Internet community to support the option. Aleksandar Kuzmanovic |
SIGCOMM | 1 |
| 2005 | Denial-of-service resilience in peer-to-peer file sharing systemsabstractPeer-to-peer (p2p) file sharing systems are characterized by highly replicated content distributed among nodes with enormous aggregate resources for storage and communication. These properties alone are not sufficient, however, to render p2p networks immune to denial-of-service (DoS) attack. In this paper, we study, by means of analytical modeling and simulation, the resilience of p2p file sharing systems against DoS attacks, in which malicious nodes respond to queries with erroneous responses. We consider the file-targeted attacks in current use in the Internet, and we introduce a new class of p2p-network-targeted attacks.In file-targeted attacks, the attacker puts a large number of corrupted versions of a single file on the network. We demonstrate that the effectiveness of these attacks is highly dependent on the clients' behavior. For the attacks to succeed over the long term, clients must be unwilling to share files, slow in removing corrupted files from their machines, and quick to give up downloading when the system is under attack.In network-targeted attacks, attackers respond to queries for any file with erroneous information. Our results indicate that these attacks are highly scalable: increasing the number of malicious nodes yields a hyperexponential decrease in system goodput, and a moderate number of attackers suffices to cause a near-collapse of the entire system. The key factors inducing this vulnerability are (i) hierarchical topologies with misbehaving "supernodes," (ii) high path-length networks in which attackers have increased opportunity to falsify control information, and (iii) power-law networks in which attackers insert themselves into high-degree points in the graph.Finally, we consider the effects of client counter-strategies such as randomized reply selection, redundant and parallel download, and reputation systems. Some counter-strategies (e.g., randomized reply selection) provide considerable immunity to attack (reducing the scaling from hyperexponential to linear), yet significantly hurt performance in the absence of an attack. Other counter-strategies yield little benefit (or penalty). In particular, reputation systems show little impact unless they operate with near perfection. Dan Dumitriu, Edward W. Knightly, Aleksandar Kuzmanovic, Ion Stoica, Willy Zwaenepoel |
SIGMETRICS | 3 |
| 2004 | A Performance vs. Trust Perspective in the Design of End-Point Congestion Control ProtocolsabstractReceiver-driven TCP protocols delegate key congestion control functions to receivers. Their goal is to exploit information available only at receivers in order to improve latency and throughput in diverse scenarios ranging from wireless access links to wireline and wireless Web browsing. Unfortunately, in contrast to today's sender-driven protocols, receiver-driven congestion control introduces an incentive for misbehavior. Namely, the primary beneficiary of a flow (the receiver of data) has both the means and incentive to manipulate the congestion control algorithm in order to obtain higher throughput or reduced latency. We study the deployability of receiver-driven TCP in environments with untrusted receivers which may tamper with the congestion control algorithm for their own benefit. Using analytical modeling and extensive simulation experiments, we show that deployment of receiver-driven TCP must strike a balance between enforcement mechanisms, which can limit performance, and complete trust of end-points, which results in vulnerability to cheaters and even DoS attackers. Aleksandar Kuzmanovic, Edward W. Knightly |
ICNP | 1 |
| 2003 | TCP-LP: A Distributed Algorithm for Low Priority Data TransferabstractService prioritization among different traffic classes is an important goal for the future Internet. Conventional approaches to solving this problem consider the existing best-effort class as the low-priority class, and attempt to develop mechanisms that provide "better-than-best-effort" service. In this paper, we explore the opposite approach, and devise a new distributed algorithm to realize a low-priority service (as compared to the existing best effort) from the network endpoints. To this end, we develop TCP Low Priority (TCP-LP), a distributed algorithm whose goal is to utilize only the excess network bandwidth as compared to the "fair share" of bandwidth as targeted by TCP. The key mechanisms unique to TCP-LP congestion control are the use of one-way packet delays for congestion indications and a TCP-transparent congestion avoidance policy. Our simulation results show that: (1) TCP-LP is largely non-intrusive to TCP traffic; (2) both single and aggregate TCP-LP flows are able to successfully utilize excess network bandwidth; moreover, multiple TCP-LP flows share excess bandwidth fairly; (3) substantial amounts of excess bandwidth are available to low-priority class, even in the presence of "greedy" TCP flows; (4) the response times of web connections in the best-effort class decrease by up to 90% when long-lived bulk data transfers use TCP-LP rather than TCP. Aleksandar Kuzmanovic, Edward W. Knightly |
INFOCOM | 1 |
| 2003 | Low-rate TCP-targeted denial of service attacks: the shrew vs. the mice and elephantsabstractDenial of Service attacks are presenting an increasing threat to the global inter-networking infrastructure. While TCP's congestion control algorithm is highly robust to diverse network conditions, its implicit assumption of end-system cooperation results in a well-known vulnerability to attack by high-rate non-responsive flows. In this paper, we investigate a class of low-rate denial of service attacks which, unlike high-rate attacks, are difficult for routers and counter-DoS mechanisms to detect. Using a combination of analytical modeling, simulations, and Internet experiments, we show that maliciously chosen low-rate DoS traffic patterns that exploit TCP's retransmission time-out mechanism can throttle TCP flows to a small fraction of their ideal rate while eluding detection. Moreover, as such attacks exploit protocol homogeneity, we study fundamental limits of the ability of a class of randomized time-out mechanisms to thwart such low-rate DoS attacks. Aleksandar Kuzmanovic, Edward W. Knightly |
SIGCOMM | 1 |
| 2003 | Measurement-Based Characterization and Classification of QoS-Enhanced SystemsabstractQuality-of-service mechanisms and differentiated service classes are increasingly available in networks and Web servers. While network and Web server clients can assess their service by measuring basic performance parameters such as packet loss and delay, such measurements do not expose the system's core QoS functionality such as multiclass service discipline. In this paper, we develop a framework and methodology for enabling network and Web server clients to assess system's multiclass mechanisms and parameters. Using hypothesis testing, maximum likelihood estimation, and empirical arrival and service rates measured across multiple time scales, we devise techniques for clients to: 1) determine the most likely service discipline among earliest deadline first, class-based weighted fair queuing, and strict priority; 2) estimate the system's parameters with high confidence; and (3) detect and parameterize non work-conserving elements such as rate limiters. We describe the important role of time scales in such a framework and identify the conditions necessary for obtaining accurate and high confidence inferences. Aleksandar Kuzmanovic, Edward W. Knightly |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2001 | Measuring Service in Multi-Class NetworksabstractQuality of service mechanisms and differentiated service classes are increasingly available in networks and servers. While network clients can assess their service by measuring basic performance parameters such as packet loss and delay, such measurements do not expose the network's core QoS functionality. We develop a framework and methodology for enabling network clients to assess a system's multi-class mechanisms and parameters. Using hypothesis testing, maximum likelihood estimation, and empirical arrival and service rates measured across multiple time scales, we devise techniques for clients to (1) determine the most likely service discipline among EDF, WFQ, and SP, (2) estimate the server's parameters with high confidence, and (3) detect and parameterize non-work-conserving elements such as rate limiters. We describe the important role of time scales in such a framework and identify the conditions necessary for obtaining accurate and high confidence inferences. Aleksandar Kuzmanovic, Edward W. Knightly |
INFOCOM | 1 |