Songqing Chen

dblp:47/3276 · DBLP profile ↗
← Back
135ranked-venue papers
16as first author
33since 2021 · last 2026
0000-0003-4650-7125ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 46 · 6 first-author · 12 since 2021Systems, architecture and hardware · 31 · 5 first-author · 7 since 2021Security and privacy · 24 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 10 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3Theory of computation · 1
YearPublicationVenuePosition
2026 Triton-Sanitizer: A Fast and Device-Agnostic Memory Sanitizer for Triton with Rich Diagnostic Context
abstract
Memory access errors remain one of the most pervasive bugs in GPU programming. Existing GPU sanitizers such as compute-sanitizer detect memory access errors by instrumenting every memory instruction in low-level IRs or binaries, which imposes high overhead and provides minimal memory access error diagnostic context for fixing problems. We present Triton-Sanitizer, the first device-agnostic memory sanitizer designed for Triton, a domain-specific language for developing portable, efficient GPU kernels for deep learning workloads. Triton-Sanitizer leverages Triton's tile-oriented semantics to construct symbolic expressions for memory addresses and masks, verifies them with an SMT solver, and selectively falls back to eager simulation for indirect accesses. This hybrid analysis enables precise detection of memory access errors without false positives while avoiding the cost of per-access instrumentation. Beyond detection, Triton-Sanitizer generates rich diagnostic reports that attribute violations to the tensors nearest to the violated addresses, track the complete call path, and expose the symbolic operations responsible for incorrect addresses. Evaluated on seven widely used open-source repositories of Triton kernels, Triton-Sanitizer uncovered 24 previously unknown memory access errors, of which 8 have already been fixed and upstreamed by us. Compared to compute-sanitizer, Triton-Sanitizer achieves speedups ranging from 1.07× to 14.66×, with an average improvement of 1.62×, demonstrating its ability to enhance performance, precision, and usability in memory access error detection.
Hao Wu 0077, Qidong Zhao, Songqing Chen, Yueming Hao, Tony C. W. Liu, Adnan Aziz, Keren Zhou 0001
ASPLOS (2)3
2026 Exploring Collaborative Immersive Visualization & Analytics for High-Dimensional Scientific Data through Domain Expert Perspectives
abstract
Cross-disciplinary teams increasingly work with high-dimensional scientific datasets, yet fragmented toolchains and limited support for shared exploration hinder collaboration. Prior immersive visualization & analytics research has emphasized individual interaction, leaving open how multi-user collaboration can be supported at scale. To fill this critical gap, we conduct semi-structured interviews with 20 domain experts from diverse academic, government, and industry backgrounds. Using deductive–inductive hybrid thematic analysis, we identify four collaboration-focused themes: workflow challenges, adoption perceptions, prospective features, and anticipated usability and ethical risks. These findings show how current ecosystems disrupt coordination and shared understanding, while highlighting opportunities for effective multi-user engagement. Our study contributes empirical insights into collaboration practices for high-dimensional scientific data visualization & analysis, offering design implications to enhance coordination, mutual awareness, and equitable participation in next-generation collaborative immersive platforms. These contributions point toward future environments enabling distributed, cross-device teamwork on high-dimensional scientific data.
Fahim Arsad Nafis, Jie Li 0064, Simon Su, Songqing Chen, Bo Han 0001
CHI4
2026 Physical Self-Supervised Learning: IMU Sensing without Manual Labels
abstract
Deep neural networks have become a promising approach for IMU-based sensing, but their scalability is fundamentally limited by costly labeled data and poor robustness to heterogeneous devices, placements, and users. Existing unsupervised and self-supervised methods reduce but do not remove this dependence, still requiring labeled data for domain adaptation and largely ignoring known physical structure. We propose physical self-supervised learning, an autoencoder-style paradigm for label-free IMU sensing. We replace the conventional neural decoder with an auto-adaptive physics decoder—a learnable family of kinematic equations that enforces explicit physical structure while adapting across environments—and adopt a hybrid two-stage IMU encoder with reconstruction in a structured latent space to mitigate sensor noise. Our framework further introduces probabilistic frequency-spatial constraints to disentangle sensor and object motion, a multi-view kinematic tree to exploit sparse physical self-supervised signals, and an uncertainty-aware formulation to handle the inherent ambiguity of IMU inference. Evaluated on inertial tracking and full-body motion capture over public datasets and realistic deployments, physical self-supervised learning reduces errors by up to 5× for tracking and 4× for motion capture in challenging generalization scenarios, consistently outperforming state-of-the-art supervised and self-supervised baselines without any labels.
Yuyang Leng, Renyuan Liu, Shaohan Hu, Peijun Zhao, Chun-Fu Chen 0001, Songqing Chen, Shuochao Yao
MobiSys6
2025 Hello, GenAI? Dissecting Human to Generative AI Calling
abstract
The rise of generative artificial intelligence (GenAI), powered by large language models, has led to the emergence of real-time, voice-based conversational applications that enable dynamic, multi-modal interactions for everyday tasks such as checking the weather or planning a trip. These human-to-GenAI calling applications blend speech processing, generative intelligence, and real-time communication, presenting new challenges in latency optimization, network infrastructure design, and resilience under load. Despite their growing popularity, little is known about the operational characteristics and performance of these applications. This paper conducts an empirical measurement of six human-to-GenAI calling applications from Google, Meta, Microsoft, and OpenAI, focusing on their input/output modalities, network behavior, latency metrics, and robustness. Our findings reveal key design choices and performance bottlenecks in these emerging applications. For example, the conversational latency often reaches several seconds, far exceeding the typical sub-second delays of human-to-human voice communication and potentially impairing interactivity. Moreover, voice-based GenAI traffic is inherently asymmetric: the uplink, carrying real-time human speech, benefits from streaming-based transmission, while the typically large downlink GenAI responses are better served through batch-based delivery.
Ruizhi Cheng, Surendra Pathak, Guowu Xie, Matteo Varvello, Songqing Chen, Bo Han 0001
IMC5
2025 From WebGL to WebGPU: A Reality Check of Browser-Based GPU Acceleration
abstract
With the rising demand for cost-effective and privacy-preserving deep learning and visualization services, service providers are increasingly turning to in-browser solutions. General-purpose computations, leveraging the graphics processing unit (GPU), are foundational to executing algorithms that power these services. Web graphics library (WebGL) is a widely adopted GPU-access application programming interface (API) designed for multidimensional rendering in browsers, while WebGPU is a newer API developed with compute-specific capabilities. Although WebGPU is a promising standard, its performance has not been systematically evaluated for general-purpose computation. This paper investigates WebGPU and WebGL for accelerating client-side computation in web browsers. By benchmarking key computational GPU kernels for 16 PolyBench and 2 CHStone functions, we measure the performance of WebGPU and WebGL across varying input sizes and algorithmic complexities. Our results show that: 1) both WebGPU and WebGL exhibit poorer performance than central processing unit (CPU)-based execution for small input data due to setup and CPU-GPU synchronization overheads, but they outperform CPU execution as the input data size increases; 2) WebGL performs better than WebGPU for small inputs, except for CPU-driven loop functions, due to its lower initial setup overhead; and 3) WebGPU outperforms WebGL for large inputs through optimized GPU thread utilization and achieves better performance for loop-driven algorithms across all input sizes by minimizing CPU-GPU data exchange. Overall, our results indicate that WebGPU is a competitive option for enhancing the execution performance of large-scale web applications.
Sthitadhi Sengupta, Nan Wu 0012, Matteo Varvello, Krish Jana, Songqing Chen, Bo Han 0001
IMC5
2025 The Decentralization Dilemma: Performance Trade-Offs in IPFS and Breakpoints
abstract
Web 3.0 is redefining the current Web (Web 2.0) with a focus on data and governance decentralization. The InterPlanetary File System (IPFS) exemplifies this shift. However, it faces a trade-off between decentralization and performance: prior studies have shown IPFS's performance degradations but fail to diagnose root causes or deliver actionable fixes.
Ruizhe Shi, Yuqi Fu, Ruizhi Cheng, Bo Han 0001, Yue Cheng 0001, Songqing Chen
IMC6
2025 NeRFCompressor: Enhancing Dynamic Scene Representation for Efficient 6-DoF Object Transportation
abstract
3D scene modeling is essential for immersive experiences in Virtual, Augmented, and Mixed Reality (VR/AR/MR) applications. Neural Radiance Fields (NeRF) have emerged as a strong alternative to traditional representations such as meshes and point clouds for 6-DoF rendering. However, maintaining high visual quality while enabling efficient transmission in dynamic environments remains a significant challenge. In this paper, we propose NeRFCompressor, a novel compression framework for dynamic scene representation using NeRF-like models. Building on tensor decomposition-based 3D reconstruction, NeRFCompressor improves transmission efficiency by leveraging existing video codecs to exploit both intra-scene and inter-scene redundancies. It maintains high QoE with minimal degradation in reconstruction quality. Experiments show that NeRFCompressor outperforms state-of-the-art methods in compressing both static and dynamic scene representations.
Jin Zhou 0006, Mufeng Zhu, Yao Liu 0001, Songqing Chen
MMSP4
2025 PIPE: Privacy-preserving 6DoF Pose Estimation for Immersive Applications
abstract
Image-based mapping and localization offer six degrees of freedom (6DoF) pose estimation for immersive applications. This is achieved by matching, on a server, 2D visual features extracted from a mobile device's camera view and 3D features stored in a map. While effective, this process may lead to privacy breaches (e.g., exposure of sensitive information captured by camera views). To tackle this crucial issue, we present PIPE, a first-of-its-kind Privacy-preserving Image-based 6DoF Pose Estimation system. The design of PIPE is motivated by our key observation that uploading only a small amount of features extracted from camera views for pose estimation could reduce privacy leakage. However, trade-offs exist between privacy preservation, system utility (i.e., pose estimation accuracy), and system performance (e.g., end-to-end latency). To balance the trade-offs, PIPE deliberately explores the feature-detection space to reduce computation latency, designs an efficient feature ranking method by judiciously utilizing map data, and optimizes feature selection by jointly considering the features' ranking and spatial distribution to improve pose estimation accuracy. Moreover, we construct a learning-based metric to quantify the extent of privacy leakage in images. Our extensive performance evaluation reveals that PIPE can effectively preserve privacy and reduce end-to-end latency by up to 22.6%, while marginally affecting pose estimation accuracy (e.g., as low as 2.7%).
Nan Wu 0012, Ruizhi Cheng, Songqing Chen, Bo Han 0001
SenSys3
2025 Centralization in the Decentralized Web: Challenges and Opportunities in IPFS Data Management
abstract
The InterPlanetary File System (IPFS) is a pioneering effort for Web 3.0, well-known for its decentralized infrastructure. However, some recent studies have shown that IPFS exhibits a high degree of centralization and has integrated centralized components for improved performance. While this change contradicts the core decentralized ethos of IPFS and introduces risks of hurting the data replication level and thus availability, it also opens some opportunities for better data management and cost savings through deduplication.
Ruizhe Shi, Ruizhi Cheng, Yuqi Fu, Bo Han 0001, Yue Cheng 0001, Songqing Chen
WWW6
2025 Dissecting User Experience of Social Virtual Reality: A Tale of Five Platforms
abstract
Social virtual reality (VR) has the potential to replace conventional online social media by offering quasi-real-world social experiences. As such, it has been extensively examined by the research community. However, existing studies fall short of providing a comprehensive understanding of how different aspects of social VR platforms interact to affect user experience. Motivated by this limitation, we conduct a user study with Oculus Quest 2 headsets and dissect the user experience on five social VR platforms. We evenly and randomly divide 42 participants into short-term (spending 10-30 minutes/platform) and long-term (spending at least 120 minutes/platform) groups. Besides employing surveys and interviews, we measure the frame rate and resolution of these platforms and explore how various factors interplay to influence the user experience of social VR. Our findings reveal that the frame rate, resolution, and interactive events of social VR platforms have a more significant impact on the experience of long-term users compared to short-term users. The scalability limitations of these platforms, as evidenced by decreased frame rates with the increasing number of concurrent users, result in an increased prevalence of motion sickness among long-term users, negatively impacting their overall experience. Moreover, the absence of highly interactive events also deteriorates their overall experience, and the low resolution combined with the lack of interactive events further decreases their sense of social presence. Additionally, our study demonstrates several common limitations negatively affecting the experience of both long-term and short-term users. For example, the harassment prevention mechanisms on all five platforms are inadequate, and being harassed has a detrimental effect on users' overall experience and sense of social presence. The avatar embodiment of investigated platforms has limited contribution to users' sense of social presence, mainly due to the lack of realism and full-body tracking. Our findings call for more research in scalability support, motion sickness relief, interactive event design, harassment prevention, and avatar development for improving social VR platforms in the future.
Ruizhi Cheng, Jie Li 0064, Songqing Chen, Bo Han 0001
Proc. ACM Hum. Comput. Interact.3
2024 A First Look at Immersive Telepresence on Apple Vision Pro
abstract
Due to the widespread adoption of "work-from-home" policies, videoconferencing applications (e.g., Zoom) have become indispensable for remote communication. However, they often lack immersiveness, leading to "Zoom fatigue" and degrading communication efficiency. The recent debut of Apple Vision Pro, a mobile headset that supports "spatial personas", offers an immersive telepresence experience. In this paper, we conduct a first-of-its-kind in-depth and empirical study to analyze the performance of immersive telepresence with FaceTime, Webex, Teams, and Zoom on Vision Pro. We find that only FaceTime provides a truly immersive experience with spatial personas, whereas others still operate 2D personas. Our measurements reveal that (1) FaceTime delivers semantic data to optimize bandwidth consumption, which is even lower than that of 2D personas for other applications, and (2) it employs visibility-aware optimizations to reduce rendering overhead. However, the scalability of FaceTime remains limited, with a simple server-allocation strategy that potentially leads to high network delay for users.
Ruizhi Cheng, Nan Wu 0012, Matteo Varvello, Eugene Chai, Songqing Chen, Bo Han 0001
IMC5
2024 Dynamic 6-DoF Volumetric Video Generation: Software Toolkit and Dataset
abstract
Volumetric video streaming has become increasingly popular in recent years due to its support of 6 degrees-of-freedom (6-DoF) exploration. There is, however, a shortage of dynamic 6-DoF content suitable for comparing the performance among heterogeneous volumetric video representations. This paper introduces a software toolkit for creating both a dataset of dynamic 6-DoF content in point clouds and a dataset for training and testing neural-based representations such as neural radiance fields (NeRF). Starting with freely available 3D assets online, our software toolkit uses the Blender Python API to generate training and testing datasets for neural-based dynamic volumetric model training. The created datasets are compliant with existing neural-based model training and rendering frameworks. The software can also construct point cloud sequences derived from synthetic dynamic 3D meshes. This further facilitates comparing point clouds and neural-based methods for volumetric video representation. We release the software toolkit along with a rich set of sequence datasets generated in compliance with the permissions granted by the original 3D asset creators. With our toolkit and dataset, we aim to facilitate research from the multimedia systems community to support practical volumetric streaming. Our software toolkit and dataset are available at: https://6-dof-dynamic-content-software.github.io/.
Mufeng Zhu, Yuan-Chun Sun, Na Li 0032, Jin Zhou 0006, Songqing Chen, Cheng-Hsin Hsu, Yao Liu 0001
MMSP5
2024 MetaFL: Privacy-preserving User Authentication in Virtual Reality with Federated Learning
abstract
The increasing popularity of virtual reality (VR) has stressed the importance of authenticating VR users while preserving their privacy. Behavioral biometrics, owing to their robustness and ease of collection, compared to traditional modes such as passwords, have become a favored authentication choice. While current approaches that utilize behavioral biometrics to train classifiers for authentication yield promising accuracy, they cause privacy breaches by sharing sensitive data with a server to train a central model. In this paper, we present MetaFL, a first-of-its-kind privacy-preserving VR authentication framework that leverages federated learning (FL) on multi-modal motion data. The design of MetaFL is motivated by our key insight that various modalities of motion data uniquely affect authentication performance for individual users and among different users. It is attributed to the fundamental challenge of privacy-preserving user authentication: users can access only their own data with limited global knowledge. To tackle this issue, MetaFL judiciously selects the most suitable modalities for each user, which is decomposed into within-user ordering and between-user selection to eliminate the complex interplay between various conflicting factors. Moreover, we develop a personalized strategy to initialize FL models, further improving authentication accuracy. Our extensive performance evaluation on six public datasets shows that MetaFL outperforms state-of-the-art FL-based models (e.g., 17--28% higher authentication accuracy), and its accuracy gap with the non-privacy-preserving central model is small (i.e., only <2%).
Ruizhi Cheng, Yuetong Wu, Ashish Kundu, Hugo Latapie, Myungjin Lee, Songqing Chen, Bo Han 0001
SenSys6
2024 ALPS: An Adaptive Learning, Priority OS Scheduler for Serverless Functions
Yuqi Fu, Ruizhe Shi, Songqing Chen, Yue Cheng 0001
USENIX ATC4
2024 Understanding Online Education in Metaverse: Systems and User Experience Perspectives
abstract
Thanks to recent advances in immersive technologies, virtual reality (VR) is becoming increasingly popular in online education, particularly in light of the rise of the Metaverse. However, there is currently no in-depth investigation of the user experience of VR-based online education and the comparison of it with video-conferencing-based counterparts. To fill these critical gaps, we conduct multiple sessions of two courses in a university with 10 and 37 participants on Mozilla Hubs (Hubs for short), a social VR platform that is deemed as one of the early prototypes of the Metaverse, and let them compare the classroom experience on Hubs with Zoom, a popular video-conferencing application. In addition to employing traditional analytical methods to understand user experience, we benefit from an end-to-end measurement study of Hubs to corroborate our findings and systematically detect its performance bottlenecks. Our study leads to the following key observations. First, the scalability issue of Hubs makes it inadequate for accommodating large courses. Second, compared to Zoom, Hubs can offer a better sense of place presence and social presence to students, thanks to its avatar-based interactions and the hand and head tracking enabled by headsets. Third, even though VR headsets help students concentrate in class, effectively utilizing learning tools through them remains a challenge.
Ruizhi Cheng, Erdem Murat, Lap-Fai Yu, Songqing Chen, Bo Han 0001
VR4
2024 HardenVR: Harassment Detection in Social Virtual Reality
abstract
Social Virtual Reality (VR) is regarded as one of the most popular VR applications since it transcends geographical barriers, allowing users to interact in simulated environments for various purposes. Despite its promising prospects, there is a growing concern about the harassment issue due to the immersive nature of social VR compared to other online social environments. Existing protections against harassment in social VR are highly limited in terms of practical effectiveness. The deficiency of studies toward understanding and preventing harassment in social VR further complicates the regulation and intervention efforts of social VR platforms in such situations. To address these challenges, we, in this paper, quantitatively investigate human interaction behaviors in social VR. More specifically, we first build a customized platform based on Mozilla Hubs, a popular social VR platform, to collect data about users’ social interaction behaviors involving harassment instances. A subsequent analysis of the collected dataset SAHARA (Social interAction beHAviors in vR with hArassment) reveals that the task of online harassment detection in social VR is complicated since it depends on not only users’ actions but also their spatial and temporal relationships. To accurately discern harassment, we propose a novel framework HardenVR (HA-Rassment DEtectioN framework for social VR). As a context-aware harassment detection framework, HardenVR employs a transformer-based model to capture relative poses and learn users’ hand actions in 6-DOF (Degree-of-Freedom). Meanwhile, multiple mechanisms, including the extra attention mechanism, distance-aware clustering method, and the sliding window, have been introduced into the model to handle challenges of data imbalance, over-fitting, and continuous detection. The design of HardenVR aims to achieve the balance between accuracy, efficiency, and cost-effectiveness for the task of harassment detection. As a starting point, HardenVR successfully learns pose information as the context to identify harassment and the experiment results show its detection accuracy as high as 98.26%.
Jin Zhou 0006, Jie Li 0064, Bo Han 0001, Fei Li 0001, Songqing Chen
VR6
2023 ScaleFlow: Efficient Deep Vision Pipeline with Closed-Loop Scale-Adaptive Inference
abstract
Deep visual data processing is underpinning many life-changing applications, such as auto-driving and smart cities. Improving the accuracy while minimizing their inference time under constrained resources has been the primary pursuit for their practical adoptions. Existing research thus has been devoted to either narrowing down the area of interest for the detection or miniaturizing the deep learning model for faster inference time. However, the former may risk missing/delaying small but important object detection, potentially leading to disastrous consequences (e.g., car accidents), while the latter often compromises the accuracy without fully utilizing intrinsic semantic information. To overcome these limitations, in this work, we propose ScaleFlow, a closed-loop scale-adaptive inference that can reduce model inference time by progressively processing vision data with increasing resolution but decreasing spatial size, achieving speedup without compromising accuracy. For this purpose, ScaleFlow refactors existing neural networks to be scale-equivariant on multiresolution data with the assistance of wavelet theory, providing predictable feature patterns on different data resolutions. Comprehensive experiments have been conducted to evaluate ScaleFlow. The results show that ScaleFlow can support anytime inference, consistently provide 1.5× to 2.2× speed up, and save around 25% ~ 45% energy consumption with < 1% accuracy loss on four embedded and edge platforms
Yuyang Leng, Renyuan Liu, Hongpeng Guo, Songqing Chen, Shuochao Yao
ACM Multimedia4
2023 An Efficient and Robust Cloud-Based Deep Learning With Knowledge Distillation
abstract
In recent years, deep neural networks have shown extraordinary power in various practical learning tasks, especially in object detection, classification, natural language processing. However, deploying such large models on resource-constrained devices or embedded systems is challenging due to their high computational cost. Efforts such as model partition, pruning, or quantization have been used at the expense of accuracy loss. Knowledge distillation is a technique that transfers model knowledge from a well-trained model (teacher) to a smaller and shallow model (student). Instead of using a learning model on the cloud, we can deploy distilled models on various edge devices, significantly reducing the computational cost, memory usage and prolonging the battery lifetime. In this work, we propose a novel neuron manifold distillation (NMD) method, where the student models imitate the teacher's output distribution and learn the feature geometry of the teacher model. In addition, to further improve the cloud-based learning system reliability, we propose a confident prediction mechanism to calibrate the model predictions. We conduct experiments with different distillation configurations over multiple datasets. Our proposed method demonstrates a consistent improvement in accuracy-speed trade-offs for the distilled model.
Zeyi Tao, Qi Xia 0003, Songqing Chen, Qun Li 0001
IEEE Trans. Cloud Comput.3
2023 Towards Software Defined Measurement in Data Centers: A Comparative Study of Designs, Implementation, and Evaluation
abstract
Cloud data centers are increasingly adopting the Software-Defined Networking (SDN) technologies for their underlying connection and communications. However, as a critical part of daily operations and management of such data centers, the network measurement is essential but has often been constrained by the available resources in the traditional network devices. Thus, how to properly balance the resource consumption while maintain timely and accurate measurement remains a challenge to data center systems. Recent advances in Software-Defined Networking (SDN) have enabled flexible and programmable network measurement, which is referred to as Software Defined Measurement (SDM). A promising trend for SDM is to conduct network traffic measurement on widely deployed Open vSwitches (OVS) in data centers. However, little attention has been paid to the design options for conducting traffic measurement on the OVS. In this study, we set to explore different designs and investigate the corresponding trade-offs among resource consumption, measurement accuracy, implementation complexity, and impact on switching speed. Through extensive experiments and comparisons, we quantitatively show the various trade-offs that the different schemes strike to balance, and demonstrate the feasibility of instrumenting OVS with monitoring capabilities. These results provide valuable insights into which design will best serve different measurement and monitoring needs.
Zili Zha, An Wang 0002, Yang Guo 0001, Songqing Chen
IEEE Trans. Cloud Comput.4
2023 Elastically Augmenting the Control-path Throughput in SDN to Deal with Internet DDoS Attacks
abstract
Distributed denial of service (DDoS) attacks have been prevalent on the Internet for decades. Albeit various defenses, they keep growing in size, frequency, and duration. The new network paradigm, Software-defined networking (SDN), is also vulnerable to DDoS attacks. SDN uses logically centralized control, bringing the advantages in maintaining a global network view and simplifying programmability. When attacks happen, the control path between the switches and their associated controllers may become congested due to their limited capacity. However, the data plane visibility of SDN provides new opportunities to defend against DDoS attacks in the cloud computing environment. To this end, we conduct measurements to evaluate the throughput of the software control agents on some of the hardware switches when they are under attacks. Then, we design a new mechanism, called Scotch , to enable the network to scale up its capability and handle the DDoS attack traffic. In our design, the congestion works as an indicator to trigger the mitigation mechanism. Scotch elastically scales up the control plane capacity by using an Open vSwitch-based overlay. Scotch takes advantage of both the high control plane capacity of a large number of vSwitches and the high data plane capacity of commodity physical switches to increase the SDN network scalability and resiliency under abnormal (e.g., DDoS attacks) traffic surges. We have implemented a prototype and experimentally evaluated Scotch . Our experiments in the small-scale lab environment and large-scale GENI testbed demonstrate that Scotch can elastically scale up the control channel bandwidth upon attacks.
Yuanjun Dai, An Wang 0002, Yang Guo 0001, Songqing Chen
ACM Trans. Internet Techn.4
2022 Are we ready for metaverse?: a measurement study of social virtual reality platforms
abstract
Social virtual reality (VR) has the potential to gradually replace traditional online social media, thanks to recent advances in consumer-grade VR devices and VR technology itself. As the vital foundation for building the Metaverse, social VR has been extensively examined by the computer graphics and HCI communities. However, there has been little systematic study dissecting the network performance of social VR, other than hype in the industry. To fill this critical gap, we conduct an in-depth measurement study of five popular social VR platforms: AltspaceVR, Horizon Worlds, Mozilla Hubs, Rec Room, and VRChat. Our experimental results reveal that all these platforms are still in their early stage and face fundamental technical challenges to realize the grand vision of Metaverse. For example, their throughput, end-to-end latency, and on-device computation resource utilization increase almost linearly with the number of users, leading to potential scalability issues. We identify the platform servers' direct forwarding of avatar data for embodying users without further processing as the main reason for the poor scalability and discuss potential solutions to address this problem. Moreover, while the visual quality of the current avatar embodiment is low and fails to provide a truly immersive experience, improving the avatar embodiment will consume more network bandwidth and further increase computation overhead and latency, making the scalability issues even more pressing.
Ruizhi Cheng, Nan Wu 0012, Matteo Varvello, Songqing Chen, Bo Han 0001
IMC4
2022 Towards Accurate Positioning in Multiuser Augmented Reality on Mobile Devices
abstract
Multiuser Augmented Reality (MuAR) is essential to implementing the vision of Metaverse for its capability to provide immersive and interactive experiences. In such experiences, peer positions are critical to understand each other’s intentions and actions so as to guarantee the smooth cooperation among users. However, we find that the explicit peer positions provided by the current practice could be incomplete and/or inaccurate in some situations, which leads to the weakened spatial awareness. To achieve the accurate peer tracking in MuAR, we propose a novel multiple sensors information fusion method, CSA (Coordinate System Alignment), to detect and correct defective relative positions by the current practice. CSA firstly formulates problem of correcting erroneous positions into an overdetermined system, and then finds the solution by applying the simulated annealing algorithm to expedite the search process. The evaluation results show that CSA’s ability to reduce errors significantly (58.3% on average) under long-term error duration, especially its advantage in reducing the relative direction errors. The result confirms the potential of CSA to provide reliable peer tracking in MuAR. Meanwhile, it does not impose extra restrictions on users’ practice with current mobile devices in experiences.
Stefano Petrangeli, Viswanathan (Vishy) Swaminathan, Fei Li 0001, Songqing Chen
ISM6
2022 Exploring Spherical Autoencoder for Spherical Video Content Processing
abstract
3D spherical content is increasingly presented in various applications (e.g., AR/MR/VR) for better users' immersiveness experience, yet today processing such spherical 3D content still mainly relies on the traditional 2D approaches after projection, leading to the distortion and/or loss of critical information. This study sets to explore methods to process spherical 3D content directly and more effectively. Using 360-degree videos as an example, we propose a novel approach called Spherical Autoencoder (SAE) for spherical video processing. Instead of projecting to a 2D space, SAE represents the 360-degree video content as a spherical object and employs encoding and decoding on the 360-degree video directly. Furthermore, to support the adoption of SAE on pervasive mobile devices that often have resource constraints, we further propose two optimizations on top of SAE.First, since the FoV (Field of View) prediction is widely studied and leveraged to transport only a portion of the content to the mobile device to save bandwidth and battery consumption, we design p-SAE, a SAE scheme with the partial view support that can utilize such FoV prediction. Second, since machine learning models are often compressed when running on mobile devices in order to reduce the processing load, which usually leads to degradation of output (e.g., video quality in SAE), we propose c-SAE by applying the compressive sensing theory into SAE to maintain the video quality when the model is compressed. Our extensive experiments show that directly incorporating and processing spherical signals is promising, and it outperforms the traditional approaches by a large margin. Both p-SAE and c-SAE show their effectiveness in delivering high quality videos (e.g., PSNR results) when used alone or combined together with model compression.
Jin Zhou 0006, Na Li 0032, Yao Liu 0001, Shuochao Yao, Songqing Chen
ACM Multimedia5
2022 A Reality Check of Positioning in Multiuser Mobile Augmented Reality: Measurement and Analysis
abstract
Multiuser Augmented Reality (MuAR) is essential to implementing the vision of Metaverse. With the pervasive mobile devices, MuAR enables multiple devices to share a common AR experience. In such experiences, the peer positions are critical to understand peers' intentions and actions so as to achieve the smooth interaction in AR. Such a spacial awareness requirement poses new challenges to MuAR. Traditionally, in AR experiences designed for the single user, the SLAM algorithm is adopted to compute self positions. However, the computed positions cannot be directly used to compute the relative positions of peer devices in MuAR, because they are computed with respect to independent coordinate systems associated with participating devices. To fill in the gap, the industry has recently proposed to implement peer tracking with the help of built-in Ultra Wideband (UWB) chip. In this work, we aim to perform a reality check on the proposed support, with the Nearby Interaction (NI) framework developed for iOS mobile devices as an example. The goal of our study is to gain an in-depth understanding about the reliability of the proposed support and identify potential issues. Through extensive measurements, we discover the peer tracking solution is not reliable sometimes, in terms of availability and accuracy. Furthermore, with regard to erroneous position reports, we present a quantitative analysis, summarizing the error types (e.g., transient errors and permanent errors) and revealing their underlying reasons. We believe the preliminary findings could help to improve the spacial awareness and enhance user experiences in MuAR.
Stefano Petrangeli, Viswanathan (Vishy) Swaminathan, Fei Li 0001, Songqing Chen
MMAsia6
2022 Preserving privacy in mobile spatial computing
abstract
Mapping and localization are the key components in mobile spatial computing to facilitate interactions between users and the digital model of the physical world. To enable localization, mobile devices keep capturing images of the real-world surroundings and uploading them to a server with spatial maps for localization. This leads to privacy concerns on the potential leakage of sensitive information in both spatial maps and localization images (e.g., when used in confidential industrial settings or our homes). Motivated by the above issues, we present a holistic research agenda in this paper for designing principled approaches to preserve privacy in spatial mapping and localization. We introduce our ongoing research, including learning-assisted noise generation to shield spatial maps, distributed architecture with intelligent aggregation to protect localization images, and end-to-end privacy preservation with fully homomorphic encryption. We also discuss the technical challenges, our preliminary results, and open research problems in those areas.
Nan Wu 0012, Ruizhi Cheng, Songqing Chen, Bo Han 0001
NOSSDAV3
2022 BinProv: Binary Code Provenance Identification without Disassembly
abstract
Provenance identification, which is essential for binary analysis, aims to uncover the specific compiler and configuration used for generating the executable. Traditionally, the existing solutions extract syntactic, structural, and semantic features from disassembled programs and employ machine learning techniques to identify the compilation provenance of binaries. However, their effectiveness heavily relies on disassembly tools (e.g., IDA Pro) and tedious feature engineering, since it is challenging to obtain accurate assembly code, particularly, from the stripped or obfuscated binaries. In addition, the features in machine learning approaches are manually selected based on the domain knowledge of one specific architecture, which cannot be applied to other architectures. In this paper, we develop an end-to-end provenance identification system BinProv, which leverages a BERT (Bidirectional Encoder Representations from Transformers) based embedding model to learn and represent the context semantics and syntax directly from the binary code. Therefore, BinProv avoids the disassembling step and manual feature selection in provenance identification. Moreover, BinProv can distinguish the compilers and the four optimization levels (O0/O1/O2/O3) by fine-tuning the classifier model with the embedding inputs for specific provenance identification tasks. Experimental results show that BinProv achieves 92.14%, 99.4%, and 99.8% accuracy at byte sequence, function, and binary levels, respectively. We further demonstrate that BinProv works well on obfuscated binary code, suggesting that BinProv is a viable approach to remarkably mitigate the disassembler dependence in future provenance identification tasks. Finally, our case studies show that BinProv can better identify compiler helper functions and improve the performance of binary code similarity detection.
Shu Wang 0004, Yunlong Xing, Pengbin Feng, Haining Wang 0001, Qi Li 0002, Songqing Chen, Kun Sun 0001
RAID7
2022 SFS: Smart OS Scheduling for Serverless Functions
abstract
Serverless computing enables a new way of building and scaling cloud applications by allowing developers to write fine-grained serverless or cloud functions. The execution duration of a cloud function is typically short-ranging from a few milliseconds to hundreds of seconds. However, due to resource contentions caused by public clouds' deep consolidation, the function execution duration may get significantly prolonged and fail to accurately account for the function's true resource usage. We observe that the function duration can be highly unpredictable with huge amplification of more than 50× for an open-source FaaS platform (OpenLambda). Our experiments show that the OS scheduling policy of cloud functions' host server can have a crucial impact on performance. The default Linux scheduler, CFS (Completely Fair Scheduler), being oblivious to workloads, frequently context-switches short functions, causing a turnaround time that is much longer than their service time. We propose SFS (Smart Function Scheduler), which works entirely in the user space and carefully orchestrates existing Linux FIFO and CFS schedulers to approximate Shortest Remaining Time First (SRTF). SFS uses two-level scheduling that seamlessly combines a new FILTER policy with Linux CFS, to trade off increased duration of long functions for significant performance improvement for short functions. We implement SFS in the Linux user space and port it to OpenLambda. Evaluation results show that SFS significantly improves short functions' duration with a small impact on relatively longer functions, compared to CFS.
Yuqi Fu, Li Liu 0045, Yue Cheng 0001, Songqing Chen
SC5
2022 Understanding Internet of Things malware by analyzing endpoints in their static artifacts
Jinchun Choi, Afsah Anwar, Abdulrahman Alabduljabbar, Hisham Alasmary, Jeffrey Spaulding, An Wang 0002, Songqing Chen, DaeHun Nyang, Amro Awad, David Mohaisen
Comput. Networks7
2022 Cleaning the NVD: Comprehensive Quality Assessment, Improvements, and Analyses
Afsah Anwar, Ahmed Abusnaina, Songqing Chen, Frank Li 0001, David Mohaisen
IEEE Trans. Dependable Secur. Comput.3
2021 SyncAttack: Double-spending in Bitcoin Without Mining Power
abstract
The existing Bitcoin security research has mainly followed the security models in [22, 35], which stipulate that an adversary controls some mining power in order to violate the blockchain consistency property (i.e., through a double-spend attack). These models, however, largely overlooked the impact of the realistic network synchronization, which can be manipulated given the permissionless nature of the network. In this paper, we revisit the security of Bitcoin blockchain by incorporating the network synchronization into the security model and evaluating that in practice. Towards this goal, we propose the ideal functionality for the Bitcoin network synchronization and specify bounds on the network outdegree and the block propagation delay in order to preserve the consistency property. By contrasting the ideal functionality against measurements, we find deteriorating network synchronization reported by Bitnodes and a notable churn rate with 10% of the nodes arriving and departing from the network daily.
Muhammad Saad 0001, Songqing Chen, David Mohaisen
CCS2
2021 Mind the Gap: Broken Promises of CPU Reservations in Containerized Multi-tenant Clouds
abstract
Containerization is becoming increasingly popular, but unfortunately, containers often fail to deliver the anticipated performance with the allocated resources. In this paper, we first demonstrate the performance variance and degradation are significant (by up to 5x) in a multi-tenant environment where containers are co-located. We then investigate the root cause of such performance degradation. Contrary to the common belief that such degradation is caused by resource contention and interference, we find that there is a gap between the amount of CPU a container reserves and actually gets. The root cause lies in the design choices of today's Linux scheduling mechanism, which we call Forced Runqueue Sharing and Phantom CPU Time. In fact, there are fundamental conflicts between the need to reserve CPU resources and Completely Fair Scheduler's work-conserving nature, and this contradiction prevents a container from fully utilizing its requested CPU resources. As a proof-of-concept, we implement a new resource configuration mechanism atop the widely used Kubernetes and Linux to demonstrate its potential benefits and shed light on future scheduler redesign. Our proof-of-concept, compared to the existing scheduler, improves the performance of both batch and interactive containerized apps by up to 5.6x and 13.7x.
Li Liu 0045, An Wang 0002, Mengbai Xiao, Yue Cheng 0001, Songqing Chen
SoCC6
2021 CE-SGD: Communication-Efficient Distributed Machine Learning
abstract
Training large-scale machine learning models usually demands a distributed approach to process the huge amount of training data efficiently. However, the high network communication cost introduced by parallel stochastic gradient descent (SGD) algorithms is a well-known bottleneck. To this end, we propose CE-SGD, a communication-efficient distributed machine learning algorithm that aggressively reduces the amount of gradient data exchanged among the training workers. CE-SGD belongs to the family of gradient sparsification schemes. CE-SGD adaptively adjusts the gradient sparsity according to the model's feedback. It also selectively transmits the gradients based on their degree of participation in the backpropagation. We mathematically prove the convergence of CE-SGD for both convex and non-convex cases and conduct a series of experiments on our CE-SGD implementation. Our experiments reveal that CE-SGD can achieve fast convergence, desirable gradient compression ratio, and high accuracy with low network bandwidth cost compared to state-of-the-art algorithms.
Zeyi Tao, Qi Xia 0003, Qun Li 0001, Songqing Chen
GLOBECOM4
2021 Root Cause Analyses for the Deteriorating Bitcoin Network Synchronization
abstract
The Bitcoin network synchronization is crucial for its security against partitioning attacks. From 2014 to 2018, the Bitcoin network size has increased, while the percentage of synchronized nodes has decreased due to block propagation delay, which increases with the network size. However, in the last few months, the network synchronization has deteriorated despite a constant network size. The change in the synchronization pattern suggests that the network size is not the only factor in place, necessitating a root cause analysis of network synchronization. In this paper, we perform a root cause analysis to study four factors that affect network synchronization: the unreachable nodes, the addressing protocol, the information relaying protocol, and the network churn. Our study reveals that the unreachable nodes size is 24x the reachable network size. We also found that the network addressing protocol does not distinguish between reachable and unreachable nodes, leading to inefficiencies due to attempts to connect with unreachable nodes/addresses. We note that the outcome of this behavior is a low success rate of the outgoing connections, which reduces the average outdegree. Through measurements, we found malicious nodes that exploit this opportunity to flood the network with unreachable addresses. We also discovered that Bitcoin follows a round-robin relaying mechanism that adds a small delay in block propagation. Finally, we observe a high churn in the Bitcoin network where ≈8 % nodes leave the network every day. In the last few months the churn among synchronized nodes has doubled, which is likely the most dominant factor in decreasing network synchronization. Consolidating our insights, we propose improvements in Bitcoin Core to increase network synchronization.
Muhammad Saad 0001, Songqing Chen, David Mohaisen
ICDCS2
2020 Statically Dissecting Internet of Things Malware: Analysis, Characterization, and Detection
Afsah Anwar, Hisham Alasmary, Jeman Park 0001, An Wang 0002, Songqing Chen, David Mohaisen
ICICS5
2020 A Data-Driven Study of DDoS Attacks and Their Dynamics
abstract
Despite continuous defense efforts, DDoS attacks are still very prevalent on the Internet. In such arms races, attackers are becoming more agile and their strategies are more sophisticated to escape from detection. Effective defenses demand in-depth understanding of such strategies. In this paper, we set to investigate the DDoS landscape from the perspective of the attackers. We focus on the dynamics of the attacking force, aiming to explore the strategies behind the scenes, if any. Our study is based on 50,704 different Internet DDoS attacks across the globe in a seven-month period. Our results indicate that attackers deliberately schedule their controlled bots in a dynamic fashion, and such dynamics can be well captured by statistical distributions. Furthermore, different botnet families exhibit similar scheduling patterns, strongly suggesting their close relationship and potential collaborations. Such collaborations are further confirmed by bots rotating in multiple families, and such rotation patterns are examined and confirmed at various levels. These findings lay a promising foundation for predicting DDoS attacks in the future and aid mitigation efforts.
An Wang 0002, Wentao Chang, Songqing Chen, David Mohaisen
IEEE Trans. Dependable Secur. Comput.3
2019 XLF: A Cross-layer Framework to Secure the Internet of Things (IoT)
abstract
The burgeoning Internet of Things (IoT) has offered unprecedented opportunities for innovations and applications that are continuously changing our life. At the same time, the large amount of pervasive IoT applications have posed paramount threats to the user's security and privacy. While a lot of efforts have been dedicated to deal with such threats from the hardware, the software, and the applications, in this paper, we argue and envision that more effective and comprehensive protection for IoT systems can only be achieved via a cross-layer approach. As such, we present our initial design of XLF, a cross-layer framework towards this goal. XLF can secure the IoT systems not only from each individual layer of device, network, and service, but also through the information aggregation and correlation of different layers.
An Wang 0002, David Mohaisen, Songqing Chen
ICDCS3
2019 Companion Paper for
abstract
This artifact includes source code, scripts and datasets required to reproduce the experimental figures in the evaluation of the MM'18 paper, which is entitled "MiniView Layout for Bandwidth-Efficient 360-Degree Video". The artifact reports the comparison results among the standard cube layout (CUBE), the equi-angular layout (EAC), and the MiniView layout (MVL) in terms of compressed video size, visual quality of views and decoding and rendering time.
Mengbai Xiao, Shuoqian Wang, Chao Zhou 0004, Li Liu 0045, Zhenhua Li 0001, Yao Liu 0001, Songqing Chen, Lucile Sassatelli, Gwendal Simon
ACM Multimedia7
2019 vCPU as a container: towards accurate CPU allocation for VMs
abstract
With our increasing reliance on cloud computing, accurate resource allocation of virtual machines (or domains) in the cloud have become more and more important. However, the current design of hypervisors (or virtual machine monitors) fails to accurately allocate resources to the domains in the virtualized environment. In this paper, we claim the root cause is that the protection scope is erroneously used as the resource scope for a domain in the current virtualization design. Such design flaw prevents the hypervisor from accurately accounting resource consumption of each domain. In this paper, using virtual CPUs as a container we propose to redefine the resource scope of a domain, so that the new resource scope is aligned with all the CPU consumption incurred by this domain. As a demonstration, we implement a novel system, called VASE (vCPU as a container), on top of the Xen hypervisor. Evaluations on our testbed have shown our proposed approach is effective in accounting system-wide CPU consumption incurred by domains, while introducing negligible overhead to the system.
Li Liu 0045, An Wang 0002, Mengbai Xiao, Yue Cheng 0001, Songqing Chen
VEE6
2019 Software-Defined Networking Enhanced Edge Computing: A Network-Centric Survey
abstract
Edge computing is burgeoning along with the rapidly increasing adoption of the Internet of Things (IoT). While there are studies on various aspects of edge computing, we find there is a lack of network perspective. In this paper, we, thus, first present an overview of how software-defined networking (SDN) and related technologies are being investigated in edge computing. Our purpose is to survey the state of the art and discuss the potential (remaining) challenges for future research. For this, we survey how SDN and related technologies are integrated to facilitate the management and operations of edge servers and various IoT devices. For the former, we review how SDN has been utilized in the access network, the core network, and the wide area network (WAN) between the edge and the cloud. For the latter, we focus on how SDN is leveraged to provide unified and programmable interfaces to manage devices. Through our discussion, we suggest that the SDN-related network support for edge computing deserves more in-depth investigations. We also identify several challenges and open issues to be addressed in the future.
An Wang 0002, Zili Zha, Yang Guo 0001, Songqing Chen
Proc. IEEE4
2018 Empirical Evaluation of the Hypervisor Scheduling on Side Channel Attacks
abstract
Along with the wide adoption of the cloud platform, various attacks also target clouds. Due to the sharing of the underlying physical resources among different virtual machines (VMs), various side-channel attacks have been demonstrated to be capable of stealing victim's secret information, such as encryption key, by monitoring the victim's access pattern to a shared hardware, such as CPU cache. Among various defense mechanisms proposed, the hypervisor scheduling based schemes shed some light on lightweight solutions that are more likely to be adopted in practice. However, scheduling is affected by several factors that have not been thoroughly investigated so far. In this study, we aim to study in-depth the impact of various factors affecting the hypervisor scheduling, with the objective to understand their impact on mitigating these side-channel attacks. Our results can not only deepen our understanding, but also provide some guidelines to design effective scheduling based defenses in the future.
Li Liu 0045, An Wang 0002, Wanyu Zang, Meng Yu 0001, Songqing Chen
ICC5
2018 BAS-360°: Exploring Spatial and Temporal Adaptability in 360-degree Videos over HTTP/2
abstract
Today, 360-degree video streaming has become a popular Internet service with the rise of affordable virtual reality (VR) technologies. However, streaming 360-degree videos suffers from the prohibitive bandwidth demand. Existing bandwidth-efficient solutions mainly focus on exploiting the inherent spatial adaptability of 360-degree videos, delivering only video content (spatially-cut tiles) in the viewer's region of interest (ROI) with higher quality. Temporal adaptability, which has been widely leveraged in HTTP streaming, has not been well exploited to select proper quality for video segments according to the bandwidth variations. When these two dimensions of adaptability are jointly considered, bitrate selection for the tiles become more complicated and challenging. The importance of a tile with a spatial coordination played at a specific time should be quantified so that we can determine how to allocate bandwidth for improving the viewer's quality of experience. Furthermore, viewer's head orientation prediction is highly variable, which makes the determination of important tiles highly dynamic. In addition, network fluctuations are very common on the Internet. To overcome these challenges, we propose Bi-Adaptive Streaming for 360-degree videos (BAS-360°). In BAS-360°, both spatial and temporal adaptabilities are explored in the bitrate selection for different tiles. The objective is to minimize the bandwidth waste by allocating bandwidth to more important tiles (the tiles that are more likely to be watched). To tackle the high variability of visual region prediction and the unpredictable network fluctuations, we employ two features provided by HTT P /2: stream termination and stream priority, to efficiently organize tile delivery. Evaluation results show that BAS-360° outperforms naive tile-based 360-degree video streaming strategies when network fluctuations or errors in viewport predictions occur.
Mengbai Xiao, Chao Zhou 0004, Viswanathan (Vishy) Swaminathan, Yao Liu 0001, Songqing Chen
INFOCOM5
2018 MiniView Layout for Bandwidth-Efficient 360-Degree Video
abstract
With the recent increase in popularity of VR devices, 360-degree video has become increasingly popular. As more users experience this new medium, it will likely see further increases in popularity as users experience its greater immersiveness compared to traditional video streams. 360-degree video streams must encode the omnidirectional view, and, with current encoding techniques, these views require significantly higher bandwidth than traditional video streams. These larger bandwidth requirements comprise the main barrier toward wider adoption by video streaming services.
Mengbai Xiao, Shuoqian Wang, Chao Zhou 0004, Li Liu 0045, Zhenhua Li 0001, Yao Liu 0001, Songqing Chen
ACM Multimedia7
2018 Shuffler: Mitigate Cross-VM Side-Channel Attacks via Hypervisor Scheduling
Li Liu 0045, An Wang 0002, Wanyu Zang, Meng Yu 0001, Menbai Xiao, Songqing Chen
SecureComm (1)6
2018 Risk-aware multi-objective optimized virtual machine placement in the cloud
abstract
Cloud computing, while becoming more and more popular as a dominant computing platform, introduces new security challenges. When virtual machines are deployed in a cloud environment, virtual machine placement strategies can significantly affect the overall security risks of the entire cloud. In recent years, the attacks are specifically designed to co-locate with target virtual machines in the cloud. The virtual machine placement without considering the security risks may put the users, or even the entire cloud, in danger. In this paper, we present a comprehensive approach to quantify the security risk of cloud environments from network, host and VM. Accordingly, we propose a Security-aware Multi-Objective Optimization based virtual machine Placement scheme (SMOOP) to seek a Pareto-optimal solution that reduces the overall security risks of a cloud, while considering workload balance, resource utilization on CPU, memory, disk, and network traffic. New placement strategies are designed and our evaluation results demonstrate their effectiveness. The security of clouds could be improved with affordable overheads. The latest VM allocation policies are further studied and integrated into our designs to defeat the co-residence attacks.
Wangyu Zang, Li Liu 0045, Songqing Chen, Meng Yu 0001
J. Comput. Secur.4
2018 Delving Into Internet DDoS Attacks by Botnets: Characterization and Analysis
abstract
Internet distributed denial of service (DDoS) attacks are prevalent but hard to defend against, partially due to the volatility of the attacking methods and patterns used by attackers. Understanding the latest DDoS attacks can provide new insights for effective defense. But most of existing understandings are based on indirect traffic measures (e.g., backscatters) or traffic seen locally. In this paper, we present an in-depth analysis based on 50 704 different Internet DDoS attacks directly observed in a seven-month period. These attacks were launched by 674 botnets from 23 different botnet families with a total of 9026 victim IPs belonging to 1074 organizations in 186 countries. Our analysis reveals several interesting findings about today's Internet DDoS attacks. Some highlights include: 1) geolocation analysis shows that the geospatial distribution of the attacking sources follows certain patterns, which enables very accurate source prediction of future attacks for most active botnet families; 2) from the target perspective, multiple attacks to the same target also exhibit strong patterns of inter-attack time interval, allowing accurate start time prediction of the next anticipated attacks from certain botnet families; and 3) there is a trend for different botnets to launch DDoS attacks targeting the same victim, simultaneously or in turn. These findings add to the existing literature on the understanding of today's Internet DDoS attacks and offer new insights for designing new defense schemes at different levels.
An Wang 0002, Wentao Chang, Songqing Chen, David Mohaisen
IEEE/ACM Trans. Netw.3
2018 A New Deep Learning-Based Food Recognition System for Dietary Assessment on An Edge Computing Service Infrastructure
abstract
Literature has indicated that accurate dietary assessment is very important for assessing the effectiveness of weight loss interventions. However, most of the existing dietary assessment methods rely on memory. With the help of pervasive mobile devices and rich cloud services, it is now possible to develop new computer-aided food recognition system for accurate dietary assessment. However, enabling this future Internet of Things-based dietary assessment imposes several fundamental challenges on algorithm development and system design. In this paper, we set to address these issues from the following two aspects: (1) to develop novel deep learning-based visual food recognition algorithms to achieve the best-in-class recognition accuracy; (2) to design a food recognition system employing edge computing-based service computing paradigm to overcome some inherent problems of traditional mobile cloud computing paradigm, such as unacceptable system latency and low battery life of mobile devices. We have conducted extensive experiments with real-world data. Our results have shown that the proposed system achieved three objectives: (1) outperforming existing work in terms of food recognition accuracy; (2) reducing response time that is equivalent to the minimum of the existing approaches; and (3) lowering energy consumption which is close to the minimum of the state-of-the-art.
Chang Liu 0033, Yu Cao 0002, Yan Luo 0001, Vinod Vokkarane, Yunsheng Ma, Songqing Chen
IEEE Trans. Serv. Comput.7
2017 Reducing Security Risks of Clouds Through Virtual Machine Placement
Wanyu Zang, Songqing Chen, Meng Yu 0001
DBSec3
2017 An Adversary-Centric Behavior Modeling of DDoS Attacks
abstract
Distributed Denial of Service (DDoS) attacks are some of the most persistent threats on the Internet today. The evolution of DDoS attacks calls for an in-depth analysis of those attacks. A better understanding of the attackers' behavior can provide insights to unveil patterns and strategies utilized by attackers. The prior art on the attackers' behavior analysis often falls in two aspects: it assumes that adversaries are static, and makes certain simplifying assumptions on their behavior, which often are not supported by real attack data. In this paper, we take a data-driven approach to designing and validating three DDoS attack models from temporal (e.g., attack magnitudes), spatial (e.g., attacker origin), and spatiotemporal (e.g., attack inter-launching time) perspectives. We design these models based on the analysis of traces consisting of more than 50,000 verified DDoS attacks from industrial mitigation operations. Each model is also validated by testing its effectiveness in accurately predicting future DDoS attacks. Comparisons against simple intuitive models further show that our models can more accurately capture the essential features of DDoS attacks.
An Wang 0002, David Mohaisen, Songqing Chen
ICDCS3
2017 vPROM: VSwitch enhanced programmable measurement in SDN
abstract
While being critical to the network management, the current state of the art in network measurement is inadequate, providing surprisingly little visibility into detailed network behaviors and often requiring high level of manual intervention to operate. Such a practice becomes increasingly ineffective as the networks grow both in size and complexity. In this paper, we propose vPROM, a vSwitch enhanced SDN programmable measurement framework that automates the measurement process, minimizes the measurement resource usage, and addresses several significant technical challenges faced by early works. vPROM leverages the SDN programmability and extends the Pyretic runtime system and OpenFlow network interface to achieve the measurement automation. The required measurement resources are minimized by only acquiring the necessary statistics, made possible with instrumented Open vSwitches1with user defined monitoring capability. By decoupling monitoring from routing, vPROM reduces the interference between the measurement applications and other applications, and eliminates the frequent involvement of the controller. A vPROM prototype is implemented with DDoS and port-scan detection applications. The performance of vPROM is evaluated and the comparison results with other existing programmable measurement approaches are also presented.
An Wang 0002, Yang Guo 0001, Songqing Chen, Fang Hao, T. V. Lakshman, Doug Montgomery, Kotikalapudi Sriram
ICNP3
2017 OpTile: Toward Optimal Tiling in 360-degree Video Streaming
abstract
360-degree videos are encoded for adaptive streaming by first projecting the spherical surface onto two-dimensional frames, then encoding these as standard video segments. During playback of these 360-degree videos, the video player renders the portion of the spherical surface in the direction of the user's view. These user viewports typically cover only a small portion of the 360 degree surface, causing much of the downloaded bandwidth to be wasted. Tile-based approaches can reduce the wasted bandwidth by cutting video spatially into motion-constrained rectangles. Streaming logic then only needs to download the tiles necessary to render the viewport seen by the user. Existing tile-based approaches cut 360-degree videos into tiles of fixed sizes. These fixed-size tiling approaches, however, suffer from reduced encoding efficiency. Tiling cuts away portions of the video that can be copied by the encoder from adjacent frames or within the current frame that are needed for effective video compression.
Mengbai Xiao, Chao Zhou 0004, Yao Liu 0001, Songqing Chen
ACM Multimedia4
2017 Understanding Adversarial Strategies from Bot Recruitment to Scheduling
Wentao Chang, David Mohaisen, An Wang 0002, Songqing Chen
SecureComm4
2016 DASH2M: Exploring HTTP/2 for Internet Streaming to Mobile Devices
abstract
Today HTTP/1.1 is the most popular vehicle for delivering Internet content, including streaming video. Standardized in 2015 with a few new features, HTTP/2 is gradually replacing HTTP 1.1 to improve user experience. Yet, how HTTP/2 can help improve the video streaming delivery has not been thoroughly investigated. In this work, we set to investigate how to utilize the new features offered by HTTP/2 for video streaming over the Internet, focusing on the streaming delivery to mobile devices as, today, more and more users watch video on their mobile devices. For this purpose, we design DASH2M, Dynamic Adaptive Streaming over HTTP/2 to Mobile Devices. DASH2M deliberately schedules the streaming content delivery by comprehensively considering the user's Quality of Experience (QoE), the dynamics of the network resources, and the power efficiency on the mobile devices. Experiments based on an implemented prototype show that DASH2M can outperform prior strategies for users' QoE while minimizing the battery power consumption on mobile devices.
Mengbai Xiao, Viswanathan (Vishy) Swaminathan, Sheng Wei 0001, Songqing Chen
ACM Multimedia4
2016 Evaluating and improving push based video streaming with HTTP/2
abstract
The sever-initiated push mechanism is one of the most prominent features in the next generation HTTP/2 protocol, having shown its capability on saving network traffic and improving the web page retrieval latency. Our prior work has investigated the server push-based mechanism for HTTP video streaming and proposed a k-push scheme, where the server pushes k video segments following the response to a request. In this study, we further conduct an analysis and evaluation of the k-push scheme in HTTP streaming. Our results uncover that the push mechanism can efficiently increase the network utilization (under certain conditions) compared to regular HTTP streaming. However the results also show that the k-push scheme deteriorates network adaptability and leads to the "over-push" problem, in which the pushed video content waste network resources due to user abandonment behaviors. To overcome these limitations, we propose a new " adaptive-push" scheme, which dynamically adjusts the parameter k to adapt to the runtime environment. To evaluate the performance of adaptive-push, we implemented a prototype system. The experimental results show that compared to k-push, adaptive-push can improve the network adaptability. Furthermore, our real-world trace based simulation results show that adaptive-push can effectively alleviate the over-push problem.
Mengbai Xiao, Viswanathan (Vishy) Swaminathan, Sheng Wei 0001, Songqing Chen
NOSSDAV4
2016 Attribution of Economic Denial of Sustainability Attacks in Public Clouds
Mohammad Karami, Songqing Chen
SecureComm2
2016 GoCAD: GPU-Assisted Online Content-Adaptive Display Power Saving for Mobile Devices in Internet Streaming
abstract
During Internet streaming, a significant portion of the battery power is always consumed by the display panel on mobile devices. To reduce the display power consumption, backlight scaling, a scheme that intelligently dims the backlight has been proposed. To maintain perceived video appearance in backlight scaling, a computationally intensive luminance compensation process is required. However, this step, if performed by the CPU as existing schemes suggest, could easily offset the power savings gained from backlight scaling. Furthermore, computing the optimal backlight scaling values requires per-frame luminance information, which is typically too energy intensive for mobile devices to compute. Thus, existing schemes require such information to be available in advance. And such an offline approach makes these schemes impractical. To address these challenges, in this paper, we design and implement GoCAD, a GPU-assisted Online Content-Adaptive Display power saving scheme for mobile devices in Internet streaming sessions. In GoCAD, we employ the mobile device's GPU rather than the CPU to reduce power consumption during the luminance compensation phase. Furthermore, we compute the optimal backlight scaling values for small batches of video frames in an online fashion using a dynamic programming algorithm. Lastly, we make novel use of the widely available video storyboard, a pre-computed set of thumbnails associated with a video, to intelligently decide whether or not to apply our backlight scaling scheme for a given video. For example, when the GPU power consumption would offset the savings from dimming the backlight, no backlight scaling is conducted. To evaluate the performance of GoCAD, we implement a prototype within an Android application and use a Monsoon power monitor to measure the real power consumption. Experiments are conducted on more than 460 randomly selected YouTube videos. Results show that GoCAD can effectively produce power savings without affecting rendered video quality.
Yao Liu 0001, Mengbai Xiao, Xin Li 0078, Mian Dong, Zhenhua Li 0001, Songqing Chen
WWW8
2016 Content-Adaptive Display Power Saving for Internet Video Applications on Mobile Devices
abstract
Backlight scaling is a technique proposed to reduce the display panel power consumption by strategically dimming the backlight. However, for mobile video applications, a computationally intensive luminance compensation step must be performed in combination with backlight scaling to maintain the perceived appearance of video frames. This step, if done by the Central Processing Unit (CPU), could easily offset the power savings via backlight dimming. Furthermore, computing the backlight scaling values requires per-frame luminance information, which is typically too energy intensive to compute on mobile devices. In this article, we propose Content-Adaptive Display (CAD) for two typical Internet mobile video applications: video streaming and real-time video communication. CAD uses the mobile device’s Graphics Processing Unit (GPU) rather than the CPU to perform luminance compensation at reduced power consumption. For video streaming where video frames are available in advance, we compute the backlight scaling schedule using a more efficient dynamic programming algorithm than existing work. For real-time video communication where video frames are generated on the fly, we propose a greedy algorithm to determine the backlight scaling at runtime. We implement CAD in one video streaming application and one real-time video call application on the Android platform and use a Monsoon power meter to measure the real power consumption. Experiment results show that CAD can save more than 10% overall power consumption for up to 55.7% videos during video streaming and up to 31.0% overall power consumption in real-time video calls.
Yao Liu 0001, Mengbai Xiao, Xin Li 0078, Mian Dong, Zhenhua Li 0001, Lei Guo 0004, Songqing Chen
ACM Trans. Multim. Comput. Commun. Appl.9
2015 Measuring Botnets in the Wild: Some New Trends
abstract
Today, botnets are still responsible for most large scale attacks on the Internet. Botnets are versatile, they remain the most powerful attack platform by constantly and continuously adopting new techniques and strategies in the arms race against various detection schemes. Thus, it is essential to understand the latest of the botnets in a timely manner so that the insights can be utilized in developing more efficient defenses. In this work, we conduct a measurement study on some of the most active botnets on the Internet based on a public dataset collected over a period of seven months by a monitoring entity. We first examine and compare the attacking capabilities of different families of today's active botnets. Our analysis clearly shows that different botnets start to collaborate when launching DDoS attacks.
Wentao Chang, David Mohaisen, An Wang 0002, Songqing Chen
AsiaCCS4
2015 UMON: flexible and fine grained traffic monitoring in open vSwitch
abstract
We study how to provide fine-grained, flexible traffic monitoring in the Open vSwitch (OVS). We argue that the existing OVS monitoring tools are neither flexible nor sufficient for supporting many monitoring applications. We propose UMON, a mechanism that decouples monitoring from forwarding, and offers flexible and fine-grained traffic stats. We describe a prototype implementation of UMON that integrates well with the OVS architecture. Finally, we evaluate the performance using the prototype, and illustrate UMON's efficiency with the example use cases such as detecting port scans.
An Wang 0002, Yang Guo 0001, Fang Hao, T. V. Lakshman, Songqing Chen
CoNEXT5
2015 Capturing DDoS Attack Dynamics Behind the Scenes
An Wang 0002, David Mohaisen, Wentao Chang, Songqing Chen
DIMVA4
2015 Delving into Internet DDoS Attacks by Botnets: Characterization and Analysis
abstract
Internet Distributed Denial of Service (DDoS) at- tacks are prevalent but hard to defend against, partially due to the volatility of the attacking methods and patterns used by attackers. Understanding the latest DDoS attacks can provide new insights for effective defense. But most of existing understandings are based on indirect traffic measures (e.g., backscatters) or traffic seen locally. In this study, we present an in-depth analysis based on 50,704 different Internet DDoS attacks directly observed in a seven-month period. These attacks were launched by 674 botnets from 23 different botnet families with a total of 9,026 victim IPs belonging to 1,074 organizations in 186 countries. Our analysis reveals several interesting findings about today's Internet DDoS attacks. Some highlights include: (1) geolocation analysis shows that the geospatial distribution of the attacking sources follows certain patterns, which enables very accurate source prediction of future attacks for most active botnet families, (2) from the target perspective, multiple attacks to the same target also exhibit strong patterns of inter-attack time interval, allowing accurate start time prediction of the next anticipated attacks from certain botnet families, (3) there is a trend for different botnets to launch DDoS attacks targeting the same victim, simultaneously or in turn. These findings add to the existing literature on the understanding of today's Internet DDoS attacks, and offer new insights for designing new defense schemes at different levels.
An Wang 0002, David Mohaisen, Wentao Chang, Songqing Chen
DSN4
2015 Reducing display power consumption for real-time video calls on mobile devices
abstract
The display subsystem of a mobile device usually consumes 38%-68% [1] of the total battery power in video streaming. Therefore, a few schemes have been designed to reduce the display power consumption. The basic idea is to dim the backlight level while properly compensating the pixel luminance to maintain image fidelity. The luminance compensation and proper backlight level calculation are computation intensive and demand per-frame luminance information. For these reasons, existing schemes only work for video-on-demand where each frame (and thus the luminance information) is available in advance. In addition, they demand additional computing resource support. Otherwise, if the computation is conducted on the mobile device, the power consumption due to such computation can easily offset the power savings from dimming the backlight. In this work, we set to investigate power saving for real-time video calls on mobile devices. Different from video-on-demand, real-time video calls are highly delay sensitive and the frame luminance information is not known in advance. Moreover, video calls often involve multiple streaming sources from multiple (≥2) participants, making it more difficult. Because there are few background changes and the frame rate is usually small in video calls, we design a Greedy Display Power saving scheme, called LCD-GDP, which utilizes the commonly available GPU on mobile devices without demanding additional support. Our design is implemented on WebRTC, a popular real-time web browser based video call standard. Experiments show that our scheme can save up to 33% power consumption in video calls without affecting the video call quality.
Mengbai Xiao, Yao Liu 0001, Lei Guo 0004, Songqing Chen
ISLPED4
2015 FAST: A fog computing assisted distributed analytics system to monitor fall for stroke mitigation
abstract
Fog computing is a recently proposed computing paradigm that extends Cloud computing and services to the edge of the network. The new features offered by fog computing (e.g., distributed analytics and edge intelligence), if successfully applied for pervasive health monitoring applications, has great potential to accelerate the discovery of early predictors and novel biomarkers to support smart care decision making in a connected health scenarios. While promising, how to design and develop real-word fog computing-based pervasive health monitoring system is still an open question. As a first step to answer this question, in this paper, we employ pervasive fall detection for stroke mitigation as a case in study. There are four major contributions in this paper: (1) to investigate and develop a set of new fall detection algorithms, including new fall detection algorithms based on acceleration magnitude values and non-linear time series analysis techniques, as well as new filtering techniques to facilitate fall detection process; (2) to design and employ a real-time fall detection system employing fog computing paradigm, which distribute the analytics throughout the network by splitting the detection task between the edge devices (e.g., smartphones attached to the user) and the server (e.g., servers in the cloud); (3) we carefully exam the special needs and constraints of stroke patients and propose patient-centered design that is minimal intrusive to patients. This type of patient-centered design is currently lacking in most of the existing work; and (4) our experiments with real-word data show that our proposed system achieves the high sensitivity (low missing rate) while it also achieves the high specificity (low false alarm rate). At the same time, the response time and energy consumption of our system are close to the minimum of the existing approaches.
Yu Cao 0002, Songqing Chen, Donald Brown
NAS2
2015 Elicit: Efficiently identify computation-intensive tasks in mobile applications for offloading
abstract
As mobile devices are battery powered and have less computing resources, plenty of research has been conducted on how to efficiently offload computing-intensive tasks in a mobile application to more powerful counterpart. However, prior research either implicitly assumes that the computing-intensive tasks are known in advance or the application developers will make special notations about them. In this paper, we design a framework Elicit to efficiently identify the computation-intensive tasks in mobile applications for offloading. Furthermore, we also consider the response time savings dynamically when deciding whether to offload a task based on the runtime system resources. A prototype of Elicit is built based on the Dalvik VM. Our evaluation with some popular Android applications from Google Play shows that Elicit can efficiently find an application's computing-intensive task and save response time and energy consumption when these tasks are offloaded.
Mohammed Anowarul Hassan, Qi Wei 0003, Songqing Chen
NAS3
2015 Content-adaptive display power saving in internet mobile streaming
abstract
Backlight scaling is a technique proposed to reduce the display panel power consumption by strategically dimming the backlight. However, for Internet streaming to mobile devices, a computationally intensive luminance compensation step must be performed in combination with backlight scaling to maintain the perceived appearance of video frames. This step, if done by the CPU, could easily offset the power savings via backlight dimming. Furthermore, computing the backlight scaling values requires per-frame luminance information, which is typically too energy intensive to compute on mobile devices.
Yao Liu 0001, Mengbai Xiao, Xin Li 0078, Mian Dong, Zhenhua Li 0001, Songqing Chen
NOSSDAV8
2015 A Quantitative Study of Video Duplicate Levels in YouTube
Yao Liu 0001, Sam Blasiak, Weijun Xiao, Zhenhua Li 0001, Songqing Chen
PAM5
2015 Introduction to the Special Issue on MMSys 2014 and NOSSDAV 2014
abstract
No abstract available.
Kuan-Ta Chen, Songqing Chen, Wei Tsang Ooi
ACM Trans. Multim. Comput. Commun. Appl.2
2014 POSTER: How Distributed Are Today's DDoS Attacks?
abstract
Today botnets are responsible for most of the DDoS attacks on the Internet. Understanding the characteristics of such DDoS attacks is critical to develop effective DDoS mitigation schemes. In this poster, we present some preliminary findings, mainly concerning the distribution of the attackers, of today's DDoS attacks. Our investigation is based on 50,704 different Internet DDoS attacks collected within a seven-month period for activities across the globe. These attacks were launched by 674 botnet generations from 23 different bonet families with a total of 9026 victim IPs belonging to 1074 organizations that are collectively located in 186 countries. We find that different from the traditional widely distributed intuition, most of these DDoS attacks are not widely distributed as the attackers are mostly from the same region, i.e., highly regionalized. We also find that different botnet families have strong target preferences in the same area as well. These findings refresh our understanding on the modern DDoS attacks.
An Wang 0002, Wentao Chang, David Mohaisen, Songqing Chen
CCS4
2014 Scotch: Elastically Scaling up SDN Control-Plane using vSwitch based Overlay
abstract
Software Defined Networks use logically centralized control due to its benefits in maintaining a global network view and in simplifying programmability. However, the use of centralized controllers can affect network performance if the control path between the switches and their associated controllers becomes a bottleneck. We find from measurements that the software control agents on some of the switches have very limited throughput. This can cause performance degradation if the switch has to handle a high traffic load, as for instance due to flash crowds or DDoS attacks. This degradation can occur even when the data plane capacity is under-utilized. The goal of our paper is to design new mechanisms to enable the network to scale up its ability to handle high control traffic loads. For this purpose, we design, implement, and experimentally evaluate Scotch, a solution that elastically scales up the control plane capacity by using a vSwitch based overlay. Scotch takes advantage of both the high control plane capacity of a large number of vSwitches and the high data plane capacity of commodity physical switches to increase the SDN network scalability and resiliency under normal (e.g., flash crowds) or abnormal (e.g., DDoS attacks) traffic surge.
An Wang 0002, Yang Guo 0001, Fang Hao, T. V. Lakshman, Songqing Chen
CoNEXT5
2014 Characterizing botnets-as-a-service
abstract
No abstract available.
Wentao Chang, An Wang 0002, David Mohaisen, Songqing Chen
SIGCOMM4
2014 A Host-Based Approach for Unknown Fast-Spreading Worm Detection and Containment
abstract
The fast-spreading worm, which immediately propagates itself after a successful infection, is becoming one of the most serious threats to today’s networked information systems. In this article, we present WormTerminator, a host-based solution for fast Internet worm detection and containment with the assistance of virtual machine techniques based on the fast-worm defining characteristic. In WormTerminator, a virtual machine cloning the host OS runs in parallel to the host OS. Thus, the virtual machine has the same set of vulnerabilities as the host. Any outgoing traffic from the host is diverted through the virtual machine. If the outgoing traffic from the host is for fast worm propagation, the virtual machine should be infected and will exhibit worm propagation pattern very quickly because a fast-spreading worm will start to propagate as soon as it successfully infects a host. To prove the concept, we have implemented a prototype of WormTerminator and have examined its effectiveness against the real Internet worm Linux/Slapper. Our empirical results confirm that WormTerminator is able to completely contain worm propagation in real-time without blocking any non-worm traffic. The major performance cost of WormTerminator is a one-time delay to the start of each outgoing normal connection for worm detection. To reduce the performance overhead, caching is utilized, through which WormTerminator will delay no more than 6% normal outgoing traffic for such detection on average.
Songqing Chen, Lei Liu 0021, Xinyuan Wang 0005, Xinwen Zhang, Zhao Zhang 0010
ACM Trans. Auton. Adapt. Syst.1
2014 Investigating Redundant Internet Video Streaming Traffic on iOS Devices: Causes and Solutions
abstract
The Internet has witnessed rapidly increasing streaming traffic to various mobile devices. In this paper, through analysis of a server-side workload and experiments in a controlled lab environment, we find that current practice has introduced a significant amount of redundant traffic. In particular, for the popular iOS based mobile devices, accessing popular Internet streaming services typically involves about 10%-70% redundant traffic. Such a practice not only over-utilizes and wastes resources on the server side and the network (cellular or Internet), but also consumes additional battery power on user's mobile devices and leads to possible monetary cost. To alleviate such a situation without changing the server side or the client side, we design and implement CStreamer that can transparently work between existing mobile clients and servers. We have implemented a prototype and installed on Amazon EC2. Experiments conducted based on this prototype show that CStreamer can completely eliminate the redundant traffic without degrading user's QoS.
Yao Liu 0001, Qi Wei 0003, Lei Guo 0004, Bo Shen 0003, Songqing Chen, Yingjie Lan
IEEE Trans. Multim.5
2013 UNIK: unsupervised social network spam detection
abstract
Social network spam increases explosively with the rapid development and wide usage of various social networks on the Internet. To timely detect spam in large social network sites, it is desirable to discover unsupervised schemes that can save the training cost of supervised schemes. In this work, we first show several limitations of existing unsupervised detection schemes. The main reason behind the limitations is that existing schemes heavily rely on spamming patterns that are constantly changing to avoid detection. Motivated by our observations, we first propose a sybil defense based spam detection scheme SD2 that remarkably outperforms existing schemes by taking the social network relationship into consideration. In order to make it highly robust in facing an increased level of spam attacks, we further design an unsupervised spam detection scheme, called UNIK. Instead of detecting spammers directly, UNIK works by deliberately removing non-spammers from the network, leveraging both the social graph and the user-link graph. The underpinning of UNIK is that while spammers constantly change their patterns to evade detection, non-spammers do not have to do so and thus have a relatively non-volatile pattern. UNIK has comparable performance to SD2 when it is applied to a large social network site, and outperforms SD2 significantly when the level of spam attacks increases. Based on detection results of UNIK, we further analyze several identified spam campaigns in this social network site. The result shows that different spammer clusters demonstrate distinct characteristics, implying the volatility of spamming patterns and the ability of UNIK to automatically extract spam signatures.
Enhua Tan, Lei Guo 0004, Songqing Chen, Xiaodong Zhang 0001, Yihong Eric Zhao
CIKM3
2013 Defeat Information Leakage from Browser Extensions via Data Obfuscation
Wentao Chang, Songqing Chen
ICICS2
2013 Effectively minimizing redundant Internet streaming traffic to iOS devices
abstract
The Internet has witnessed rapidly increasing streaming traffic to various mobile devices. In this paper, we find that for the popular iOS based mobile devices, accessing popular Internet streaming services typically involves about 10% - 70% unnecessary redundant traffic. Such a practice not only overutilizes and wastes resources on the server side and the network (cellular or Internet), but also consumes additional battery power on users' mobile devices and leads to possible monetary cost. To alleviate such a situation without changing the server side or the iOS, we design and implement a CStreamer prototype that can transparently work between existing iOS devices and media servers. We also build a CStreamer iOS App to enable end users to access Internet streaming services via CStreamer. Experiments conducted based on this prototype running on Amazon EC2 show that CStreamer can completely eliminate the redundant traffic without degrading user's QoS.
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Bo Shen 0003, Songqing Chen
INFOCOM5
2013 A Comparative Study of Android and iOS for Accessing Internet Streaming Services
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Bo Shen 0003, Songqing Chen
PAM5
2013 Using adaptively coupled models and high-performance computing for enabling the computability of dust storm forecasting
abstract
Forecasting dust storms for large geographical areas with high resolution poses great challenges for scientific and computational research. Limitations of computing power and the scalability of parallel systems preclude an immediate solution to such challenges. This article reports our research on using adaptively coupled models to resolve the computational challenges and enable the computability of dust storm forecasting by dividing the large geographical domain into multiple subdomains based on spatiotemporal distributions of the dust storm. A dust storm model (Eta-8bin) performs a quick forecasting with low resolution (22 km) to identify potential hotspots with high dust concentration. A finer model, non-hydrostatic mesoscale model (NMM-dust) performs high-resolution (3 km) forecasting over the much smaller hotspots in parallel to reduce computational requirements and computing time. We also adopted spatiotemporal principles among computing resources and subdomains to optimize parallel systems and improve the performance of high-resolution NMM-dust model. This research enabled the computability of high-resolution, large-area dust storm forecasting using the adaptively coupled execution of the two models Eta-8bin and NMM-dust.
Qunying Huang, Chaowei Phil Yang, Karl Benedict, Abdelmounaam Rezgui, Jibo Xie, Jizhe Xia, Songqing Chen
Int. J. Geogr. Inf. Sci.7
2013 Measurement and Analysis of an Internet Streaming Service to Mobile Devices
abstract
Receiving Internet streaming services on various mobile devices is getting increasingly popular, and cloud platforms have also been gradually employed for delivering streaming services to mobile devices. While a number of studies have been conducted at the client side to understand and characterize Internet mobile streaming delivery, little is known about the server side, particularly for the recent cloud-based Internet mobile streaming delivery. In this work, we aim to investigate the Internet mobile streaming service at the server side. For this purpose, we have collected a 4-month server-side log on the cloud (with 1,002 TB delivered video traffic) from a top Internet mobile streaming service provider serving worldwide mobile users. Through trace analysis, we find that 1) a major challenge for providing Internet mobile streaming services is rooted from the mobile device hardware and software heterogeneity. In this workload, we find over 3,400 different hardware models with more than 100 different screen resolutions running 14 different mobile OS and three audio codecs and four video codecs. 2) To deal with the device heterogeneity, CPU-intensive transcoding is used on the cloud to customize the video to the appropriate versions at runtime for different devices. A video clip could be transcoded into more than 40 different versions to serve requests from different devices. 3) Compared to videos in traditional Internet streaming, mobile streaming videos are typically of much smaller size (a median of 1.68 MBytes) and shorter duration (a median of 2.7 minutes). Furthermore, the daily mobile user accesses are more skewed following a Zipf-like distribution but users' interests also quickly shift. Considering the huge demand of CPU cycles for online transcoding, we further examine server-side caching to reduce the total CPU cycle demand from the cloud. We show that a policy considering different versions of a video altogether outperforms other intuitive ones when the cache size is limited.
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Bo Shen 0003, Songqing Chen, Yingjie Lan
IEEE Trans. Parallel Distributed Syst.5
2012 An experimental study of open-source cloud platforms for dust storm forecasting
abstract
Cloud computing is becoming a viable computing solution for scientific research and several open-source cloud solutions are available to support scientific studies. However, little has been done to systematically investigate the performance of these solutions in supporting scientific pursuits. Taking dust storm forecasting as an example, we test three popular open-source cloud solutions, namely Eucalyptus, OpenNebula, and CloudStack, on the same hardware and compare against a bare cluster. We find that: (1) compared to the bare cluster, a cloud has about 10% virtualization and management overhead when one virtual machine is used. Overhead increases when more virtual machines are used. Leveraging more virtual resources would not necessarily yield better performance. (2) For computing- and communication-intensive dust storm forecasting, the performance overhead is mainly due to virtualized network rather than virtualized computing resources when more than one virtual machine is involved. (3) Compared to Eucalyptus and CloudStack, OpenNebula provides better support for dust storm forecasting with relatively better performance. The results can provide some insights for scientific community in adopting these open-source cloud solutions.
Qunying Huang, Jizhe Xia, Chaowei Phil Yang, Kai Liu 0017, Jing Li 0029, Zhipeng Gui, Mohammed Anowarul Hassan, Songqing Chen
SIGSPATIAL/GIS8
2012 Spammer Behavior Analysis and Detection in User Generated Content on Social Networks
abstract
Spam content is surging with an explosive increase of user generated content (UGC) on the Internet. Spammers often insert popular keywords or simply copy and paste recent articles from the Web with spam links inserted, attempting to disable content-based detection. In order to effectively detect spam in user generated content, we first conduct a comprehensive analysis of spamming activities on a large commercial UGC site in 325 days covering over 6 million posts and nearly 400 thousand users. Our analysis shows that UGC spammers exhibit unique non-textual patterns, such as posting activities, advertised spam link metrics, and spam hosting behaviors. Based on these non-textual features, we show via several classification methods that a high detection rate could be achieved offline. These results further motivate us to develop a runtime scheme, BARS, to detect spam posts based on these spamming patterns. The experimental results demonstrate the effectiveness and robustness of BARS.
Enhua Tan, Lei Guo 0004, Songqing Chen, Xiaodong Zhang 0001, Yihong Eric Zhao
ICDCS3
2012 A server's perspective of Internet streaming delivery to mobile devices
abstract
Receiving Internet streaming services on various mobile devices is getting more and more popular. To understand and better support Internet streaming delivery to mobile devices, a number of studies have been conducted. However, existing studies have mainly focused on the client side resource consumption and streaming quality. So far, little is known about the server side, which is the key for providing successful mobile streaming services. In this work, we set to investigate the Internet mobile streaming service at the server side. For this purpose, we have collected a one-month server log (with 212 TB delivered video traffic) from a top Internet mobile streaming service provider serving worldwide mobile users. Through trace analysis, we find that (1) a major challenge for providing Internet mobile streaming services is rooted from the mobile device hardware and software heterogeneity. In this workload, we find over 2800 different hardware models with about 100 different screen resolutions running 14 different mobile OS and 3 audio codecs and 4 video codecs. (2) To deal with the device heterogeneity, transcoding is used to customize the video to the appropriate versions at runtime for different devices. A video clip could be transcoded into more than 40 different versions in order to serve requests from different devices. (3) Compared to videos in traditional Internet streaming, mobile streaming videos are typically of much smaller size (a median of 1.68 MBytes) and shorter duration (a median of 2.7 minutes). Furthermore, the daily mobile user accesses are more skewed following a Zipf-like distribution but users' interests also quickly shift. Considering the huge demand of CPU cycles for online transcoding, we further examine server-side caching in order to reduce CPU cycle demand. We show that a policy considering different versions of a video altogether outperforms other intuitive ones when the cache size is limited.
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Bo Shen 0003, Songqing Chen
INFOCOM5
2012 Chrome Extensions: Threat Analysis and Countermeasures
Lei Liu 0021, Xinwen Zhang, Guanhua Yan, Songqing Chen
NDSS4
2012 ACM/Springer Mobile Networks and Applications (MONET) Special Issue on "Collaborative Computing: Networking, Applications and Worksharing"
Songqing Chen, Le Gruenwald, James B. D. Joshi, Karl Aberer
Mob. Networks Appl.1
2012 Building an efficient transcoding overlay for P2P streaming to heterogeneous devices
abstract
With the increasing deployment of Internet P2P/overlay streaming systems, more and more clients use mobile devices, such as smart phones and PDAs, to access these Internet streaming services. Compared to wired desktops, mobile devices normally have a smaller screen size, a less color depth, and lower bandwidth and thus cannot correctly and effectively render and display the data streamed to desktops. To address this problem, in this paper, we propose PAT (Peer-Assisted Transcoding) to enable effective online transcoding in P2P/overlay streaming. PAT has the following unique features. First, it leverages active peer cooperation without demanding infrastructure support such as transcoding servers. Second, as online transcoding is computationally intensive while the various devices used by participating clients may have limited computing power and related resources (e.g., battery, bandwidth), an additional overlay, called metadata overlay, is constructed to instantly share the intermediate transcoding result of a transcoding procedure with other transcoding nodes to minimize the total computing overhead in the system. The experimental results collected within a realistically simulated testbed show that by consuming 6% extra bandwidth, PAT could save up to 58% CPU cycles for online transcoding.
Dongyu Liu, Fei Li 0001, Bo Shen 0003, Songqing Chen
ACM Trans. Multim. Comput. Commun. Appl.4
2011 RatBot: Anti-enumeration Peer-to-Peer Botnets
Guanhua Yan, Songqing Chen, Stephan J. Eidenbenz
ISC2
2011 An empirical evaluation of battery power consumption for streaming data transmission to mobile devices
abstract
Internet streaming applications are becoming increasingly popular on mobile devices. However, receiving streaming services on mobile devices is often constrained by their limited battery power supply. Various techniques have been proposed to save battery power consumption on mobile devices, mainly focusing on how much data to transmit and how to transmit.
Yao Liu 0001, Lei Guo 0004, Fei Li 0001, Songqing Chen
ACM Multimedia4
2011 BlueStreaming: towards power-efficient internet P2P streaming to mobile devices
abstract
P2P streaming applications are very popular on the Internet today. However, a mobile device in P2P streaming not only needs to continuously receive streaming data from other peers for its playback, but also needs to continuously exchange control information (e.g., buffermaps and file chunk requests) with neighboring peers and upload the downloaded streaming data to them. These lead to excessive battery power consumption on the mobile device.
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Yang Guo 0001, Songqing Chen
ACM Multimedia5
2011 A measurement study of resource utilization in internet mobile streaming
abstract
The pervasive usage of mobile devices and wireless networking support have enabled more and more Internet stream- ing services to all kinds of heterogeneous mobile devices. However, Internet mobile streaming services are challenged by the inherently limited on-device resources, device heterogeneity, and the bulk amount of streaming data.
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Songqing Chen
NOSSDAV4
2011 Some special minimum k-geodetically connected graphs
Yingjie Lan, Songqing Chen
Discret. Appl. Math.2
2010 TopBT: A Topology-Aware and Infrastructure-Independent BitTorrent Client
abstract
BitTorrent (BT) has carried out a significant and continuously increasing portion of Internet traffic. While several designs have been recently proposed and implemented to improve the resource utilization by bridging the application layer (overlay) and the network layer (underlay), these designs are largely dependent on Internet infrastructures, such as ISPs and CDNs. In addition, they also demand large-scale deployments of their systems to work effectively. Consequently, they require multiefforts far beyond individual users' ability to be widely used in the Internet. In this paper, aiming at building an infrastructure-independent user-level facility, we present our design, implementation, and evaluation of a topology-aware BT system, called TopBT, to significantly improve the overall Internet resource utilization without degrading user downloading performance. The unique feature of TopBT client lies in that a TopBT client actively discovers network proximities (to connected peers), and uses both proximities and transmission rates to maintain fast downloading while reducing the transmitting distance of the BT traffic and thus the Internet traffic. As a result, a TopBT client neither requires feeds from major Internet infrastructures, such as ISPs or CDNs, nor requires large-scale deployment of other TopBT clients on the Internet to work effectively. We have implemented TopBT based on widely used open-source BT client code base, and made the software publicly available. By deploying TopBT and other BitTorrent clients on hundreds of Internet hosts, we show that on average TopBT can reduce about 25% download traffic while achieving a 15% faster download speed compared to several prevalent BT clients. TopBT has been widely used in the Internet by many users all over the world.
Shansi Ren, Enhua Tan, Songqing Chen, Lei Guo 0004, Xiaodong Zhang 0001
INFOCOM4
2010 Online learning approaches in maximizing weighted throughput
abstract
Motivated by providing quality-of-service for next generation IP-based networks, we design algorithms to schedule packets with values and deadlines. Packets arrive over time; each packet has a non-negative value and an integer deadline. In each time step, at most one packet can be sent. Packets can be dropped at any time before they are sent. The objective is to maximize the total value gained by delivering packets no later than their respective deadlines. This model is the well-studied bounded-delay model (Hajek. CISS 2001. Kesselman et al. SICOMP 2004) which extensive competitive online algorithms have been developed for. In a generalization of this model, the success of delivering a packets in each time step depends on the reliability of the communication channel. In this paper, we apply online learning approaches on this model as well as a few of its variants. We design online learning algorithms and analyze their performance theoretically in terms of external regret. We also measure these algorithms' performance experimentally. We conclude that no online learning algorithms have a constant regret. Our online learning algorithms outperform the competitive algorithms for algorithmic simplicity and running complexity. However, in general, this online learning algorithms work no worse than the best known competitive online algorithm for maximizing weighted throughput in practice.
Zhi Zhang 0010, Fei Li 0001, Songqing Chen
IPCCC3
2010 Reducing data request contentions for improved streaming quality
abstract
In P2P assisted multi-channel live streaming systems, it is commonly believed that in unpopular channels, quality degradation is due to the small number of participating peers with almost-the-same set of available data; this phenomena prevents effective data exchanges among peers themselves and automatically leads to data request contentions once a new data chunk becomes available. In popular programs, our measurement on PPLive for a continuous three-month period at various locations also shows numerous occurrences of quality degradation because of the even higher ratio (up to 190%) of repetitive data requests for the same data chunks.
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Songqing Chen
NOSSDAV4
2010 An Application-Level Data Transparent Authentication Scheme without Communication Overhead
abstract
With abundant aggregate network bandwidth, continuous data streams are commonly used in scientific and commercial applications. Correspondingly, there is an increasing demand of authenticating these data streams. Existing strategies explore data stream authentication by using message authentication codes (MACs) on a certain number of data packets (a data block) to generate a message digest, then either embedding the digest into the original data, or sending the digest out-of-band to the receiver. Embedding approaches inevitably change the original data, which is not acceptable under some circumstances (e.g., when sensitive information is included in the data). Sending the digest out-of-band incurs additional communication overhead, which consumes more critical resources (e.g., power in wireless devices for receiving information) besides network bandwidth. In this paper, we propose a novel strategy, DaTA, which effectively authenticates data streams by selectively adjusting some interpacket delay. This authentication scheme requires no change to the original data and no additional communication overhead. Modeling-based analysis and experiments conducted on an implemented prototype system in an LAN and over the Internet show that our proposed scheme is efficient and practical.
Songqing Chen, Shiping Chen 0003, Xinyuan Wang 0005, Zhao Zhang 0010, Sushil Jajodia
IEEE Trans. Computers1
2009 Malyzer: Defeating Anti-detection for Application-Level Malware Analysis
Lei Liu 0021, Songqing Chen
ACNS2
2009 Exploitation and threat analysis of open mobile devices
abstract
The increasingly open environment of mobile computing systems such as PDAs and smartphones brings rich applications and services to mobile users. Accompanied with this trend is the growing malicious activities against these mobile systems, such as information leakage, service stealing, and power exhaustion. Besides the threats posed against individual mobile users, these unveiled mobile devices also open the door for more serious damage such as disabling critical public cyber physical systems that are connected to the mobile/wireless infrastructure. The impact of such attacks, however, has not been fully recognized.
Lei Liu 0021, Xinwen Zhang, Guanhua Yan, Songqing Chen
ANCS4
2009 Run-Time Detection of Malwares via Dynamic Control-Flow Inspection
abstract
Conventional approach of detecting malwares relies on static scanning of malware signature. However, it may not work on the malwares that use software protection methods such as encryption and packing with run-time decryption and unpacking. We propose a hardware-assisted malware detection system that detects malwares during program run time to complement the conventional approach. It searches for control flow-based signature of malware during program execution, therefore bypassing the protection method used by those malwares. A new hardware design is used to assist the collection of control flow information. We have implemented and evaluated a prototype system on top of a full-system simulator based on the Intel x86 architecture. The experimental results show that the system can successfully distinguish all 30 malware variants and other benign programs that we have randomly collected, and that the overall run-time performance overhead is negligible. In short, the study demonstrates that it is a viable approach to detect malware in run time using control flow-based signature.
Yong-Joon Park, Zhao Zhang 0010, Songqing Chen
ASAP3
2009 A Case Study of Traffic Locality in Internet P2P Live Streaming Systems
abstract
With the ever-increasing P2P Internet traffic, recently much attention has been paid to the topology mismatch between the P2P overlay and the underlying network due to the large amount of cross-ISP traffic. Mainly focusing on BitTorrent-like file sharing systems, several recent studies have demonstrated how to efficiently bridge the overlay and the underlying network by leveraging the existing infrastructure, such as CDN services or developing new application-ISP interfaces, such as P4P. However, so far the traffic locality in existing P2P live streaming systems has not been well studied. In this work, taking PPLive as an example, we examine traffic locality in Internet P2P streaming systems. Our measurement results on both popular and unpopular channels from various locations show that current PPLive traffic is highly localized at the ISP level. In particular, we find: (1) a PPLive peer mainly obtains peer lists referred by its connected neighbors (rather than tracker servers) and up to 90% of listed peers are from the same ISP as the requesting peer; (2) the major portion of the streaming traffic received by a requesting peer (up to 88% in popular channels) is served by peers in the same ISP as the requestor; (3) the top 10\% of the connected peers provide most (about 70%) of the requested streaming data and these top peers have smaller RTT to the requesting peer. Our study reveals that without using any topology information or demanding any infrastructure support, PPLive achieves such high ISP level traffic locality spontaneously with its decentralized, latency based, neighbor referral peer selection strategy. These findings provide some new insights for better understanding and optimizing the network- and user-level performance in practical P2P live streaming systems.
Yao Liu 0001, Lei Guo 0004, Fei Li 0001, Songqing Chen
ICDCS4
2009 Towards Optimal Resource Utilization in Heterogeneous P2P Streaming
abstract
Though plenty of research has been conducted to improve Internet P2P streaming quality perceived by end-users, little has been known about the upper bounds of achievable performance with available resources so that different designs could compare against. On the other hand, the current practice has shown increasing demand of server capacities in P2P-assisted streaming systems in order to maintain high-quality streaming to end-users. Both research and practice call for a design that can optimally utilize available peer resources. In the paper, we first present a new design, aiming to reveal the best achievable throughput for heterogeneous P2P streaming systems. We measure the performance gaps between various designs and this optimal resource allocation. Through extensive simulations, we find out that several typical existing designs have not fully exploited the potential of system resources. However, the control overhead prohibits the adoption of this optimal approach. Then, we design a hybrid system in trading off the cost of assignment and utilization of resources. This hybrid approach has a proved theoretical bound on efficiency of utilization. Simulation results show that compared with the optimal resource allocation, our proposed hybrid design can achieve near-optimal (up to 90%) utilization while only use much less (below 4%) control overhead. Our results provide a basis for both server capacity planning in current P2P-assisted streaming practice and future protocol designs.
Dongyu Liu, Fei Li 0001, Songqing Chen
ICDCS3
2009 CUBS: Coordinated Upload Bandwidth Sharing in Residential Networks
abstract
Millions of residential users are widely served by cable or DSL connections with modest upload bandwidth and relatively high download bandwidth. For the increasingly important and demanding P2P applications such as VoIP, BitTorrent, and Internet streaming, stable or high upload bandwidth is required. Inadequate upload bandwidth degrades the performance of these applications among residential users. On the other hand, our Internet measurements show that plenty of idle upload bandwidth (from 50% to 80%) is always available in a local residential network. Based on this observation, we propose a system prototype to Coordinate Upload Bandwidth Sharing (CUBS) among neighboring residential users. Specifically, the idle upload bandwidth of neighbors can be used upon a request from a demanding user. Since it has become a common practice to deploy wireless access points in a residential user's home, we have built CUBS by leveraging the support from the wireless networks. In CUBS, to discover and manage idle bandwidth, a localized overlay is constructed by the cooperative users. CUBS is application independent as the bandwidth sharing is implemented at the network layer. CUBS is also ISP transparent because the sharing of neighbors' bandwidth does not demand any additional bandwidth supplies. We have evaluated the CUBS system prototype with experiments on Internet. The experimental results demonstrate that CUBS can effectively improve the performance of upload intensive applications by more than 30%.
Enhua Tan, Lei Guo 0004, Songqing Chen, Xiaodong Zhang 0001
ICNP3
2009 Analyzing patterns of user content generation in online social networks
abstract
Various online social networks (OSNs) have been developed rapidly on the Internet. Researchers have analyzed different properties of such OSNs, mainly focusing on the formation and evolution of the networks as well as the information propagation over the networks. In knowledge-sharing OSNs, such as blogs and question answering systems, issues on how users participate in the network and how users "generate/contribute" knowledge are vital to the sustained and healthy growth of the networks. However, related discussions have not been reported in the research literature.
Lei Guo 0004, Enhua Tan, Songqing Chen, Xiaodong Zhang 0001, Yihong Eric Zhao
KDD3
2009 VirusMeter: Preventing Your Cellphone from Spies
Lei Liu 0021, Guanhua Yan, Xinwen Zhang, Songqing Chen
RAID4
2008 BotTracer: Execution-Based Bot-Like Malware Detection
Lei Liu 0021, Songqing Chen, Guanhua Yan, Zhao Zhang 0010
ISC2
2008 Dynamic Balancing of Packet Filtering Workloads on Distributed Firewalls
abstract
Firewalls are widely deployed nowadays to enforce security policies of enterprise networks. While having played crucial roles in securing these networks, firewalls themselves are subject to performance limitations. An overloaded firewall can cause severe damage to the protected enterprise network, because any legitimate communication through it is either degraded or even completely severed. In this paper, we address how to dynamically balance packet filtering workloads on distributed firewalls efficiently in large enterprise networks. We model dynamic load balancing on distributed firewalls as a minimax optimization problem, and show that it is strongly NP-complete even if we eliminate all precedence relationships among policy rules by rule rewriting. Accordingly, we propose a light-weight rule distribution scheme that quickly balances workloads among all firewalls. Our scheme is adaptive to incoming traffic. Moreover, dynamically placing and ordering policy rules on distributed firewalls reduces the probability that attackers successfully infer the rule distribution. Experimental results show that using a commodity PC, our approach can reduce the peak firewall workload in distributed firewall systems by 40% within less than five minutes, compared against alternative solutions that only optimize rule ordering on individual firewalls.
Guanhua Yan, Songqing Chen, Stephan J. Eidenbenz
IWQoS2
2008 The stretched exponential distribution of internet media access patterns
abstract
The commonly agreed Zipf-like access pattern of Web workloads is mainly based on Internet measurements when text-based content dominated the Web traffic. However, with dramatic increase of media traffic on the Internet, the inconsistency between the access patterns of media objects and the Zipf model has been observed in a number of studies. An insightful understanding of media access patterns is essential to guide Internet system design and management, including resource provisioning and performance optimizations.
Lei Guo 0004, Enhua Tan, Songqing Chen, Xiaodong Zhang 0001
PODC3
2008 Modeling and Optimization of Meta-Caching Assisted Transcoding
abstract
The increase of aggregate Internet bandwidth and the rapid development of 3G wireless networks demand efficient delivery of multimedia objects to all types of wireless devices. To handle requests from wireless devices at runtime, the transcoding-enabled caching proxy has been proposed to save transcoded versions to reduce the intensive computing demanded by online transcoding. Constrained by available CPU and storage, existing transcoding-enabled caching schemes always selectively cache certain transcoded versions, expecting that many future requests can be served from the cache. But such schemes treat the transcoder as a black box, leaving no room for flexible control of joint resource management between CPU and storage. In this paper, we first introduce the idea of meta-caching by looking into a transcoding procedure. Instead of caching certain selected transcoded versions in full, meta-caching identifies intermediate transcoding steps from which certain intermediate results (calledmetadata) can be cached so that a fully transcoded version can be easily produced from the metadata with a small amount of CPU cycles. Achieving big saving in caching space with possibly small sacrifice on CPU load, the proposed meta-caching scheme provides a unique method to balance the utilization of CPU and storage resources at the proxy. We further construct a model to analyze the meta-caching scheme. Based on the analysis, we proposeAMTrac,AdaptiveMeta-caching forTranscoding, which adaptively applies meta-caching based on the client request patterns and available resources. Experimental results show that AMTrac can significantly improve the system throughput over existing approaches.
Dongyu Liu, Songqing Chen, Bo Shen 0003
IEEE Trans. Multim.2
2008 Achieving simultaneous distribution control and privacy protection for Internet media delivery
abstract
Massive Internet media distribution demands prolonged continuous consumption of networking and disk bandwidths in large capacity. Many proxy-based Internet media distribution algorithms and systems have been proposed, implemented, and evaluated to address the scalability and performance issue. However, few of them have been used in practice, since two important issues are not satisfactorily addressed. First, existing proxy-based media distribution architectures lack an efficient media distribution control mechanism. Without copyright protection, content providers are hesitant to use proxy-based fast distribution techniques. Second, little has been done to protect client privacy during content accesses on the Internet. Straightforward solutions to address these two issues independently lead to conflicts. For example, to enforce distribution control, only legitimate users should be granted access rights. However, this normally discloses more information (such as which object the client is accessing) other than the client identity, which conflicts with the client's desire for privacy protection. In this article, we propose a unified proxy-based media distribution protocol to effectively address these two problems simultaneously. We further design a set of new algorithms in a cooperative proxy environment where our proposed scheme works efficiently and practically. Simulation-based experiments are conducted to extensively evaluate the proposed system. Preliminary results demonstrate the effectiveness of our proposed strategy.
Songqing Chen, Shiping Chen 0003, Huiping Guo, Bo Shen 0003, Sushil Jajodia
ACM Trans. Multim. Comput. Commun. Appl.1
2007 SecureBus: towards application-transparent trusted computing with mandatory access control
abstract
The increasing number of software-based attacks has attracted substantial efforts to prevent applications from malicious interference. For example, Trusted Computing (TC) technologies have been recently proposed to provide strong isolation on application platforms. On the other hand, today pervasively available computing cycles and data resources have enabled various distributed applications that require collaboration among different application processes. These two conflicting trends grow in parallel. While much existing research focuses on one of these two aspects, a few authors have considered simultaneously providing strong isolation as well as collaboration convenience, particularly in the TC environment. However, none of these schemes is transparent. That is, they require modifications either of legacy applications or the underlying Operating System (OS).In this paper, we propose the SecureBus (SB) architecture, aiming to provide strong isolation and flexible controlled information flow and communication between processes at runtime. Since SB is application and OS transparent, existing applications can run without changes to commodity OS's. Furthermore, SB enables the enforcement of general access control policies, which is required but difficult to achieve for typical legacy applications. To study its feasibility and performance overhead, we have implemented a prototype system based on User-Mode Linux. Our experimental results show that SB can effectively achieve its design goals.
Xinwen Zhang, Michael J. Covington, Songqing Chen, Ravi S. Sandhu
AsiaCCS3
2007 SCAP: Smart Caching inWireless Access Points to Improve P2P Streaming
abstract
The increasing number of wireless users in Internet P2P applications causes two new performance problems due to the requirement of uploading the downloaded traffic for other peers, limited bandwidth of wireless communications, and resource competition between the access point and wireless stations. First, an active P2P wireless user can significantly reduce the downloading throughput of other wireless users in the WLAN. Second, the slowdown of a P2P wireless user communication can also delay its relay and data sharing service for other dependent wired/wireless peers. In order to address these problems, in this paper, we propose an efficient caching mechanism called SCAP (Smart Caching in Access Points). Conducting intensive Internet measurements on representative P2P streaming applications, we observe a high percentage of duplicated data packets in successive downloading and uploading data streams. Through duplication detection and caching at the access point, these duplicated packets can be compressed so that the uploading traffic in the WLAN is significantly reduced. Our prototype-based experimental evaluation demonstrates that by effectively reducing the redundant P2P traffic in the WLAN, SCAP improves the throughput of the WLAN by up to 88% and reduces the response delay to other Internet users meanwhile.
Enhua Tan, Lei Guo 0004, Songqing Chen, Xiaodong Zhang 0001
ICDCS3
2007 PSM-throttling: Minimizing Energy Consumption for Bulk Data Communications in WLANs
abstract
While the 802.11 power saving mode (PSM) and its enhancements can reduce power consumption by putting the wireless network interface (WNI) into sleep as much as possible, they either require additional infrastructure support, or may degrade the transmission throughput and cause additional transmission delay. These schemes are not suitable for long and bulk data transmissions with strict QoS requirements on wireless devices. With increasingly abundant bandwidth available on the Internet, we have observed that TCP congestion control is often not a constraint of bulk data transmissions as bandwidth throttling is widely used in practice. In this paper, instead of further manipulating the trade-off between the power saving and the incurred delay, we effectively explore the power saving potential by considering the bandwidth throttling on streaming/downloading servers. We propose an application-independent protocol, called PSM-throttling. With a quick detection on the TCP flow throughput, a client can identify bandwidth throttling connections with a low cost Since the throttling enables us to reshape the TCP traffic into periodic bursts with the same average throughput as the server transmission rate, the client can accurately predict the arriving time of packets and turn on/off the WNI accordingly. PSM-throttling can minimize power consumption on TCP-based bulk traffic by effectively utilizing available Internet bandwidth without degrading the application's performance perceived by the user. Furthermore, PSM-throttling is client-centric, and does not need any additional infrastructure support. Our lab-environment and Internet-based evaluation results show that PSM-throttling can effectively improve energy savings (by up to 75%) and/or the QoS for a broad types of TCP-based applications, including streaming, pseudo streaming, and large file downloading, over existing PSM-like methods.
Enhua Tan, Lei Guo 0004, Songqing Chen, Xiaodong Zhang 0001
ICNP3
2007 Does internet media traffic really follow Zipf-like distribution?
abstract
No abstract available.
Lei Guo 0004, Enhua Tan, Songqing Chen, Xiaodong Zhang 0001
SIGMETRICS3
2007 A performance study of BitTorrent-like peer-to-peer systems
abstract
This paper presents a performance study of BitTorrent-like P2P systems by modeling, based on extensive measurements and trace analysis. Existing studies on BitTorrent systems are single-torrent based and usually assume the process of request arrivals to a torrent is Poisson-like. However, in reality, most BitTorrent peers participate in multiple torrents and file popularity changes over time. Our study of representative BitTorrent traffic provides insights into the evolution of single-torrent systems and several new findings regarding the limitations of BitTorrent systems: (1) Due to the exponentially decreasing peer arrival rate in a torrent, the service availability of the corresponding file becomes poor quickly, and eventually it is hard to locate and download this file. (2) Client performance in the BitTorrent-like system is unstable, and fluctuates significantly with the changes of the number of online peers. (3) Existing systems could provide unfair services to peers, where a peer with a higher downloading speed tends to download more and upload less. Motivated by the analysis and modeling results, we have further proposed a graph based model to study interactions among multiple torrents. Our model quantitatively demonstrates that inter-torrent collaboration is much more effective than stimulating seeds to serve longer for addressing the service unavailability in BitTorrent systems. An architecture for inter-torrent collaboration under an exchange based instant incentive mechanism is also discussed and evaluated by simulations.
Lei Guo 0004, Songqing Chen, Enhua Tan, Xiaoning Ding, Xiaodong Zhang 0001
IEEE J. Sel. Areas Commun.2
2007 Cooperative Relay Service in a Wireless LAN
abstract
As a family of wireless local area network (WLAN) protocols between physical layer and higher layer protocols, IEEE 802.11 has to accommodate the features and requirements of both ends. However, current practice has addressed the problems of these two layers separately and is far from satisfactory. On one end, due to varying channel conditions, WLANs have to provide multiple physical channel rates to support various signal qualities. A low channel rate station not only suffers low throughput, but also significantly degrades the throughput of other stations. On the other end, the power saving mechanism of 802.11 is ineffective in TCP-based communications, in which the wireless network interface (WNI) has to stay awake to quickly acknowledge senders, and hence, the energy is wasted on channel listening during idle awake time. In this paper, considering the needs of both ends, we utilize the idle communication power of the WNI to provide a Cooperative Relay Service (CRS) for WLANs with multiple channel rates. We characterize energy efficiency as energy per bit, instead of energy per second. In CRS, a high channel rate station relays data frames as a proxy between its neighboring stations with low channel rates and the Access Point, improving their throughput and energy efficiency. Different from traditional relaying approaches, CRS compensates a proxy for the energy consumed in data forwarding. The proxy obtains additional channel access time from its clients, leading to the increase of its own throughput without compromising its energy efficiency. Extensive experiments are conducted through a prototype implementation and ns-2 simulations to evaluate our proposed CRS. The experimental results show that CRS achieves significant performance improvements for both low and high channel rate stations
Lei Guo 0004, Xiaoning Ding, Haining Wang 0001, Qun Li 0001, Songqing Chen, Xiaodong Zhang 0001
IEEE J. Sel. Areas Commun.5
2007 SProxy: A Caching Infrastructure to Support Internet Streaming
abstract
Many algorithmic efforts have been made to address technical issues in designing a streaming media caching proxy. Typical of those are segment-based caching approaches that efficiently cache large media objects in segments which reduces the startup latency while ensuring continuous streaming. However, few systems have been practically implemented and deployed. The implementation and deployment efforts are hindered by several factors: 1) streaming of media content in complicated data formats is difficult; 2) typical streaming protocols such as RTP often run on UDP; in practice, UDP traffic is likely to be blocked by firewalls at the client side due to security considerations; and 3) coordination between caching discrete object segments and streaming continuous media data is challenging. To address these problems, we have designed and implemented a segment-based streaming media proxy, called SProxy. This proxy system has the following merits. First, SProxy leverages existing Internet infrastructure to address the flash crowd. The content server is now free of the streaming duty while hosting streaming content through a regular Web server. Thus, UDP based streaming traffic from SProxy suffers less dropping and no blocking. Second, SProxy streams and caches media objects in small segments determined by the object popularity, causing very low startup latency, and significantly reducing network traffic. Finally, prefetching techniques are used to pro-actively preload uncached segments that are likely to be used soon, thus providing continuous streaming. SProxy has been extensively tested and we show that it provides high quality streaming delivery in both local area networks and wide area networks (e.g., between Japan and the U.S.).
Songqing Chen, Bo Shen 0003, Susie J. Wee, Xiaodong Zhang 0001
IEEE Trans. Multim.1
2006 V-COPS: A Vulnerability-Based Cooperative Alert Distribution System
abstract
The efficiency of promptly releasing security alerts of established analysis centers has been greatly challenged by the continuous emergence of various large scale network attacks, such as worms. With a limited number of sensors deployed over the Internet and a long attack verification period, when the alert is released by analysis centers, the best time to stop the attack may have passed. On the other hand, (1) most of the past large scale attacks targeted known vulnerabilities, and (2) today numerous Internet systems have integrated detection tools, such as virus detection software and intrusion detection systems (IDS), the power of which could be harnessed to defend against large scale attacks. In this paper, we propose V-COPS - a vulnerability-based cooperative alert distribution system, by leveraging existing independent local attack detection systems. V-COPS is capable of promptly propagating genuine alerts with critical vulnerability information, based on which relevant stakeholders can take preventive actions in time. Extensive analysis and experiments have been performed to study the performance of V-COPS. The preliminary results show V-COPS is effective
Shiping Chen 0003, Dongyu Liu, Songqing Chen, Sushil Jajodia
ACSAC3
2006 WormTerminator: an effective containment of unknown and polymorphic fast spreading worms
abstract
The fast spreading worm is becoming one of the most serious threats to today's networked information systems. A fast spreading worm could infect hundreds of thousands of hosts within a few minutes. In order to stop a fast spreading worm, we need the capability to detect and contain worms automatically in real-time. While signature based worm detection and containment are effective in detecting and containing known worms, they are inherently ineffective against previously unknown worms and polymorphic worms. Existing traffic anomaly pattern based approaches have the potential to detect and/or contain previously unknown and polymorphic worms, but they either impose too much constraint on normal traffic or allow too much infectious worm traffic to go out to the Internet before an unknown or polymorphic worm can be detected.In this paper, we present WormTerminator, which can detect and completely contain, at least in theory, almost all fast spreading worms in real-time while blocking virtually no normal traffic. WormTerminator detects and contains the fast spreading worm based on its defining characteristic -- a fast spreading worm will start to infect others as soon as it successfully infects one host. WormTerminator also exploits the observation that a fast spreading worm keeps exploiting the same set of vulnerabilities when infecting new machines. To prove the concept, we have implemented a prototype of WormTerminator and have examined its effectiveness against the real Internet worm Linux/Slapper.
Songqing Chen, Xinyuan Wang 0005, Lei Liu 0021, Xinwen Zhang
ANCS1
2006 A Case for Internet Streaming via Web Servers
abstract
Hosting Internet streaming services has its unique challenges. Aiming at making Internet streaming services be widely and easily adopted in practice, in this paper, we have designed and implemented a system, called SProxy that can leverage existing Internet infrastructure to free the streaming content providers so that they only need to host streaming content through a regular Web server. SProxy has been extensively tested and evaluated and it provides high quality streaming delivery in both local area networks and wide area networks (e.g. between Japan and US)
Songqing Chen, Bo Shen 0003, Wai-tian Tan, Susie J. Wee, Xiaodong Zhang 0001
ICME1
2006 Delving into internet streaming media delivery: a quality and resource utilization perspective
abstract
Modern Internet streaming services have utilized various techniques to improve the quality of streaming media delivery. Despite the characterization of media access patterns and user behaviors in many measurement studies, few studies have focused on the streaming techniques themselves, particularly on the quality of streaming experiences they offer end users and on the resources of the media systems that they consume. In order to gain insights into current streaming services techniques and thus provide guidance on designing resource-efficient and high quality streaming media systems, we have collected a large streaming media workload from thousands of broadband home users and business users hosted by a major ISP, and analyzed the most commonly used streaming techniques such as automatic protocol switch, Fast Streaming, MBR encoding and rate adaptation. Our measurement and analysis results show that with these techniques, current streaming systems these techniques tend to over-utilize CPU and bandwidth resources to provide better services to end users, which may not be a desirable and effective is not necessary the best way to improve the quality of streaming media delivery. Motivated by these results, we propose and evaluate a coordination mechanism that effectively takes advantage of both Fast Streaming and rate adaptation to better utilize the server and Internet resources for streaming quality improvement.
Lei Guo 0004, Enhua Tan, Songqing Chen, Oliver Spatscheck, Xiaodong Zhang 0001
Internet Measurement Conference3
2006 Exploiting Idle Communication Power to Improve Wireless Network Performance and Energy Efficiency
abstract
Abstract — As a family of wireless local area network (WLAN) protocols between physical layer and higher-layer protocols, IEEE 802.11 has to accommodate the features and requirements of both ends. However, current practice has addressed the problems separately and is far from being satisfactory. On the one end, due to varying channel conditions, WLANs have to provide multiple data channel rates to support various bit error rates. A low channel rate station not only suffers low throughput itself, but also significantly degrades the throughput of other stations. On the other end, TCP is not energy efficient running on 802.11. This is because a wireless network interface (WNI) has to stay awake to generate timely acknowledgments during a TCP session, and hence, the energy consumed during idle awake time is wasted for channel listening. In this paper, considering the needs of both ends, we utilize the idle communication power of the WNI to improve the throughput and energy efficiency of stations in WLANs supporting multiple channel rates. We characterize the energy efficiency as energy per bit, instead of energy per second. Based on modeling and analysis, we propose a data forwarding mechanism and an energy-aware channel allocation mechanism. In such a system, a high channel rate station relays data frames between its neighboring stations with low channel rates and Access Point, improving their throughput and energy efficiency. Different from traditional relaying approaches, our scheme compensates for the energy consumption for data forwarding. The forwarding station gets additional channel access time from its beneficiaries, leading to the increase of its own throughput without compromising its energy efficiency. We implement a prototype of our proposed system and evaluate it through extensive experiments. Our results show significant performance improvements for both low and high channel rate stations. I.
Lei Guo 0004, Xiaoning Ding, Haining Wang 0001, Qun Li 0001, Songqing Chen, Xiaodong Zhang 0001
INFOCOM5
2006 Efficient Proxy-Based Internet Media Distribution Control and Privacy Protection Infrastructure
abstract
Massive Internet media distribution demands pro longed continuous consumption of networking and disk band widths in large capacity. Many proxy-based Internet media distribution algorithms and systems have been proposed, implemented, and evaluated to address the scalability issue. However, few of them have been used in practice, since two important issues are not satisfactorily addressed. First, existing proxy-based media distribution architectures lack an efficient media distribution control mechanism. Without protection on the Internet, content providers are hesitant to use existing fast distribution techniques. Second, little has been done to protect client privacy during client accesses. Straightforward solutions to address these two issues independently lead to conflicts. For example, to enforce distribution control, only legitimate users should be granted access rights. However, this normally discloses more information (such as which object the client is accessing) other than the client identity, which conflicts with the client's desire for privacy protection. In this paper, we propose a unified proxy-based media distribution protocol to effectively address these two problems simultaneously. We further design a set of new algorithms for cooperative proxies where our proposed scheme works practically. Simulation results show that our proposed strategy is efficient
Songqing Chen, Shiping Chen 0003, Huiping Guo, Bo Shen 0003, Sushil Jajodia
IWQoS1
2006 AMTrac: adaptive meta-caching for transcoding
abstract
The increase of aggregate Internet bandwidth and the rapid development of 3G wireless networks demand efficient delivery of multimedia objects to all types of wireless devices. To handle requests from wireless devices at runtime, the transcode-enabled caching proxy has been proposed and a lot of research has been conducted to study online transcoding. Since transcoding is a CPU-intensive task, the transcoded versions can be saved to reduce the CPU load for future requests. However, extensively caching all transcoded results can quickly exhaust cache space. Constrained by available CPU and storage, existing transcode-enabled caching schemes always selectively cache certain transcoded versions, expecting that many future requests can be served from the cache while leaving CPU cycles for online transcoding for other requests. But such schemes treat the transcoder as a black box, leaving little room for flexible control of joint resource management between CPU and storage. In this paper, we first introduce the idea of meta-caching by looking into a transcoding procedure. Instead of caching certain selected transcoded versions in full, meta-caching identifies intermediate transcoding steps from which certain intermediate results (called metadata) can be cached so that a fully transcoded version can be easily produced from the metadata with a small amount of CPU cycles. Achieving big saving in caching space with possibly small sacrifice on CPU load, the proposed meta-caching scheme provides a unique method to balance the utilization of CPU and storage resources at the proxy. We further construct a model to analyze the meta-caching scheme. Based on modeling results, we propose AMTrac, Adaptive Meta-caching for Transcoding, which adaptively applies meta-caching based on the client request pattern and available resources. Experimental results show that our proposed AMTrac can significantly improve the system throughput over existing approaches.
Dongyu Liu, Songqing Chen, Bo Shen 0003
NOSSDAV2
2006 Design and Evaluation of a Scalable and Reliable P2P Assisted Proxy for On-Demand Streaming Media Delivery
abstract
To efficiently deliver streaming media, researchers have developed technical solutions that fall into three categories, each of which has its merits and limitations. Infrastructure-based CDNs with dedicated network bandwidths and hardware supports can provide high-quality streaming services, but at a high cost. Server-based proxies are cost-effective but not scalable due to the limited proxy capacity in storage and bandwidth, and its centralized control also brings a single point of failure. Client-based P2P networks are scalable, but do not guarantee high-quality, streaming service due to the transient nature of peers. To address these limitations, we present a novel and efficient design of a scalable and reliable media proxy system assisted by P2P networks, called PROP. In the PROP system, the clients' machines in an intranet are self-organized into a structured P2P system to provide a large media storage and to actively participate in the streaming media delivery, where the proxy is also embedded as an important member to ensure the quality of streaming service. The coordination and collaboration in the system are efficiently conducted by our P2P management structure and replacement policies. Our system has the following merits: 1) It addresses both the scalability problem in centralized proxy systems and the unreliable service concern by only relying on the P2P sharing of clients. 2) The proposed content locating scheme can timely serve the demanded media data and fairly dispatch media streaming tasks in appropriate granularity across the system. 3) Based on the modeling and analysis, we propose global replacement policies for proxy and clients, which well balance the demand and supply of streaming data in the system, achieving a high utilization of peers' cache. We have comparatively evaluated our system through trace-driven simulations with synthetic workloads and with a real-life workload extracted from the media server logs in an enterprise network, which shows our design significantly improves the quality of media streaming and the system scalability.
Lei Guo 0004, Songqing Chen, Xiaodong Zhang 0001
IEEE Trans. Knowl. Data Eng.2
2006 Segment-based streaming media proxy: modeling and optimization
abstract
Researchers often use segment-based proxy caching strategies to deliver streaming media by partially caching media objects. The existing strategies mainly consider increasing the byte hit ratio and/or reducing the client perceived startup latency (denoted by the metric delayed startup ratio). However, these efforts do not guarantee continuous media delivery because the to-be-viewed object segments may not be cached in the proxy when they are demanded. The potential consequence is playback jitter at the client side due to proxy delay in fetching the uncached segments, which we call proxy jitter. Thus, for the best interests of clients, a correct model for streaming proxy system design should aim to minimize proxy jitter subject to reducing the delayed startup ratio and increasing the byte hit ratio. However, we have observed two major pairs of conflicting interests inherent in this model: (1) one between improving the byte hit ratio and reducing proxy jitter, and (2) the other between improving the byte hit ratio and reducing the delayed startup ratio. In this study, we first propose and analyze prefetching methods for in-time prefetching of uncached segments, which provides insights into the first pair of conflicting interests. Second, to address the second pair of the conflicting interests, we build a general model to analyze the performance tradeoff between the second pair of conflicting performance objectives. Finally, considering our main objective of minimizing proxy jitter and optimizing the two tradeoffs, we propose a new streaming proxy system called Hyper Proxy. Synthetic and real workloads are used to evaluate our system. The performance results show that Hyper Proxy generates minimum proxy jitter with a low delayed startup ratio and a small decrease of byte hit ratio compared with existing schemes.
Songqing Chen, Bo Shen 0003, Susie J. Wee, Xiaodong Zhang 0001
IEEE Trans. Multim.1
2005 DISC: Dynamic Interleaved Segment Caching for Interactive Streaming
abstract
Streaming media objects have become widely used on the Internet, and the demand of interactive requests to these objects has increased dramatically. Typical interactive requests include fast forward and direct jumps. Unfortunately, most of existing streaming proxies are designed for sequential accesses, and only a few solutions have been proposed to maintain additional data structures in the proxy to support some interactive operations (such as fast forward) other than jumps, which are among the most common interactive requests from the clients. Focusing on interactive accesses, in this paper, we present an analysis of streaming media workload collected from thousands of broadband users hosted by a major ISP. Our analysis shows that jump accesses (48%) and pauses (51%) are the dominant client interactive requests and that jump accesses often suffer serious delays due to slow buffering through the network. To support jump accesses effectively, we further propose a novel caching algorithm - DISC (dynamic interleaved segment caching), which trades cache performance for response time to client interactive requests. In this algorithm, segments of a media object are cached dynamically according to client access patterns. DISC can support direct jumps efficiently while ensuring timely prefetching of uncached segments for sequential accesses. Trace-driven simulations demonstrate that DISC outperforms other caching schemes significantly for interactive requests with only a small degradation in cache performance
Lei Guo 0004, Songqing Chen, Xiaodong Zhang 0001
ICDCS2
2005 Measurements, Analysis, and Modeling of BitTorrent-like Systems
Lei Guo 0004, Songqing Chen, Enhua Tan, Xiaoning Ding, Xiaodong Zhang 0001
Internet Measurement Conference2
2005 Analysis of multimedia workloads with implications for internet streaming
abstract
In this paper, we study the media workload collected from a large number of commercial Web sites hosted by a major ISP and that collected from a large group of home users connected to the Internet via a well-known cable company. Some of our key findings are: (1) Surprisingly, the majority of media contents are still delivered via downloading from Web servers. (2) A substantial percentage of media downloading connections are aborted before completion due to the long waiting time. (3) A hybrid approach, pseudo streaming, is used by clients to imitate real streaming. (4) The mismatch between the downloading rate and the client playback speed in pseudo streaming is common, which either causes frequent playback delays to the clients, or unnecessary traffic to the Internet. (5) Compared with streaming, downloading and pseudo streaming are neither bandwidth efficient nor performance effective. To address this problem, we propose the design of AutoStream, an innovative system that can provide additional previewing and streaming services automatically for media objects hosted on standard Web sites in server farms at the client's will.
Lei Guo 0004, Songqing Chen, Xiaodong Zhang 0001
WWW2
2005 Fast proxy delivery of multiple streaming sessions in shared running buffers
abstract
With the falling price of memory, an increasing number of multimedia servers and proxies are now equipped with a large memory space. Caching media objects in the memory of a proxy helps to reduce the network traffic, the disk I/O bandwidth requirement, and the data delivery latency. The running buffer approach and its alternatives are representative techniques to caching streaming data in the memory. There are two limits in the existing techniques. First, although multiple running buffers for the same media object co-exist in a given processing period, data sharing among multiple buffers is not considered. Second, user access patterns are not insightfully considered in the buffer management. In this paper, we propose two techniques based on shared running buffers in the proxy to address these limits. Considering user access patterns and characteristics of the requested media objects, our techniques adaptively allocate memory buffers to fully utilize the currently buffered data of streaming sessions, with the aim to reduce both the server load and the network traffic. Experimentally comparing with several existing techniques, we show that the proposed techniques achieve significant performance improvement by effectively using the shared running buffers.
Songqing Chen, Bo Shen 0003, Yong Yan 0003, Sujoy Basu, Xiaodong Zhang 0001
IEEE Trans. Multim.1
2004 SRB: Shared Running Buffers in Proxy to Exploit Memory Locality of Multiple Streaming Media Sessions
abstract
With the falling price of the memory, an increasing number of multimedia servers and proxies are now equipped with a large DRAM memory space. Caching media objects in the memory of a proxy helps to reduce network traffic, disk I/O bandwidth requirement, and data delivery latency. The running buffer approach and its alternatives are representative techniques to cache streaming data in the memory. However, there are two limits in the existing techniques. First, although multiple running buffers for the same media object co-exist in a given processing period, data sharing among the multiple buffers is not considered. Second, user access patterns are not insightfully considered in the buffer management. In this paper, we propose two techniques based on shared running buffers (SRB) in the proxy to address these limits. Considering user access patterns and characteristics of the requested media objects, our techniques adoptively allocate memory buffers to fully utilize the currently buffered data of streaming sessions, with the aim to reduce both the server load and the network traffic. Experimentally comparing with several existing techniques, we show that the proposed techniques have achieved significant performance improvement by effectively using the shared running buffers.
Songqing Chen, Bo Shen 0003, Yong Yan 0003, Sujoy Basu, Xiaodong Zhang 0001
ICDCS1
2004 PROP: A Scalable and Reliable P2P Assisted Proxy Streaming System
abstract
The demand of delivering streaming media content in the Internet has become increasingly high for scientific, educational, and commercial applications. Three representative technologies have been developed for this purpose, each of which has its merits and serious limitations. Infrastructure-based CDNs with dedicated network bandwidths and powerful media replicas can provide high quality streaming services but at a high cost. Server-based proxies are cost-effective but not scalable due to the limited proxy capacity and its centralized control. Client-based P2P networks are scalable but do not guarantee high quality streaming service due to the transient nature of peers. To address these limitations, we present a novel and efficient design of a scalable and reliable media proxy system supported by P2P networks. This system is called PROP abbreviated from our technical theme of "collaborating and coordinating PROxy and its P2P clients". Our objective is to address both scalability and reliability issues of streaming media delivery in a cost-effective way. In the PROP system, the clients' machines in an intranet are self-organized into a structured P2P system to provide a large media storage and to actively participate in the streaming media delivery, where the proxy is also embedded as an important member to ensure quality of streaming service. The coordination and collaboration in the system are efficiently conducted by our P2P management structure and replacement policies. We have comparatively evaluated our system by trace-driven simulations with synthetic workloads and with a real-life workload trace extracted from the media server logs in an enterprise network. The results show that our design significantly improves the quality of media streaming and the system scalability.
Lei Guo 0004, Songqing Chen, Shansi Ren, Xin Chen 0034, Song Jiang 0001
ICDCS2
2004 Designs of High Quality Streaming Proxy Systems
abstract
Researchers often use segment-based proxy caching strategies to deliver streaming media by partially caching media objects. The existing strategies mainly consider increasing the byte hit ratio and/or reducing the client perceived startup latency (denoted by the metric delayed startup ratio). However, these efforts do not guarantee continuous media delivery because the to-be-viewed object segments may not be cached in the proxy when they are demanded. The potential consequence is playback jitter at the client side due to proxy delay in fetching the uncached segments, which we call proxy jitter. Thus, for the best interests of clients, a correct model for streaming proxy system design should aim to minimize proxy jitter subject to reducing the delayed startup ratio and increasing the byte hit ratio. However, we have observed two major pairs of conflicting interests inherent in this model: (1) one between improving the byte hit ratio and reducing proxy jitter, and (2) the other between improving the byte hit ratio and reducing the delayed startup ratio. In this study, we first propose an active prefetching method for in-time prefetching of uncached segments, which provides insights into the first pair of conflicting interests. Second, we further improve our lazy-segmentation scheme which effectively addresses the second pair of the conflicting interests. Finally, considering our main objective of minimizing proxy jitter and optimizing the two trade-offs, we propose a new streaming proxy system called Hyper Proxy by effectively coordinating both prefetching and segmentation techniques. Synthetic and real workloads are used to systematically evaluate our system. The performance results show that the hyper proxy system generates minimum proxy jitter with a low delayed startup ratio and a small decrease of byte hit ratio compared with existing schemes.
Songqing Chen, Bo Shen 0003, Susie J. Wee, Xiaodong Zhang 0001
INFOCOM1
2004 Enforcing direct communications between clients and Web servers to improve proxy performance and security
abstract
Abstract The amount of dynamic Web contents and secured e‐commerce transactions has been dramatically increasing on the Internet, where proxy servers between clients and Web servers are commonly used for the purpose of sharing commonly accessed data and reducing Internet traffic. A significant and unnecessary Web access delay is caused by the overhead in proxy servers to process two types of accesses, namely dynamic Web contents and secured transactions, not only increasing response time, but also raising some security concerns. Conducting experiments on Squid proxy 2.3STABLE4, we have quantified the unnecessary processing overhead to show its significant impact on increased client access response times. We have also analyzed the technical difficulties in eliminating or reducing the processing overhead and the security loopholes based on the existing proxy structure. In order to address these performance and security concerns, we propose a simple but effective technique from the client side that adds a detector interfacing with a browser. With this detector, a standard browser, such as the Netscape/Mozilla, will have simple detective and scheduling functions, called a detective browser. Upon an Internet request from a user, the detective browser can immediately determine whether the requested content is dynamic or secured. If so, the browser will bypass the proxy and forward the request directly to the Web server; otherwise, the request will be processed through the proxy. We implemented a detective browser prototype in Mozilla version 0.9.7, and tested its functionality and effectiveness. Since we have simply moved the necessary detective functions from a proxy server to a browser, the detective browser introduces little overhead to Internet accessing, and our software can be patched to existing browsers easily. Copyright © 2004 John Wiley & Sons, Ltd.
Songqing Chen, Xiaodong Zhang 0001
Softw. Pract. Exp.1
2004 Building a Large and Efficient Hybrid Peer-to-Peer Internet Caching System
abstract
Proxy hit ratios tend to decrease as the demand and supply of Web contents are becoming more diverse. By case studies, we quantitatively confirm this trend and observe significant document duplications among a proxy and its client browsers' caches. One reason behind this trend is that the client/server Web caching model does not support direct resource sharing among clients, causing the Web contents and the network bandwidths among clients to be relatively underutilized. To address these limits and improve Web caching performance, we have extensively enhanced and deployed our browsers-aware framework, a peer-to-peer Web caching management scheme. We make the browsers and their proxy share the contents to exploit the neglected but rich data locality in browsers and reduce document duplications among the proxy and browsers' caches to effectively utilize the Web contents and network bandwidth among clients. The objective of our scheme is to improve the scalability of proxy-based caching both in the number of connected clients and in the diversity of Web documents. We show that building such a caching system with considerations of sharing contents among clients, minimizing document duplications, and achieving data integrity and communication anonymity is not only feasible but also highly effective.
Li Xiao 0001, Xiaodong Zhang 0001, Artur Andrzejak 0001, Songqing Chen
IEEE Trans. Knowl. Data Eng.4
2004 Adaptive Memory Allocations in Clusters to Handle Unexpectedly Large Data-Intensive Jobs
abstract
In a cluster system with dynamic load sharing support, a job submission or migration to a workstation is determined by the availability of CPU and memory resources of the workstation at the time (L. Xiao et al., 2002). In such a system, a small number of running jobs with unexpectedly large memory allocation requirements may significantly increase the queuing delay times of the rest of jobs with normal memory requirements, slowing down execution of each individual job and decreasing the system throughput. We call this phenomenon the job blocking problem because the big jobs block the execution pace of majority jobs in the cluster. Since the memory demand of jobs may not be known in advance and may change dynamically, the possibility of unsuitable job submissions/migrations to cause the blocking problem is high, and existing load sharing schemes are unable to effectively handle this problem. We propose two schemes to address this problem. The first scheme, network RAM supported load sharing, combines job migrations with network RAM, which uses remote execution to initially allocate a job to the most lightly loaded workstation and, if necessary, network RAM to provide a global memory space for the job larger than it would be available otherwise. This scheme has the merits of both job migrations and network RAM. Our experiments show its effectiveness and scalability. However, this scheme requires a network RAM facility in the cluster, which may cause additional overhead and increase cluster network traffic. In order to address this limit, we propose a second scheme, memory reservation, incorporated with dynamic load sharing, which adaptively reserves a small set of workstations to provide special services to the jobs demanding large memory allocations. As soon as the blocking problem is resolved by the memory reservation scheme, the system will adaptively switch back to the normal load sharing state. Both schemes target on handling large data-intensive jobs in clusters, and are mutually complementary. The network RAM supported load sharing scheme can fully utilize the cluster global memory space, while the memory reservation scheme has the advantage of simple implementations and low overhead. Thus, they both can be effective alternatives, and practically deployed in cluster computing under different system conditions.
Li Xiao 0001, Songqing Chen, Xiaodong Zhang 0001
IEEE Trans. Parallel Distributed Syst.2
2003 Adaptive and lazy segmentation based proxy caching for streaming media delivery
abstract
Streaming media objects are often cached in segments. Previous segment-based caching strategies cache segments with constant or exponentially increasing lengths and typically favor caching the beginning segments of media objects. However, these strategies typically do not consider the fact that most accesses are targeted toward a few popular objects. In this paper, we argue that neither the use of a predefined segment length nor the favorable caching of the beginning segments is the best caching strategy for reducing network traffic. We propose an adaptive and lazy segmentation based caching mechanism by delaying the segmentation as late as possible and determining the segment length based on the client access behaviors in real time. In addition, the admission and eviction of segments are carried out adaptively based on an accurate utility function. The proposed method is evaluated by simulations using traces including one from actual enterprise server logs. Simulation results indicate that our proposed method achieves a 30% reduction in network traffic. The utility functions of the replacement policy are also evaluated with different variations to show its accuracy.
Songqing Chen, Bo Shen 0003, Susie J. Wee, Xiaodong Zhang 0001
NOSSDAV1
2002 Adaptive and Virtual Reconfigurations for Effective Dynamic Job Scheduling in Cluster Systems
abstract
In a cluster system with dynamic load sharing support, a job submission or migration to a workstation is determined by the availability of CPU and memory resources of the workstation at the time. In such a system, a small number of running jobs with unexpectedly large memory allocation requirements may significantly increase the queuing delay times of the rest of jobs with normal memory requirements, slowing down executions of individual jobs and decreasing the system throughput. We call this phenomenon as the job blocking problem because the big jobs block the execution pace of majority jobs in the cluster. We propose a software method incorporating with dynamic load sharing, which adaptively reserves a small set of workstations through virtual cluster reconfiguration to provide special services to the jobs demanding large memory allocations. This policy implies the principle of shortest-remaining-processing-time policy. As soon as the blocking problem is resolved by the reconfiguration, the system will adaptively switch back to the normal load sharing state. We present three contributions in this study. (1) the conditions to cause the job blocking problem; (2) the adaptive software method in a dynamic load sharing system; and (3) trace-driven simulations. We show that our method can effectively improve the cluster computing performance by quickly resolving the job blocking problem. The effectiveness and performance insights are also analytically verified.
Songqing Chen, Li Xiao 0001, Xiaodong Zhang 0001
ICDCS1
2002 Dynamic Cluster Resource Allocations for Jobs with Known and Unknown Memory Demands
abstract
The cluster system we consider for load sharing is a compute farm which is a pool of networked server nodes providing high-performance computing for CPU-intensive, memory-intensive, and I/O active jobs in a batch mode. Existing resource management systems mainly target at balancing the usage of CPU loads among server nodes. With the rapid advancement of CPU chips, memory and disk access speed improvements significantly lag behind advancement of CPU speed, increasing the penalty for data movement, such as page faults and I/O operations, relative to normal CPU operations. Aiming at reducing the memory resource contention caused by page faults and I/O activities, we have developed and examined load sharing policies by considering effective usage of global memory in addition to CPU load balancing in clusters. We study two types of application workloads: 1) Memory demands are known in advance or are predictable and 2) memory demands are unknown and dynamically changed during execution. Besides using workload traces with known memory demands, we have also made kernel instrumentation to collect different types of workload execution traces to capture dynamic memory access patterns. Conducting different groups of trace-driven simulations, we show that our proposed policies can effectively improve overall job execution performance by well utilizing both CPU and memory resources with known and unknown memory demands.
Li Xiao 0001, Songqing Chen, Xiaodong Zhang 0001
IEEE Trans. Parallel Distributed Syst.2
2001 Dynamic Load Sharing with Unknown Memory Demands in Clusters
abstract
A compute farm is a pool of clustered workstations to provide high performance computing services for CPU-intensive, memory-intensive, and I/O active jobs in a batch mode. Existing load sharing schemes with memory considerations assume jobs' memory demand sizes are known in advance or predictable based on users' hints. This assumption can greatly simplify the designs and implementations of load sharing schemes, but is not desirable in practice. In order to address this concern, we present three new results and contributions in this study. Conducting Linux kernel instrumentation, we have collected different types of workload execution traces to quantitatively characterize job interactions, and modeled page fault behavior as a function of the overloaded memory sizes and the amount of jobs' I/O activities. Based on experimental results and collected dynamic system information, we have built a simulation model which accurately emulates the memory system operations and job migrations with virtual memory considerations. We have proposed a memory-centric load sharing scheme and its variations to effectively process dynamic memory, allocation demands, aiming at minimizing execution time of each individual job by dynamically migrating and remotely submitting jobs to eliminate or reduce page faults and to reduce the queuing time for CPU services. Conducting trace-driven simulations, we have examined these load sharing policies to show their effectiveness.
Songqing Chen, Li Xiao 0001, Xiaodong Zhang 0001
ICDCS1