Cong Wang 0014

dblp:18/2771-14 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0001-6429-8799ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 4 first-author · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Taming Imbalance and Complexity in Resilient WAN Traffic Engineering
abstract
The rapid expansion of global cloud infrastructures and increasing network traffic volume and dynamicity have led to a rise in research on developing scalable and resilient Traffic Engineering (TE) solutions for Wide Area Networks (WANs). Despite recent advancements, striking the right balance between network availability, computational complexity, and resource utilization remains a significant challenge.This paper presents empirical findings that highlight the inherent traffic demand imbalance and link utilization heterogeneity that current TE solutions have overlooked. We then define two new performance metrics, namely critical link set and network criticality, that jointly represent these heterogeneities. We further introduce an efficient extension to the tunnel-based resilient TE algorithm to be adaptive to traffic profiles. An extensive simulation study on representative WAN topologies demonstrates the substantial performance enhancements achieved by our holistic solution approach.
Yufeng Xin, Sajith Sasidharan, Cong Wang 0014, Mert Cevik
ICCCN3
2024 DISTRI: Development and Integration of Simulation Tools for Resilient Infrastructure
abstract
In contemporary scientific research, data acquisition and analysis platforms have grown increasingly complex, often spanning multiple facilities with diverse internal structures. Efficiently managing the interactions between job scheduling, resource allocation, and networking across these distributed systems requires a robust simulation framework. However, existing simulators fall short in capturing the detailed interactions necessary for comprehensive analysis of large-scale distributed environments. To address this gap, we introduce DISTRI, a versatile framework specifically designed for the development and testing of distributed multi-facility workflows. DISTRI allows for customizable facility configurations and includes built-in support for distributed, resilient scheduling and resource management, alongside detailed network simulation for data communication. Key features of DISTRI encompass inter- and intra-facility resource management, agent-based distributed scheduling, and extensive performance metrics logging for both resource and network management. By providing these essential tools, DISTRI enables thorough analysis and optimization, thereby advancing research in the resilience and efficiency of multi-facility systems.
Imtiaz Mahmud, Pawel Zuk, Cong Wang 0014, Mariam Kiran, Kesheng Wu, Komal Thareja, Raghavan Krishnan, Anirban Mandal, Ewa Deelman
IEEE Big Data3
2024 Large Language Models for Anomaly Detection in Computational Workflows: From Supervised Fine-Tuning to In-Context Learning
abstract
Anomaly detection in computational workflows is critical for ensuring system reliability and security. However, traditional rule-based methods struggle to detect novel anomalies. This paper leverages large language models (LLMs) for workflow anomaly detection by exploiting their ability to learn complex data patterns. Two approaches are investigated: (1) supervised fine-tuning (SFT), where pretrained LLMs are fine-tuned on labeled data for sentence classification to identify anomalies, and (2) in-context learning (ICL), where prompts containing task descriptions and examples guide LLMs in few-shot anomaly detection without fine-tuning. The paper evaluates the performance, efficiency, and generalization of SFT models and explores zeroshot and few-shot ICL prompts and interpretability enhancement via chain-of-thought prompting. Experiments across multiple workflow datasets demonstrate the promising potential of LLMs for effective anomaly detection in complex executions.
George Papadimitriou 0002, Raghavan Krishnan, Pawel Zuk, Prasanna Balaprakash, Cong Wang 0014, Anirban Mandal, Ewa Deelman
SC6
2022 Automating Edge-to-cloud Workflows for Science: Traversing the Edge-to-cloud Continuum with Pegasus
abstract
In this paper, we describe how we extended the Pegasus Workflow Management System to support edge-to-cloud workflows in an automated fashion. We discuss how Pegasus and HTCondor (its job scheduler) work together to enable this automation. We use HTCondor to form heterogeneous pools of compute resources and Pegasus to plan the workflow onto these resources and manage containers and data movement for executing workflows in hybrid edge-cloud environments. We then show how Pegasus can be used to evaluate the execution of workflows running on edge only, cloud only, and edge-cloud hybrid environments. Using the Chameleon Cloud testbed to set up and configure an edge-cloud environment, we use Pegasus to benchmark the executions of one synthetic workflow and two production workflows: CASA-Wind and the Ocean Observatories Initiative Orcasound workflow, all of which derive their data from edge devices. We present the performance impact on workflow runs of job and data placement strategies employed by Pegasus when configured to run in the above three execution environments. Results show that the synthetic workflow performs best in an edge only environment, while the CASA - Wind and Orcasound workflows see significant improvements in overall makespan when run in a cloud only environment. The results demonstrate that Pegasus can be used to automate edge-to-cloud science workflows and the workflow provenance data collection capabilities of the Pegasus monitoring daemon enable computer scientists to conduct edge-to-cloud research.
Ryan Tanaka, George Papadimitriou 0002, Sai Charan Viswanath, Cong Wang 0014, Eric Lyons 0001, Komal Thareja, Chengyi Qu, Alicia Esquivel Morel, Ewa Deelman, Anirban Mandal, Prasad Calyam, Michael Zink
CCGRID4
2022 Towards Production Deployment of a SDX Control Framework
abstract
Developing a distributed controller system for pro-duction software defined networks (SDN) requires substantial en-gineering effort to satisfy the stringent performance, adaptability, availability and security requirements that are well beyond the basic function development. The softwarization nature of SDN provides the opportunity to leverage the recent advancements in software engineering and system automation. In this paper we present our recent work to develop and deploy a wide-area SDN control framework prototype towards production deployment and operation. We primarily focus on three critical areas (1) enhancement of the software functions and the substrate virtualization configurations to fully support advanced connection services and fault tolerance in both data plane and control plane (2) a high-fidelity testing pipeline con-sisting of unit tests emulation and a testbed and (3) a continuous integration and continuous deployment (CI/CD) pipeline. Our experience proved that the presented environment and process greatly increased the software quality and development efficiency which would ultimately lead to a reliable and continuous deploy-ment and automation of the targeted production SDN network.
Mert Cevik, Michael J. Stealey, Cong Wang 0014, Jeronimo Bezerra, Julio Ibarra, Vasilka Chergarova, Heidi Morgan, Yufeng Xin
ICCCN3
2021 Mining Workflows for Anomalous Data Transfers
abstract
Modern scientific workflows are data-driven and are often executed on distributed, heterogeneous, high-performance computing infrastructures. Anomalies and failures in the work-flow execution cause loss of scientific productivity and inefficient use of the infrastructure. Hence, detecting, diagnosing, and mitigating these anomalies are immensely important for reliable and performant scientific workflows. Since these workflows rely heavily on high-performance network transfers that require strict QoS constraints, accurately detecting anomalous network performance is crucial to ensure reliable and efficient workflow execution. To address this challenge, we have developed X-FLASH, a network anomaly detection tool for faulty TCP workflow transfers. X-FLASH incorporates novel hyperparameter tuning and data mining approaches for improving the performance of the machine learning algorithms to accurately classify the anomalous TCP packets. X-FLASH leverages XGBoost as an ensemble model and couples XGBoost with a sequential optimizer, FLASH, borrowed from search-based Software Engineering to learn the optimal model parameters. X-FLASH found configurations that outperformed the existing approach up to 28%, 29%, and 40% relatively for F-measure, G-score, and recall in less than 30 evaluations. From (1) large improvement and (2) simple tuning, we recommend future research to have additional tuning study as a new standard, at least in the area of scientific workflow anomaly detection.
Huy Tu, George Papadimitriou 0002, Mariam Kiran, Cong Wang 0014, Anirban Mandal, Ewa Deelman, Tim Menzies
MSR4
2021 End-to-end online performance data capture and analysis for scientific workflows
George Papadimitriou 0002, Cong Wang 0014, Karan Vahi, Rafael Ferreira da Silva, Anirban Mandal, Zhengchun Liu, Rajiv Mayani, Mats Rynge, Mariam Kiran, Vickie E. Lynch, Rajkumar Kettimuthu, Ewa Deelman, Jeffrey S. Vetter, Ian T. Foster
Future Gener. Comput. Syst.2
2020 Detecting anomalous packets in network transfers: investigations using PCA, autoencoder and isolation forest in TCP
Mariam Kiran, Cong Wang 0014, George Papadimitriou 0002, Anirban Mandal, Ewa Deelman
Mach. Learn.2
2019 Toward a Dynamic Network-Centric Distributed Cloud Platform for Scientific Workflows: A Case Study for Adaptive Weather Sensing
abstract
Computational science today depends on complex, data-intensive applications operating on datasets from a variety of scientific instruments. A major challenge is the integration of data into the scientist's workflow. Recent advances in dynamic, networked cloud resources provide the building blocks to construct reconfigurable, end-to-end infrastructure that can increase scientific productivity. However, applications have not adequately taken advantage of these advanced capabilities. In this work, we have developed a novel network-centric platform that enables high-performance, adaptive data flows and coordinated access to distributed cloud resources and data repositories for atmospheric scientists. We demonstrate the effectiveness of our approach by evaluating time-critical, adaptive weather sensing workflows, which utilize advanced networked infrastructure to ingest live weather data from radars and compute data products used for timely response to weather events. The workflows are orchestrated by the Pegasus workflow management system and were chosen because of their diverse resource requirements. We show that our approach results in timely processing of Nowcast workflows under different infrastructure configurations and network conditions. We also show how workflow task clustering choices affect throughput of an ensemble of Nowcast workflows with improved turnaround times. Additionally, we find that using our network-centric platform powered by advanced layer2 networking techniques results in faster, more reliable data throughput, makes cloud resources easier to provision, and the workflows easier to configure for operational use and automation.
Eric Lyons 0001, Anirban Mandal, George Papadimitriou 0002, Cong Wang 0014, Komal Thareja, Paul Ruth, Juan J. Villalobos, Ivan Rodero, Ewa Deelman, Michael Zink
eScience4
2019 COMET: Distributed Metadata Service for Multi-cloud Experiments
abstract
A majority of today’s cloud services are independently operated by individual cloud service providers. In this approach, the locations of cloud resources are strictly constrained by the distribution of cloud service providers’ sites. As the popularity and scale of cloud services increase, we believe this traditional paradigm is about to change toward further federated services, a.k.a., multi-cloud, due to the improved performance, reduced cost of compute, storage and network resources, as well as increased user demands. In this paper, we present COMET, a lightweight, distributed storage system for managing metadata on large scale, federated cloud infrastructure providers, end users, and their applications (e.g. HTCondor Cluster or Hadoop Cluster). We showcase use case from NSF’s, Chameleon, ExoGENI and JetStream research cloud testbeds to show the effectiveness of COMET design and deployment.
Komal Thareja, Cong Wang 0014, Paul Ruth, Anirban Mandal, Ilya Baldin, Michael J. Stealey
ICNP2
2019 Sustainable Cloud Encoding for Adaptive Bitrate Streaming over CDNs
abstract
Video streaming is the most popular application on today's Internet. Millions of people around the globe access video contents using various end user devices, such as smart phones, tablets, laptops, and TVs. To meet the requirements of different end user devices and variable network conditions, videos need to be encoded into different quality versions before delivery to the clients. Such large-scale encoding tasks consume significant amounts of energy. In this paper, we investigate to what extent the realtime video encoding clouds can be powered by renewable energy sources. We show that video encoding tasks are suitable for execution on clouds that are powered by a combination of renewable and grid energy sources. With the use of our power management policies, grid energy usage can be reduced by 73-83%, which leads to electricity cost reductions of 14-28% compared to unlimited non-renewable power.
Cong Wang 0014, Michael Zink
LANMAN1
2017 Design and Analysis of QoE-Aware Quality Adaptation for DASH: A Spectrum-Based Approach
abstract
The dynamics of the application-layer-based control loop of dynamic adaptive streaming over HTTP (DASH) make video bitrate selection for DASH a difficult problem. In this work, we provide a DASH quality adaptation algorithm, named SQUAD, that is specifically tailored to provide a high quality of experience (QoE). We review and provide new insights into the challenges for DASH rate estimation. We found that in addition to the ON-OFF behavior of DASH clients, there exists a discrepancy in the timescales that form the basis of the rate estimates across (i) different video segments and (ii) the rate control loops of DASH and Transmission Control Protocol (TCP). With these observations in mind, we design SQUAD aiming to maximize the average quality bitrate while minimizing the quality variations. We test our implementation of SQUAD together with a number of different quality adaptation algorithms under various conditions in the Global Environment for Networking Innovation testbed, as well as, in a series of measurements over the public Internet. Through a measurement study, we show that by sacrificing little to nothing in average quality bitrate, SQUAD can provide significantlygt ; better QoE in terms of quality switching and magnitude. In addition, we show that retransmission of higher-quality segments that were originally received in low-quality is feasible and improves the QoE.
Cong Wang 0014, Divyashri Bhat, Amr Rizk, Michael Zink
ACM Trans. Multim. Comput. Commun. Appl.1
2016 SQUAD: a spectrum-based quality adaptation for dynamic adaptive streaming over HTTP
abstract
The application-layer based control loops of dynamic adaptive streaming over HTTP (DASH) make video bitrate selection a complex problem. In this work, we review and present new insights into the challenges of DASH rate adaptation. We identify several critical issues that contribute to the degradation of DASH performance with respect to the rate control loops of DASH and TCP. We then introduce a novel DASH quality adaptation algorithm SQUAD, which is specifically designed to ensure high quality of experience (QoE). We implement and test our algorithm together with a number of state-of-the-art quality adaptation algorithms. Through extensive experiments on both testbed and cross-Atlantic Internet scenarios, we show that by sacrificing little to none in average quality bitrate, SQUAD provides significantly better QoE in terms of number and magnitude of quality switches.
Cong Wang 0014, Amr Rizk, Michael Zink
MMSys1
2015 VHub: Single-stage virtual network mapping through hub location
Shashank Shanbhag, Arun Reddy Kandoor, Cong Wang 0014, Ramgopal R. Mettu, Tilman Wolf
Comput. Networks3
2014 On the Feasibility of DASH Streaming in the Cloud
abstract
As shown in recent studies, video streaming is by far the biggest category of backbone Internet traffic in the US. As a measure to reduce the cost of highly over-provisioned physical infrastructures while remaining the quality of video services, many streaming service providers started to use cloud services where physical resources can be dynamically allocated based on current demand. This paper characterizes the performance of Dynamic Adaptive Streaming over HTTP (DASH), a new MPEG standard on adaptive streaming, in the cloud. We seek to answer the following questions that are critical to content providers that are hosting video in clouds: Which data center is the best to host videos? Does geographical distance matter? What type of instance is best suitable depending on different needs? How to efficiently solve the trade-off between performance and cost? The measurement methods and results presented in this paper can be easily expanded into other VoD services, and they allow us to i) characterize DASH behavior when streaming from the cloud; ii) identify the key factors that influence the DASH performance; and iii) suggest improvements for related services.
Cong Wang 0014, Michael Zink
NOSSDAV1
2012 Virtual network mapping with traffic matrices
abstract
Network virtualization is a core technology in next-generation networks to overcome the ossification problem that is observed in the current Internet. The key idea of network virtualization is to split physical network resources into multiple logical networks, each supporting different network services and functionalities. One of the key challenges for virtual network infrastructure providers is to efficiently allocate network resources based on virtual network requests, which is referred to as the virtual network mapping problem. While several algorithms have been proposed previously to solve this mapping problem, their effectiveness is limited since virtual requests specify the internal topology of the virtual network. In this paper, we argue that such internal topologies lead to unnecessary constraints and less efficient solutions. Instead, we propose an alternate formulation of the problem that represents virtual network requests as traffic matrices. We provide a solutions to solving this traffic-matrix-based mapping problem using a mixed integer programming formulation. Our simulation results show that our approach can map significantly more virtual network requests on a physical network infrastructure than previous mapping algorithms and thus improves the efficient use of networking resources in virtual networks.
Cong Wang 0014, Shashank Shanbhag, Tilman Wolf
ICC1
2011 Virtual Network Mapping with Traffic Matrices
abstract
Network virtualization allows multiple logical networks to coexist on the same physical infrastructure. A key problem that needs to be solved in this context is management of substrate network resources, in particular when mapping virtual network requests to the substrate. Several efficient algorithms have been proposed to provide effective solutions, but they assume that the virtual network request topology is given. In this paper, we show that the use of these topology-based requests may limit the efficiency of physical network usage. As an alternative, we propose a problem formulation and mapping algorithm that is based on traffic matrices to specify virtual network requests.
Cong Wang 0014, Tilman Wolf
ANCS1