Shanchieh Jay Yang

dblp:26/2191 · also Shan-Chieh Yang, Shanchieh Yang · DBLP profile ↗
← Back
44ranked-venue papers
5as first author
16since 2021 · last 2026
0009-0004-5503-2082ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 15 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 14 · 1 first-author · 4 since 2021Computer networks · 10 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 6 since 2021Systems, architecture and hardware · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 How Can You Tell if Your Large Language Model Could Be a Closet Antisemite? An Explainability-Based Audit Framework for Implicit Bias
abstract
Auditing large language models (LLMs) for biases is an ongoing and dynamic process, resembling a proverbial cat-and-mouse game. As researchers identify new vulnerabilities in LLMs, guardrails are updated to address them, prompting the need for innovative approaches to audit the increasingly fortified LLMs for biases. This paper makes three contributions. First, it introduces a scalable, explainable framework to measure biases against various identity groups across multiple open large language models. Second, it conducts a bias audit considering five well-known open LLMs and demonstrates their bias inclinations towards several historically disadvantaged groups. Our audit reveals disturbing antisemitic, Islamophobic, and xenophobic biases present in several well-known LLMs. Finally, we release a dataset of 1,000 probes curated under the supervision of an expert social scientist that can facilitate similar audits.
Arka Dutta 0001, Reza Fayyazi, Shanchieh Jay Yang, Ashiqur R. KhudaBukhsh
AAAI3
2024 Adversary Tactic Driven Scenario and Terrain Generation with Partial Infrastructure Specification
abstract
Diverse, accurate, and up-to-date training environments are essential for training cybersecurity experts and autonomous systems. However, preparation of their content is time-consuming and requires experts to provide detailed specifications. In this paper, we explore the challenges of automated generation of the content (composed of scenarios and terrains) for these environments.
Ádám Ruman, Martin Drasar, Lukás Sadlek, Shanchieh Jay Yang, Pavel Celeda
ARES4
2024 The 4th Workshop on Artificial Intelligence-enabled Cybersecurity Analytics
abstract
Cybersecurity remains a grand societal challenge. Large and constantly changing attack surfaces are non-trivial to protect against malicious actors. Entities like the United States and the European Union have recently emphasized the value of Artificial Intelligence (AI) for advancing cybersecurity. For example, the National Science Foundation has called for AI systems that can enhance cyber threat intelligence, detect new and evolving threats, and analyze massive troves of cybersecurity data. The 4th Workshop on Artificial Intelligence-enabled Cybersecurity Analytics (co-located with ACM KDD) sought to make significant and novel contributions within these relevant topics. Submissions were reviewed by highly qualified AI for cybersecurity researchers and practitioners spanning academia and private industry firms.
Steven Ullman, Benjamin Ampel, Sagar Samtani, Shanchieh Jay Yang, Hsinchun Chen
KDD4
2023 DUCK: A Drone-Urban Cyber-Defense Framework Based on Pareto-Optimal Deontic Logic Agents
abstract
Drone based terrorist attacks are increasing daily. It is not expected to be long before drones are used to carry out terror attacks in urban areas. We have developed the DUCK multi-agent testbed that security agencies can use to simulate drone-based attacks by diverse actors and develop a combination of surveillance camera, drone, and cyber defenses against them.
Tonmoay Deb, Jürgen Dix, Mingi Jeong, Cristian Molinaro, Andrea Pugliese 0001, Alberto Quattrini Li, Eugene Santos Jr., V. S. Subrahmanian, Shanchieh Jay Yang, Youzhi Zhang 0001
AAAI9
2023 Unraveling Network-Based Pivoting Maneuvers: Empirical Insights and Challenges
Martin Husák, Shanchieh Jay Yang, Joseph Khoury, Dorde Klisura, Elias Bou-Harb
ICDF2C (2)2
2023 The 3rd Workshop on Artificial Intelligence-enabled Cybersecurity Analytics
abstract
Artificial Intelligence (AI) has gripped modern society as a viable approach to revolutionize operational capabilities across multiple industries. One critical application area that could stand to benefit from the capabilities of AI is cybersecurity. Increasingly, federal funding agencies such as the National Science Foundation are calling for enhanced AI-enabled analytics capabilities to improve cyber threat intelligence, cyber defense generation, and more. To this end, this half-day workshop, not in its third year at ACM KDD, sought to attain significant contributions related to various aspects of AI-enabled cybersecurity analytics. This workshop received a record number of submissions. Submissions were reviewed by a highly-qualified, interdisciplinary group of AI for cybersecurity researchers and practitioners spanning academia and private industry firms.
Sagar Samtani, Shanchieh Jay Yang, Hsinchun Chen
KDD2
2022 Non-cooperative Learning for Robust Spectrum Sharing in Connected Vehicles with Malicious Agents
abstract
Multi-agent reinforcement learning (MARL) has pre-viously been employed for efficient spectrum sharing among co-operative connected vehicles. However, we show in this paper that existing MARL models are not robust against non-cooperative or malicious agents (vehicles) whose spectrum selection strategy may cause congestion and reduce the spectrum utilization. For example, a selfish (non-cooperative) agent aims to only maximize its own spectrum utilization, irrespective of the overall system efficiency and spectrum availability to others. We investigate and analyze the MARL-based spectrum sharing problem in connected vehicles including vehicles (agents) with selfish or sabotage strategies. We then develop a theoretical framework to consider the selfish agent, and study various adversarial scenarios (including attacks with disruptive goals) via simulations. Our robust MARL approach where “robust” agents are trained to be prepared for selfish agents in testing phase achieves more resiliency in the presence of a selfish agent and even a sabotage one; achieving 6.7%~20% and 50.7% ~ 138% higher unicast throughput and broadcast delivery success rate over regular benign agents, respectively.
Hanif Rahbari, Shanchieh Jay Yang, Li-Chun Wang 0001
GLOBECOM3
2022 Discovery of Rare yet Co-occurring Actions with Temporal Characteristics in Episodic Cyberattack Streams
abstract
The large number of streaming intrusion alerts make it challenging for security analysts to quickly identify attack patterns. This is especially difficult since critical alerts often occur rarely for traditional pattern mining algorithms to be effective. Recognizing the attack speed as an inherent indicator of differing cyber attacks, this work aggregates alerts into attack episodes that have distinct attack speeds, and finds attack actions regularly co-occurring within the same episode. This enables a novel use of SPADE – a temporal pattern mining algorithm – to extract consistent co-occurrences of alert signatures that are indicative of attack actions that follow each other. The proposed system, R-CAD, extracts not only the co-occurring patterns but also the temporal characteristics of the co-occurrences, giving the ‘strong rules’ indicative of critical and repeated attack behaviors. Through the use of a real-world dataset, we demonstrate that R-CAD helps reduce from the overwhelming volume and variety of intrusion alerts to a manageable set of co-occurring strong rules. We show specific rules that reveal how critical attack actions follow one another and in what attack speed.
Gordon Werner, Shanchieh Jay Yang
ICCCN2
2022 ACM KDD AI4Cyber/MLHat: Workshop on AI-enabled Cybersecurity Analytics and Deployable Defense
abstract
Federal funding agencies and industry entities are seeking innovative approaches to address the ever-growing cybersecurity crisis. Increasingly, numerous cybersecurity thought leaders are indicating that Artificial Intelligence (AI)-enabled analytics can help tackle key cybersecurity tasks and deploy defenses. This half-day workshop, co-located with ACM KDD, sought to attain significant research contributions to various aspects of AI-enabled analytics for cybersecurity applications and deployable defense solutions from academics and practitioners. This workshop was a joint workshop of the 2021 AI-enabled Cybersecurity Analytics and 2021 International Workshop on Deployable Machine Learning for Security Defense. As such, we developed an interdisciplinary Program Committee with significant experience in various aspects of AI, cybersecurity, and/or deployable defense.
Sagar Samtani, Gang Wang 0011, Ali Ahmadzadeh, Arridhana Ciptadi, Shanchieh Jay Yang, Hsinchun Chen
KDD5
2022 Alert-Driven Attack Graph Generation Using S-PDFA
abstract
Ideal cyber threat intelligence (CTI) includes insights into attacker strategies that are specific to a network under observation. Such CTI currently requires extensive expert input for obtaining, assessing, and correlating system vulnerabilities into a graphical representation, often referred to as an attack graph (AG). Instead of deriving AGs based on system vulnerabilities, this work advocates the direct use of intrusion alerts. We propose SAGE, an explainable sequence learning pipeline that automatically constructs AGs from intrusion alerts without a priori expert knowledge. SAGE exploits the temporal and probabilistic dependence between alerts in a suffix-based probabilistic deterministic finite automaton (S-PDFA)-a model that brings infrequent severe alerts into the spotlight and summarizes paths leading to them. Attack graphs are extracted from the model on a per-victim, per-objective basis. SAGE is thoroughly evaluated on three open-source intrusion alert datasets collected through security testing competitions in order to analyze distributed multi-stage attacks. SAGE compresses over 330k alerts into 93 AGs that show how specific attacks transpired. The AGs are succinct, interpretable, and provide directly relevant insights into strategic differences and fingerprintable paths. They even show that attackers tend to follow shorter paths after they have discovered a longer one in 84.5% of the cases.
Azqa Nadeem, Sicco Verwer, Stephen Moskal, Shanchieh Jay Yang
IEEE Trans. Dependable Secur. Comput.4
2021 On the Evaluation of Sequential Machine Learning for Network Intrusion Detection
abstract
Recent advances in deep learning renewed the research interests in machine learning for Network Intrusion Detection Systems (NIDS). Specifically, attention has been given to sequential learning models, due to their ability to extract the temporal characteristics of network traffic flows (NetFlows), and use them for NIDS tasks. However, the applications of these sequential models often consist of transferring and adapting methodologies directly from other fields, without an in-depth investigation on how to leverage the specific circumstances of cybersecurity scenarios; moreover, there is a lack of comprehensive studies on sequential models that rely on NetFlow data, which presents significant advantages over traditional full packet captures. We tackle this problem in this paper. We propose a detailed methodology to extract temporal sequences of NetFlows that denote patterns of malicious activities. Then, we apply this methodology to compare the efficacy of sequential learning models against traditional static learning models. In particular, we perform a fair comparison of a ‘sequential’ Long Short-Term Memory (LSTM) against a ‘static’ Feedforward Neural Networks (FNN) in distinct environments represented by two well-known datasets for NIDS: the CICIDS2017 and the CTU13. Our results highlight that LSTM achieves comparable performance to FNN in the CICIDS2017 with over 99.5% F1-score; while obtaining superior performance in the CTU13, with 95.7% F1-score against 91.5%. This paper thus paves the way to future applications of sequential learning models for NIDS.
Andrea Corsini, Shanchieh Jay Yang, Giovanni Apruzzese
ARES2
2021 Enabling Visual Analytics via Alert-driven Attack Graphs
abstract
Attack graphs (AG) are a popular area of research that display all the paths an attacker can exploit to penetrate a network. Existing techniques for AG generation rely heavily on expert input regarding vulnerabilities and network topology. In this work, we advocate the use of AGs that are built directly using the actions observed through intrusion alerts, without prior expert input. We have developed an unsupervised visual analytics system, called SAGE, to learn alert-driven attack graphs. We show how these AGs (i) enable forensic analysis of prior attacks, and (ii) enable proactive defense by providing relevant threat intelligence regarding attacker strategies. We believe that alert-driven AGs can play a key role in AI-enabled cyber threat intelligence as they open up new avenues for attacker strategy analysis whilst reducing analyst workload.
Azqa Nadeem, Sicco Verwer, Stephen Moskal, Shanchieh Jay Yang
CCS4
2021 Near real-time intrusion alert aggregation using concept-based learning
abstract
Intrusion detection systems generate a large number of streaming alerts. It can be overwhelming for analysts to quickly and effectively find related alerts stemmed from correlated attack actions. What if fast arriving alerts could be automatically processed with no prior knowledge to find related actions in near real-time? The Concept Learning for Intrusion Event Aggregation in Realtime (CLEAR) system aims to learn and update an evolving set of temporal 'concepts,' each consisting of aggregates of related alerts that exhibit similar statistical arrival patterns. With no training data, the system constructs the concepts in near real-time from statistically similar alert aggregates. Tracked concepts are then applied to incoming alerts for fast and high-fidelity aggregation. The concepts learned by CLEAR are significantly more unique and invariant when compared to those learned by alternative drift detection methods. Furthermore, it provides insights for how specific individual, or co-occuring, alerts arrive with distinct and consistent temporal patterns.
Gordon Werner, Shanchieh Jay Yang, Katie McConky
CF2
2021 Towards an Efficient Detection of Pivoting Activity
Martin Husák, Giovanni Apruzzese, Shanchieh Jay Yang, Gordon Werner
IM3
2021 ACM KDD AI4Cyber: The 1st Workshop on Artificial Intelligence-enabled Cybersecurity Analytics
abstract
Despite significant contributions to various aspects of cybersecurity, cyber-attacks remain on the unfortunate rise. Increasingly, internationally recognized entities such as the National Science Foundation and National Science & Technology Council have noted Artificial Intelligence can help analyze billions of log files, Dark Web data, malware, and other data sources to help execute fundamental cybersecurity tasks. Our objective for the 1st Workshop on Artificial Intelligence-enabled Cybersecurity Analytics (half-day; co-located with ACM KDD) was to gather academic and practitioners to contribute recent work pertaining to AI-enabled cybersecurity analytics. We composed an outstanding, inter-disciplinary Program Committee with significant expertise in various aspects of AI-enabled Cybersecurity Analytics to evaluate the submitted work. Significant contributions to the half-day workshop were made in the areas of CTI, vulnerability assessment, and malware analysis.
Sagar Samtani, Shanchieh Jay Yang, Hsinchun Chen
KDD2
2021 SAGE: Intrusion Alert-driven Attack Graph Extractor
abstract
Attack graphs (AG) are used to assess pathways availed by cyber adversaries to penetrate a network. State-of-the-art approaches for AG generation focus mostly on deriving dependencies between system vulnerabilities based on network scans and expert knowledge. In real-world operations however, it is costly and ineffective to rely on constant vulnerability scanning and expert-crafted AGs. We propose to automatically learn AGs based on actions observed through intrusion alerts, without prior expert knowledge. Specifically, we develop an unsupervised sequence learning system, SAGE, that leverages the temporal and probabilistic dependence between alerts in a suffix-based probabilistic deterministic finite automaton (S-PDFA) – a model that accentuates infrequent severe alerts and summarizes paths leading to them. AGs are then derived from the S-PDFA on a per-objective, per-victim basis. Tested with intrusion alerts collected through Collegiate Penetration Testing Competition, SAGE compresses over 330k alerts into 93 AGs. These AGs reflect the strategies used by the participating teams. The AGs are succinct, interpretable, and capture behavioral dynamics, e.g., that attackers will often follow shorter paths to re-exploit objectives.
Azqa Nadeem, Sicco Verwer, Shanchieh Jay Yang
VizSec3
2020 SoK: contemporary issues and challenges to enable cyber situational awareness for network security
abstract
Cyber situational awareness is an essential part of cyber defense that allows the cybersecurity operators to cope with the complexity of today's networks and threat landscape. Perceiving and comprehending the situation allow the operator to project upcoming events and make strategic decisions. In this paper, we recapitulate the fundamentals of cyber situational awareness and highlight its unique characteristics in comparison to generic situational awareness known from other fields. Subsequently, we provide an overview of existing research and trends in publishing on the topic, introduce front research groups, and highlight the impact of cyber situational awareness research. Further, we propose an updated taxonomy and enumeration of the components used for achieving cyber situational awareness. The updated taxonomy conforms to the widely-accepted three-level definition of cyber situational awareness and newly includes the projection level. Finally, we identify and discuss contemporary research and operational challenges, such as the need to cope with rising volume, velocity, and variety of cybersecurity data and the need to provide cybersecurity operators with the right data at the right time and increase their value through visualization.
Martin Husák, Tomás Jirsík, Shanchieh Jay Yang
ARES3
2020 Session-level Adversary Intent-Driven Cyberattack Simulator
abstract
Recognizing the need for proactive analysis of cyber adversary behavior, this paper presents a new event-driven simulation model and implementation to reveal the efforts needed by attackers who have various entry points into a network. Unlike previous models which focus on the impact of attackers' actions on the defender's infrastructure, this work focuses on the attackers' strategies and actions. By operating on a request-response session level, our model provides an abstraction of how the network infrastructure reacts to access credentials the adversary might have obtained through a variety of strategies. We present the current capabilities of the simulator by showing three variants of Bronze Butler APT on a network with different user access levels.
Martin Drasar, Stephen Moskal, Shanchieh Jay Yang, Pavol Zat'ko
DS-RT3
2019 Experimental Evaluation of Jamming Threat in LoRaWAN
abstract
LoRaWAN is a promising solution of Low-power wide area network (LPWAN) operating in unlicensed spectrum to support long range wireless services for Internet of Things (IoTs). However, with the growth of IoT devices deployed in a fixed geographic area, the immunity to interference on the communication increases significantly. Attackers may utilize such situation to jam packet transmission by emitting RF interference signal at the same time when a LoRa end node is sending data to the LoRa gateway. As a consequence, the transmission of the LoRa end node would fail due to collision which would in turn reduce the network performance. In this paper, we implement a LoRa jammer on commercial LoRa devices by modifying the open source and figure out the proper setting of the jammer through three scenarios aiming to evaluate the influence of LoRa transmission configuration on jamming performance. Specifically, the impact of non-orthogonality of LoRa transmission on jamming effect is investigated. Possible countermeasures for LoRaWAN are then presented to alleviate the jamming attacks.
Chin-Ya Huang, Ching-Wei Lin, Ray-Guang Cheng, Shanchieh Jay Yang, Shiann-Tsong Sheu
VTC Spring4
2019 ASSERT: attack synthesis and separation with entropy redistribution towards predictive cyber defense
abstract
The sophistication of cyberattacks penetrating into enterprise networks has called for predictive defense beyond intrusion detection, where different attack strategies can be analyzed and used to anticipate next malicious actions, especially the unusual ones. Unfortunately, traditional predictive analytics or machine learning techniques that require training data of known attack strategies are not practical, given the scarcity of representative data and the evolving nature of cyberattacks. This paper describes the design and evaluation of a novel automated system, ASSERT, which continuously synthesizes and separates cyberattack behavior models to enable better prediction of future actions. It takes streaming malicious event evidences as inputs, abstracts them to edge-based behavior aggregates, and associates the edges to attack models, where each represents a unique and collective attack behavior. It follows a dynamic Bayesian-based model generation approach to determine when a new attack behavior is present, and creates new attack models by maximizing a cluster validity index. ASSERT generates empirical attack models by separating evidences and use the generated models to predict unseen future incidents. It continuously evaluates the quality of the model separation and triggers a re-clustering process when needed. Through the use of 2017 National Collegiate Penetration Testing Competition data, this work demonstrates the effectiveness of ASSERT in terms of the quality of the generated empirical models and the predictability of future actions using the models.
Ahmet Okutan, Shanchieh Jay Yang
Cybersecur.2
2018 Extracting and Evaluating Similar and Unique Cyber Attack Strategies from Intrusion Alerts
abstract
Intrusion detection system (IDS) is an integral part of computer networks to monitor and detect threats. However, the alerts raised by these systems are often overwhelming to security analysts, making it difficult to uncover the steps an attacker took to compromise one or more systems in the network. This work presents a novel approach that aggregates IDS alerts and forms sequences of attack activities and their corresponding probabilistic models. This allows comparison of attack sequences to offer insights for unique as well as similar attack behaviors. We aggregate alerts by performing a Gaussian filter on specific alert attributes and model attackers using a suffix-based probabilistic model. We compare sequences generated from ten independent attacking teams with similar objectives demonstrating how our process uncovers similarities and uniqueness between the attacks that was not obvious. The sequences revealed by our process creates meaningful sequences that offers insights on how the attacking teams exploit a network.
Stephen Moskal, Shanchieh Jay Yang, Michael E. Kuhl
ISI2
2018 Leveraging Intra-Day Temporal Variations to Predict Daily Cyberattack Activity
abstract
Cyber attacks against organizations are occurring with increasing regularity. Defensive systems are in place that can detect malicious traffic within a network. However, these systems can only provide analysis after malicious activity has occurred. What if one can forecast the number of cyberattacks expected for a future day with reasonable accuracies? This paper investigates the use of Auto-Regressive Integrated Moving Average (ARIMA) models to forecast daily counts of different cyberattack types against multiple targets. Smaller measurement periods are used to better capture temporal trends in attack data and increase forecasting accuracy, reducing error by over 14% compared to naive predictions based on average historical occurrence rates. Aggregation techniques are employed to construct a daily forecast using a number of smaller predictions, providing over 11% more accuracy than standard ARIMA models based on daily counts. Temporal intensity variations are leveraged as regressors to further improve model accuracy by over 11% compared to aggregated forecasts. The ARIMA with intensity-based regressors were put into testing to perform predictions up to 7 days in advance, and achieved over 15% improvement over the baseline. ARIMA is able to reduce forecasting error compared to naive approaches, showing that cyber incidents do not occur completely randomly and could be captured and modeled with statistical time series forecasting techniques.
Gordon Werner, Shanchieh Jay Yang, Katie McConky
ISI2
2018 Forecasting cyberattacks with incomplete, imbalanced, and insignificant data
abstract
Having the ability to forecast cyberattacks before they happen will unquestionably change the landscape of cyber warfare and cyber crime. This work predicts specific types of attacks on a potential victim network before the actual malicious actions take place. The challenge to forecasting cyberattacks is to extract relevant and reliable signals to treat sporadic and seemingly random acts of adversaries. This paper builds on multi-faceted machine learning solutions and develops an integrated system to transform large volumes of public data to aggregate signals with imputation that are relevant and predictive of cyber incidents. A comprehensive analysis of the individual parts and the integrated whole demonstrates the effectiveness and trade-offs of the proposed approach. Using 16-months of reported cyber incidents by an anonymized victim organization, the integrated approach achieves up to 87%, 90%, and 96% AUC for forecasting endpoint-malware, malicious-destination, and malicious-email attacks, respectively. When assessed month-by-month, the proposed approach shows robustness to perform consistently well, achieving F -Measure between 0.6 and 1.0. The framework also enables an examination of which unconventional signals are meaningful for cyberattack forecasting.
Ahmet Okutan, Gordon Werner, Shanchieh Jay Yang, Katie McConky
Cybersecur.3
2017 POSTER: Cyber Attack Prediction of Threats from Unconventional Resources (CAPTURE)
abstract
This paper outlines the design, implementation and evaluation of CAPTURE - a novel automated, continuously working cyber attack forecast system. It uses a broad range of unconventional signals from various public and private data sources and a set of signals forecasted via the Auto-Regressive Integrated Moving Average (ARIMA) model. While generating signals, auto cross correlation is used to find out the optimum signal aggregation and lead times. Generated signals are used to train a Bayesian classifier against the ground truth of each attack type. We show that it is possible to forecast future cyber incidents using CAPTURE and the consideration of the lead time could improve forecast performance.
Ahmet Okutan, Gordon Werner, Katie McConky, Shanchieh Jay Yang
CCS4
2017 Modeling Information Sharing Behavior on Q&A Forums
Biru Cui, Shanchieh Jay Yang, Christopher Homan
PAKDD (2)2
2015 Privacy Sensitive Resource Access Monitoring for Android Systems
abstract
Existing works have studied how to collect and analyze human usage of mobile devices, to aid in further understanding of human behavior. Typical data collection utilizes applications or background services installed on the mobile device with user permission to collect user usage data via accelerometer, call logs, location, Wi-Fi transmission, etc. through a data tainting process. Built on the existing work, this research developed a system called Panorama (Privacy-sensitive Resource Access Monitoring for Android Systems) to collect application behavior instead of user behavior. The goal is to provide the means to analyze how background services access mobile resources, and potentially to identify suspicious applications that access sensitive user information. Panorama tracks the access of mobile resources in real time and enhances the concept of taint tracking. Each identified user privacy-sensitive resource is tagged and marked for tracking. The result is a dynamic, real-time tool that monitors the process flow of applications. This paper presents the development of Panorama and a set of analysis with respect to a variety of legitimate application behaviors.
Leah Zhao, Neil Wong Hon Chan, Shanchieh Jay Yang, Roy W. Melton
ICCCN3
2014 Non-independent Cascade Formation: Temporal and Spatial Effects
abstract
Determining cascade size and the factors affecting cascade size are two fundamental research problems in social network analysis. The commonly considered independent cascade model, when applied to social networks such as Digg, produces a phase-transition phenomenon where the cascade is either very small or very large. This phenomenon can be explained based on the concept of Giant Propagation Component (GPC). The GPC is defined as a maximally connected component, such that, by applying the independent cascade model, once any node of the component is infected, most of the remaining nodes in the component will eventually become infected with a high probability. While GPC exists in social networks, the phase-transition phenomenon, is not observed in the actual cascade size distribution when the information propagation is due to actions such as ``like'' or ``dig''.
Biru Cui, Shanchieh Jay Yang, Christopher Homan
CIKM2
2014 Probabilistic Inference for Obfuscated Network Attack Sequences
abstract
Facing diverse network attack strategies and overwhelming alters, much work has been devoted to correlate observed malicious events to pre-defined scenarios, attempting to deduce the attack plans based on expert models of how network attacks may transpire. Sophisticated attackers can, however, employ a number of obfuscation techniques to confuse the alert correlation engine or classifier. Recognizing the need for a systematic analysis of the impact of attack obfuscation, this paper models attack strategies as general finite order Markov models, and treats obfuscated observations as noises. Taking into account that only finite observation window and limited computational time can be afforded, this work develops an algorithm to efficiently inference on the joint distribution of clean and obfuscated attack sequences. The inference algorithm recovers the optimal match of obfuscated sequences to attack models, and enables a systematic and quantitative analysis on the impact of obfuscation on attack classification.
Haitao Du, Shanchieh Jay Yang
DSN2
2013 Introduction to the special section on social computing, behavioral-cultural modeling, and prediction
abstract
No abstract available.
Shanchieh Jay Yang, Dana S. Nau, John J. Salerno
ACM Trans. Intell. Syst. Technol.1
2011 Optimizing collection requirements through analysis of plausible impact
Khiem Tong, Shanchieh Jay Yang, Moises Sudit, Jared Holsopple
FUSION2
2011 Characterizing Transition Behaviors in Internet Attack Sequences
abstract
Cyber attacks from the Internet often span over multiple ports and multiple hosts. This work hypothesizes that there are distinct sequential patterns revealing hacking behavior. A feature called Attack Transition Action (ATA) is defined to represent the changes on attacked destinations and ports over time. The simplicity of the feature enables the development of a probabilistic model, revealing higher order transitions hidden within the attack sequences. The model trained with a real-world attack dataset uncovers several natural clusters of Internet attack behaviors. The discovered behavior patterns are explained with representative hacking strategies. Our systematic modeling and analysis provides an effective means to characterize classes of Internet attacks.
Haitao Du, Shanchieh Jay Yang
ICCCN2
2010 Clustering of multistage cyber attacks using significant services
Chris Murphy, Shanchieh Jay Yang
FUSION2
2010 Issues and challenges in higher level fusion: Threat/impact assessment and intent modeling (a panel summary)
John S. Salerno, Moises Sudit, Shanchieh Jay Yang, George P. Tadda, Ivan Kadar, Jared Holsopple
FUSION3
2010 Toward Ensemble Characterization and Projection of Multistage Cyber Attacks
abstract
With expanding network infrastructures, increasing vulnerabilities and uncertain malicious activities, cyber security research has begun to provide situation assessment beyond Intrusion Detection Systems (IDSs). A key goal of cyber situation assessment is to efficiently and effectively project the likely future targets of ongoing multistage attacks. This work presents two ensemble techniques that combine real-time projection algorithms modeling the behavior, capability, and opportunity of malicious activities in a network. Sugeno fuzzy inference system and Transferable Belief Model are used to combine supporting evidence and resolve conflicts between the algorithm outputs. The two ensemble techniques are analyzed and compared using simulated attack datasets generated for varying network environments and attack parameters. The results are discussed to reveal the benefits and limitations of individual algorithms and ensemble techniques.
Haitao Du, Daniel F. Liu, Jared Holsopple, Shanchieh Jay Yang
ICCCN4
2009 Toward unsupervised classification of non-uniform cyber attack tracks
Haitao Du, Chris Murphy, Jordan Bean, Shanchieh Jay Yang
FUSION4
2008 Real-time fusion and Projection of network intrusion activity
Stephen R. Byers, Shanchieh Jay Yang
FUSION2
2008 FuSIA: Future Situation and Impact Awareness
Jared Holsopple, Shanchieh Jay Yang
FUSION2
2008 Intrusion activity projection for cyber situational awareness
abstract
Previous works in the area of network security have emphasized the creation of intrusion detection systems (IDSs) to flag malicious network traffic and computer usage. Raw IDS data may be correlated and form attack tracks, each of which consists of ordered collections of alerts belonging to a single multi-stage attack. Assessing an attack track in its early stage may reveal the attackerpsilas capability and behavior trends, leading to projections of future intrusion activities. Behavior trends are captured via variable length Markov models (VLMM) without predetermined attack plans. A virtual terrain schema is developed to model network and system configurations, and used to estimate critical elements and vulnerabilities exposed to each attacker given his/her progress. Experimental results show promises for these proactive measures in ensuring continuous and critical cyber operations.
Shanchieh Jay Yang, Stephen R. Byers, Jared Holsopple, Brian Argauer, Daniel S. Fava
ISI1
2008 Projecting Cyberattacks Through Variable-Length Markov Models
abstract
Previous works in the area of network security have emphasized the creation of intrusion detection systems (IDSs) to flag malicious network traffic and computer usage, and the development of algorithms to analyze IDS alerts. One possible byproduct of correlating raw IDS data are attack tracks, which consist of ordered collections of alerts belonging to a single multistage attack. This paper presents a variable-length Markov model (VLMM) that captures the sequential properties of attack tracks, allowing for the prediction of likely future actions on ongoing attacks. The proposed approach is able to adapt to newly observed attack sequences without requiring specific network information. Simulation results are presented to demonstrate the performance of VLMM predictors and their adaptiveness to new attack scenarios.
Daniel S. Fava, Stephen R. Byers, Shanchieh Jay Yang
IEEE Trans. Inf. Forensics Secur.3
2007 Terrain and behavior modeling for projecting multistage cyber attacks
abstract
Contributions from the information fusion community have enabled comprehensible traces of intrusion alerts occurring on computer networks. Traced or tracked cyber attacks are the bases for threat projection in this work. Due to its complexity, we separate threat projection into two sub-tasks: predicting likely next targets and predicting attacker behavior. A virtual cyber terrain is proposed for identifying likely targets. Overlaying traced alerts onto the cyber terrain reveals exposed vulnerabilities, services, and hosts. Meanwhile, a novel attempt to extract cyber attack behavior is discussed. Leveraging traditional work on prediction and compression, this work identities behavior patterns from traced cyber attack data. The extracted behavior patterns are expected to further refine projections deduced from the cyber terrain.
Daniel S. Fava, Jared Holsopple, Shanchieh Jay Yang, Brian Argauer
FUSION3
2004 Enhancing both network and user performance for networks supporting best effort traffic
abstract
With a view on improving user-perceived performance on networks supporting best effort flows, e.g., multimedia/data file transfers, we propose a family of bandwidth allocation criteria that depends on the residual work of on-going transfers. Analysis and simulations show that allocating bandwidth in this fashion can significantly improve the user-perceived delay, bit transmission delay, and throughput over traditional approaches, e.g., by 58% on an 80% loaded linear network. A simple implementation based on TCP Reno, exemplifies how one might approach practically realizing such gains. We discuss several other advantages of incorporating such differentiation at the transport level. In particular we make the case that favoring small transfers combined with user impatience or peak rate constraints, both of which are natural mechanisms for users to express the utility of completing transfers, offers a lightweight approach to achieving good overall network goodput and/or utility for best effort networks.
Shanchieh Jay Yang, Gustavo de Veciana
IEEE/ACM Trans. Netw.1
2002 Size-based Adaptive Bandwidth Allocation: Optimizing the Average QoS for Elastic Flows
abstract
With a view on improving user perceived performance on networks supporting elastic flows, e.g., multimedia/data file transfers, we identify the key properties that an online dynamic bandwidth allocation policy should have. We then propose a family of bandwidth allocation criteria which depends on the residual work of on-going transfers. Analysis and simulations show that allocating bandwidth in this fashion can improve the user perceived average bit transmission delay (BTD), i.e., delay/flow size, by up to 70% at 80% traffic load over traditional approaches. A simple implementation based upon TCP Reno, exemplifies how one might approach practically realizing such gains. Further studies on simple network topologies show that as the penetration of the proposed transport mechanism increases, users will have the proper incentives to upgrade from TCP Reno, and that the overall performance is better for all users once the penetration exceeds 20%.
Shanchieh Jay Yang, Gustavo de Veciana
INFOCOM1
2001 Bandwidth sharing: the role of user impatience
abstract
Empirical work has shown that up to 20% of the volume transferred on data networks might correspond to 'aborted' connections, i.e., badput. With this in mind, we propose two generic models that capture a variety of user impatience behaviors and investigate their impact on user perceived and actual system performance achieved by various bandwidth sharing schemes. Our study suggests that differentiated bandwidth allocation based on job size, rather than using traditional fair share allocations, results in a more 'graceful' performance degradation and, particularly in the presence of impatient users, leads to better network efficiency as well as user perceived performance.
Shanchieh Jay Yang, Gustavo de Veciana
GLOBECOM1
1997 Reactive Bandwidth Arbitration for Priority and Multicasting Control in ATM Switching
abstract
A new service scheduling scheme, reactive bandwidth arbitration (RBA), is proposed as an effective way to arbitrate the contention among priority classes or between unicast and multicast services in ATM switching. The RBA scheme integrates two previously proposed concepts: bandwidth allocation and reactive arbitration. In, RBA, guaranteed bandwidth is allocated to a connection at call set up based on the traffic characteristics and service requirements, while the arbitration of cell delivery takes into account the queue status to react to traffic fluctuation. Through simulation, we found that by bandwidth allocation, a set of queue threshold levels can be obtained to provide a good approximation to the desired delay performance. Furthermore, around this initial configuration, a linear relationship between the queue threshold and the resulting delay performance can be established to fine-tune the configuration parameters for the desired delay performance. Similar effects were observed as the RBA scheme was applied to arbitrate the contention between unicast and multicast connections in a shared buffer ATM switch. The proposed RBA scheme was incorporated in a queue manager chip for an 8/spl times/8 shared buffer ATM switch with four priority classes per port and link rate at 622 Mbps. The chip has 130 k gates in a chip area of 137.88 mm/sup 2/ and operates at 25 MHz.
Shu-Jen Fang, Shanchieh Jay Yang, C. Bernard Shung
ICC (3)3