VLDB 2026 Research / reviewers in the wild / expert
Sagar Samtani
dblp:165/9230
· DBLP profile ↗
45ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0002-4513-805XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 30 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Computer networks · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VoiceFormer: Fusing Non-Acoustic Motion Sensors for High-Fidelity Voice Synthesis in Mobile DevicesabstractWith the popularity of mobile devices, a variety of motion sensors are integrated to enhance the user experience. Although existing studies demonstrated that non-acoustic motion sensors can be attacked by adversaries, they overlook the limited sampling frequencies of motion sensors (e.g., < 500 Hz) in mobile devices and are evaluated in the controlled laboratory settings. In this article, we explore a new attack model on non-acoustic motion sensors based on the off-the-shelf mobile devices. We propose a general framework named VoiceFormer to synthesize high-fidelity speeches based on the vibrations of accelerometers and gyroscopes with a low sampling frequency. Specifically, in VoiceFormer , we introduce a signal alignment approach to remove the time offsets between two nonsynchronous signals, and leverage Time Interleaved Analog-Digital-Conversion (TI-ADC) to generate a high-frequency synthetic signal (e.g., > 8 KHz) based on the vibration signals of accelerometers and gyroscopes on the same motherboard. To synthesize the high-fidelity acoustic waveforms, we propose a wavelet-based generative adversarial network to learn the spatiotemporal latent mapping between vibrations and original speech signals. Extensive experimental results demonstrate the feasibility of voice synthesis by spying the low-frequency non-acoustic motion sensors in off-the-shelf mobile devices. VoiceFormer shows impressive performance in the synthesized acoustical signals with a Mean Opinion Score of 3.38. Although there are significant differences of mobile devices in hardware settings, VoiceFormer shows robust performance in synthesizing intelligible voice signals. Our results suggest that eavesdropping an off-the-shelf mobile device remotely by fusing non-acoustic sensors is feasible. Xiaokai Yan, Yunji Liang, Lei Liu 0073, Sagar Samtani, Bin Guo 0001, Zhiwen Yu 0001 |
ACM Trans. Priv. Secur. | 4 |
| 2026 | Corner Case Detection and Generation for Autonomous Driving: An OverviewabstractSafety concerns remain one of the most significant obstacles to the large-scale deployment and continued advancement of autonomous driving (AD) systems. A major underlying cause of many safety-related incidents in AD systems is suboptimal or erroneous decision-making when the vehicle encounters corner cases (CCs)—rare, unexpected, or extreme situations that fall outside typical operating conditions. Although recent advances in artificial intelligence have driven substantial progress in both autonomous driving and corner-case research, the field still lacks a coherent conceptual foundation and a systematic, widely accepted categorization of CCs. In this survey, we address this gap by offering a structured review of the existing corner-case literature along three key dimensions: understanding, detection, and generation. The key contribution is a three-level classification of corner cases—spanning data-level, model-level, and semantic-level CCs—that clarifies and disambiguates competing definitions and perspectives on AD corner cases. Building on this framework, we examine simulator-based methods for corner-case selection and data generation, and we identify open challenges, promising directions, and potential solutions to more effectively handle corner cases in autonomous driving systems. Yunji Liang, Junteng Liu, Xiaokai Yan, Xiaolong Zheng 0001, Lei Tang 0002, Luwen Huangfu, Sagar Samtani, Zhiwen Yu 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2026 | Hard Sample Mining: A New Paradigm of Efficient and Robust Model TrainingabstractOver the past two decades, deep learning (DL) has achieved unprecedented breakthroughs across diverse application domains spanning computer vision (CV) to natural language processing (NLP). However, despite significant advances in computational resources and algorithmic frameworks, the training of deep neural networks continues to present formidable challenges due to persistent issues of training inefficiency and inherent data distribution biases. Recent years have witnessed the emergence of hard sample mining (HSM) as a promising paradigm to mitigate training inefficiencies and enhance model robustness through representative sample selection. Although HSM is reshaping contemporary AI research, its critical role in enabling efficient and robust model training has not yet been systematically explored. This article presents a comprehensive survey of HSM methodologies by: 1) establishing unified definitions of hard samples through rigorous sample complexity quantification criteria; 2) proposing a systematic taxonomy of HSM approaches with in-depth technical analysis; and 3) identifying pivotal research frontiers in this evolving field. This survey not only consolidates the foundations of HSM but also provides a roadmap for advancing efficient, robust, and generalizable deep learning models. Lei Liu 0073, Yunji Liang, Xiaokai Yan, Luwen Huangfu, Sagar Samtani, Zhiwen Yu 0001, Yanyong Zhang, Daniel Dajun Zeng |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2026 | A Hierarchical Hard Negative Sampling Strategy for Robust Out-of-Distribution Object DetectionabstractOut-of-distribution (OOD) detection is crucial for deploying models in open-world environments. This process aims to mitigate the issue of overconfident predictions, which is a common problem for models designed for closed-domain tasks when they encounter OOD data. Recent progress in OOD detection has shown that integrating auxiliary datasets during model training can greatly enhance OOD detection performance. However, existing methods tend to heavily depend on these auxiliary datasets to establish the decision boundary for in-distribution (ID) data, while not adequately addressing OOD object detection in safety–critical applications. In this article, we propose object-level OOD detection and introduce a hierarchical hard negative sampling (HNS) strategy that does not necessitate auxiliary data. Specifically, we offer a new metric that strategically considers difficult negatives near the decision boundary between inter-class and intra-class instances. Inspired by adversarial thinking, we sample outliers for each class, ensuring that the negative samples capture both diversity and informative traits. We conducted comprehensive experiments on three public datasets. The results demonstrate that HNS performs superiorly in object-level OOD detection, even without auxiliary datasets. The source code can be accessed at: https://github.com/Aurevior1/HNS . Junteng Liu, Zizhe Wang, Yunji Liang, Sagar Samtani, Lei Tang 0002, Zhiwen Yu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | What Do Machine Learning Researchers Mean by "Reproducible"?abstractThe concern that Artificial Intelligence (AI) and Machine Learning (ML) are entering a "reproducibility crisis" has spurred significant research in the past few years. Yet with each paper, it is often unclear what someone means by "reproducibility". Our work attempts to clarify the scope of "reproducibility" as displayed by the community at large. In doing so, we propose to refine the research to eight general topic areas. In this light, we see that each of these areas contains many works that do not advertise themselves as being about "reproducibility", in part because they go back decades before the matter came to broader attention. Edward Raff, Michel Benaroch, Sagar Samtani, Andrew L. Farris |
AAAI | 3 |
| 2024 | The 4th Workshop on Artificial Intelligence-enabled Cybersecurity AnalyticsabstractCybersecurity remains a grand societal challenge. Large and constantly changing attack surfaces are non-trivial to protect against malicious actors. Entities like the United States and the European Union have recently emphasized the value of Artificial Intelligence (AI) for advancing cybersecurity. For example, the National Science Foundation has called for AI systems that can enhance cyber threat intelligence, detect new and evolving threats, and analyze massive troves of cybersecurity data. The 4th Workshop on Artificial Intelligence-enabled Cybersecurity Analytics (co-located with ACM KDD) sought to make significant and novel contributions within these relevant topics. Submissions were reviewed by highly qualified AI for cybersecurity researchers and practitioners spanning academia and private industry firms. Steven Ullman, Benjamin Ampel, Sagar Samtani, Shanchieh Jay Yang, Hsinchun Chen |
KDD | 3 |
| 2024 | Understand your shady neighborhood: An approach for detecting and investigating hacker communities
Dalyapraz Manatova, Charles Devries, Sagar Samtani |
Decis. Support Syst. | 3 |
| 2024 | Learning Entangled Interactions of Complex Causality via Self-Paced Contrastive LearningabstractLearning causality from large-scale text corpora is an important task with numerous applications—for example, in finance, biology, medicine, and scientific discovery. Prior studies have focused mainly on simple causality, which only includes one cause-effect pair. However, causality is notoriously difficult to understand and analyze because of multiple cause spans and their entangled interactions. To detect complex causality, we propose a self-paced contrastive learning model, namely N2NCause, to learn entangled interactions between multiple spans. Specifically, N2NCause introduces data enhancement operations to convert implicit expressions into explicit expressions with the most rational causal connectives for the synthesis of positive samples and to invert the directed connection between a cause-effect pair for the synthesis of negative samples. To learn the semantic dependency and causal direction of positive and negative samples, self-paced contrastive learning is proposed to learn the entangled interactions among spans, including the interaction direction and interaction field. We evaluated the performance of N2NCause in three cause-effect detection tasks. The experimental results show that, with the least data annotation efforts, N2NCause demonstrates competitive performance in detecting simple cause-effect relations, and it is superior to existing solutions for the detection of complex causality. Yunji Liang, Lei Liu 0073, Luwen Huangfu, Sagar Samtani, Zhiwen Yu 0001, Daniel Dajun Zeng |
ACM Trans. Knowl. Discov. Data | 4 |
| 2023 | Disrupting Ransomware Actors on the Bitcoin Blockchain: A Graph Embedding ApproachabstractRansomware is a growing problem and significant threat to cybersecurity in the United States. One primary vector for ransomware payments is the Bitcoin network. Network science techniques are a potential approach to analyze ransomware payment networks to discover salient ransomware actors. In this study, we propose a design framework for labeling nodes in a ransomware payment network and identifying key ransomware Bitcoin addresses that can be targeted for disruption. By leveraging semi-supervised graph embedding methodology and updating the loss function of a prevailing algorithm, GraphSAGE, to manage dataset imbalance, we identify key wallets in our ransomware network. We demonstrate the utility of our approach with a case study identifying a Bitcoin wallet that has been reported as a ransomware actor as recently as December 2021 and has transferred over $450 million in Bitcoin. Benjamin Ampel, Kaeli Otto, Sagar Samtani, Hsinchun Chen |
ISI | 3 |
| 2023 | Mapping Exploit Code on Paste Sites to the MITRE ATT&CK Framework: A Multi-label Transformer ApproachabstractCyber-criminals often use information-sharing platforms such as paste sites (e.g., Pastebin) to share vast amounts of malicious text content, such as exploit source code. Careful analysis of malicious paste site content can provide Cyber Threat Intelligence (CTI) about potential threats. In this research, we propose a Convolutional BiLSTM Transformer multi-label classification method that automatically maps paste site exploit source code to the MITRE ATT&CK framework to identify adversarial techniques in support of proactive CTI. The Convolutional BiLSTM Transformer combines a convolutional neural network layer placed before a Transformer block, a concatenated pooling from a global max pooling and global average, and a BiLSTM pair-wise function within the Transformer to capture word and sequence orders. We conducted an multi-label classification experiment where our proposed Convolutional BiLSTM Transformer model achieved state-of-the-art results in terms of accuracy, recall, F1-score, and hamming loss. The results of a case study showed the tactics and tools that are used by malicious actors on paste sites. Benjamin Ampel, Tala Vahedi, Sagar Samtani, Hsinchun Chen |
ISI | 3 |
| 2023 | Assessing the Vulnerabilities of the Open-Source Artificial Intelligence (AI) Landscape: A Large-Scale Analysis of the Hugging Face PlatformabstractArtificial Intelligence (AI) has rapidly proliferated as a critical disruptive technology in the 21st century. Hugging Face hosts pre-trained models, facilitating the sharing and use of open-source code. Hugging Face has been used by 22,000+ organizations, including Intel and Microsoft, with 2.6+ billion model downloads. While Hugging Face democratizes access to AI models, these models may contain unknown security vulnerabilities. In this research, we automatically collect models from Hugging Face, link them to their underlying code bases on GitHub, and perform a large-scale vulnerability assessment of these repositories. Through our approaches, we collected about 110,000 models from Hugging Face and over 29,000 GitHub repositories. Our vulnerability assessment revealed a larger percentage (35.98%) of high-severity vulnerabilities compared to low-severity vulnerabilities (6.79%). This trend in severity levels contradicts the results of severities detected in repositories forked from root repositories and searched repositories. Given that many of the vulnerabilities reside in fundamental AI repositories such as Transformers, the results of this vulnerability assessment have significant implications for supply chain software security and AI risk management more broadly. Adhishree Kathikar, Aishwarya Nair, Ben Lazarine, Agrim Sachdeva, Sagar Samtani |
ISI | 5 |
| 2023 | The 3rd Workshop on Artificial Intelligence-enabled Cybersecurity AnalyticsabstractArtificial Intelligence (AI) has gripped modern society as a viable approach to revolutionize operational capabilities across multiple industries. One critical application area that could stand to benefit from the capabilities of AI is cybersecurity. Increasingly, federal funding agencies such as the National Science Foundation are calling for enhanced AI-enabled analytics capabilities to improve cyber threat intelligence, cyber defense generation, and more. To this end, this half-day workshop, not in its third year at ACM KDD, sought to attain significant contributions related to various aspects of AI-enabled cybersecurity analytics. This workshop received a record number of submissions. Submissions were reviewed by a highly-qualified, interdisciplinary group of AI for cybersecurity researchers and practitioners spanning academia and private industry firms. Sagar Samtani, Shanchieh Jay Yang, Hsinchun Chen |
KDD | 1 |
| 2023 | A deep interpretable representation learning method for speech emotion recognition
Erkang Jing, Ye-Zheng Liu 0001, Yidong Chai, Jianshan Sun, Sagar Samtani, Yuan-Chun Jiang, Yang Qian 0001 |
Inf. Process. Manag. | 5 |
| 2023 | An Escalated Eavesdropping Attack on Mobile Devices via Low-Resolution Vibration SignalsabstractWith the global prevalence of mobile devices, concerns about mobile devices regarding privacy breaches and data leakage are rising. Although sensor permissions are required for mobile applications to access outputs of built-in sensors, motion sensors (e.g., accelerometer and gyroscope) can be visited directly without permission requirement. Extant studies have shown that motion sensors may cause breaches of confidential information, such as passwords, digits, and voice-based commands, but whether it is possible to synthesize intelligible speech waveforms from low-resolution motion sensors has been understudied. In this article, we present an escalated side-channel attack of built-in speakers by synthesizing intelligible speech waveforms from low-resolution vibration signals. Opposite to traditional classification problems, we formulate this task as a generative problem and introduce an end-to-end synthesis framework dubbed asAccMyrinxto eavesdrop on the speaker via the low-resolution vibration signals. InAccMyrinx, we introduce the data alignment solution to provide the pair-wise voice-vibration sequences and present wavelet-based MelGAN (WMelGAN) with multi-scale time-frequency domain discriminators to generate intelligible acoustic waveforms. We conducted intensive experiments and demonstrated the feasibility of synthesizing the intelligible acoustic signals from low-resolution solid-borne vibration signals. Compared with existing synthesis solutions, our proposed solution outperforms the baselines in both subject and object metrics with the smoothed word error rate of 42.67% and the Mel-Cepstral distortion of 0.298. In addition, the quality of synthetic speeches could be impacted by several factors, including gender, speech rate, volume, and sampling frequency. Yunji Liang, Yuchen Qin, Qi Li 0048, Xiaokai Yan, Luwen Huangfu, Sagar Samtani, Bin Guo 0001, Zhiwen Yu 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2023 | Additive Feature Attribution Explainable Methods to Craft Adversarial Attacks for Text Classification and Text RegressionabstractDeep learning (DL) models have significantly improved the performance of text classification and text regression tasks. However, DL models are often strikingly vulnerable to adversarial attacks. Many researchers have aimed to develop adversarial attacks against DL models in realistic black-box settings (i.e., assuming no model knowledge is accessible to attackers). These attacks typically operate with a two-phase framework: (1) sensitivity estimation through gradient-based or deletion-based methods to evaluate the sensitivity of each token to the prediction of the target model, and (2) perturbation execution to craft adversarial examples based on the estimated token sensitivity. However, gradient-based and deletion-based methods used to estimate sensitivity often face issues of capturing token directionality and overlapping token sensitivities, respectively. In this study, we propose a novel eXplanation-based method for Adversarial Text Attacks (XATA) that leverages additive feature attribution explainable methods, namely LIME or SHAP, to measure the sensitivity of input tokens when crafting black-box adversarial attacks on DL models performing text classification or text regression. We evaluated XATA's attack performance on DL models executing text classification on the IMDB Movie Review, Yelp Reviews-Polarity, and Amazon Reviews-Polarity datasets and DL models conducting text regression on the My Personality, Drug Review, and CommonLit Readability datasets. The proposed XATA outperformed the existing gradient-based and deletion-based adversarial attack baselines in both tasks. These findings indicate that the ever-growing research focused on improving the explainability of DL models with additive feature attribution explainable methods can provide attackers with weapons to launch targeted adversarial attacks. Yidong Chai, Ruicheng Liang, Sagar Samtani, Hongyi Zhu 0001, Meng Wang 0001, Ye-Zheng Liu 0001, Yuan-Chun Jiang |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Enforcement of Laws and Privacy Preferences in Modern Computing SystemsabstractModern civilization is highly dependent on computing systems, touching all aspects of business, government, and individual life. At the same time, there has been an increase in laws and privacy preferences whose implementation and effectiveness depend on software. Whereas organizations and individuals have been expected to comply with laws and regulations, now computing systems must also be compliant and accountable. Computing systems need to be designed with privacy preferences and legal statutes in mind, and should be adaptable to change. Murat Kantarcioglu, Barbara Carminati, Sagar Samtani, Sudip Mittal, Maanak Gupta |
CODASPY | 3 |
| 2022 | ACM KDD AI4Cyber/MLHat: Workshop on AI-enabled Cybersecurity Analytics and Deployable DefenseabstractFederal funding agencies and industry entities are seeking innovative approaches to address the ever-growing cybersecurity crisis. Increasingly, numerous cybersecurity thought leaders are indicating that Artificial Intelligence (AI)-enabled analytics can help tackle key cybersecurity tasks and deploy defenses. This half-day workshop, co-located with ACM KDD, sought to attain significant research contributions to various aspects of AI-enabled analytics for cybersecurity applications and deployable defense solutions from academics and practitioners. This workshop was a joint workshop of the 2021 AI-enabled Cybersecurity Analytics and 2021 International Workshop on Deployable Machine Learning for Security Defense. As such, we developed an interdisciplinary Program Committee with significant experience in various aspects of AI, cybersecurity, and/or deployable defense. Sagar Samtani, Gang Wang 0011, Ali Ahmadzadeh, Arridhana Ciptadi, Shanchieh Jay Yang, Hsinchun Chen |
KDD | 1 |
| 2022 | Explainable Artificial Intelligence for Cyber Threat Intelligence (XAI-CTI)abstractThe papers in this special section focus on explainable artificial intelligence for cyber threat intelligence. Despite concerted efforts from industry, academia, and government on improving cybersecurity capabilities, cyber-threats such as ransomware, fake news, advanced malware, and others, continue to exact a substantial toll on modern infrastructure and day-to-day societal operations. To help combat the ever-growing quantity and severity of cyber-threats, many organizations are adopting Cyber Threat Intelligence (CTI). At its core, CTI is a data-driven process that aims to identify emerging threats and key threat actors to help enable effective cybersecurity decision-making. Sagar Samtani, Hsinchun Chen, Murat Kantarcioglu, Bhavani Thuraisingham |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2021 | AI for Security and Security for AIabstractOn one side, the security industry has successfully adopted some AI-based techniques. Use varies from mitigating denial of service attacks, forensics, intrusion detection systems, homeland security, critical infrastructures protection, sensitive information leakage, access control, and malware detection. On the other side, we see the rise of Adversarial AI. Here the core idea is to subvert AI systems for fun and profit. The methods utilized for the production of AI systems are systematically vulnerable to a new class of vulnerabilities. Adversaries are exploiting these vulnerabilities to alter AI system behavior to serve a malicious end goal. This panel discusses some of these aspects. Elisa Bertino, Murat Kantarcioglu, Cuneyt Gurcan Akcora, Sagar Samtani, Sudip Mittal, Maanak Gupta |
CODASPY | 4 |
| 2021 | Exploring the Evolution of Exploit-Sharing Hackers: An Unsupervised Graph Embedding ApproachabstractCybercrime was estimated to cost the global economy $945 billion in 2020. Increasingly, law enforcement agencies are using social network analysis (SNA) to identify key hackers from Dark Web hacker forums for targeted investigations. However, past approaches have primarily focused on analyzing key hackers at a single point in time and use a hacker’s structural features only. In this study, we propose a novel Hacker Evolution Identification Framework to identify how hackers evolve within hacker forums. The proposed framework has two novelties in its design. First, the framework captures features such as user statistics, node-level metrics, lexical measures, and post style, when representing each hacker with unsupervised graph embedding methods. Second, the framework incorporates mechanisms to align embedding spaces across multiple time-spells of data to facilitate analysis of how hackers evolve over time. Two experiments were conducted to assess the performance of prevailing graph embedding algorithms and nodal feature variations in the task of graph reconstruction in five time-spells. Results of our experiments indicate that Text-Associated Deep-Walk (TADW) with all of the proposed nodal features outperforms methods without nodal features in terms of Mean Average Precision in each time-spell. We illustrate the potential practical utility of the proposed framework with a case study on an English forum with 51,612 posts. The results produced by the framework in this case study identified key hackers posting piracy assets. Kaeli Otto, Benjamin Ampel, Sagar Samtani, Hongyi Zhu 0001, Hsinchun Chen |
ISI | 3 |
| 2021 | Identifying and Categorizing Malicious Content on Paste Sites: A Neural Topic Modeling ApproachabstractMalicious cyber activities impose substantial costs on the U.S. economy and global markets. Cyber-criminals often use information-sharing social media platforms such as paste sites (e.g., Pastebin) to share vast amounts of plain text content related to Personally Identifiable Information (PII), credit card numbers, exploit code, malware, and other sensitive content. Paste sites can provide targeted Cyber Threat Intelligence (CTI) about potential threats and prior breaches. In this research, we propose a novel Bidirectional Encoder Representation from Transformers (BERT) with Latent Dirichlet Allocation (LDA) model to categorize pastes automatically. Our proposed BERT-LDA model leverages a neural network transformer architecture to capture sequential dependencies when representing each sentence in a paste. BERT-LDA replaces the Bag-of-Words (BoW) approach in the conventional LDA with a Bag-of-Labels (BoL) that encompasses class labels at the sequence level. We compared the performance of the proposed BERT-LDA against the conventional LDA and BERT-LDA variants (e.g., GPT2-LDA) on 4,254,453 pastes from three paste sites. Experiment results indicate that the proposed BERT-LDA outperformed the standard LDA and each BERT-LDA variant in terms of perplexity on each paste site. Results of our BERT-LDA case study suggest that significant content relating to hacker community activities, malicious code, network and website vulnerabilities, and PII are shared on paste sites. The insights provided by this study could be used by organizations to proactively mitigate potential damage on their infrastructure. Tala Vahedi, Benjamin Ampel, Sagar Samtani, Hsinchun Chen |
ISI | 3 |
| 2021 | ACM KDD AI4Cyber: The 1st Workshop on Artificial Intelligence-enabled Cybersecurity AnalyticsabstractDespite significant contributions to various aspects of cybersecurity, cyber-attacks remain on the unfortunate rise. Increasingly, internationally recognized entities such as the National Science Foundation and National Science & Technology Council have noted Artificial Intelligence can help analyze billions of log files, Dark Web data, malware, and other data sources to help execute fundamental cybersecurity tasks. Our objective for the 1st Workshop on Artificial Intelligence-enabled Cybersecurity Analytics (half-day; co-located with ACM KDD) was to gather academic and practitioners to contribute recent work pertaining to AI-enabled cybersecurity analytics. We composed an outstanding, inter-disciplinary Program Committee with significant expertise in various aspects of AI-enabled Cybersecurity Analytics to evaluate the submitted work. Significant contributions to the half-day workshop were made in the areas of CTI, vulnerability assessment, and malware analysis. Sagar Samtani, Shanchieh Jay Yang, Hsinchun Chen |
KDD | 1 |
| 2021 | Fusion of heterogeneous attention mechanisms in multi-view convolutional neural network for text classification
Yunji Liang, Bin Guo 0001, Zhiwen Yu 0001, Xiaolong Zheng 0001, Sagar Samtani, Daniel Dajun Zeng |
Inf. Sci. | 6 |
| 2021 | Energy-efficient Collaborative Sensing: Learning the Latent Correlations of Heterogeneous SensorsabstractWith the proliferation of Internet of Things (IoT) devices in the consumer market, the unprecedented sensing capability of IoT devices makes it possible to develop advanced sensing and complex inference tasks by leveraging heterogeneous sensors embedded in IoT devices. However, the limited power supply and the restricted computation capability make it challenging to conduct seamless sensing and continuous inference tasks on resource-constrained devices. How to conduct energy-efficient sensing and perform rich-sensor inference tasks on IoT devices is crucial for the success of IoT applications. Therefore, we propose a novel energy-efficient collaborative sensing framework to optimize the energy consumption of IoT devices. Specifically, we explore the latent correlations among heterogeneous sensors via an attention mechanism in temporal convolutional network to quantify the dependency among sensors, and characterize the heterogeneous sensors in terms of energy consumption to categorize them into low-power sensors and energy-intensive sensors . Finally, to decrease the sampling frequency of energy-intensive sensors , we propose a multi-task learning strategy to predict the statuses of energy-intensive sensors based on the low-power sensors . To evaluate the performance of the proposed collaborative sensing framework, we develop a mobile application to collect concurrent heterogeneous data streams from all sensors embedded in Huawei Mate 8. The experimental results show that latent correlation learning is greatly helpful to understand the latent correlations among heterogeneous streams, and it is feasible to predict the statuses of energy-intensive sensors by low-power sensors with high accuracy and fast convergence. In terms of energy consumption, the proposed collaborative sensing framework is able to preserve the energy consumption of IoT devices by nearly 50% for continuous data acquisition tasks. Yunji Liang, Zhiwen Yu 0001, Bin Guo 0001, Xiaolong Zheng 0001, Sagar Samtani |
ACM Trans. Sens. Networks | 6 |
| 2020 | Labeling Hacker Exploits for Proactive Cyber Threat Intelligence: A Deep Transfer Learning ApproachabstractWith the rapid development of new technologies, vulnerabilities are at an all-time high. Companies are investing in developing Cyber Threat Intelligence (CTI) to counteract these new vulnerabilities. However, this CTI is generally reactive based on internal data. Hacker forums can provide proactive CTI value through automated analysis of new trends and exploits. One way to identify exploits is by analyzing the source code that is posted on these forums. These source code snippets are often noisy and unlabeled, making standard data labeling techniques ineffective. This study aims to design a novel framework for the automated collection and categorization of hacker forum exploit source code. We propose a deep transfer learning framework, the Deep Transfer Learning for Exploit Labeling (DTL-EL). DTL-EL leverages the learned representation from professional labeled exploits to better generalize to hacker forum exploits. This model classifies the collected hacker forum exploits into eight predefined categories for proactive and timely CTI. The results of this study indicate that DTL-EL outperforms other prominent models in hacker forum literature. Benjamin Ampel, Sagar Samtani, Hongyi Zhu 0001, Steven Ullman, Hsinchun Chen |
ISI | 2 |
| 2020 | Identifying Vulnerable GitHub Repositories and Users in Scientific Cyberinfrastructure: An Unsupervised Graph Embedding ApproachabstractThe scientific cyberinfrastructure community heavily relies on public internet-based systems (e.g., GitHub) to share resources and collaborate. GitHub is one of the most powerful and popular systems for open source collaboration that allows users to share and work on projects in a public space for accelerated development and deployment. Monitoring GitHub for exposed vulnerabilities can save financial cost and prevent misuse and attacks of cyberinfrastructure. Vulnerability scanners that can interface with GitHub directly can be leveraged to conduct such monitoring. This research aims to proactively identify vulnerable communities within scientific cyberinfrastructure. We use social network analysis to construct graphs representing the relationships amongst users and repositories. We leverage prevailing unsupervised graph embedding algorithms to generate graph embeddings that capture the network attributes and nodal features of our repository and user graphs. This enables the clustering of public cyberinfrastructure repositories and users that have similar network attributes and vulnerabilities. Results of this research find that major scientific cyberinfrastructures have vulnerabilities pertaining to secret leakage and insecure coding practices for high-impact genomics research. These results can help organizations address their vulnerable repositories and users in a targeted manner. Ben Lazarine, Sagar Samtani, Mark W. Patton, Hongyi Zhu 0001, Steven Ullman, Benjamin Ampel, Hsinchun Chen |
ISI | 2 |
| 2020 | Smart Vulnerability Assessment for Scientific Cyberinfrastructure: An Unsupervised Graph Embedding ApproachabstractThe accelerated growth of computing technologies has provided interdisciplinary teams a platform for producing innovative research at an unprecedented speed. Advanced scientific cyberinfrastructures, in particular, provide data storage, applications, software, and other resources to facilitate the development of critical scientific discoveries. Users of these environments often rely on custom developed virtual machine (VM) images that are comprised of a diverse array of open source applications. These can include vulnerabilities undetectable by conventional vulnerability scanners. This research aims to identify the installed applications, their vulnerabilities, and how they vary across images in scientific cyberinfrastructure. We propose a novel unsupervised graph embedding framework that captures relationships between applications, as well as vulnerabilities identified on corresponding GitHub repositories. This embedding is used to cluster images with similar applications and vulnerabilities. We evaluate cluster quality using Silhouette, Calinski-Harabasz, and Davies-Bouldin indices, and application vulnerabilities through inspection of selected clusters. Results reveal that images pertaining to genomics research in our research testbed are at greater risk of high-severity shell spawning and data validation vulnerabilities. Steven Ullman, Sagar Samtani, Ben Lazarine, Hongyi Zhu 0001, Benjamin Ampel, Mark W. Patton, Hsinchun Chen |
ISI | 2 |
| 2020 | On data-driven curation, learning, and analysis for inferring evolving internet-of-Things (IoT) botnets in the wild
Morteza Safaei Pour, Antonio Mangino, Kurt Friday, Matthias Rathbun, Elias Bou-Harb, Farkhund Iqbal, Sagar Samtani, Jorge Crichigno, Nasir Ghani |
Comput. Secur. | 7 |
| 2020 | Behavioral Biometrics for Continuous Authentication in the Internet-of-Things Era: An Artificial Intelligence PerspectiveabstractIn the Internet-of-Things (IoT) era, user authentication is essential to ensure the security of connected devices and the customization of passive services. However, conventional knowledge-based and physiological biometric-based authentication systems (e.g., password, face recognition, and fingerprints) are susceptible to shoulder surfing attacks, smudge attacks, and heat attacks. The powerful sensing capabilities of IoT devices, including smartphones, wearables, robots, and autonomous vehicles enable continuous authentication (CA) based on behavioral biometrics. The artificial intelligence (AI) approaches hold significant promise in sifting through large volumes of heterogeneous biometrics data to offer unprecedented user authentication and user identification capabilities. In this survey article, we outline the nature of CA in IoT applications, highlight the key behavioral signals, and summarize the extant solutions from an AI perspective. Based on our systematic and comprehensive analysis, we discuss the challenges and promising future directions to guide the next generation of AI-based CA research. Yunji Liang, Sagar Samtani, Bin Guo 0001, Zhiwen Yu 0001 |
IEEE Internet Things J. | 2 |
| 2020 | Proactively Identifying Emerging Hacker Threats from the Dark Web: A Diachronic Graph Embedding Framework (D-GEF)abstractCybersecurity experts have appraised the total global cost of malicious hacking activities to be $450 billion annually. Cyber Threat Intelligence (CTI) has emerged as a viable approach to combat this societal issue. However, existing processes are criticized as inherently reactive to known threats. To combat these concerns, CTI experts have suggested proactively examining emerging threats in the vast, international online hacker community. In this study, we aim to develop proactive CTI capabilities by exploring online hacker forums to identify emerging threats in terms of popularity and tool functionality. To achieve these goals, we create a novel Diachronic Graph Embedding Framework (D-GEF). D-GEF operates on a Graph-of-Words (GoW) representation of hacker forum text to generate word embeddings in an unsupervised manner. Semantic displacement measures adopted from diachronic linguistics literature identify how terminology evolves. A series of benchmark experiments illustrate D-GEF's ability to generate higher quality than state-of-the-art word embedding models (e.g., word2vec) in tasks pertaining to semantic analogy, clustering, and threat classification. D-GEF's practical utility is illustrated with in-depth case studies on web application and denial of service threats targeting PHP and Windows technologies, respectively. We also discuss the implications of the proposed framework for strategic, operational, and tactical CTI scenarios. All datasets and code are publicly released to facilitate scientific reproducibility and extensions of this work. Sagar Samtani, Hongyi Zhu 0001, Hsinchun Chen |
ACM Trans. Priv. Secur. | 1 |
| 2019 | Dark-Net Ecosystem Cyber-Threat Intelligence (CTI) ToolabstractThe frequency and costs of cyber-attacks are increasing each year. By the end of 2019, the total cost of data breaches is expected to reach $2.1 trillion through the ever-growing online presence of enterprises and their consumers. The tools to perform these attacks and the breached data can often be purchased within the Dark-net. Many of the threat actors within this realm use its various platforms to broker, discuss, and strategize these cyber-threat assets. To combat these attacks, researchers are developing Cyber-Threat Intelligence (CTI) tools to proactively monitor the ever-growing online hacker community. This paper will detail the creation and use of a CTI tool that leverages a social network to identify cyber-threats across major Dark-net data sources. Through this network, emerging threats can be quickly identified so proactive or reactive security measures can be implemented. Nolan Arnold, Reza Ebrahimi 0001, Ben Lazarine, Mark W. Patton, Hsinchun Chen, Sagar Samtani |
ISI | 7 |
| 2019 | Identifying High-Impact Opioid Products and Key Sellers in Dark Net Marketplaces: An Interpretable Text Analytics ApproachabstractAs the Internet based applications become more and more ubiquitous, drug retailing on Dark Net Marketplaces (DNMs) has raised public health and law enforcement concerns due to its highly accessible and anonymous nature. To combat illegal drug transaction among DNMs, authorities often require agents to impersonate DNM customers in order to identify key actors within the community. This process can be costly in time and resource. Research in DNMs have been conducted to provide better understanding of DNM characteristics and drug sellers' behavior. Built upon the existing work, researchers can further leverage predictive analytics techniques to take proactive measures and reduce the associated costs. To this end, we propose a systematic analytical approach to identify key opioid sellers in DNMs. Utilizing machine learning and text analysis, this research provides prediction of high-impact opioid products in two major DNMs. Through linking the high-impact products and their sellers, we then identify the key opioid sellers among the communities. This work intends to help law enforcement authorities to formulate strategies by providing specific targets within the DNMs and reduce the time and resources required for prosecuting and eliminating the criminals from the market. Po-Yi Du, Reza Ebrahimi 0001, Hsinchun Chen, Randall A. Brown, Sagar Samtani |
ISI | 6 |
| 2018 | Identifying, Collecting, and Presenting Hacker Community Data: Forums, IRC, Carding Shops, and DNMsabstractCyber-attacks cost the global economy over $450 billion annually. To combat this issue, researchers and practitioners put enormous efforts into developing Cyber Threat Intelligence, or the process of identifying emerging threats and key hackers. However, the reliance on internal network data to has resulted in inherently reactive intelligence. CTI experts have urged the importance of proactively studying the large, ever-evolving online hacker community. Despite their CTI value, collecting data from hacker community platforms is a non-trivial task. In this paper, we summarize our efforts in systematically identifying and automatically collecting a large-scale of hacker forums, carding shops, Internet-Relay-Chat, and Dark Net Marketplaces. We also present our efforts to provide this data to the larger CTI community via the AZSecure Hacker Assets Portal (www.azsecure-hap.com). With our methodology, we collected 102 platforms for a total of 43,981,647 records. To the best of our knowledge, this compilation of hacker community data is the largest such collection in academia. Po-Yi Du, Reza Ebrahimi 0001, Sagar Samtani, Ben Lazarine, Nolan Arnold, Rachael Dunn, Sandeep Suntwal, Guadalupe Angeles, Robert Schweitzer, Hsinchun Chen |
ISI | 4 |
| 2018 | Detecting Cyber Threats in Non-English Dark Net Markets: A Cross-Lingual Transfer Learning ApproachabstractRecent advances in proactive cyber threat intelligence rely on early detection of cyber threats in hacker communities. Dark Net Markets (DNMs) are growing platforms in hacker community that provide hackers with highly- specialized tools and products which may not be found in other platforms. While text classification techniques have been used for cyber threat detection in English DNMs, the task is hindered in non-English platforms due to the language barrier and lack of ground-truth data. Current approaches use monolingual models on machine translated data to overcome these challenges. However, the translation errors can deteriorate the classification results. The abundance of data in English DNMs can be leveraged in learning non-English threats without using machine translation. In this study, we show that a deep cross-lingual model that can jointly learn the common language representation from two languages, significantly outperforms a monolingual model learned on machine translated data for identifying cyber threats in non-English DNMs. Unlike most studies, our approach does not require any external data source such as bilingual word embeddings or bilingual lexicons. Our experiments on Russian DNMs show that this approach can achieve better performance than state-of-the-art methods for non-English cyber threat detection in malicious hacker community. Reza Ebrahimi 0001, Mihai Surdeanu, Sagar Samtani, Hsinchun Chen |
ISI | 3 |
| 2018 | Vulnerability Assessment, Remediation, and Automated Reporting: Case Studies of Higher Education InstitutionsabstractScientific advances of higher education institutions make them attractive targets for malicious cyber-attacks. Modern scanners such as Nessus and Burp can pinpoint an organization's vulnerabilities for subsequent mitigation. However, the remediation reports generated from the tools often cause significant information overload while failing to provide actionable solutions. Consequently, higher education institutions lack the appropriate knowledge to improve their cybersecurity posture. In this study, we conduct a large-scale vulnerability assessment of 272 higher education institutions. From the results, we identified vulnerabilities that fail to provide comprehensive remediation strategies. Selected flaws are recreated and remediated in a virtual environment to develop enhanced, automated reporting mechanisms that provide succinct reports to enable the efficient vulnerability remediation. Our enhanced reports address 27.80% of vulnerabilities found in scanned higher education institutions. Christopher R. Harrell, Mark W. Patton, Hsinchun Chen, Sagar Samtani |
ISI | 4 |
| 2018 | Benchmarking Vulnerability Assessment Tools for Enhanced Cyber-Physical System (CPS) ResiliencyabstractCyber-Physical Systems (CPSs) are engineered systems seamlessly integrating computational algorithms and physical components. CPS advances offer numerous benefits to domains such as health, transportation, smart homes and manufacturing. Despite these advances, the overall cybersecurity posture of CPS devices remains unclear. In this paper, we provide knowledge on how to improve CPS resiliency by evaluating and comparing the accuracy, and scalability of two popular vulnerability assessment tools, Nessus and OpenVAS. Accuracy and suitability are evaluated with a diverse sample of pre-defined vulnerabilities in Industrial Control Systems (ICS), smart cars, smart home devices, and a smart water system. Scalability is evaluated using a large-scale vulnerability assessment of 1,000 Internet accessible CPS devices found on Shodan, the search engine for the Internet of Things (IoT). Assessment results indicate several CPS devices from major vendors suffer from critical vulnerabilities such as unsupported operating systems, OpenSSH vulnerabilities allowing unauthorized information disclosure, and PHP vulnerabilities susceptible to denial of service attacks. Emma McMahon, Mark W. Patton, Sagar Samtani, Hsinchun Chen |
ISI | 3 |
| 2018 | Incremental Hacker Forum Exploit Collection and Classification for Proactive Cyber Threat Intelligence: An Exploratory StudyabstractCyber threats have emerged as a key societal concern. To counter the growing threat of cyber-attacks, organizations, in recent years, have begun investing heavily in developing Cyber Threat Intelligence (CTI). Fundamentally a data driven process, many organizations have traditionally collected and analyzed data from internal log files, resulting in reactive CTI. The online hacker community can offer significant proactive CTI value by alerting organizations to threats they were not previously aware of. Amongst various platforms, forums provide the richest metadata, data permanence, and tens of thousands of freely available Tools, Techniques, and Procedures (TTP). However, forums often employ anti-crawling measures such as authentication, throttling, and obfuscation. Such limitations have restricted many researchers to batch collections. This exploratory study aims to (1) design a novel web crawler augmented with numerous anti-crawling countermeasures to collect hacker exploits on an ongoing basis, (2) employ a state-of-the-art deep learning approach, Long Short-Term Memory (LSTM) Recurrent Neural Network (RNN), to automatically classify exploits into pre-defined categories on-the-fly, and (3) develop interactive visualizations enabling CTI practitioners and researchers to explore collected exploits for proactive, timely CTI. The results of this study indicate, among other findings, that system and network exploits are shared significantly more than other exploit types. Sagar Samtani, Mark W. Patton, Hsinchun Chen |
ISI | 2 |
| 2017 | Benchmarking vulnerability scanners: An experiment on SCADA devices and scientific instrumentsabstractCybersecurity is a critical concern in society today. One common avenue of attack for malicious hackers is exploiting vulnerable websites. It is estimated that there are over one million websites that are attacked daily. Two emerging targets of such attacks are Supervisory Control and Data Acquisition (SCADA) devices and scientific instruments. Vulnerability assessment tools can help provide owners of these devices with the knowledge on how to protect their infrastructure. However, owners face difficulties in identifying which tools are ideal for their assessments. This research aims to benchmark two state-of-the-art vulnerability assessment tools, Nessus and Burp Suite, in the context of SCADA devices and scientific instruments. We specifically focus on identifying the accuracy, scalability, and vulnerability results of the scans. Results of our study indicate that both tools together can provide a comprehensive assessment of the vulnerabilities in SCADA devices and scientific instruments. Malaka El, Emma McMahon, Sagar Samtani, Mark W. Patton, Hsinchun Chen |
ISI | 3 |
| 2017 | Identifying mobile malware and key threat actors in online hacker forums for proactive cyber threat intelligenceabstractCyber-attacks are constantly increasing and can prove difficult to mitigate, even with proper cybersecurity controls. Currently, cyber threat intelligence (CTI) efforts focus on internal threat feeds such as antivirus and system logs. While this approach is valuable, it is reactive in nature as it relies on activity which has already occurred. CTI experts have argued that an actionable CTI program should also provide external, open information relevant to the organization. By finding information about malicious hackers prior to an attack, organizations can provide enhanced CTI and better protect their infrastructure. Hacker forums can provide a rich data source in this regard. This research aims to proactively identify mobile malware and associated key authors. Specifically, we use a state-of-the-art neural network architecture, recurrent neural networks, to identify mobile malware attachments followed by social network analysis techniques to determine key hackers disseminating the mobile malware. Results of this study indicate that many identified attachments are zipped Android apps made by threat actors holding administrative positions in hacker forums. Our identified mobile malware attachments are consistent with some of the emerging mobile malware concerns as highlighted by industry leaders. John Grisham, Sagar Samtani, Mark W. Patton, Hsinchun Chen |
ISI | 2 |
| 2017 | Assessing medical device vulnerabilities on the Internet of ThingsabstractInternet enabled medical devices offer patients with a level of convenience. In recent years, the healthcare industry has seen a surge in the number of cyber-attacks. Given the potentially fatal impact of a compromised medical device, this study aims to identify vulnerabilities of medical devices. Our approach uses Shodan to obtain a large collection of IP addresses that will be passed through Nessus to verify if any vulnerabilities exist. We determined some devices manufactured by primary vendors such as Omron Corporation, FORA, Roche, and Bionet contain serious vulnerabilities such as Dropbear SSH Server and MS17-010. These allow remote execution of code and authentication bypassing potentially giving attackers control of their systems. Emma McMahon, Malaka El, Sagar Samtani, Mark W. Patton, Hsinchun Chen |
ISI | 4 |
| 2017 | Identifying vulnerabilities of consumer Internet of Things (IoT) devices: A scalable approachabstractThe Internet of Things becomes more defined year after year. Companies are looking for novel ways to implement various smart capabilities into their products that increase interaction between users and other network devices. While many smart devices offer greater convenience and value, they also present new security vulnerabilities that can have a detrimental effect on consumer privacy. Given the societal impact of IoT device vulnerabilities, this study aims to perform a large-scale vulnerability assessment of consumer IoT devices exposed on the Internet. Specifically, Shodan is used to collect a large testbed of consumer IoT devices which are then passed through Nessus to determine whether potential vulnerabilities exist. Results of this study indicate that a significant number of consumer IoT devices are vulnerable to exploits that can compromise user information and privacy. Emma McMahon, Sagar Samtani, Mark W. Patton, Hsinchun Chen |
ISI | 3 |
| 2016 | Using social network analysis to identify key hackers for keylogging tools in hacker forumsabstractCyber-attacks are critical cybersecurity concerns across the world. Catching malicious hackers prior to a cyber-attack can save significant financial cost as well as avoid devastating cyber-attacks. Current methods of identifying and reprimanding hackers generally occurs after an attack and is reactive in nature. This research aims to proactively identify key hackers who are creating and disseminating malicious tools within hacker forums. Specifically, we utilize social network analysis techniques to systematically identify key hackers for keylogging tools within a large English hacker forum. Results of this study indicate that many key hackers are the most senior, longest tenured participants of their community. Sagar Samtani, Hsinchun Chen |
ISI | 1 |
| 2016 | AZSecure Hacker Assets Portal: Cyber threat intelligence and malware analysisabstractCyber threats pose grave national security dangers to the US. Many cyber-attacks today are executed with ever-growing collection of malicious tools. Cyber threat intelligence (CTI) and malware analysis portals aim to provide knowledge and tools to help prevent and mitigate attacks. However, current CTI and malware analysis portals and techniques have been criticized for being too reactive as they rely on data collected from past cyber-attacks. Online hacker forums provide a novel source of data that can inform a proactive CTI and malware portal. This research demonstrates the AZSecure Hacker Assets Portal. This portal collects and analyzes malicious assets directly from the largely untapped and rich data source of online hacker communities by utilizing state-of-the-art machine learning techniques. This paper explores the development and evolution of the AZSecure Hacker Assets Portal. We also present key portal functionalities such as asset searching, browsing, and downloading, source code visualizations and code comparison analytics, and an interactive CTI dashboard. Sagar Samtani, Kory Chinn, Cathy Larson, Hsinchun Chen |
ISI | 1 |
| 2016 | Identifying SCADA vulnerabilities using passive and active vulnerability assessment techniquesabstractCritical infrastructure such as power plants, oil refineries, and sewage are at the core of modern society. Supervisory Control and Data Acquisition (SCADA) systems were designed to allow human operators supervise, maintain, and control critical infrastructure. Recent years has seen an increase in connectivity of SCADA systems to the Internet. While this connectivity provides an increased level of convenience, it also increases their susceptibility to cyber-attacks. Given the potentially severe ramifications of exploiting SCADA systems, the purpose of this study is to utilize passive and active vulnerability assessment techniques to identify the vulnerabilities of Internet enabled SCADA systems. Specifically, we collect a large testbed of SCADA devices from Shodan, a search engine for the IoT, and assess their vulnerabilities with Nessus and against the National Vulnerability Database (NVD). Results of this study indicate that many SCADA systems from major vendors such as Rockwell Automation and Siemens are vulnerable to default credential, man-in-the-middle, and SSH exploit attacks. Sagar Samtani, Shuo Yu 0002, Hongyi Zhu 0001, Mark W. Patton, Hsinchun Chen |
ISI | 1 |
| 2015 | Exploring hacker assets in underground forumsabstractMany large companies today face the risk of data breaches via malicious software, compromising their business. These types of attacks are usually executed using hacker assets. Researching hacker assets within underground communities can help identify the tools which may be used in a cyberattack, provide knowledge on how to implement and use such assets and assist in organizing tools in a manner conducive to ethical reuse and education. This study aims to understand the functions and characteristics of assets in hacker forums by applying classification and topic modeling techniques. This research contributes to hacker literature by gaining a deeper understanding of hacker assets in well-known forums and organizing them in a fashion conducive to educational reuse. Additionally, companies can apply our framework to forums of their choosing to extract their assets and appropriate functions. Sagar Samtani, Ryan Chinn, Hsinchun Chen |
ISI | 1 |