EDBT 2026 Demo / reviewers in the wild / expert
Kevin Leach
dblp:133/3698
· DBLP profile ↗
55ranked-venue papers
3as first author
34since 2021 · last 2026
0000-0002-4001-3442ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 19 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 18 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Systems, architecture and hardware · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EyeMulator: Improving Code Language Models by Mimicking Human Visual AttentionabstractYifan Zhang, Chen Huang, Yueke Zhang, Jiahao Zhang, Toby Jia-Jun Li, Collin McMillan, Kevin Leach, Yu Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yifan Zhang 0013, Chen Huang 0006, Yueke Zhang, Toby Jia-Jun Li, Collin McMillan, Kevin Leach, Yu Huang 0015 |
ACL (1) | 7 |
| 2026 | EyeLayer: Integrating Human Attention Patterns into LLM-Based Code SummarizationabstractCode summarization is the task of generating natural language descriptions of source code, which is critical for software comprehension and maintenance. While large language models (LLMs) have achieved remarkable progress on this task, an open question remains: can human expertise in code understanding further guide and enhance these models? We propose EyeLayer, a lightweight attention-augmentation module that incorporates human eye-gaze patterns, as a proxy of human expertise, into LLM-based code summarization. EyeLayer models human attention during code reading via a Multimodal Gaussian Mixture, redistributing token embeddings based on learned parameters \((\mu _i, \sigma _i^2)\) that capture where and how intensively developers focus. This design enables learning generalizable attention priors from eye-tracking data and incorporating them into LLMs seamlessly, without disturbing existing representations. We evaluate EyeLayer across diverse model families (i.e., LLaMA-3.2, Qwen3, and CodeBERT) covering different scales and architectures. EyeLayer consistently outperforms strong fine-tuning baselines across standard metrics, achieving gains of up to 13.17% on BLEU-4. These results demonstrate that human gaze patterns encode complementary attention signals that enhance the semantic focus of LLMs and transfer effectively across diverse models for code summarization. Yifan Zhang 0013, Kevin Leach, Yu Huang 0015 |
ICPC | 3 |
| 2026 | SUGAR: A Sweeter Spot for Generative Unlearning of Many IdentitiesabstractRecent advances in 3D-aware generative models have enabled high-fidelity image synthesis of human identities. However, this progress raises urgent questions around user consent and the ability to remove specific individuals from a model’s output space. We address this by introducing SUGAR, a framework for scalable generative un-learning that enables the removal of many identities (simultaneously or sequentially) without retraining the entire model. Rather than projecting unwanted identities to unrealistic outputs or relying on static template faces, SUGAR learns a personalized surrogate latent for each identity, diverting reconstructions to visually coherent alternatives while preserving the model’s quality and diversity. We further introduce a continual utility preservation objective that guards against degradation as more identities are forgotten. SUGAR achieves state-of-the-art performance in removing up to 200 identities, while delivering up to a 700% improvement in retention utility compared to existing baselines. Our code is publicly available at https://github.com/judydnguyen/SUGAR-Generative-Unlearn. Dung Thuy Nguyen, Preston Robinette, Eli Jiang, Taylor T. Johnson, Kevin Leach |
WACV | 6 |
| 2026 | Building Confidential Accelerator Computing Environment for Arm CCA
Chenxu Wang 0005, Fengwei Zhang, Yunjie Deng 0001, Kevin Leach, Jiannong Cao 0001, Zhenyu Ning, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | ComCat: Expertise-Guided Context Generation to Enhance Code ComprehensionabstractSoftware maintenance constitutes a substantial portion of the total lifetime costs of software, with a significant portion attributed to code comprehension. Software comprehension is eased by documentation such as comments that summarize and explain code. We present ComCat , an approach to automate comment generation by augmenting Large Language Models (LLMs) with expertise-guided context to target the annotation of source code with comments that improve comprehension. Our approach enables the selection of the most relevant and informative comments for a given snippet or file containing source code. We develop the ComCat pipeline to comment C/C++ files by (1) automatically identifying suitable locations in which to place comments, (2) predicting the most helpful type of comment for each location, and (3) generating a comment based on the selected location and comment type. In a human subject evaluation, we demonstrate that ComCat -generated comments significantly improve developer code comprehension across three indicative software engineering tasks by up to 13% for 80% of participants. In addition, we demonstrate that ComCat -generated comments are at least as accurate and readable as human-generated comments and are preferred over standard ChatGPT-generated comments for up to 92% of snippets of code. Furthermore, we develop and release a dataset containing source code snippets, human-written comments, and human-annotated comment categories. ComCat leverages LLMs to offer a significant improvement in code comprehension across a variety of human software engineering tasks. Skyler Grandel, Scott Thomas Andersen, Yu Huang 0015, Kevin Leach |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2025 | Spurious Cues in RVL-CDIP and Tobacco3482 Document Classification: The Case of ID CodesabstractRVL-CDIP and Tobacco3482 are commonly used document classification benchmarks, but recent work on explainability has revealed that ID codes stamped on the documents in these datasets may be used by machine learning models to learn shortcuts on the classification task. In this paper, we present an in-depth investigation into the influence and impact of these ID codes on model performance. We annotate ID codes in documents from RVL-CDIP and Tobacco3482 and find that shallow learning models can achieve classification accuracy scores of roughly 40% on RVL-CDIP and 60% on Tobacco3482 using only features derived from the ID codes. We also find that a state-of-the-art document classifier sees a performance drop of 11 accuracy points on RVL-CDIP when ID codes are removed from the data. Finally, we train an ID code detection model in order to remove ID codes from RVL-CDIP and Tobacco3482 and make this data publicly available. Stefan Larson, Sharad Duwal, Brian Vilnrotter, Gayatri Chakkithara, Vedant Padwal, Kevin Leach |
DocEng | 6 |
| 2025 | Document Classification using File NamesabstractRapid document classification is critical in several time-sensitive applications like digital forensics and large-scale media classification. Traditional approaches that rely on heavy-duty deep learning models fall short due to high inference times over vast input datasets and computational resources associated with analyzing whole documents. In this paper, we present a method using lightweight supervised learning models, combined with a TF-IDF feature extraction-based tokenization method, to accurately and efficiently classify documents based solely on their file names, which substantially reduces inference time. Experiments on two datasets introduced in this paper show that our file name classifiers correctly predict more than 90% of in-scope documents with 99.63% and 96.57% accuracy while being 442x faster than more complex models such as DiT. Our results demonstrate that incorporating lightweight file name classification as a front-end to document analysis pipelines can efficiently process vast document datasets in critical scenarios, enabling fast and more reliable document classification. Stefan Larson, Kevin Leach |
DocEng | 3 |
| 2025 | A Human Study of Automatically Generated Decompiler AnnotationsabstractReverse engineering is a crucial technique in software security, enabling professionals to analyze malware, identify vulnerabilities, and patch legacy software without access to source code. Although decompilers attempt to reconstruct high-level code from binaries, essential information, such as variable names and types, is often dissimilar from the original version, hindering readability and comprehension.Recent advancements have employed AI to enhance decompiler output by recovering original variable names and types. Traditional evaluation of recovery techniques relies on measuring similarity between original and recovered names, assuming that higher similarity enhances readability. However, studies suggest that these "intrinsic" metrics may not accurately predict "extrinsic" outcomes like user comprehension or task performance, revealing a gap in understanding readability and cognitive load in reverse engineering.This paper presents an extrinsic evaluation of machine-generated variable and type names, focusing on their impact on reverse engineers’ comprehension of decompiled code. We conducted a user study with 40 participants—including students and professionals—to assess code comprehension both with and without AI-generated variable and type name assistance. Our findings indicate a lack of correlation between traditional machine learning metrics and actual comprehension gains, highlighting limitations in current evaluation techniques. Despite this, participants showed a preference for AI-augmented decompiler outputs. These insights contribute to understanding the effectiveness of automatic recovery techniques in enhancing reverse engineering tasks and underscore the need for comprehensive, user-centered evaluation frameworks. Skyler Grandel, Jeremy Lacomis, Edward J. Schwartz, Bogdan Vasilescu, Claire Le Goues, Kevin Leach |
DSN | 7 |
| 2025 | MalMixer: Few-Shot Malware Classification with Retrieval-Augmented Semi-Supervised LearningabstractRecent growth and proliferation of malware have tested practitioners’ ability to promptly classify new samples according to malware families. In contrast to labor-intensive reverse engineering efforts, machine learning approaches have demonstrated increased speed and accuracy. However, most existing deep-learning malware family classifiers must be calibrated using a large number of samples that are painstakingly manually analyzed before training. Furthermore, as novel malware samples arise that are beyond the scope of the training set, additional reverse engineering effort must be employed to update the training set. The sheer volume of new samples found in the wild creates substantial pressure on practitioners’ ability to reverse engineer enough malware to adequately train modern classifiers.In this paper, we present MALMIXER, a malware family classifier using semi-supervised learning that achieves high accuracy with sparse training data. We present a domain-knowledge-aware data augmentation technique for malware feature representations, enhancing few-shot performance of semi-supervised malware family classification. We show that MALMIXER achieves state-of-the-art performance in few-shot malware family classification settings. Our research confirms the feasibility and effectiveness of lightweight, domain-knowledge-aware data augmentation methods for malware features and shows the capabilities of similar semi-supervised classifiers in addressing malware classification issues. Yifan Zhang 0013, Yu Huang 0015, Kevin Leach |
EuroS&P | 4 |
| 2025 | PARDON: Privacy-Aware and Robust Federated Domain GeneralizationabstractWhile Federated Learning (FL) shows promise in preserving privacy and enabling collaborative learning, most current solutions concentrate on private data collected from a single domain. Yet, a substantial performance degradation on unseen domains arises when data among clients is drawn from diverse domains (i.e., domain shift). However, existing Federated Domain Generalization (FedDG) methods are typically designed under the assumption that each client has access to the complete dataset of a single domain. This assumption hinders their performance in real-world FL scenarios, which are characterized by domain-based heterogeneity—where data from a single domain is distributed heterogeneously across clients—and client sampling, where only a subset of clients participate in each training round.In addition, certain methods enable information sharing among clients, raising privacy concerns as this information could be used to reconstruct sensitive private data. To overcome this limitation, we present PARDON, a novel FedDG paradigm designed to robustly handle more complicated domain distributions between clients while ensuring security. PARDON facilitates client learning across domains by extracting an interpolative style from abstracted local styles obtained from each client and using contrastive learning. This approach provides each client with a multi-domain representation and an unbiased convergent target. Empirical results on multiple datasets, including PACS, Office-Home, and IWildCam, demonstrate PARDON’s superiority over state-of-the-art methods. Notably, our method outperforms state-of-the-art techniques by a margin ranging from 3.64 to 57.22% in terms of accuracy on unseen domains. Our code is available at https://github.com/judydnguyen/PARDON-FedDG. Dung Thuy Nguyen, Taylor T. Johnson, Kevin Leach |
ICDCS | 3 |
| 2025 | Who's Pushing the Code? An Exploration of GitHub ImpersonationabstractGitHub is one of the largest open-source software (OSS) communities for software development and collaboration. Impersonation in the OSS communities refers to the malicious act of assuming another user's identity, often aiming to gain unauthorized access to code, manipulate project outcomes, or spread misinformation. With several recent real-world attacks resulting from impersonation, this issue is becoming more and more concerning within the OSS community. We present the first exploration of the impact of impersonation in GitHub. Specifically, we conduct structured interviews with 17 real-world OSS contributors about their perception of impersonation and corresponding mitigations. Our study reveals that, in general, GitHub users lack awareness of impersonation and underestimate the severity of its implications. After witnessing a demo of impersonation, they show significant concern for the OSS community. Meanwhile, we also demonstrate that the current best practices (i.e., commit signing) that might mitigate impersonation must be improved to encourage use and adoption. We also present and discuss participant perceptions of potential ways to mitigate GitHub impersonation. We collect a dataset comprising 12.5 million commits to investigate the current status of impersonation. Interestingly, we find out that currently impersonation cannot be easily detected. We observe that existing commit histories treat impersonation behavior identically to pull request events, resulting in a lack of detection methods for impersonation. Yueke Zhang, Anda Liang, Pamela J. Wisniewski, Fengwei Zhang, Kevin Leach, Yu Huang 0015 |
ICSE | 6 |
| 2025 | Optimizing Code Runtime Performance Through Context-Aware Retrieval-Augmented GenerationabstractOptimizing software performance through automated code refinement offers a promising avenue for enhancing execution speed and efficiency. Despite recent advancements in LLMs, a significant gap remains in their ability to perform indepth program analysis. This study introduces AutoPatch, an in-context learning approach designed to bridge this gap by enabling LLMs to automatically generate optimized code. Inspired by how programmers learn and apply knowledge to optimize software, AutoPatch incorporates three key components: (1) an analogy-driven framework to align LLM optimization with human cognitive processes, (2) a unified approach that integrates historical code examples and CFG analysis for context-aware learning, and (3) an automated pipeline for generating optimized code through in-context prompting. Experimental results demonstrate that AutoPatch achieves a$\mathbf{7. 3 \%}$improvement in execution efficiency over GPT-4o across common generated executable code, highlighting its potential to advance automated program runtime optimization. Manish Acharya, Yifan Zhang 0013, Kevin Leach, Yu Huang 0015 |
ICPC | 3 |
| 2025 | CodeACT-R: A Cognitive Simulation Framework for Human Attention in Code ReadingabstractReading code is a fundamental activity in both software engineering and computer science education. Understanding the cognitive processes involved in reading code is crucial for identifying effective cognitive strategies, which can inform teaching methods and tooling support for developers. However, collecting large human subject eye tracking datasets, especially for programming tasks, is often costly and time-consuming, limiting its scalability and applicability. To address this issue, we present CodeACT-R, the first cognitive simulation framework tailored for code reading, based on the well-established Adaptive Control of Thought—Rational (ACT-R) architecture from cognitive science. CodeACT-R simulates how humans read code and requires only a small, manageable amount of human data to initiate the simulator design, offering a cost-effective and scalable alternative to traditional data collection methods like eye tracking.Specifically, we first collected real human visual attention data from 48 programmers reading code using eye tracking. These data were then used to develop CodeACT-R, enabling the simulation of human-like code reading behaviors. Our evaluation demonstrates that CodeACT-R is capable of simulating visual attention patterns (i.e., scanpaths) that closely resemble real-world human attention patterns, also accounting for up to 87% of observed pattern variations. Yueke Zhang, Zihan Fang 0001, J. Gregory Trafton, Daniel Levin 0001, Kevin Leach, Yu Huang 0015 |
ASE | 5 |
| 2025 | PBP: Post-training Backdoor Purification for Malware Classifiers
Dung Thuy Nguyen, Ngoc N. Tran, Taylor T. Johnson, Kevin Leach |
NDSS | 4 |
| 2024 | Generating Hard-Negative Out-of-Scope Data with ChatGPT for Intent ClassificationabstractIntent classifiers must be able to distinguish when a user’s utterance does not belong to any supported intent to avoid producing incorrect and unrelated system responses. Although out-of-scope (OOS) detection for intent classifiers has been studied, previous work has not yet studied changes in classifier performance against hard-negative out-of-scope utterances (i.e., inputs that share common features with in-scope data, but are actually out-of-scope). We present an automated technique to generate hard-negative OOS data using ChatGPT. We use our technique to build five new hard-negative OOS datasets, and evaluate each against three benchmark intent classifiers. We show that classifiers struggle to correctly identify hard-negative OOS utterances more than general OOS utterances. Finally, we show that incorporating hard-negative OOS data for training improves model robustness when detecting hard-negative OOS data and general OOS data. Our technique, datasets, and evaluation address an important void in the field, offering a straightforward and inexpensive way to collect hard-negative OOS data and improve intent classifiers’ robustness. Stefan Larson, Kevin Leach |
LREC/COLING | 3 |
| 2024 | De-Identification of Sensitive Personal Data in Datasets Derived from IIT-CDIPabstractStefan Larson, Nicole Cornehl Lima, Santiago Pedroza Diaz, Amogh Manoj Joshi, Siddharth Betala, Jamiu Tunde Suleiman, Yash Mathur, Kaushal Kumar Prajapati, Ramla Alakraa, Junjie Shen, Temi Okotore, Kevin Leach. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Stefan Larson, Nicole Lima, Santiago Diaz, Amogh Manoj Joshi, Siddharth Betala, Jamiu Suleiman, Yash Mathur, Kaushal Prajapati, Ramla Alakraa, Junjie Shen 0011, Temi Okotore, Kevin Leach |
EMNLP | 12 |
| 2024 | Breaking the Flow: A Study of Interruptions During Software Engineering ActivitiesabstractIn software engineering, interruptions during tasks can have significant implications for productivity and well-being. While previous studies have investigated the effect of interruptions on productivity, to the best of our knowledge, no prior work has yet distinguished the effect of different types of interruptions on software engineering activities. Yimeng Ma, Yu Huang 0015, Kevin Leach |
ICSE | 3 |
| 2024 | ChatGPT Giving Relationship Advice - How Reliable Is It?abstractIn the evolving realm of natural language processing (NLP), generative AI models like ChatGPT are increasingly utilized across various applications. Among the possible purposes, many people are considering asking ChatGPT for relationship advice. However, the lack of in-depth examination of ChatGPT's response quality could be concerning when it is used for personal topics like mental health issues and intimate relationship problems. In these topics, a piece of misleading advice could cause harmful repercussions. In response to people's growing interest in using ChatGPT as a relationship advisor, our research evaluates ChatGPT's proficiency in discerning relationship advice. Specifically, we investigate its alignment with human judgements. We conducted our analysis with 13,138 Reddit posts about intimate relationship problems to examine the overall alignment. Furthermore, we investigate ChatGPT's consistency in judging intimate relationship advice by re-prompting identical queries. Our results indicate a significant disparity between ChatGPT and human judgments, with the model displaying inconsistency in its own decisions. Our findings emphasize the need for comprehensive insights into ChatGPT's mechanisms for intimacy problems and future improvements in its proficiency in helping people's relationship struggles. Haonan Hou, Kevin Leach, Yu Huang 0015 |
ICWSM | 2 |
| 2024 | Do Machines and Humans Focus on Similar Code? Exploring Explainability of Large Language Models in Code SummarizationabstractRecent language models have demonstrated proficiency in summarizing source code. However, as in many other domains of machine learning, language models of code lack sufficient explainability --- informally, we lack a formulaic or intuitive understanding of what and how models learn from code. Explainability of language models can be partially provided if, as the models learn to produce higher-quality code summaries, they also align in deeming the same code parts important as those identified by human programmers. In this paper, we report negative results from our investigation of explainability of language models in code summarization through the lens of human comprehension. We measure human focus on code using eye-tracking metrics such as fixation counts and duration in code summarization tasks. To approximate language model focus, we employ a state-of-the-art model-agnostic, black-box, perturbation-based approach, SHAP (SHapley Additive exPlanations), to identify which code tokens influence that generation of summaries. Using these settings, we find no statistically significant relationship between language models' focus and human programmers' attention. Furthermore, alignment between model and human foci in this setting does not seem to dictate the quality of the LLM-generated summaries. Our study highlights an inability to align human focus with SHAP-based model focus measures. This result calls for future investigation of multiple open questions for explainable language models for code summarization and software engineering tasks in general, including the training mechanisms of language models for code, whether there is an alignment between human and model attention on code, whether human attention can improve the development of language models, and what other model focus measures are appropriate for improving explainability. Yifan Zhang 0013, Zachary Karas, Collin McMillan, Kevin Leach, Yu Huang 0015 |
ICPC | 5 |
| 2024 | CAGE: Complementing Arm CCA with GPU Extensions
Chenxu Wang 0005, Fengwei Zhang, Yunjie Deng 0001, Kevin Leach, Jiannong Cao 0001, Zhenyu Ning, Shoumeng Yan, Zhengyu He |
NDSS | 4 |
| 2024 | Reducing Malware Analysis Overhead With CoveringsabstractThere is a substantial and growing body of malware samples that evade automated analysis and detection tools. Malware may measure fingerprints (“artifacts”) of the underlying analysis tool or environment, and change their behavior when such artifacts are detected. While analysis tools can mitigate artifacts to reduce exposure, such concealment is expensive and limits scalable automated malware analysis. However, not every sample checks for every type of artifact—analysis efficiency can be improved by mitigating only those artifacts most likely to be used by a sample. Using that insight, we proposeMimosa, a system that identifies a small set of “covering” configurations that collectively and efficiently defeat most malware samples in a corpus.Mimosaidentifies a set of configurations that maximize analysis throughput and detection accuracy while minimizing manual effort, enabling scalable automation for analyzing stealthy malware. We evaluate our approach against a benchmark of 1535 meticulously labeled stealthy malware samples. We further test our approach on an additional set of 1221 stealthy malware samples and successfully analyze nearly 99% of them using only 2 VM backends.Mimosaprovides a practical, tunable method for efficiently deploying malware analysis resources. Michael Sandborn, Zach Stoebner, Westley Weimer, Stephanie Forrest, Ryan E. Dougherty, Jules White, Kevin Leach |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2024 | Building a Lightweight Trusted Execution Environment for Arm GPUsabstractA wide range of Arm endpoints leverage integrated and discrete GPUs to accelerate computation. However, Arm GPU security has not been explored by the community. Existing work has used Trusted Execution Environments (TEEs) to address GPU security concerns on Intel-based platforms, but there are numerous architectural differences that lead to novel technical challenges in deploying TEEs for Arm GPUs. There is a need for generalizable and efficient Arm-based GPU security mechanisms. To address these problems, we presentStrongBox, the first GPU TEE for secured general computation on Arm endpoints.StrongBoxprovides an isolated execution environment by ensuring exclusive access to GPU. Our approach is based in part on a dynamic, fine-grained memory protection policy as Arm-based GPUs typically share a unified memory with the CPU. Furthermore,StrongBoxreduces runtime overhead from the redundant security introspection operations. We also design an effective defense mechanism withinsecure worldto protect the confidential GPU computation. Our design leverages the widely-deployed Arm TrustZone and generic Arm features, without hardware modification or architectural changes. We prototypeStrongBoxusing an off-the-shelf Arm Mali GPU and perform an extensive evaluation. Results show thatStrongBoxsuccessfully ensures GPU computation security with a low (4.70%–15.26%) overhead. Chenxu Wang 0005, Yunjie Deng 0001, Zhenyu Ning, Kevin Leach, Jin Li 0002, Shoumeng Yan, Zhengyu He, Jiannong Cao 0001, Fengwei Zhang |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | Hardware-Assisted Live Kernel Function Updating on Intel PlatformsabstractTraditional kernel updates such as perfective maintenance and vulnerability patching requires shutting the system down, disrupting continuous execution of applications. Enterprises and researchers have proposed various live updating techniques to patch the kernel with lower downtime to reduce the loss of useful uptime. However, existing kernel live update techniques either rely on specific support from the target OS, or are deployed in virtualized environments (i.e., systems running in virtual machines). In this paper we presentKShot, a hardware-assisted live and secure kernel function update mechanism for native operating systems. By leveraging x86 SMM and Intel SGX,KShotruns in hardware-assisted Trusted Execution Environments and updates kernel functions at the binary-level without relying on the underlying OS support. We demonstrate the applicability ofKShotby successfully patching critical kernel vulnerabilities, upgrading base kernel functions and drivers nearly instantly and transparently. Our experimental results show thatKShotincurs merely 70 microseconds downtime to update a one kilobyte binary and 18 MB memory overhead. Lei Zhou 0023, Fengwei Zhang, Kevin Leach, Xuhua Ding, Zhenyu Ning, Guojun Wang 0001, Jidong Xiao |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | On Evaluation of Document Classifiers using RVL-CDIPabstractThe RVL-CDIP benchmark is widely used for measuring performance on the task of document classification.Despite its widespread use, we reveal several undesirable characteristics of the RVL-CDIP benchmark.These include (1) substantial amounts of label noise, which we estimate to be 8.1% (ranging between 1.6% to 16.9% per document category); (2) presence of many ambiguous or multi-label documents; (3) a large overlap between test and train splits, which can inflate model performance metrics; and (4) presence of sensitive personally-identifiable information like US Social Security numbers (SSNs).We argue that there is a risk in using RVL-CDIP for benchmarking document classifiers, as its limited scope, presence of errors (state-of-the-art models now achieve accuracy error rates that are within our estimated label error rate), and lack of diversity make it less than ideal for benchmarking.We further advocate for the creation of a new document classification benchmark, and provide recommendations for what characteristics such a resource should include. Model (Reported by) Modality Accuracy Stefan Larson, Gordon Lim, Kevin Leach |
EACL | 3 |
| 2023 | Leveraging Evidence Theory to Improve Fault Localization: An Exploratory Study
Yueke Zhang, Kevin Leach, Yu Huang 0015 |
ESEM | 2 |
| 2023 | Revisiting Deep Learning for Variable Type RecoveryabstractCompiled binary executables are often the only available artifact in reverse engineering, malware analysis, and software systems maintenance. Unfortunately, the lack of semantic information like variable types makes comprehending binaries difficult. In efforts to improve the comprehensibility of binaries, researchers have recently used machine learning techniques to predict semantic information contained in the original source code. Chen et al. implemented DIRTY, a Transformer-based Encoder-Decoder architecture capable of augmenting decompiled code with variable names and types by leveraging decompiler output tokens and variable size information. Chen et al. were able to demonstrate a substantial increase in name and type extraction accuracy on Hex-Rays decompiler outputs compared to existing static analysis and AI-based techniques. We extend the original DIRTY results by re-training the DIRTY model on a dataset produced by the open-source Ghidra decompiler. Although Chen et al. concluded that Ghidra was not a suitable decompiler candidate due to its difficulty in parsing and incorporating DWARF symbols during analysis, we demonstrate that straightforward parsing of variable data generated by Ghidra results in similar retyping performance. We hope this work inspires further interest and adoption of the Ghidra decompiler for use in research projects. Kevin Cao, Kevin Leach |
ICPC | 2 |
| 2023 | Revisiting Lightweight Compiler Provenance Recovery on ARM BinariesabstractA binary’s behavior is greatly influenced by how the compiler builds its source code. Although most compiler configuration details are abstracted away during compilation, recovering them is useful for reverse engineering and program comprehension tasks on unknown binaries, such as code similarity detection. We observe that previous work has thoroughly explored this on x86-64 binaries. However, there has been limited investigation of ARM binaries, which are increasingly prevalent.In this paper, we extend previous work with a shallow-learning model that efficiently and accurately recovers compiler configuration properties for ARM binaries. We apply opcode and register-derived features, that have previously been effective on x86-64 binaries, to ARM binaries. Furthermore, we compare this work with Pizzolotto et al., a recent architecture-agnostic model that uses deep learning, whose dataset and code are available.We observe that the lightweight features are reproducible on ARM binaries. We achieve over 99% accuracy, on par with state-of-the-art deep learning approaches, while achieving a 583-times speedup during training and 3,826-times speedup during inference. Finally, we also discuss findings of overfitting that was previously undetected in prior work. Jason Kim 0007, Daniel Genkin, Kevin Leach |
ICPC | 3 |
| 2023 | A Four-Year Study of Student Contributions to OSS vs. OSS4SG with a Lightweight InterventionabstractModern software engineering practice and training increasingly rely on Open Source Software (OSS). The recent growth in demand for professional software engineers has led to increased contributions to, and usage of, OSS. However, there is limited understanding of the factors affecting how developers, and how new or student developers in particular, decide which OSS projects to contribute to, a process critical to OSS sustainability, access, adoption, and growth. To better understand OSS contributions from the developers of tomorrow, we conducted a four-year study with 1,361 students investigating the life cycle of their contributions (from project selection to pull request acceptance). During the study, we also delivered a lightweight intervention to promote the awareness of open source projects for social good (OSS4SG), OSS projects that have positive impacts in other domains. Using both quantitative and qualitative methods, we analyze student experience reports and the pull requests they submit. Compared to general OSS projects, we find significant differences in project selection (𝑝 < 0.0001, effect size = 0.84), student motivation (𝑝 < 0.01, effect size = 0.13), and increased pull-request acceptance rates for OSS4SG contributions. We also find that our intervention correlates with increased student contributions to OSS4SG (𝑝 < 0.0001, effect size = 0.38). Finally, we analyze correlations of factors such as gender or working with a partner. Our findings may help improve the experience for new developers participating in OSS4SG and the quality of their contributions. We also hope our work helps educators, project leaders, and contributors to build a mutually-beneficial framework for the future growth of OSS4SG. Zihan Fang 0001, Madeline Endres, Thomas Zimmermann 0001, Denae Ford, Westley Weimer, Kevin Leach, Yu Huang 0015 |
ESEC/SIGSOFT FSE | 6 |
| 2022 | StrongBox: A GPU TEE on Arm EndpointsabstractA wide range of Arm endpoints leverage integrated and discrete GPUs to accelerate computation such as image processing and numerical processing applications. However, in spite of these important use cases, Arm GPU security has yet to be scrutinized by the community. By exploiting vulnerabilities in the kernel, attackers can directly access sensitive data used during GPU computing, such as personally-identifiable image data in computer vision tasks. Existing work has used Trusted Execution Environments (TEEs) to address GPU security concerns on Intel-based platforms, while there are numerous architectural differences that lead to novel technical challenges in deploying TEEs for Arm GPUs. In addition, extant Arm-based GPU defenses are intended for secure machine learning, and lack generality. There is a need for generalizable and efficient Arm-based GPU security mechanisms. Yunjie Deng 0001, Chenxu Wang 0005, Shunchang Yu, Shiqing Liu, Zhenyu Ning, Kevin Leach, Jin Li 0002, Shoumeng Yan, Zhengyu He, Jiannong Cao 0001, Fengwei Zhang |
CCS | 6 |
| 2022 | START: A Framework for Trusted and Resilient Autonomous Vehicles (Practical Experience Report)abstractFrom delivering groceries and vital medical supplies to driving trucks and passenger vehicles, society is becoming increasingly reliant on autonomous vehicles (AVs), It is therefore vital that these systems be resilient to adversarial actions, perform mission-critical functions despite known and unknown vulnerabilities, and protect and repair themselves during or after operational failures and cyber-attacks. While techniques have been proposed to address individual aspects of software resilience, vulnerability assessment, automated repair, and invariant detection, there is no approach that provides end-to-end trusted and resilient mission operation and repair on AVs. In this paper, we describe our experience of building START,11Software Techniques for Automated Resilience and Trust a framework that provides increased resilience, accurate vul-nerability assessment, and trustworthy post-repair operation in autonomous vehicles. We combine techniques from binary analysis and rewriting, runtime monitoring and verification, auto-mated program repair, and invariant detection that cooperatively detect and eliminate a swath of software security vulnerabilities in cyberphysical systems. We evaluate our framework using an autonomous vehicle simulation platform, demonstrating its holistic applicability to AVs. Kevin Leach, Christopher Steven Timperley, Kevin Angstadt, Anh Nguyen-Tuong, Jason Hiser, Aaron Paulos, Partha P. Pal, Patrick Hurley, Carl Thomas, Jack W. Davidson, Stephanie Forrest, Claire Le Goues, Westley Weimer |
ISSRE | 1 |
| 2022 | Evaluating Out-of-Distribution Performance on Document Image ClassifiersabstractThe ability of a document classifier to handle inputs that are drawn from a distribution different from the training distribution is crucial for robust deployment and generalizability. The RVL-CDIP corpus is the de facto standard benchmark for document classification, yet to our knowledge all studies that use this corpus do not include evaluation on out-of-distribution documents. In this paper, we curate and release a new out-of-distribution benchmark for evaluating out-of-distribution performance for document classifiers. Our new out-of-distribution benchmark consists of two types of documents: those that are not part of any of the 16 in-domain RVL-CDIP categories (RVL-CDIP-O), and those that are one of the 16 in-domain categories yet are drawn from a distribution different from that of the original RVL-CDIP dataset (RVL-CDIP-N). While prior work on document classification for in-domain RVL-CDIP documents reports high accuracy scores, we find that these models exhibit accuracy drops of between roughly 15-30% on our new out-of-domain RVL-CDIP-N benchmark, and further struggle to distinguish between in-domain RVL-CDIP-N and out-of-domain RVL-CDIP-O inputs. Our new benchmark provides researchers with a valuable new resource for analyzing out-of-distribution performance on document classifiers. Stefan Larson, Gordon Lim, Yutong Ai, David Kuang, Kevin Leach |
NeurIPS | 5 |
| 2022 | Redwood: Using Collision Detection to Grow a Large-Scale Intent Classification DatasetabstractDialog systems must be capable of incorporating new skills via updates over time in order to reflect new use cases or deployment scenarios.Similarly, developers of such ML-driven systems need to be able to add new training data to an already-existing dataset to support these new skills.In intent classification systems, problems can arise if training data for a new skill's intent overlaps semantically with an alreadyexisting intent.We call such cases collisions.This paper introduces the task of intent collision detection between multiple datasets for the purposes of growing a system's skillset.We introduce several methods for detecting collisions, and evaluate our methods on real datasets that exhibit collisions.To highlight the need for intent collision detection, we show that model performance suffers if new data is added in such a way that does not arbitrate colliding intents.Finally, we use collision detection to construct and benchmark a new dataset, Redwood, which is composed of 451 intent categories from 13 original intent classification datasets, making it the largest publicly available intent classification benchmark. Stefan Larson, Kevin Leach |
SIGDIAL | 2 |
| 2021 | A Coprocessor-Based Introspection Framework Via Intel Management EngineabstractDuring the past decade, virtualization-based (e.g., virtual machine introspection) and hardware-assisted approaches (e.g., x86 SMM and ARM TrustZone) have been used to defend against low-level malware such as rootkits. However, these approaches either require a large Trusted Computing Base (TCB) or they must share CPU time with the operating system, disrupting normal execution. In this article, we propose an introspection framework called Nighthawk that transparently checks system integrity and monitor the runtime state of target system. Nighthawk leverages the Intel Management Engine (IME), a co-processor that runs in isolation from the main CPU. By using the IME, our approach has a minimal TCB and incurs negligible overhead on the host system on a suite of indicative benchmarks. We use Nighthawk to introspect the system software and firmware of a host system at runtime. The experimental results show that Nighthawk can detect real-world attacks against the OS, hypervisors, and System Management Mode while mitigating several classes of evasive attacks. Additionally, Nighthawk can monitor the runtime state of host system against the suspicious applications running in target machine. Lei Zhou 0023, Fengwei Zhang, Jidong Xiao, Kevin Leach, Westley Weimer, Xuhua Ding, Guojun Wang 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2021 | Toward an Objective Measure of Developers' Cognitive ActivitiesabstractUnderstanding how developers carry out different computer science activities with objective measures can help to improve productivity and guide the use and development of supporting tools in software engineering. In this article, we present two controlled experiments involving 112 students to explore multiple computing activities (code comprehension, code review, and data structure manipulations) using three different objective measures including neuroimaging (functional near-infrared spectroscopy (fNIRS) and functional magnetic resonance imaging (fMRI)) and eye tracking. By examining code review and prose review using fMRI, we find that the neural representations of programming languages vs. natural languages are distinct. We can classify which task a participant is undertaking based solely on brain activity, and those task distinctions are modulated by expertise. We leverage insights from the psychological notion of spatial ability to decode the neural representations of several fundamental data structures and their manipulations using fMRI, fNIRS, and eye tracking. We examine list, array, tree, and mental rotation tasks and find that data structure and spatial operations use the same focal regions of the brain but to different degrees: they are related but distinct neural tasks. We demonstrate best practices and describe the implication and tradeoffs between fMRI, fNIRS, eye tracking, and self-reporting for software engineering research. Zohreh Sharafi, Yu Huang 0015, Kevin Leach, Westley Weimer |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2020 | Inconsistencies in Crowdsourced Slot-Filling Annotations: A Typology and Identification MethodsabstractSlot-filling models in task-driven dialog systems rely on carefully annotated training data.However, annotations by crowd workers are often inconsistent or contain errors.Simple solutions like manually checking annotations or having multiple workers label each sample are expensive and waste effort on samples that are correct.If we can identify inconsistencies, we can focus effort where it is needed.Toward this end, we define six inconsistency types in slot-filling annotations.Using three new noisy crowd-annotated datasets, we show that a wide range of inconsistencies occur and can impact system performance if not addressed.We then introduce automatic methods of identifying inconsistencies.Experiments on our new datasets show that these methods effectively reveal inconsistencies in data, though there is further scope for improvement. Stefan Larson, Adrian Cheung, Anish Mahendran, Kevin Leach, Jonathan K. Kummerfeld |
COLING | 4 |
| 2020 | KShot: Live Kernel Patching with SMM and SGXabstractLive kernel patching is an increasingly common trend in operating system distributions, enabling dynamic updates to include new features or to fix vulnerabilities without having to reboot the system. Patching the kernel at runtime lowers downtime and reduces the loss of useful state from running applications. However, existing kernel live patching techniques (1) rely on specific support from the target operating system, and (2) admit patch failures resulting from kernel faults. We present KSHOT, a kernel live patching mechanism based on x86 SMM and Intel SGX that focuses on patching Linux kernel security vulnerabilities. Our patching processes are protected by hardware-assisted Trusted Execution Environments. We demonstrate that our technique can successfully patch vulnerable kernel functions at the binary-level without support from the underlying OS and regardless of whether the kernel patching mechanism is compromised. We demonstrate the applicability of KSHOT by successfully patching 30 critical indicative kernel vulnerabilities. Lei Zhou 0023, Fengwei Zhang, Jinghui Liao, Zhenyu Ning, Jidong Xiao, Kevin Leach, Westley Weimer, Guojun Wang 0001 |
DSN | 6 |
| 2020 | Iterative Feature Mining for Constraint-Based Data Collection to Increase Data Diversity and Model RobustnessabstractStefan Larson, Anthony Zheng, Anish Mahendran, Rishi Tekriwal, Adrian Cheung, Eric Guldan, Kevin Leach, Jonathan K. Kummerfeld. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Stefan Larson, Anthony Zheng, Anish Mahendran, Rishi Tekriwal, Adrian Cheung, Eric Guldan, Kevin Leach, Jonathan K. Kummerfeld |
EMNLP (1) | 7 |
| 2020 | Neurological divide: an fMRI study of prose and code writingabstractSoftware engineering involves writing new code or editing existing code. Recent efforts have investigated the neural processes associated with reading and comprehending code --- however, we lack a thorough understanding of the human cognitive processes underlying code writing. While prose reading and writing have been studied thoroughly, that same scrutiny has not been applied to code writing. In this paper, we leverage functional brain imaging to investigate neural representations of code writing in comparison to prose writing. We present the first human study in which participants wrote code and prose while undergoing a functional magnetic resonance imaging (fMRI) brain scan, making use of a full-sized fMRI-safe QWERTY keyboard. Ryan Krueger, Yu Huang 0015, Tyler Santander, Westley Weimer, Kevin Leach |
ICSE | 6 |
| 2020 | A Human Study of Comprehension and Code SummarizationabstractSoftware developers spend a great deal of time reading and understanding code that is poorly-documented, written by other developers, or developed using differing styles. During the past decade, researchers have investigated techniques for automatically documenting code to improve comprehensibility. In particular, recent advances in deep learning have led to sophisticated summary generation techniques that convert functions or methods to simple English strings that succinctly describe that code's behavior. However, automatic summarization techniques are assessed using internal metrics such as BLEU scores, which measure natural language properties in translational models, or ROUGE scores, which measure overlap with human-written text. Unfortunately, these metrics do not necessarily capture how machine-generated code summaries actually affect human comprehension or developer productivity. Sean Stapleton, Yashmeet Gambhir, Alexander LeClair, Zachary Eberhart, Westley Weimer, Kevin Leach, Yu Huang 0015 |
ICPC | 6 |
| 2020 | Data Query Language and Corpus Tools for Slot-Filling and Intent Classification DataabstractTypical machine learning approaches to developing task-oriented dialog systems require the collection and management of large amounts of training data, especially for the tasks of intent classification and slot-filling. Managing this data can be cumbersome without dedicated tools to help the dialog system designer understand the nature of the data. This paper presents a toolkit for analyzing slot-filling and intent classification corpora. We present a toolkit that includes (1) a new lightweight and readable data and file format for intent classification and slot-filling corpora, (2) a new query language for searching intent classification and slot-filling corpora, and (3) tools for understanding the structure and makeup for such corpora. We apply our toolkit to several well-known NLU datasets, and demonstrate that our toolkit can be used to uncover interesting and surprising insights. By releasing our toolkit to the research community, we hope to enable others to develop more robust and intelligent slot-filling and intent classification models. Stefan Larson, Eric Guldan, Kevin Leach |
LREC | 3 |
| 2020 | Biases and differences in code review using medical imaging and eye-tracking: genders, humans, and machinesabstractCode review is a critical step in modern software quality assurance, yet it is vulnerable to human biases. Previous studies have clarified the extent of the problem, particularly regarding biases against the authors of code,but no consensus understanding has emerged. Advances in medical imaging are increasingly applied to software engineering, supporting grounded neurobiological explorations of computing activities, including the review, reading, and writing of source code. In this paper, we present the results of a controlled experiment using both medical imaging and also eye tracking to investigate the neurological correlates of biases and differences between genders of humans and machines (e.g., automated program repair tools) in code review. We find that men and women conduct code reviews differently, in ways that are measurable and supported by behavioral, eye-tracking and medical imaging data. We also find biases in how humans review code as a function of its apparent author, when controlling for code quality. In addition to advancing our fundamental understanding of how cognitive biases relate to the code review process, the results may inform subsequent training and tool design to reduce bias. Yu Huang 0015, Kevin Leach, Zohreh Sharafi, Nicholas McKay, Tyler Santander, Westley Weimer |
ESEC/SIGSOFT FSE | 2 |
| 2019 | An Evaluation Dataset for Intent Classification and Out-of-Scope PredictionabstractStefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang, Jason Mars. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Stefan Larson, Anish Mahendran, Joseph Peper, Christopher Clarke, Andrew Lee 0001, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael Laurenzano, Lingjia Tang, Jason Mars |
EMNLP/IJCNLP (1) | 8 |
| 2019 | Nighthawk: Transparent System Introspection from Ring -3
Lei Zhou 0023, Jidong Xiao, Kevin Leach, Westley Weimer, Fengwei Zhang, Guojun Wang 0001 |
ESORICS (2) | 3 |
| 2019 | Distilling neural representations of data structure manipulation using fMRI and fNIRSabstractData structures permeate many aspects of software engineering, but their associated human cognitive processes are not thoroughly understood. We leverage medical imaging and insights from the psychological notion of spatial ability to decode the neural representations of several fundamental data structures and their manipulations. In a human study involving 76 participants, we examine list, array, tree, and mental rotation tasks using both functional near-infrared spectroscopy (fNIRS) and functional magnetic resonance imaging (fMRI). We find a nuanced relationship: data structure and spatial operations use the same focal regions of the brain but to different degrees. They are related but distinct neural tasks. In addition, more difficult computer science problems induce higher cognitive load than do problems of pure spatial reasoning. Finally, while fNIRS is less expensive and more permissive, there are some computing-relevant brain regions that only fMRI can reach. Yu Huang 0015, Ryan Krueger, Tyler Santander, Xiao-Su Hu, Kevin Leach, Westley Weimer |
ICSE | 6 |
| 2018 | Towards Transparent DebuggingabstractTraditional malware analysis relies on virtualization or emulation technology to run samples in a confined environment, and to analyze malicious activities by instrumenting code execution. However, virtual machines and emulators inevitably create artifacts in the execution environment, making these approaches vulnerable to detection or subversion. In this paper, we present MALT, a debugging framework that employs System Management Mode, a CPU mode in the x86 architecture, to transparently study armored malware. MALT does not depend on virtualization or emulation and thus is immune to threats targeting such environments. Our approach reduces the attack surface at the software level, and advances state-of-the-art debugging transparency. MALT embodies various debugging functions, including register/memory accesses, breakpoints, and seven stepping modes. Additionally, MALT restores the system to a clean state after a debugging session. We implemented a prototype of MALT on two physical machines, and we conducted experiments by testing an array of existing anti-virtualization, anti-emulation, and packing techniques against MALT. The experimental results show that our prototype remains transparent and undetected against the samples. Furthermore, debugging and restoration introduce moderate but manageable overheads on both Windows and Linux platforms. Fengwei Zhang, Kevin Leach, Angelos Stavrou, Haining Wang 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2017 | Scotch: Combining Software Guard Extensions and System Management Mode to Monitor Cloud Resource Usage
Kevin Leach, Fengwei Zhang, Westley Weimer |
RAID | 1 |
| 2017 | Daehr: A Discriminant Analysis Framework for Electronic Health Record Data and an Application to Early Detection of Mental Health DisordersabstractElectronic health records (EHR) provide a rich source of temporal data that present a unique opportunity to characterize disease patterns and risk of imminent disease. While many data-mining tools have been adopted for EHR-based disease early detection, linear discriminant analysis (LDA) is one of the most commonly used statistical methods. However, it is difficult to train an accurate LDA model for early disease diagnosis when too few patients are known to have the target disease. Furthermore, EHR data are heterogeneous with significant noise. In such cases, the covariance matrices used in LDA are usually singular and estimated with a large variance. This article presents Daehr , an extension of the LDA framework using electronic health record data to address these issues. Beyond existing LDA analyzers, we propose Daehr to (1) eliminate the data noise caused by the manual encoding of EHR data and (2) lower the variance of parameter (covariance matrices) estimation for LDA models when only a few patients’ EHR are available for training. To achieve these two goals, we designed an iterative algorithm to improve the covariance matrix estimation with embedded data-noise/parameter-variance reduction for LDA. We evaluated Daehr extensively using the College Health Surveillance Network, a large, real-world EHR dataset. Specifically, our experiments compared the performance of LDA to three baselines (i.e., LDA and its derivatives) in identifying college students at high risk for mental health disorders from 23 U.S. universities. Experimental results demonstrate Daehr significantly outperforms the three baselines by achieving 1.4%--19.4% higher accuracy and a 7.5%--43.5% higher F1-score. Haoyi Xiong, Jinghe Zhang, Yu Huang 0015, Kevin Leach, Laura E. Barnes |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2016 | Assessing social anxiety using gps trajectories and point-of-interest dataabstractMental health problems are highly prevalent and appear to be increasing in frequency and severity among the college student population. The upsurge in mobile and wearable wireless technologies capable of intense, longitudinal tracking of individuals, provide valuable opportunities to examine temporal patterns and dynamic interactions of key variables in mental health research. In this paper, we present a feasibility study leveraging non-invasive mobile sensing technology to passively assess college students' social anxiety, one of the most common disorders in the college student population. We have first developed a smartphone application to continuously track GPS locations of college students, then we built an analytic infrastructure to collect the GPS trajectories and finally we analyzed student behaviors (e.g. studying or staying at home) using Point-Of-Interest (POI). The whole framework supports intense, longitudinal, dynamic tracking of college students to evaluate how their anxiety and behaviors change in the college campus environment. The collected data provides critical information about how students' social anxiety levels and their mobility patterns are correlated. Our primary analysis based on 18 college students demonstrated that social anxiety level is significantly correlated with places students' visited and location transitions. Yu Huang 0015, Haoyi Xiong, Kevin Leach, Philip Chow, Karl C. Fua, Bethany A. Teachman, Laura E. Barnes |
UbiComp | 3 |
| 2016 | LO-PHI: Low-Observable Physical Host Instrumentation for Malware Analysis
Chad Spensky, Hongyi Hu, Kevin Leach |
NDSS | 3 |
| 2016 | Towards Transparent IntrospectionabstractThere is a growing need for the dynamic analysis of sensitive systems thatdo not support traditional debugging or emulation environments. Analysiscan alter program behavior, necessitating transparency. For example, asthe cat and mouse game between malware authors and malware analystsprogresses, malicious software can increasingly detect and confounddebuggers. Analysts must understand variable values, stack traces, andfactors influencing dynamic behavior, but recent malware samples leverageany piece of information or artifact available that signals the presence ofa debugger or emulator. In this work, we advance the state-of-the-art for transparent programanalysis by introducing a low-artifact introspection technique. Ourapproach uses hardware-assisted live memory snapshots ofprocess execution on native targets (e.g., x86 processors), coupledwith static reasoning about programs. We produce high-fidelity data and control flow information with minimaldetectable artifacts that could influence benign subject behavior or beleveraged for anti-analysis. We evaluate our system using two hardwareimplementations (x86-supported System Management Mode and PCI-basedSlotScreamer devices) and two software configurations (benign and evasiveprograms). We also analyze the theoretical and practical limitations of our technique. We discuss an expert case study in which we apply our technique to amalware reverse engineering task. Finally, we present results of a human study in which 30 participantsperformed debugging tasks using information provided by our approach, ourtool was as useful as a gdb baseline, but applies transparently. Our dynamic analysis approach permitstransparent introspection to access previously-unavailable informationabout a process's internal state with minimal instrumentation artifacts. Kevin Leach, Chad Spensky, Westley Weimer, Fengwei Zhang |
SANER | 1 |
| 2015 | M-SEQ: Early detection of anxiety and depression via temporal orders of diagnoses in electronic health dataabstractAccording to a 2014 Spring American College Health Association Survey, almost 50% of college students reported feeling things were hopeless and that it was difficult to function within the last 12 months. More than 80% reported feeling overwhelmed and exhausted by their responsibilities. This critical subpopulation of Americans is facing significant levels of mental health disorders, challenging colleges to provide accessible and high quality behavioral health care. However, psychiatric disorders are frequently unrecognized in primary care settings, posing physical, emotional, economic, and social burdens to patients and others. Towards the goal of earlier identification and treatment of mental health disorders, this paper proposes M-SEQ, an early detection framework for anxiety/depression using electronic health data from primary care visit sequences. Specifically, compared to existing methods that predict a future disease state using frequency of diagnoses in a patient's medical history, we hypothesize that future disease might also be correlated with the temporal orders of diagnoses. Thus, M-SEQ first discovers a set of diagnosis codes that are discriminative of anxiety/depression, and then extracts each diagnosis pair from each patient's health record to represent the temporal orders of diagnoses. Further, it incorporates the extracted temporal order information with the existing representation to predict whether a patient is at risk of anxiety/depression. We evaluate M-SEQ using the electronic health record (EHR) data of 213,112 college students from 10 schools participating in the College Health Surveillance Network (CHSN) from January 1, 2011 through December 31, 2014. The experimental results shows that our framework can detect a future diagnosis of anxiety and depression based on the primary care visit data up to 3 months in advance, with approximately 1%-4.5% higher accuracy, compared to baseline methods using frequency of diagnoses. Jinghe Zhang, Haoyi Xiong, Yu Huang 0015, Kevin Leach, Laura E. Barnes |
IEEE BigData | 5 |
| 2015 | TrustLogin: Securing Password-Login on Commodity Operating SystemsabstractWith the increasing prevalence of Web 2.0 and cloud computing, password-based logins play an increasingly important role on user-end systems. We use passwords to authenticate ourselves to countless applications and services. However, login credentials can be easily stolen by attackers. In this paper, we present a framework, TrustLogin, to secure password-based logins on commodity operating systems. TrustLogin leverages System Management Mode to protect the login credentials from malware even when OS is compromised. TrustLogin does not modify any system software in either client or server and is transparent to users, applications, and servers. We conduct two study cases of the framework on legacy and secure applications, and the experimental results demonstrate that TrustLogin is able to protect login credentials from real-world keyloggers on Windows and Linux platforms. TrustLogin is robust against spoofing attacks. Moreover, the experimental results also show TrustLogin introduces a low overhead with the tested applications. Fengwei Zhang, Kevin Leach, Haining Wang 0001, Angelos Stavrou |
AsiaCCS | 2 |
| 2015 | Using Hardware Features for Increased Debugging TransparencyabstractWith the rapid proliferation of malware attacks on the Internet, understanding these malicious behaviors plays a critical role in crafting effective defense. Advanced malware analysis relies on virtualization or emulation technology to run samples in a confined environment, and to analyze malicious activities by instrumenting code execution. However, virtual machines and emulators inevitably create artifacts in the execution environment, making these approaches vulnerable to detection or subversion. In this paper, we present MALT, a debugging framework that employs System Management Mode, a CPU mode in the x86 architecture, to transparently study armored malware. MALT does not depend on virtualization or emulation and thus is immune to threats targeting such environments. Our approach reduces the attack surface at the software level, and advances state-of-the-art debugging transparency. MALT embodies various debugging functions, including register/memory accesses, breakpoints, and four stepping modes. We implemented a prototype of MALT on two physical machines, and we conducted experiments by testing an array of existing anti-virtualization, anti-emulation, and packing techniques against MALT. The experimental results show that our prototype remains transparent and undetected against the samples. Furthermore, our prototype of MALT introduces moderate but manageable overheads on both Windows and Linux platforms. Fengwei Zhang, Kevin Leach, Angelos Stavrou, Haining Wang 0001, Kun Sun 0001 |
IEEE Symposium on Security and Privacy | 2 |
| 2014 | A Framework to Secure Peripherals at Runtime
Fengwei Zhang, Haining Wang 0001, Kevin Leach, Angelos Stavrou |
ESORICS (1) | 3 |
| 2013 | SPECTRE: A dependable introspection framework via System Management ModeabstractVirtual Machine Introspection (VMI) systems have been widely adopted for malware detection and analysis. VMI systems use hypervisor technology for system introspection and to expose malicious activity. However, recent malware can detect the presence of virtualization or corrupt the hypervisor state thus avoiding detection. We introduce SPECTRE, a hardware-assisted dependability framework that leverages System Management Mode (SMM) to inspect the state of a system. Contrary to VMI, our trusted code base is limited to BIOS and the SMM implementations. SPECTRE is capable of transparently and quickly examining all layers of running system code including a hypervisor, the OS, and user level applications. We demonstrate several use cases of SPECTRE including heap spray, heap overflow, and rootkit detection using real-world attacks on Windows and Linux platforms. In our experiments, full inspection with SPECTRE is 100 times faster than similar VMI systems because there is no performance overhead due to virtualization. Fengwei Zhang, Kevin Leach, Kun Sun 0001, Angelos Stavrou |
DSN | 2 |